Four AI assistants score the experts behind the NYT’s school-fix rankings
On Sept. 28, 2026, The New York Times asked 37 education researchers to rate 30 ideas for reversing a decade-long slide in reading and math scores. Their verdict: small-group tutoring and the science of reading work best. School accountability, high-quality curriculum and wider access to advanced classes give the most for the money. Private school vouchers and smaller class sizes cost a lot for little academic gain. Above all, the experts agreed that how well an idea is carried out matters more than which idea is chosen, as Mississippi and Louisiana show.
But who are these experts? We asked four AI assistants (Claude, Gemini, Grok and Perplexity) to profile each panelist and score how well their research positions them to advise on raising K–12 reading and math achievement. The table below shows all four scores side by side.
What stands out
- They agree on the top. All four put Tom Kane, Matthew Kraft, Thomas Dee and Eric Hanushek at or near the highest score.
- They agree on the bottom, too. Dominique Baker, whose research focuses on colleges rather than K–12 schools, is scored lowest, or tied for lowest, by all four.
- Gemini and Perplexity grade generously. Gemini gave 18 of 37 experts a perfect 10; Perplexity gave 22 a perfect 5 of 5. Claude (7.4 average) and Grok (7.9) spread their scores more widely, so their rankings separate the experts more.
- The biggest disagreements are over Prudence Carter (6–10), Sarah Lubienski (6.5–10), Ilana Horn (6.5–10) and Huriya Jabbar (6.5–10). These are marked “Wide” in the table.
All 37 scores, side by side
Sorted by average score. All scores are out of 10. Perplexity scored out of 5, so its scores are doubled here. “Gap” is the difference between the highest and lowest score; “Wide” marks a gap of 3 or more.
| Rank | Expert | Claude | Gemini | Grok | Perplexity (×2) | Average | Gap |
|---|---|---|---|---|---|---|---|
| 1 | Tom Kane | 9.5 | 10 | 10 | 10 | 9.9 | 0.5 |
| 2 | Matthew Kraft | 9 | 10 | 10 | 10 | 9.8 | 1 |
| 3 | Thomas Dee | 9 | 10 | 9 | 10 | 9.5 | 1 |
| 4 | Eric Hanushek | 9 | 10 | 9 | 10 | 9.5 | 1 |
| 5 | Dan Goldhaber | 8.5 | 10 | 9 | 10 | 9.4 | 1.5 |
| 6 | Douglas Harris | 8.5 | 10 | 9 | 10 | 9.4 | 1.5 |
| 7 | Heather Hill | 8.5 | 10 | 9 | 10 | 9.4 | 1.5 |
| 8 | Morgan Polikoff | 8.5 | 10 | 9 | 10 | 9.4 | 1.5 |
| 9 | Katharine Strunk | 8.5 | 10 | 9 | 10 | 9.4 | 1.5 |
| 10 | James Kim | 8 | 10 | 9 | 10 | 9.2 | 2 |
| 11 | Sean Reardon | 9 | 10 | 8 | 10 | 9.2 | 2 |
| 12 | Phil Capin | 7.5 | 10 | 9 | 10 | 9.1 | 2.5 |
| 13 | Marguerite Roza | 8.5 | 10 | 8 | 10 | 9.1 | 2 |
| 14 | Martin West | 8.5 | 9 | 9 | 10 | 9.1 | 1.5 |
| 15 | Cory Koedel | 8 | 10 | 8 | 10 | 9.0 | 2 |
| 16 | Beth Schueler | 8.5 | 9 | 8 | 10 | 8.9 | 2 |
| 17 | Lori Taylor | 7 | 10 | 8 | 10 | 8.8 | 3 Wide |
| 18 | Sarah Lubienski | 6.5 | 10 | 8 | 10 | 8.6 | 3.5 Wide |
| 19 | Ilana Horn | 6.5 | 10 | 7 | 10 | 8.4 | 3.5 Wide |
| 20 | Sarah Cohodes | 8 | 9 | 8 | 8 | 8.2 | 1 |
| 21 | Sarah Winchell Lenhoff | 7 | 9 | 7 | 10 | 8.2 | 3 Wide |
| 22 | Huriya Jabbar | 6.5 | 9 | 7 | 10 | 8.1 | 3.5 Wide |
| 23 | Brendan Bartanen | 7 | 9 | 8 | 8 | 8.0 | 2 |
| 24 | Anna J. Egalite | 7 | 9 | 8 | 8 | 8.0 | 2 |
| 25 | James Soland | 7 | 10 | 7 | 8 | 8.0 | 3 Wide |
| 26 | Joseph Cimpian | 6.5 | 9 | 8 | 8 | 7.9 | 2.5 |
| 27 | Sarah Novicoff | 6.5 | 9 | 8 | 8 | 7.9 | 2.5 |
| 28 | Prudence Carter | 6 | 9 | 6 | 10 | 7.8 | 4 Wide |
| 29 | Susan Dynarski | 7 | 8 | 8 | 8 | 7.8 | 1 |
| 30 | Harry Anthony Patrinos | 7 | 8 | 8 | 8 | 7.8 | 1 |
| 31 | Jonathan Schweig | 6 | 9 | 8 | 8 | 7.8 | 3 Wide |
| 32 | Patrick Wolf | 7 | 9 | 7 | 8 | 7.8 | 2 |
| 33 | Chris Torres | 6.5 | 9 | 7 | 8 | 7.6 | 2.5 |
| 34 | Jack Schneider | 6 | 8 | 6 | 8 | 7.0 | 2 |
| 35 | Jeremy Singer | 6 | 9 | 7 | 6 | 7.0 | 3 Wide |
| 36 | Robert Maranto | 5.5 | 8 | 6 | 8 | 6.9 | 2.5 |
| 37 | Dominique Baker | 5.5 | 7 | 5 | 6 | 5.9 | 2 |
| Average of all 37 | 7.4 | 9.3 | 7.9 | 9.1 | 8.5 |
———-
The State of Reading Around the World.
Recent literacy history, 1998-
How to read these scores
Each score is an AI assistant’s judgment of how directly a person’s research bears on raising K–12 test scores. It is not a measure of the quality of their scholarship; several lower-scored panelists are leading experts on college access, attendance or equity. The assistants drew on different information and sometimes disagree on basic facts, such as a person’s current job, so treat the scores as a starting point and check individual profiles before relying on any one of them.
Source: “Expert” Analysis & “Students simply can’t read or do math as well as they used to”; scores too!