The program improved achievement; AI’s contribution could not be isolated.
This two-year Tennessee trial involved 18 middle schools and 53 school–grade groups. Its preregistered primary analysis used 6,902 student-period observations over five assessment periods. TCAP results remained pending.
What was compared
School–grade groups randomly assigned
Khan Academy + Khanmigo
Usual remediation
Estimated assignment effect
+1.26
MAP mathematics national percentile-rank points per assessment period
Standard error: 0.60
Whisker: ±1 standard error; not a confidence interval.
0: no difference
A few reading rules matter here. A national percentile rank describes a position relative to other students; it is not a percentage increase in marks. A student-period is also not a unique student: the same person can contribute observations in several assessment periods. The standard error indicates estimation uncertainty. A positive estimate does not mean every child improved by the same amount.
What the development team measures
A separate AIED paper by Khanmigo’s development team shows the intermediate measures used in product testing. For exercises, its primary metric is correctness on the next item attempted without further AI help in the same exercise and session. For open-ended tutoring, it measures cognitive engagement, such as explaining and reasoning. Turns per conversation are a separate, secondary measure. The team describes live experiments typically lasting one or two weeks. This is a report of product-development methods, not an independent replication of the school trial’s learning results. Development-team evaluation paper
Connecting these evaluations requires separating the decisions they inform. A school buyer wants to know whether the complete instructional arrangement improves on the existing use of money and time. A developer wants to know whether changing an explanation changes the next attempt. Both questions matter, but success on one cannot substitute for an answer to the other. Solving the next question and solving a similar question a week later require different checks.
An evaluation plan should also state the decisions its results will inform. Would a school proceed if next-question accuracy improved but a month-later assessment showed no difference? If results were promising only under one teacher’s supervision, what would need to change before expansion? Setting those decision rules first can prevent a purchasing decision from following whichever numbers a product makes easiest to display.
The same questions apply to new features
The trial ended in May 2026. Subsequently, on August 27, Google announced Gemini-powered interactive diagrams in Khanmigo, including explanations that respond when a student moves a vertex or line segment. That trial cannot establish this new feature’s effects. The announcement reported early signs of better engagement and subsequent answers, without presenting a quantified long-term learning effect or comparison design. Google’s product announcement
V’s view
If I were reviewing a school purchase, I would ask for an evaluation plan before settling for a polished demonstration. I would want a comparison group continuing the existing instruction, and a design that can also vary AI support while keeping the problem bank and instructional time comparable. I would record who intervenes when a student struggles and how much additional teacher time the arrangement requires.
The final check should happen outside the chat window: unfamiliar questions with help switched off, followed by a delayed assessment. I would look at the results for the originally defined student group, rather than a selection of enthusiastic users. The report I would want reaches beyond how much conversation happened to what students can now do independently, for whom, and under which teaching conditions.