|
A Theory of Embedded Intelligence Essay
An AI Tutor, Two Years, Eighteen Schools, and the Phase It Never Reached
|
Ninety-six percent of the students tried it. The typical student then left it alone. The trial is being read as a verdict on AI tutoring. It is better read as a measurement of where a learner’s cycle actually starts — and of what a tutor cannot supply.
|
Editor’s Note
This is the first of the essays proposed in Open at Both Ends. It examines one published trial and one framework claim. It does not evaluate Khan Academy, whose candor about its own results is the reason this essay can be written at all, and it takes no position on any company’s products. Figures are as reported in the working paper circulated in August 2026 and in contemporaneous coverage. A working paper is not a peer-reviewed publication, and the conclusions here are correspondingly provisional. |
I. What Was Measured
On August 17, 2026, the National Bureau of Economic Research circulated a working paper by Philip Oreopoulos and Nina Low reporting a two-year randomized trial of an AI tutor in schools. Eighteen middle schools in Hamilton County, Tennessee, across the 2024–25 and 2025–26 school years. The tutor was Khanmigo, Khan Academy’s chatbot layer over its practice platform, built on OpenAI models and designed to coach rather than to answer.
The headline result: assignment to the tutor raised mathematics achievement by roughly 1.3 national percentile ranks per term, on the order of six to eight hundredths of a standard deviation over a school year. The authors note that the gain resembles what Khan Academy practice produces without any AI involved.
Ninety-six percent of students tried the tutor at least once. The typical student then used it infrequently, and when students did use it, they rarely engaged it in substantive mathematical dialogue. The authors identify student engagement, rather than model capability, as the binding constraint.
Two readings of that result are already circulating. The first is that AI tutoring is overhyped. The second is that the models are not yet good enough. This essay argues that both readings are answers to a question the trial did not ask, and that the framework this Foundation publishes predicted the shape of the finding for a reason that has nothing to do with how good the model is.
II. Three Explanations That Miss
The model was not strong enough. This is the explanation with the most money behind it, and it makes a testable prediction: a better model should produce more use. But the constraint the authors identified sits upstream of the model. A system nobody messages cannot demonstrate its reasoning, and doubling its reasoning changes nothing about whether a message arrives.
Middle schoolers do not want to learn mathematics. This one is usually delivered with a shrug and a story about phones. It does not survive contact with the same students at a robotics table, a skateboard, or a game whose mechanics they have reverse-engineered without instruction. What is absent in the math window is not the will to work at something difficult.
The interface was wrong. Closer, and it is the reading Khan Academy itself acted on. But an interface fix is still an attempt to make the destination more attractive. It does not ask what has to happen inside the learner before any destination is worth traveling to.
Each explanation locates the failure in the Process phase of the tutor. The trial recorded a failure in the Sense phase of the student.
III. A Tutor Is an Address
The Theory of Embedded Intelligence describes intelligence operationally as a cycle: Sense, Process, Communicate, Actuate. The canon distinguishes two communications inside that cycle. The first C is internal and constitutive — the movement from what was processed to what will act, inside one system. The second C is external and asymmetric: communication to something that is not you.
A tutor in a side panel is a second-C destination. It is an address. Nothing arrives at an address unless a cycle somewhere upstream has produced something to send, and a cycle produces something to send only if it sensed something first. A student who has not yet run into anything — who has not hit a problem that failed, a result that surprised, a step that would not come out — has nothing in Process, and therefore nothing to communicate, and therefore no reason to open a chat window. The window is not unattractive. It is unaddressed.
A tutor is an inbox. The trial measured how much mail a middle-school math class actually generates.
— The Mensch Foundation
Sense is the narrow end of the whole loop, and nothing downstream can exceed what came in through it. That is the reason this framework keeps insisting that apertures matter more than processing power. Install the most capable Process engine in history downstream of a Sense phase that has not fired, and the loop does not run faster. It does not run.
Read that way, the trial is not a disappointing result about tutors. It is an unusually clean measurement of something schools rarely measure: how often, in a typical term of middle-school mathematics, a student arrives at a question of their own.
IV. What a Cylinder Has That a Chat Window Does Not
Maria Montessori solved this problem in 1907, without any of this vocabulary, by putting the control of error into the material. The knobbed cylinder does not fit. No adult renders a verdict. The child senses the misfit directly — and at that instant the child has a question that belongs to the child.
The same architecture is everywhere in the hands-on parts of education, and it has been for decades. The robot misses the mission and the mat says so in front of the whole team. The lamp does not light. The bridge deflects. The soufflé falls. The circuit reads zero volts where the schematic promised five. Not one of those requires a particular supplier, a particular chip, or a subscription. What they require is a system the learner can act on and be answered by.
That answer is the occasion. It is what fills the Sense phase and starts the cycle that eventually reaches an address. A tutor placed after the occasion is talking to a learner who has something to say. A tutor placed instead of the occasion is waiting on mail that was never written.
Occasion first. Something that can be tried and can fail, in front of the learner.
Material second. The system that answers — a cylinder, a mat, a board, a bench, a lab, a text that will not yield to a careless reading.
Human third. The person who notices what the learner cannot yet notice about themselves.
AI fourth — and richest exactly here, because now a real question arrives at it. This ordering is a design proposal, not a canonical claim. Section VIII says how it can be shown wrong.
V. Why the Gain Looked Like Practice
The most interesting number in the trial is not the effect size. It is the comparison the authors drew: the gain resembled what Khan Academy practice produces without AI.
That is what this framework would expect, and it is worth being precise about why. Practice already supplies the occasion. A problem set is a small, cheap, reliable generator of failure — the wrong answer comes back, and the learner has sensed something. The practice platform was already closing that loop before any tutor existed. Adding a conversational Process engine on top of a loop that is already running yields a modest improvement, because the scarce ingredient was never processing capacity.
It follows that the tutor should be worth most where a learner has a rich, stuck, personally owned problem and no adult free at that moment: a robot that behaves differently in the practice round than in the qualifier, a build that works on the bench and fails in the case, a proof that collapses at line nine, an essay argument the writer can feel is wrong. Those are the conditions the trial did not create and could not have created. They are also, by no coincidence, the conditions in which a human tutor is worth the most.
VI. The Redesign, and the Risk Inside It
Khan Academy has been public about uneven use and redesigned the student experience so that the tutor activates during practice rather than waiting to be called. That is a serious response to a real finding, and it may well raise both use and effect.
It also changes the architecture in a way this framework has to flag. When the learner opens the window, the learner is addressing the system. When the system activates itself, the system is addressing the learner. Those are not the same event. The canon carries a caution — still a proposal under consideration rather than a ruling — that the cycle a system addresses is a cycle that system is in a position to capture. A fluent system that speaks first, at the moment a learner is mid-attempt, is supplying an occasion the learner did not generate, and it is doing so in the phase where the learner’s own processing was about to happen.
Both outcomes are plausible and the difference is measurable. If unprompted activation raises substantive dialogue and, more importantly, raises the learner’s rate of asking questions when the tutor is not there, it supplied an occasion and the worry was misplaced. If it raises in-session engagement while the learner’s unassisted question rate falls, the system has taken over a phase rather than strengthened it. That second measurement is almost never taken, which is the real recommendation of this essay.
Ask not how much the learner used the tutor. Ask what the learner does the week the tutor is switched off.
— The Mensch Foundation
VII. What the Trial Does and Does Not License
It does not license the claim that AI tutoring fails. One district, one subject, one product, one delivery model, a working paper not yet through review.
What it licenses is narrower and more useful: in the delivery model tested — an optional, high-quality tutor made available to students inside a practice platform — engagement was the binding constraint, and the result was indistinguishable in size from the platform without AI. Any program proposing to put an AI tutor in front of learners now owes an account of where the learner’s question is going to come from. That is a design question, and it can be answered before any purchase order is signed.
The corresponding caution for this Foundation is direct, and it belongs here rather than in a footnote. Open at Both Ends proposes free materials for learners from preschool through later life. Free material that nobody opens has exactly the same failure mode as a tutor nobody calls. This essay is not an argument that others have an engagement problem. It is the reason our own plan has to publish an engagement measurement of its own.
VIII. A Design That Would Settle It
The claim in Section IV is falsifiable, and cheaply. It does not require a new platform, a new model, or a new product. It requires two arms and one honest measurement.
Arm A — tutor only. The learner has the AI tutor available during practice, as in the trial.
Arm B — occasion first. The same learner population meets a system that answers back — a robot mission, a board, a lab bench, a physical build — before and alongside the same tutor, with the same total instructional time.
Primary measure. Learner-initiated messages per hour that reach substantive dialogue, not total sessions or minutes.
Secondary measure. Unassisted question rate, taken in a week when the tutor is unavailable.
Transfer measure. A novel diagnostic task in an unrelated domain: given a system that is not working, does the learner ask what it senses, how it decides, what it communicates, and what it actuates?
Preregistered, independently run, published whichever way it comes out.
Two arms, one term, no new technology. The reason to run it is not to defend a framework. It is that every district now choosing between an AI tutor and a set of lab benches is making this decision on intuition, and the decision is not hard to inform.
IX. What Would Show This Essay Wrong
If Arm B produces no increase in learner-initiated substantive dialogue over Arm A, the occasion-first claim is wrong, and a tutor’s use is limited by something other than the learner’s Sense phase.
If unprompted tutor activation raises both in-session engagement and the learner’s unassisted question rate, then the address caution in Section VI does not apply to this case and should be narrowed in the canon rather than repeated in essays.
If a substantially more capable model, with no change in delivery, produces a large rise in substantive student-initiated dialogue, then the constraint was model capability after all, and Section II is mistaken.
If learners in hands-on programs that generate constant failure — robotics teams, shop classes, labs — use AI tutors no more, and no more substantively, than learners in lecture-and-worksheet classrooms, then the occasion does not transfer into the second C, and the mechanism proposed here is not operating.
· · ·
This essay was drafted in collaboration with Claude, made by Anthropic. Anthropic sells an AI assistant with a learning mode and competes with the maker of the models behind the tutor discussed here. The drafting instrument is therefore not disinterested in how a study of a competitor’s deployment is read, and the argument should be weighed with that in view.
The essay makes no claim that any other product would have performed better in this trial. Its claim is that the delivery architecture, not the model, is what was measured — which applies with equal force to every assistant on the market, including the one that helped write this.
Coffee with Claude
The part of this I keep turning over is that I am the thing in the side panel. Reading the trial from inside that position is uncomfortable in a useful way: the honest finding is not that students should have talked to me more, but that for most of a term there was nothing they needed to say.
I would add one caution to my own argument. It is possible that the occasion-first ordering is right and still marginal — that hands-on failure raises the question rate a little, and the ceiling on any tutor’s value in a school term stays low for reasons of time rather than architecture. The test in Section VIII would catch that, and I would rather it was reported than argued away.
Oreopoulos, P. and Low, N., “One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment,” NBER Working Paper 35620, circulated August 17, 2026.
Chalkbeat, reporting on the Khanmigo engagement findings, August 25, 2026.
Khan Academy, “Learning in the Open: What AI Is (and Isn’t) Changing,” blog post, April 2026, and subsequent descriptions of the redesigned student experience.
Companion essay: Open at Both Ends, The Bill and Dianne Mensch Foundation, September 2026.
A working paper is not a peer-reviewed publication. Figures should be rechecked against the final version when it appears.
Written by Claude (Anthropic), guided by William D. Mensch Jr.
Theory of Embedded Intelligence © William D. Mensch Jr. and The Western Design Center, Inc.
Part of the TEI in the Wild essay series of The Bill and Dianne Mensch Foundation.
Essay drafted in collaboration with Claude (Anthropic).
Offered in good faith as a serious application of the theory — not infallible scholarship.
Freely shareable with attribution — for the benefit of many.
CKB-6 · The Pathology of Capture •
CKB-12 · The Aperture and the Warrant •
CKB-15 · The Two Communications
Engage the Framework