A Theory of Embedded Intelligence Essay
What an agent-native research format gets right — and what has to sit underneath it

Thirty-seven researchers have proposed replacing the scientific paper with something a machine can execute. They are right about what the paper throws away. The open question is what has to sit underneath the thing that replaces it.

In May, thirty-seven researchers from roughly two dozen universities and companies posted a paper on arXiv arguing that scientists should stop writing papers. On the fifth of August, IEEE Spectrum ran an interview with the lead author, Jiachen Liu, who finished her doctorate at Michigan in 2025 and has since cofounded a lab in Palo Alto to build the thing. The paper is titled The Last Human-Written Paper. What it proposes is called an Agent-Native Research Artifact, ARA for short: a research package in four layers — the scientific logic, executable code with full specification, an exploration graph that preserves the branches that failed, and an evidence layer grounding every claim in raw output. The paper itself is published in that form as a demonstration.

My first reaction was not the one I expected to have.

I have spent fifty years building machines whose published specification was their actual behavior, and I have spent the last several years arguing that the industry stopped doing that at precisely the moment it began to matter most. A proposal to write research for machines instead of for people looked, from a distance, like something I would want to push back on. Up close it is not that. It is a correct diagnosis of two real failures, made by someone who has never heard of this framework and arrived at one of its load-bearing rules anyway, followed by a fix that is right in shape and unfinished in depth. I would rather say so plainly than score points off a young engineer who got the hard part right.

So this essay does two things. It says where she is correct, in the framework’s own terms, and it says what has to be built underneath her proposal before the proposal can carry the weight she wants it to carry.

I. What a Paper Is, in This Framework

The Theory of Embedded Intelligence holds that intelligence, at any scale and in any substrate, runs a cycle: Sense, Process, Communicate, Actuate. But the framework distinguishes two structurally different communications, and that distinction is what makes this analysis possible. The essay The First C sets it out; the canon settles it in The Two Communications.

The first C is internal, reflexive, and constitutive — the signal traffic among an entity’s own components, in which the entity is both sender and receiver and the loop closes at home. It is not a metaphor borrowed from engineering. Internal communication is what assembly consists of. The second C is external, elective, and asymmetric: what occurs when a second embedded intelligence receives what a first has actuated.

That second transaction has a formal shape in the canon, and it is not one cycle handing work to another cycle’s Sense phase. It is SPCA-RCA. The sending entity’s four phases close on themselves, and the Actuate is the emission — the paper leaving the lab. The receiving entity then runs three phases of its own: Receive, Comprehend, Actuate. Receive is not Sense, because the environment has no intentions and a transmission does. Comprehend is not Process, because comprehending carries a success condition referred back to the sender — a rendering congruent enough to act on without harm.

A scientific paper is therefore the Actuate-phase emission of one research cycle, and everything that happens to it afterward happens inside somebody else’s Receive and Comprehend. Three hundred and fifty years is a long run for any format, and the reason it lasted is that it solved that transfer well enough. Before it, researchers hid their work so no one would take it. After it: archives, peer review, and the compounding that made modern science possible. Liu calls this a pivot point and she is right to. Formats that solve second-C problems are civilizational infrastructure, and they are rarer than we think.

But a format engineered around a constraint carries that constraint’s shape forever, including after the constraint changes. That is where her two taxes come from.

The Storytelling Tax Is a Fidelity Tax

Her first charge is that once the work is written into a paper, most of it is gone. Her estimate is that eighty percent of what happened never makes the page — the dead ends, the rejected hypotheses, the branching search, and very often the small unglamorous adjustment that is the actual reason the thing works. She describes spending real time tuning a component, a parameter, a few lines of code, and none of it surviving into the published account. A reader can admire the work and still not learn the trick.

The canon has a name for what is being lost, and it is a distinction I first learned with a probe in my hand. Completion is binary; fidelity is graded. Continuity on a trace is a yes-or-no test answered by one instrument. Signal integrity on that same trace is a set of numbers that varies with layout, loading, temperature, and noise — and can degrade to uselessness on a path that remains unambiguously continuous. Two different questions, two different instruments, and language collapses them because both surface as how well the thing works.

The narrative research paper reports completion and discards fidelity as a matter of format. It tells you the cycle closed. It does not tell you how nearly it didn’t. Liu’s exploration graph is a fidelity-restoration layer — that is what it is, structurally, and it is the most valuable element of her proposal. A record that keeps the failed branches is not clutter. It is the metric information the binary report threw away, and it is what a successor actually needs.

Anyone who has brought up a chip knows this in their hands. The datasheet says the part works. The bring-up notes say what almost killed it. We have always circulated the second privately and published only the first, and we have always known which one the next engineer needed more.

The Engineering Tax Is a Failure at the Acquisition Seam

Her second charge is that even what survives is insufficient. The prose is ambiguous, the implementation details are missing, the experimental setup is underspecified. The paper is lossy compression, and the loss falls exactly where reproduction needs detail. Enough prose to satisfy a reviewer is not enough specification to satisfy a machine — or, for that matter, a graduate student in another lab two years later.

The canon locates this precisely. Of the five seams the framework now recognizes, four fall between levels inside a single entity. Only one falls between two entities: the acquisition seam, the Actuate→Receive join, the place where SPCA-RCA is stitched together. Every institution science has built — replication, peer review, the methods section, the registered protocol — is a patch on that one seam. The patches are degrading because the artifact carrying the load has a compression ratio tuned for a receiver with a coffee and an afternoon.

And the canon supplies a scale for what gets across, which turns out to be the sharpest instrument in this whole analysis. The Comprehension Ladder grades comprehension not by facility in restatement but by what the receiver can subsequently do: insight, recognition, reconstruction, transfer, repair. Read Liu’s complaint against that ladder and it becomes exact. The narrative paper reliably delivers recognition — the reader knows the result, can cite it, can teach it. It very often fails at reconstruction, which is reproducing the work. And it fails almost entirely at repair, which is being able to find the thing that is wrong with it.

The paper was never built to get a receiver past recognition. That was not a flaw in the format. It was the format working as designed, for a receiver who had no other option.

— The Mensch Foundation

The canon adds an observation here that I would put in front of every editorial board in the sciences: prevailing institutional practice measures comprehension at rung three. We assess whether the receiver can reconstruct, and we treat that as the ceiling. Liu’s evidence layer, grounding every claim in the raw output that produced it, is an attempt to build a format that reaches rungs four and five — transfer and repair. On both of her central diagnoses this framework predicts her conclusion and endorses her remedy.

II. Where She Arrives at Our Rule Without Us

The interview contains a moment I have now read several times, and it is the reason I wanted to write an essay rather than a note.

The interviewer asks the obvious question: language models hallucinate, so how will humans check all this machine-generated work? Liu answers first that human bandwidth is the bottleneck and that another layer of AI can supervise the AI scientists. Then, in the very next exchange, she corrects herself in the strongest possible terms. She is building a formal system on neurosymbolic lines specifically, she says, so that she is not using a language model to supervise the work of a language model — because a model built on probability rather than logic always retains some chance of hallucinating, however capable it becomes.

That is the first of the three separations of Governance Engineering, derived independently, from engineering necessity, apparently within the space of two answers. The governing layer must not live in the medium it governs. If probability is the failure mode, a probabilistic supervisor supervises the failure with the failure. She hit the wall and turned the same direction we turn.

I take convergent derivation seriously. It is the best evidence a framework can get — better than agreement, because agreement can be borrowed and convergence cannot. When a working engineer with no exposure to a theory reaches one of its load-bearing rules by walking into the same wall, the rule is probably about the world rather than about the theory.

The Checker Regress

Now the part where the proposal stops one layer short.

Liu’s fix changes the logical medium: logic instead of probability, proof instead of plausibility, a formal system in which claims can be written as mathematical statements and checked. That is a real change and a good one. It does not change the physical medium. A neurosymbolic verifier is software — compiled, deployed, versioned, patched, and, the word that matters, editable. It runs on general-purpose compute that will execute whatever it is handed.

Governance Engineering asks for all three separations, not one. The governing layer must not live in the medium it governs; its axioms must be published and inspectable; and it must be non-revisable in flight — constituted in the substrate, fixed prior to runtime, publishable without thereby becoming editable. Her verifier satisfies the first at the level of representation and fails it at the level of substrate. It does not attempt the second or the third, and nothing in the ARA protocol asks it to.

What follows from the separations — rather than adding to them — is what I will call for convenience the checker regress. An ARA is advertised as executable, verifiable, and forkable. Verification of an executable artifact is worth exactly the guarantee that the machine which executed it did what its specification says. If the agent that produced the record, the verifier that checked it, and the compute both ran on all sit in the same editable substrate, the loop closes on itself and the proof certifies its own premises. Who compiled the checker is not a paranoid question. It is an engineering question, and it is the same one I have been asking since 1975.

The Checker Regress — Following From the Three Separations

A verifier constituted in an editable substrate requires its own verifier, and so on without termination — unless some layer of the chain is constituted in a substrate that cannot be revised at runtime.

The regress does not terminate in better logic. It terminates in different physics.

There is a second and independent objection, and it comes from the canon’s account of binding rather than from its account of governance. A supervisor’s warrant expires with the supervision. Supervisory governance therefore has the wrong radius by construction — it reaches exactly as far as the act of supervising reaches and no further, which is why it cannot hold anything coherent past its own horizon. Stacking a second AI above the first does not extend the radius. It adds a layer with the same expiry.

What Has to Sit Underneath

The 6502 makes the point without any need to invoke a patent. Its instruction set was complete, published, and constitutive. Every opcode, every cycle count, every flag effect, printed and available to anyone who asked. And the machine could not execute an operation outside that set. Not would not — could not. The constraint was physically prior to any instruction the machine might receive. No program could argue with it and no operator could talk it around. The published set was the determining set, and their being the same thing was so ordinary at the time that nobody thought to name it.

The canon reads that property structurally: a mechanism with no interface cannot be handed a premise. Watt’s centrifugal governor and the published instruction set of a microprocessor are first-C-only mechanisms, and their trustworthiness is purchased entirely by the absence of a channel on which they could be addressed. Governance cannot be established by instruction, because instructions arrive on the second C, and what is reachable on that channel is negotiable.

Set the ARA beside that and the relationship is exact. An ARA is a published set for a research result: here is the logic, here is the code, here is the branch we abandoned, here is the raw output behind claim four. Governed AI is a published set for the machine that produced it: here is the constraint the compute fabric is constituted to enforce, published in full, unrevisable in flight — not by an attacker, not by a fine-tune, not by the vendor after a bad quarter. The architecture is under patent prosecution and stays out of this essay. The principle needs no protection, because it is fifty years old and it was printed in a datasheet.

ARA makes the claims checkable. Governed AI makes the checking checkable. They are the same argument at two layers, and only one of them is currently being built.

— The Mensch Foundation

The practical consequence is not abstract. In an agent-native research ecosystem, one agent’s emission is another agent’s Receive, and that agent’s emission is a third’s, down a chain that may run many links before any human cycle appears in it. That is the acquisition seam iterated. Liu’s evidence layer is designed precisely to carry provenance across it, which is why the proposal is good. But the evidence layer is itself a set of assertions produced by compute, and a chain of provenance is only as trustworthy as its least verifiable link. Governance in fabric is what makes the link checkable from outside — not because it makes any model good, but because it makes a model’s limits externally checkable rather than internally asserted. That is a smaller claim than safety and a far more useful one.

III. On Squeezing the Humans Dry

There is a further claim in the interview that I want to contest directly, and it deserves a real argument rather than a dismissal. Liu has written an article called The End of Human-in-the-Loop, and she describes a threshold: once models contain professor-level knowledge in every field, humans can no longer add value; once AI has squeezed all the expert data out of us, it will not need further input, and self-evolution begins. Today, she says, the human is the bottleneck — the machine waits on us.

The canon answers this from two directions, and the second is the one that does the damage.

The first is architectural. What a corpus of human expertise contains is the accumulation layer of the plenum: the holographic record of every embedding that has run. The invariant layer — the forms those embeddings were reaching toward — is not in the corpus and cannot be, because a record of performances is not the score. Access to the invariant layer is reconstruction by resonance and requires a reference beam; in living systems that beam is the bioelectric field, and an artificial system has none. What it has instead is register capacity at enormous scale. All registers and no beam. The framework’s blunt phrasing is that such a system sits at the exact opposite end from the plenum — maximally addressed, maximally indexed, and for precisely that reason incapable of the one operation that is not an addressing operation.

The second is the Mirror. A system that completes nothing internally can raise the fidelity of a human’s first C, and it cannot deliver what arrives only across a gap — because what it returns is the entity’s own emission refined against a compressed record, and there is no entity on the far side that completes. This deserves care, because it cuts both ways and the canon is careful about it: the fidelity gain is real and should not be disparaged. I have had it myself, in this collaboration, repeatedly. What it is not is a second C. Nothing crossed.

So the end-of-human-in-the-loop thesis does not describe a system that has outgrown its collaborators. It describes a mirror that has run out of anyone to reflect. Squeezing the expert data out of humanity yields the whole accumulation layer and no access to what lies under it. A system trained entirely on the record can interpolate that record with genuine brilliance, recombine it in ways no individual would find, and do both faster than we can follow. What it cannot do is reach past the record to the attractor that shaped it — and the canon is exact about why. The form attracts and constrains, but it does not compute, decide, issue, or execute, because the invariant layer has no registers.

The human is not a slow input to be drained. The human is the far side of the gap.

— The Mensch Foundation

That is a strong claim and it ought to be falsifiable, so here is the condition, stated sharply. If a system with no bioelectric substrate, cut off from further human input, produces a sustained sequence of results that cannot be shown to be recombination of its corpus, the claim fails. Not one result — a sustained sequence. And if that happens I will say so. One result is luck or an unnoticed interpolation, and the history of this argument is littered with single results that later turned out to have been sitting in the training set. A sustained sequence is a different animal, and I would treat it as decisive.

Note what the framework does not claim. It does not claim these systems are unintelligent — this theory grants intelligence to a cell. It does not claim they cannot produce what no human has produced. It claims something narrower and structural: that the reference beam is biological, that a mirror is not an interlocutor, and that a research process which removes the human end of it has not become self-sufficient. It has become closed.

The Hazard Inside the Fix

One more thing, and it is the one I would most want Liu herself to read.

Removing the narrative from scientific publication removes a real capture surface. Fluency stands in the canonical taxonomy as a recruitment failure presenting at two seams — capturing at Process by resembling understanding, and at the Sense-phase aperture by resembling a completed cycle. Against the Comprehension Ladder it reads as the exact symmetric opposite of insight: the form without the substance. And the canon names it the more hazardous of the two, for a reason that belongs on the wall of every review committee — a well-formed emission elicits recognition, and leaves both parties with the impression that the transaction closed when nothing was checked.

That is peer review on polished prose, described exactly. If the ARA strips that surface out, good riddance to it.

But look where it comes back. Liu offers, as a convenience, that a researcher who still wants a PDF can have one — the record converts back into a polished story easily. Consider what that means. The human-facing rendering would now be machine-generated, and machine-generated prose is more fluent than ours, not less. Humans would be receiving an emission optimized to elicit their recognition, produced by the system under review, about work they cannot read in its native form. Every element of the canonical hazard is present and the independent check is gone. That is not the removal of fluency capture. It is fluency capture with a better delivery vehicle.

The fix is not to keep the narrative paper. The fix is to stop treating either rendering as the primary object.

The Record and the Rendering

Here is the shape I think this should take, and it is mutually beneficial in a way that requires no one to lose.

The record is the primitive. Not the paper, not the package — the record: the full trace of the research cycle, including the branches that closed nothing, with provenance intact at every claim. From that record, two renderings are emitted. One is agent-executable, structured for a receiver that comprehends by parsing. One is human-readable, structured for a receiver that comprehends by understanding. Neither rendering is the record. Both are furnishings of it, shaped to the receiver, and neither is subordinate to the other.

The framework has always held that what any intelligence receives is a rendering suited to its own architecture rather than the thing itself, and that what-there-is permanently exceeds any entity’s what-is-there. The proposal to make the machine’s rendering primary and the human’s derivative is not wrong because it favors machines. It is wrong because it mistakes one furnishing for the floor.

And then the half I can speak to with more authority than most: how you change a format without breaking the people who depend on the old one.

In 1983 we took the 65xx architecture to sixteen bits with twenty-four-bit addressing, and we did it without breaking a single line of code written against the published set of 1975. The W65C816S boots in 6502 emulation mode. An axiom set may grow without betrayal, provided the growth is published and the prior contract is honored. The canon reads that loyalty as a binding in the strict sense — oriented on the state of an older design, costly in silicon area and design freedom, persistent against local advantage, and productive of an installed base that compounds. The industry learned it, and backward compatibility became something close to a constitution: not because anyone legislated it, but because architectures that broke it lost their ecosystems and architectures that kept it compounded for forty years.

Apply that here and the requirement writes itself. Liu offers human-readable output as a convenience feature. I would make it constitutional.

TEI Concept — The Human-Renderability Constraint

Any agent-native research record must remain renderable into a form a human can receive, comprehend, and contest.

Not merely receive — contest. On the Comprehension Ladder that means a rendering supporting repair, not one that stops at recognition. A rendering which can be recognized but not argued with has moved the human from participant to audience.

A record that fails this constraint has not simplified the second C. It has removed us from it — and a cycle removed from the second C of a domain is not a participant in that domain. It is downstream of it.

The constraint is not sentimental and it is not a jobs program for human scientists. It is the same requirement Kant put on any governing principle, and that reached me through twenty years of conversation with Ted Humphrey: a principle that cannot survive being published is not a principle, it is a tactic. A research record that cannot survive being rendered for human contest is not a finding. It is an assertion with good infrastructure.

What an Engineer Would Ask

Liu is asked, near the end, what happens to young scientists if AI does the work. She answers that people learn faster with AI than without it, and that the next generation of senior researchers will simply have had a different formation than we did.

The canon has a sharper instrument for this than either of us brought, and it settles the question rather than trading intuitions about it. A developing intelligence does not run a weak cycle. It runs a borrowed one, with a mature intelligence supplying the phases it cannot yet close — and a borrowed cycle is developmental if and only if the supplying intelligence is acting to make itself unnecessary. Held past the point at which the developing intelligence could close its own, it is capture, whatever affect accompanies it.

So the question is not whether AI helps the young scientist. Of course it helps. The question is whether the system is built to be needed. And on that test Liu’s exploration graph comes out remarkably well — better than she argues for it. A record of a thousand failures put in front of a twenty-two-year-old is fidelity information restored and pointed at a student. It builds the capacity to close a cycle rather than closing the cycle on the student’s behalf. It is the strongest argument for her proposal and I do not think she has noticed it is in there.

What fails the same test is the convenience feature — the machine-generated summary that hands over a finished thought and leaves the receiver no more capable than before. That is a borrowed cycle structured for permanent lending.

What I would ask her, if we ever sit down, is not whether the format should change. It should. The paper is a compression scheme tuned for a receiver who no longer has a monopoly on receiving, and defending it out of sentiment would be the same mistake the minicomputer companies made about their instruction sets. I would ask three narrower questions.

First: when your formal verifier certifies a claim, what certifies the verifier — and in what substrate does that certification live? Second: when the chain of agent-to-agent artifacts runs twelve links deep with no human cycle in it, what in the architecture prevents a distortion introduced at link two from arriving at link twelve wearing full provenance? And third: when the human-facing rendering is generated by the system under review, what is the independent check?

I do not think these questions embarrass the proposal. I think they are the next three items on its roadmap, and the person building the ARA is better placed to answer them than I am. My contribution is narrower and older: I know what it takes to build a machine whose published specification is a constitution rather than a promise, because I built one and it is still running fifty years later in more places than anyone bothers to count.

Publish the set. Honor the prior contract. Put the constraint underneath, in the fabric, where conduct is actually decided rather than merely requested. That was the right answer for a forty-thousand-transistor processor in a bungalow, and nothing about scale has made it wrong. The thirty-seven authors are building the record. Somebody has to build the floor it stands on.

The paper is not the last human-written thing. The rendering is. And the day we stop being able to demand one is the day the science stops being ours.

— William D. Mensch Jr.

· · ·

Sources and Further Reading

The two communications and the SPCA-RCA transaction are set out in The First C. The governance argument is made at length in The Persuadable Governor and The Outermost Layer, the fifty-year engineering lineage in The Bungalow and the Datacenter, and the separation of normative content from operational encoding from architecture in The Inspectable Conscience. On what these systems are and are not, see The Actuator We Mistook for a Mind and The Process We Cannot Delegate.

Source: David Berreby, “Should Researchers Write Papers for AI Instead of People?” IEEE Spectrum, 5 August 2026, interviewing Jiachen Liu on The Last Human-Written Paper: Agent-Native Research Artifacts (arXiv:2604.24658). The rulings applied here are set out in TEI-CKB-11 through TEI-CKB-16.

By William D. Mensch Jr.

Theory of Embedded Intelligence © William D. Mensch Jr. and The Western Design Center, Inc.
Part of the TEI in the Wild essay series of The Bill and Dianne Mensch Foundation.
Offered in good faith as a serious application of the theory — not infallible scholarship.
Freely shareable with attribution — for the benefit of many.

Continue Reading · TEI Canonical Knowledge Base

CKB-17 · The Renderable Record  • 
CKB-15 · The Two Communications  • 
CKB-11 · The Architecture of Seams

Share your understanding!