Two Rivers on AI · Ch6. What Dario Amodei Measures That Yi Zeng Cannot ← Ch5 Ch7 →
Txt Low Med High
PART ONE — Two Source Springs
Chapter 6

What Dario Amodei Measures That Yi Zeng Cannot

Page 1 · What Dario Amodei Measures
Mechanistic Interpretability
Mechanistic Interpretability

Now the mirror move.

Interpretability Problem
Interpretability Problem

Dario Amodei trained as a biophysicist. His doctorate at Princeton concerned the collective behavior of populations of neurons — how networks of cells generate coherent patterns of activity from individually noisy components. The postdoctoral work at Stanford continued in the same direction. Then a transition to machine learning, then OpenAI, then the founding of Anthropic with his sister Daniela in 2021 with a small group of OpenAI alumni who shared a specific concern about the trajectory of frontier AI capability.

The biophysics inheritance matters. Neuroscience teaches a particular set of habits. You cannot reason a network into legibility. You instrument it. You apply perturbations and observe responses. You build statistical models of activity patterns and check whether your model predicts what you see next. You assume the system contains structure you do not yet understand, and you accept that the structure will only emerge through measurement. Intuition about behavior is not a substitute for empirical characterization of mechanism. A neuron that fires in a certain pattern in response to a stimulus is not "trying" to do anything. It is doing something. The only way to know what is to record from it under controlled conditions.

You assume the system contains structure you do not yet understand, and you accept that the structure will only emerge through measurement.

Amodei transferred this methodology into AI. The 2020 scaling laws paper, which he co-authored at OpenAI, is in its essential structure a biophysics paper applied to language models. The authors did not philosophize about whether scale would matter. They measured what happened when you varied compute, parameters, and data, and fit empirical curves to the relationship. The result was the most consequential finding in the recent history of artificial intelligence. It was consequential precisely because it was treated as an empirical claim with a measurable shape rather than as a speculation about cognition.

· · ·
Page 2 · What Dario Amodei Measures
Deceptive Alignment
Deceptive Alignment

The methodology shapes everything Amodei has built since. The refusal to use "AGI" — replaced by "powerful AI" defined in capability thresholds you can in principle measure. The Responsible Scaling Policy — capability thresholds tied to specific safety measures the company commits to in advance, publicly, so the future institution is constrained by the past institution under conditions outsiders can verify. The mechanistic interpretability program — the bet that the circuits, features, and reasoning pathways inside trained models can be mapped, that the systems can be understood, not merely operated. The institutional inventory frame of the "compressed twenty-first century" essay — what institutions need to exist to receive the capability advancement that the technology is going to deliver on the timeline it is going to deliver on, whether anyone is ready or not.

Responsible Scaling Policy
Responsible Scaling Policy

All of these are the same move. Take the philosophical question. Convert it into an empirical question with measurable content. Build the instrument that produces the measurement. Commit, in advance, to acting on what the measurement reveals. Publish the commitment so that the commitment is checkable. The methodology is the biophysicist's discipline applied to a technology whose tempo has compressed to a degree the biological sciences never had to manage.

Now place Amodei next to Yi Zeng.

Dario Amodei
"The alignment problem is not just a technical challenge. It is the question of whether we can build systems that reliably do what we actually want — not what we said, not what we measured, but what we meant. [VERIFY]"
Anthropic blog, 'Core Views on AI Safety' · 2023

Yi Zeng's framework is rigorous in a different way. Yi Zeng is asking, in his BrainCog program at the Chinese Academy of Sciences, what consciousness might require architecturally — what spiking dynamics, what theory-of-mind modules, what mechanisms for embodied self-modeling, what kinds of social cognition modules might constitute the substrate on which something morally relevant could emerge. The framework is grounded in neuroscience but oriented toward a philosophical question — what makes a system the kind of system that has a perspective, that can be wronged, that earns moral consideration through its capacity to feel and intend and care.

The substantive disagreement with Amodei is not about the technology. The two researchers are working in adjacent technical traditions. The disagreement is about what the technology is for and what kind of question we are supposed to be asking of it.

· · ·
Page 3 · What Dario Amodei Measures
Assumption Of Alignment
Assumption Of Alignment

Yi Zeng's claim — that benevolence (ren) and harmony (he) must be engineering constraints, not retrofitted ethics — places the substantive philosophical question at the center of the engineering work. The question of what the system is supposed to be — what kind of being it is, in relation to what kinds of beings we are — is the engineering question. The technical work flows from a prior substantive commitment about what the good is.

Yi Zeng
"AI safety in the Chinese context cannot be separated from the question of social harmony. A system that is technically aligned but socially disruptive is not safe — it has simply moved the problem. [VERIFY]"
Brain-Inspired Intelligence: From Neuroscience to AI Safety · 2022

Amodei's claim — that "powerful AI" can be operationalized by capability thresholds, that safety measures can be specified procedurally, that interpretability provides the load-bearing bet for governance — places the procedural philosophical question at the center. The question of what the system is supposed to be is bracketed. The question is what the system is, measurably, doing — and whether the institutional response can be calibrated to evidence rather than to either fear or enthusiasm.

This is a real difference. It is not merely a difference of vocabulary. It is a difference about whether the substantive question of what AI is for can be addressed through procedural means, or whether procedural means are inadequate to a substantive question that has to be addressed substantively if it is to be addressed at all.

The question is what the system is, measurably, doing — and whether the institutional response can be calibrated to evidence rather than to either fear or enthusiasm.

Amodei would respond — has responded, in essays and testimony and the institutional architecture of Anthropic — that the substantive question is not avoidable. The Responsible Scaling Policy embeds substantive commitments about what kinds of capabilities require what kinds of safety measures. Constitutional AI embeds substantive commitments about what kinds of character a frontier model should be trained to express. The interpretability program embeds the substantive commitment that an AI we cannot understand cannot be governed in any morally serious sense. The procedural apparatus is the means by which the substantive commitments get operationalized at the tempo the technology requires.

· · ·
Page 4 · What Dario Amodei Measures
Transparent Ai
Transparent Ai

Yi Zeng would respond — has responded, in research papers and editorial work and the architecture of BrainCog — that the procedural apparatus is insufficient because it cannot reach the depth of what the substantive question requires. Specifying capability thresholds does not tell you what the system is for. Constitutional AI written in the language of Western individualism cannot encode the relational obligations the Confucian frame insists on. Interpretability tools that map circuits in a trained model cannot, by themselves, tell you whether the architecture is capable of ren, of harmony, of the wisdom that the Confucian tradition has always understood as more fundamental than intelligence.

Both researchers are right about what they see and partial about what the other supplies.

What Amodei measures that Yi Zeng cannot is the procedural discipline that allows a frontier lab to operate inside the tempo of capability advancement without abdicating governance entirely. The Chinese AI ecosystem has not produced an equivalent of the Responsible Scaling Policy — has not produced a public, published, capability-threshold-linked commitment framework that other labs can adopt and that regulators can codify. The procedural apparatus is a Western innovation. It is, at the moment, the load-bearing institutional response to the tempo problem. Without it, frontier capability advances without any structural mechanism for tying deployment decisions to safety evidence.

What Yi Zeng sees that Amodei does not is that the procedural apparatus, however well-designed, leaves the substantive question unanswered. What is the system for? What kind of being is it in relation to what kinds of beings we are? What capacities does sustained engagement with the system cultivate or atrophy in the humans on the other side? What civilizational trajectory is the deployment producing, evaluated by substantive criteria about what kind of civilization is worth becoming? The Responsible Scaling Policy cannot answer these questions. Constitutional AI gestures at them but does not, in its current form, embed the depth of substantive philosophical commitment they require.

· · ·
Page 5 · What Dario Amodei Measures
Epistemic Justice Ai
Epistemic Justice Ai

This is not a competition between the two frameworks. It is a recognition that they are addressing different parts of the same problem, and that neither part is sufficient alone. The procedural apparatus needs the substantive commitment to give the procedures their content. The substantive commitment needs the procedural apparatus to operationalize itself at the tempo the technology requires. The procedural apparatus without substantive commitment becomes empty proceduralism — accountability theater that satisfies the form of governance without doing the work. The substantive commitment without the procedural apparatus becomes wisdom literature — important, deep, philosophically rigorous, and operationally inert at the speed at which capability is advancing.

The two researchers, both of whom are deeply serious, are looking at the same problem from opposite ends. Amodei is looking at the procedural end. Yi Zeng is looking at the substantive end. The space between them is the space the global AI governance project has not yet been built into. The book that contains this chapter is one attempt to make the space visible.

What this chapter has tried to do is name, precisely, what each researcher can see that the other cannot. The naming is the prerequisite for the conversation that has not yet happened. The conversation is the prerequisite for the institutional response that the compressed century is going to require, on a timeline shorter than either tradition alone is going to be able to construct.

The next chapter steps back to the bridge between them, before the rest of the book walks each river individually.

· · ·
Dario Amodei
Further Reading From The Orange Pill Cycle · Related Thinkers
2 voices alongside this chapter — click to meet them
Continue · Chapter 7
The Bridge Before the Rivers Split
← Prev 0%
Ch6 Next →