Once capable AI can leave the publisher's custody, safety has to become bigger than the model.
I was talking with Rob, my operations manager, about AI this morning.
The conversation started with Audrey Tang's poem and the response I had written to it. Rob had read both. We were talking about intelligence, autonomy, and what happens when increasingly capable systems become available to almost anybody.
Then he made a comparison that was funny for about three seconds.
He brought up The Anarchist Cookbook.
People used to get worked up about that book because it collected dangerous information in one place. Rob's point was that the argument now feels almost quaint. Today somebody can download an open-weight model, run it privately, modify it, weaken or remove behavioral safeguards, and interact with it for hours.
The comparison stuck with me because the important difference is not simply that AI contains more information.
The cookbook could not answer back.
It could not notice that you misunderstood something and explain it differently. It could not critique your plan, translate the material, adapt to your skill level, identify missing steps, challenge your assumptions, or continue working with you as the problem changed.
Static information can be dangerous.
Interactive intelligence changes the economics of using information.
That is the part I think we are still underestimating.
- 01ReferenceStatic information
- 02Interactive intelligenceExplain · adapt · critique · translate
- 03AgentUse tools · pursue intermediate steps
- 04Automated scaleRepeat · vary · distribute
The progression describes increasing interaction and agency. It does not claim every step produces large harmful uplift.
Intelligence left the API
For the first several years of the modern generative-AI era, it was easy to think about AI safety mostly as a provider problem.
The provider owned the model. You reached it through an application or API. The provider could apply safety training, monitoring, account controls, abuse detection, rate limits, and product policy. If something went wrong, it could update the serving stack, change the model, revoke access, or deploy another safeguard.
That model still matters.
It just does not describe the whole world anymore.
Open-weight AI changes the custody boundary.
Once model weights are released, another person can possess the capability. They can copy it, preserve it, modify it, fine-tune it, quantize it, run it locally, or distribute another copy somewhere the original publisher never sees.
The 2026 International AI Safety Report makes the key point directly: already-distributed open weights cannot be recalled wholesale. A hosting platform can remove a file. Governments can regulate distribution or possession. Developers can publish safer successors. Those actions can matter a lot.
But nobody can force every retained private copy to disappear or accept the next safety update.
That is a different security property.
Software capability does not behave like a recalled automobile or a contaminated batch of medicine. There is no guaranteed call-home mechanism. No universal patch cycle. No switch the original developer can flip that changes every copy sitting on somebody else's hardware.
That is what I mean by the title.
I am not claiming intelligence is mystical or alive. I mean learned capability can now be embodied in software artifacts that are cheap to copy once released.
Once that capability leaves custody, some portion of it is going to stay outside custody.
Hosted provider custody
- Provider-held model
- Monitored service boundary
- Account / API access
- Serving stack can be updated
- Access can be revoked
Strong provider control over the service does not make hosted systems immune to misuse.
Released-weight custody
- Weights leave publisher custody
- Copies can be retained
- Downstream modification is possible
- Private execution may be unobserved
- No universal forced update of retained copies
Distribution can still be restricted or disrupted. The bounded claim is that publisher-level technical recall is incomplete after private retention.
The paradox is real
That does not make open-weight AI bad.
In fact, many of the reasons it is valuable are tied directly to the loss of centralized control.
Open weights enable independent research, wider scrutiny, red teaming, customization, local privacy, competition, sovereign operation, and access for institutions that cannot build frontier models themselves. They reduce dependence on a handful of companies deciding who may inspect or operate powerful systems.
I want those benefits.
I also think we should be honest about the tradeoff.
You cannot celebrate decentralized custody and simultaneously assume centralized behavioral control remains complete.
The same property that gives researchers and users more freedom also gives downstream operators more ability to modify safeguards and operate outside the original provider's monitoring boundary.
That does not mean safeguards are pointless.
Quite the opposite.
Researchers are making progress on defenses that are harder to remove than ordinary refusal training. If future systems can genuinely suppress or avoid learning certain dangerous capabilities rather than merely refusing to discuss them, that could materially reduce risk.
Good.
We should want that work to succeed.
But even excellent model safeguards do not change the custody fact. A released copy remains a released copy.
So model behavior cannot carry the entire burden of safety by itself.
Knowing is not doing
Rob and I talked about the frightening examples humans predict almost immediately when a new technology appears: cyberattack, fraud, propaganda, biological weapons, bombs.
This is where the argument needs discipline.
Access to an AI model does not erase the physical world.
Biology is a useful example. Current evidence shows that AI can improve performance on some biology knowledge, planning, and proxy tasks. The evidence is much less decisive when inexperienced people have to perform difficult real laboratory work end to end.
That distinction matters.
Knowledge is not laboratory skill.
A plan is not equipment.
An explanation is not procurement.
Reasoning assistance is not physical execution.
AI can reduce informational, analytical, planning, and troubleshooting friction without eliminating tacit expertise, access controls, scientific difficulty, materials, operational failure, and every other barrier between an idea and a real-world outcome.
We should worry about the friction that disappears without pretending all friction has disappeared.
That is a stronger argument because it survives contact with evidence.
- InformationWhat is known?
- ReasoningCan the plan be adapted?
- ToolsCan digital actions be executed?
- AccessAre materials, systems, credentials, or equipment available?
- ExecutionCan the real-world outcome actually be produced?
AI can reduce informational and analytical friction without eliminating tacit expertise, physical constraints, access controls, scientific difficulty, or operational failure.
The human damage may be more ordinary
The spectacular scenarios attract attention because catastrophe makes better television.
Rob raised another possibility that may be closer to what these systems already change.
One person can generate a manifesto and produce hundreds or thousands of variations around it.
That sounds less dramatic than a bioweapon. It may be more immediately relevant.
Generative AI makes it cheap to produce persuasive language, rebuttals, translations, targeted explanations, alternate framings, and endless variations of the same underlying claim.
And we have empirical evidence that AI-generated and conversational messages can shift attitudes. The effects vary by topic, model, interaction, factuality, prior belief, and distribution. The evidence does not justify fantasies about a machine hypnotizing an electorate.
But the production economics have changed.
One motivated person no longer needs a writing staff, translation team, research assistant, copy editor, debate partner, and a small army of interns to produce industrial quantities of coherent messaging.
The intelligence can participate in the production loop.
Automation can participate in the distribution loop.
Information became cheap with the internet.
Reasoning assistance is becoming cheap with AI.
Tools and agents make action cheaper.
Automation supplies scale.
Those are different changes. Their combination matters.
- 01 · Before releaseReduce and measure dangerous capability
Capability evaluation · full-access testing · stronger safeguards · deliberate release decisions
- 02 · Consequential interfacesSeparate intelligence from authority
Identity · credentials · exact targets · least privilege · policy · evidence · revocation · recovery
- 03 · Beyond one control planeBuild institutional resilience
Law · distribution response · critical-infrastructure hardening · domain controls · detection · incident response
Safety has to become bigger than the model
My first instinct was to say safety has to move outside the model.
That was too neat.
If a malicious actor owns the model, hardware, tools, network, and target environment, my control plane does not descend from the heavens and revoke anything. Architecture does not become supernatural because somebody drew a clean diagram.
The more defensible answer has three layers.
The first is before release.
Model developers should measure dangerous capabilities, test safeguards under full-access conditions, make those safeguards harder to remove, and make deliberate decisions about what capabilities they release and when. The harder downstream control becomes, the more valuable effective upstream safeguards are.
The second is where intelligence meets consequential systems.
This is the part the AI industry still tends to frame too narrowly.
A model may know how to perform an action.
That does not mean it should possess the authority to perform it.
An agent may conclude that Production deployment is the correct next step. That does not mean it gets a Production credential.
Software may know how to modify a database. That does not mean it gets write access.
Identity, credentials, capability boundaries, exact targets, execution environments, policy, revocation, evidence, and recovery can all exist outside the model.
That does not make the intelligence safe.
It makes the surrounding system less dependent on the intelligence behaving perfectly.
The third layer is outside infrastructure any one organization controls.
That is where law, distribution policy, critical-infrastructure hardening, domain controls, detection, incident response, public institutions, and social resilience matter.
A societal proliferation problem cannot be solved entirely by an enterprise authorization service.
Some problems really do require society. Humanity remains annoyingly distributed.
What GhostMesh taught me from the opposite direction
The conversation with Rob helped me understand something about the architecture we have been building.
GhostMesh did not begin as a response to open-weight AI.
We were trying to let increasingly capable agents do useful work without handing them permanent authority over everything they could reach.
We kept backing into the same rules.
The worker is temporary.
Authority survives outside the worker.
Credentials are scoped.
Consequential actions cross explicit boundaries.
Permission expires.
Execution leaves evidence.
A model can propose an action without authorizing itself to perform the action.
We eventually expressed one part of that as:
Identity is not authority.
The open-weight problem exposes an even deeper version:
Intelligence is not authority.
That distinction becomes more important as intelligence becomes cheaper and easier to possess.
When only a handful of institutions held strong machine intelligence, controlling access to the intelligence itself could perform a meaningful part of the safety function.
As useful intelligence becomes copyable, that assumption weakens.
So consequential systems need boundaries that remain meaningful regardless of which model is doing the reasoning.
Possessing the capability to reason about an action should not automatically confer the capability to perform the action.
We already understand this for humans.
I know how a bank transfer works. That does not mean the bank lets me move anybody's money.
Machine intelligence deserves at least the same maturity of authority design.
This architecture applies where an organization or institution controls the consequential interface. It is not universal containment for an actor who owns the model, hardware, tools, target, and environment.
Control where intelligence meets consequence
I do not know where the capability curve ends.
Nobody does.
I do not know how much of today's frightening speculation becomes ordinary engineering reality, how much remains constrained by physical-world friction, or how effective future tamper-resistant safeguards become.
Those are empirical questions. We should keep measuring them instead of filling uncertainty with either panic or optimism.
But one architectural fact is already difficult to escape.
Once useful intelligence can be copied, modified, preserved, and operated privately, we cannot make the complete safety model depend on the assumption that somebody still controls the model.
We should continue making models safer.
We should be deliberate about release.
We should build safeguards that survive more than a polite jailbreak attempt.
And wherever intelligence reaches systems that matter, we should separate its ability to reason from its authority to act.
Rob's Anarchist Cookbook comparison sounded like a throwaway line when he said it.
It wasn't.
The old fear was that dangerous knowledge could escape into the world.
The new fact is that something more interactive than a book can leave custody too.
The cookbook could not answer back.
The model can.
And after the model has been copied, the original publisher may not be able to call it back.
So the durable safety rule cannot simply be: make intelligence behave.
It has to be stronger than that.
Intelligence is not authority.
And where intelligence meets consequence, authority has to belong to something we can still control.
Check the work, not just the conclusion.
Public research, authority, lineage, and author testimony are labeled separately. Sources can corroborate, challenge, or bound the argument; they do not replace Tony Malott's judgment.
Take the complete artifact with you.
The deterministic package contains a self-contained offline article, the exact public-route snapshot, canonical public metadata, receipt, source text when available, plain-text context, claim ledger, source records, and a member-hash manifest.
Sources, authority, and lineage
Each record states the role it plays. Research support and governance provenance are not treated as interchangeable.
SharePlane Platform Issue #634
Governs thesis, Rob attribution, evidence gate, manuscript v2, and Creative Lock.
Governs thesis, Rob attribution, evidence gate, manuscript v2, and Creative Lock.
Open sourceSharePlane Platform Issue #635 owner acceptance and release authority
Binds owner visual acceptance, exact Development identity, merge authority, and one-mutation Production budget.
Binds owner visual acceptance, exact Development identity, merge authority, and one-mutation Production budget.
Open sourceInternational AI Safety Report 2026
Supports bounded non-recall, open-weight risk/benefit, safeguard-removal, defense-in-depth, and mixed biological-uplift claims.
Supports bounded non-recall, open-weight risk/benefit, safeguard-removal, defense-in-depth, and mixed biological-uplift claims.
Open sourceManaging risks from increasingly capable open-weight AI systems
Supports open-weight innovation benefits, irreversible distribution concerns, oversight limitations, and safeguard-removal risk.
Supports open-weight innovation benefits, irreversible distribution concerns, oversight limitations, and safeguard-removal risk.
Open sourceDeep ignorance: filtering pretraining data builds tamper-resistant safeguards into open-weight LLMs
Counterevidence showing materially stronger tamper resistance is possible while preserving the need for defense in depth.
Counterevidence showing materially stronger tamper resistance is possible while preserving the need for defense in depth.
Open sourceLLM-generated persuasive messages produce measurable attitude change
Supports bounded claims that AI-generated messages can measurably influence attitudes.
Supports bounded claims that AI-generated messages can measurably influence attitudes.
Open sourceWhat is asserted—and how it is bounded
Research, author analysis, and personal testimony remain distinct. Supporting links and caveats stay attached to each claim.
After open model weights are publicly distributed and retained privately, the original publisher cannot universally recall or force-update every copy.
Boundary Distribution channels can be removed, possession regulated, and access made harder; the claim is about universal technical recall by the publisher.
Open-weight access changes the safeguard attack surface, while stronger tamper-resistant safeguards can materially reduce risk.
Boundary The Work does not claim safeguards are futile or that every open model is unsafe.
AI-generated and conversational messages can measurably shift attitudes in controlled studies, while real-world effect sizes depend on deployment context.
Boundary Generation scale is not proof of mass persuasion effectiveness.
Possession of intelligence should not automatically confer authority over consequential systems.
Boundary This control-plane principle applies where an organization or institution controls consequential interfaces; it is not universal containment for sovereign malicious environments.
Public boundary. Rob is credited as the intellectual catalyst for the discussion and Anarchist Cookbook comparison. Tony Malott remains the accountable author and owner of the resulting synthesis, architecture argument, and publication judgment.
Continue the thinking
Each connection explains why the next work belongs here. The graph records the edge; this layer makes it useful to a reader.