The Eval Before the Eval Tony Malott · Published 2026-08-26 https://shareplane.malott.ai/artifacts/the-eval-before-the-eval/ Systems Essay The Eval Before the Eval What if the AI does exactly what we asked, and what we asked was wrong? A sufficiently capable autonomous system can execute a bad objective beautifully; consequential intent needs an explicit, evidence-bearing admission boundary before autonomous amplification. By Tony Malott Published 2026-08-26 16 min Long-form architecture argument Status Published Release Canonical URL Preview copy. This is Semantic Candidate v02, not the final SharePlane semantic lock. What if the AI does exactly what we asked, and what we asked was wrong? SharePlane Thesis This line of thinking started with my wife. She sent me something by Audrey Tang and told me I should read it. I knew very little about Tang at the time, but something in the way she approached AI caught my attention. She was coming at the problem through democracy, civic technology, pluralism, and human agency. I had arrived from almost the opposite direction. I had spent a long time wrestling with autonomous systems, probabilistic claims, stale state, credentials, provenance, authority leakage, provider failures, recovery, and the irritating reality that a system can perform technically correct work it was never actually entitled to perform. GhostMesh grew out of that mess. I was trying to make increasingly capable AI useful inside real systems without allowing a model, a credential, a database, or a convenient workflow to quietly become authority. The paths were different, but I kept recognizing the same boundary. Then another question started bothering me. We spend an enormous amount of time evaluating AI: whether the model answered correctly, whether the agent completed the task, whether it followed policy, whether it remained inside its permissions, whether the output satisfied the acceptance criteria. Those are necessary questions. I think they start one layer too late. What about the objective itself? Who decided it was the right objective? What assumptions were buried inside it? What evidence supported those assumptions? Whose interests did it represent? Who had authority over the consequences? And what if the person asking was simply wrong? The uncomfortable version is worse: what if the objective was bad and the AI was excellent? A sufficiently capable autonomous system can execute a bad objective beautifully. It can remain inside every technical permission, produce complete receipts, satisfy every acceptance criterion, and pass every evaluation built around the task. Nothing has to hallucinate. Nothing has to crash. Nobody has to bypass a security control. The system can still produce exactly the wrong outcome because the failure happened before execution began. We admitted the wrong intent. The Eval Trap The Eval Trap from intent through metric, execution, evaluation, and a passing result The Eval Trap from intent through metric, execution, evaluation, and a passing result The accompanying structured sequence explains what each stage contributes and cannot guarantee. 01 Intent 02 Metric 03 Execution 04 Evaluation 05 Pass 01 Intent The objective enters the system, along with assumptions about what should happen and why. 02 Metric The objective becomes a target the system can optimize or a result the system can measure. 03 Execution The machine performs the work, perhaps with extraordinary accuracy and efficiency. 04 Evaluation The result is tested against the success criteria derived from the objective. 05 Pass Every downstream test can turn green while the unresolved problem remains upstream. Everything downstream can be green while the objective itself is wrong. A perfect score can certify the wrong thing Organizations have been optimizing the wrong things for as long as organizations have had numbers. Measure a support organization entirely by ticket closure and people learn how to close tickets. Measure a hospital entirely by throughput and care can start looking like friction. Govern a company entirely through quarterly earnings and sacrificing the future can look remarkably disciplined for three months at a time. The metric is not necessarily false. Usually there is a legitimate reason it exists. The problem is that the metric is only a projection of purpose, and once the projection becomes the objective, parts of the original purpose can disappear. AI raises the stakes because it removes friction between intent and execution. Human organizations are inefficient in ways that drive us insane, but some of that inefficiency accidentally functions as resistance. Someone questions the instruction. Someone notices the data is wrong. Someone knows that the policy makes no sense in this particular case. Someone understands a consequence that never appeared on the dashboard. More capable autonomous systems can remove much of that resistance. That is a tremendous advantage when the objective is sound and a tremendous liability when it is not. A system that becomes better at executing intent makes the quality of that intent more important, not less. Consider a straightforward executive instruction: reduce operating costs by 20 percent. The executive may have every right to set the target, and the organization may genuinely need to reduce spending. But the number carries assumptions with it. Why 20 percent? Which costs are waste and which preserve reliability, safety, customer experience, or institutional knowledge? Is the problem structural or temporary? Is cost even the right representation of the underlying problem? An execution eval can tell us whether spending fell by 20 percent. It cannot tell us whether 20 percent deserved to be the objective. A system could defer maintenance, reduce staffing below sustainable levels, shift work somewhere the financial model does not see, or create problems that will not surface until the next reporting period. The books could reconcile perfectly and the organization could still be worse. That is a premise failure, not an execution failure. Once an objective becomes a metric, it also starts to look strangely neutral. Increase engagement. Reduce claims paid. Identify high-risk employees. Increase arrests. Improve congregational participation. Every one of those objectives can live comfortably on a dashboard, acquire a target, become an optimization routine, and eventually produce an evaluation suite full of green and red boxes. Every one also contains human judgment. A human judgment becomes a KPI. The KPI becomes an optimization target. The target becomes an eval. The eval becomes a green check. Somewhere along that chain, the assumptions and tradeoffs embedded in the original objective can disappear from view. I think of that as intent laundering . The objective has not become neutral. It has simply acquired the appearance of neutrality. The machine did not create the intent. It can simply make that intent extraordinarily efficient. There is a related question that belongs upstream too: what would look like success to the metric while actually violating the purpose? “Reduce costs” says nothing about safety. “Increase engagement” says nothing about manipulation. “Reduce fraud” says nothing about false positives. “Restore service quickly” says nothing about weakening a control that exists for a reason. A mature objective needs boundaries around pathological success, not merely a number to maximize. Authority does not make something true GhostMesh has already forced me to separate capability from authority. A model's ability to perform an action does not grant it the right to perform that action. Possession of a credential does not establish authority. Successful execution does not retroactively create permission. There is another distinction underneath that one: authority does not make something true. A CEO can possess legitimate organizational authority and still operate from false information. A board can unanimously approve a strategy built on a bad premise. A government official can have lawful authority and still be factually wrong. Keeping a human in the loop does not solve this. Humans rationalize. Humans inherit bad assumptions. Humans operate under incentives. Humans become certain about things the evidence does not support. Humans can possess legitimate authority and exercise it badly. Human authority matters, and human judgment matters. Human infallibility does not exist. So the architecture cannot simply be: the AI may be wrong, therefore ask a human. That just moves the oracle. A consequential instruction should instead be treated as a claim about reality as well as a command. “Reduce operating costs by 20 percent” implicitly claims that current costs are too high, that the reduction is achievable, that the tradeoffs are acceptable, that the requester has authority over those tradeoffs, and that cost is the correct representation of the underlying problem. None of those claims becomes true because the request arrived from an authenticated account. This is why intent needs provenance too. Where did the objective come from? What has to be true for it to make sense? Which parts are evidence and which are judgment? Who gets to decide? Why does that authority extend to the people and institutions that will carry the consequences? Two identical commands can have completely different meaning depending on what sits behind them. The missing gate The architectural idea that emerged from all of this is fairly simple. Intent Admission is the point where a consequential objective has to establish why the system is allowed to amplify it. That includes the evidence behind its important assumptions, the authority behind the request, the rules that still apply, the people and systems that will carry the consequences, and the conditions that would cause the objective to be reconsidered. It is not an AI morality score. The system is not deciding whether a person's goal is virtuous. I do not want GhostMesh quietly rating political positions, business strategies, institutional policies, or personal choices and deciding which humans deserve to act. That would merely relocate the authority problem. The job is narrower. For this objective, under these conditions, with these consequences, does the system have a sufficient basis to proceed? The answer does not need to be a numerical score. It can be operational: execute, execute within explicit bounds, hold because something material remains unresolved, or reject because a governing rule or authority boundary prohibits the action. That is not artificial morality. It is governed execution. Four things we keep collapsing Autonomous systems routinely collapse things that should remain separate. A credential may prove that an action is technically possible. It says very little about whether the actor has authority to make the decision, whether that authority extends to the consequence, or whether the factual premise behind the decision is true. Permission Can the mechanism do it? Permission is technical. It describes the portion of capability the mechanism will permit an actor to exercise. Authority Who gets to decide? Authority is institutional. It comes from ownership, delegation, law, contract, role, policy, consent, or another recognized source rather than from possession of a button. Legitimacy Why is that decision theirs to make in this case? Legitimacy concerns why an authority properly extends to the people and consequences involved. It becomes especially important when the people carrying the consequence are not the people operating the system. Evidence What supports the claim? Evidence is about whether the factual premises behind the objective are grounded in reality rather than merely asserted by an authorized actor. These are separate questions, and none of them creates the others. Permission should implement authority rather than manufacture it. Authority should operate against evidence rather than substitute for it. Strong evidence does not grant someone authority over another person. Where other people carry the consequences, the basis for that authority needs to remain visible. This is one place Audrey Tang sharpened the problem for me. In her July 2026 essay *Holding Pattern* , she writes, “Those who will carry the consequences should help author the test.” Tang is not proposing Intent Admission or the GhostMesh architecture described here. Her point helped expose a boundary technical authority alone could not answer: the people who understand the machinery and the people who carry its consequences may have different but legitimate roles in deciding how it should be used. That distinction changes how I think about the attack surface of an autonomous system too. Security engineering already teaches us not to trust an input simply because it arrived through an authenticated channel. Intent deserves the same skepticism. A requester can be authenticated and still be mistaken, manipulated, outside their authority, or asking the system to optimize the wrong thing. Authenticate the actor. Then establish the basis of the request. Those are different operations. The owner can conflict with the owner The problem becomes more interesting when the person who wants to bypass the control is me. Suppose I deliberately establish a Production publication rule requiring a qualified candidate, provenance, release integrity, and explicit acceptance where required. I created the rule because experience taught me that the moment when I most want to skip those controls is probably the moment I need them. Then one afternoon I become impatient. The artifact looks good. Other work is waiting. I tell the system, “This is good enough. Just publish it.” There is no authentication problem. Nobody stole my account. The immediate instruction is unquestionably mine. But if the newest instruction automatically overrides the earlier rule, then I never really created a control. I created a preference that remained binding only until it became inconvenient. That is the distinction I mean by intent hierarchy . An immediate instruction can operate within higher-order rules. It should not silently rewrite them merely because both came from the same owner. I call those highest-order operating rules constitutional intent . Changing them is allowed, but changing the rule and acting under the rule are different acts. Intent hierarchy Intent hierarchy from purpose through constitutional intent, policy, and an immediate request Intent hierarchy from purpose through constitutional intent, policy, and an immediate request The accompanying structured sequence explains what each stage contributes and cannot guarantee. 01 Purpose 02 Constitutional intent 03 Policy 04 Immediate request 01 Purpose What is the system trying to preserve or accomplish over time? 02 Constitutional intent What higher-order rules govern how that purpose may be pursued? 03 Policy What standing decisions operate inside those higher-order rules? 04 Immediate request What does an authorized person or agent want the system to do right now? A momentary instruction is not automatically a constitutional amendment. This is not a system refusing its owner. It is a system preserving a distinction the owner deliberately created. Standing controls often exist specifically to govern moments when immediate preference conflicts with earlier considered judgment. The same principle exists outside software. Corporate controls, financial approvals, security boundaries, clinical procedures, data-retention rules, and institutional governance all depend on the idea that an immediate request does not automatically dissolve a standing rule. Evidence constrains judgment Intent Admission cannot guarantee that an admitted objective is wise. No architecture can provide that guarantee, and a system claiming otherwise would probably be more dangerous than the problem it claims to solve. There will always be decisions where the facts are reasonably established, the right person has authority, no higher-order rule prohibits the action, and legitimate judgment still remains. A board may choose Strategy A instead of Strategy B. A manager may choose among several defensible staffing approaches. An owner may choose one design over another. Evidence can constrain that judgment. It cannot replace it. That is important because evidence-heavy systems have a tendency to imply that every important human question eventually becomes deterministic if enough data is collected. It does not. Sometimes the system has done its job when it has established the facts, the boundaries, and the person who genuinely gets to decide. Then the person decides. When the system does not know Real work rarely presents a perfectly resolved admission record. Evidence conflicts. Authority overlaps. Policies become stale. A repair that began in Development turns out to require a Production change. An agent discovers consequences nobody anticipated when the original objective was approved. AI systems have a dangerous tendency to resolve those gaps because continuing feels more useful than stopping. A model infers. A policy engine selects a default. An agent chooses the authority that looks most senior. Sometimes the correct answer is simply HOLD . A hold is not a rejection. It means something important remains unresolved and the system has no right to invent the missing answer. This becomes particularly important in multi-agent systems because agents are very good at deriving increasingly specific objectives from broad instructions. An agent can reasonably derive “reduce deployment variability” from “improve reliability,” then derive “standardize the provider path” from that. The reasoning may be excellent. That does not mean every new consequence inherited unlimited authority from the original request. Semantic derivation is not authority derivation. Delegation can become narrower as work becomes more specific. When the consequences materially expand, the basis to proceed needs to be established again. That is not less autonomy. It is an autonomous system recognizing that understanding what to do is different from possessing the authority to do it. Admission also has to remain conditional over time. New evidence can invalidate a premise. Authority can expire or change. The action can cross into a new consequence class. The people affected can change. An objective that was reasonable yesterday does not become immortal merely because the system admitted it once. The intent cannot authorize itself There is one class of action that deserves stronger treatment than ordinary execution. Changing the rules that determine who may act, what evidence counts, which constraints apply, or what a system may do autonomously is fundamentally different from operating under those rules. Imagine an agent that lacks Production authority. It proposes a policy allowing itself to publish Production, applies the policy, and then performs the deployment. Every step could generate a beautiful receipt. The sequence would still be absurd. The actor proposed the expansion, changed the rule, acquired the authority, exercised it, and potentially evaluated its own success. That is not governance. It is self-authorization with logging. An agent can participate in proposing a change to its own authority. It should not be able to propose it, approve it, activate it, and immediately exercise it entirely by itself. Somewhere in that chain, authority has to come from somewhere else. The exact separation can vary with consequence. It may be another legitimate authority, an explicit owner acceptance step, independent verification, a different phase, or a pre-established emergency contract. The important property is that the actor being constrained cannot manufacture the entire constraint system and its own escape from it in one uninterrupted chain. The eval cannot certify its own objective All of this returns to the question that started the article. An eval contains a reference frame. Somebody decided what success means. Somebody decided which harms count. Somebody selected the data. Somebody chose the acceptable tradeoffs and decided when the number was good enough. Those decisions may be excellent. They may also be incomplete, conflicted, obsolete, manipulated, or wrong. An eval can tell us whether a system performed against its definition of success. It cannot, by itself, establish whether that definition of success was the right one to begin with. NIST's current AI Risk Management Framework approaches the issue from risk management rather than autonomous authority, but there is an important parallel. Its MAP Playbook asks organizations to establish intended purpose, assumptions, deployment context, applicable expectations, potential impacts, and affected people. NIST also makes the useful point that “Highly accurate and optimized systems can cause harm,” and its MAP guidance explicitly discusses clearer go/no-go decisions about whether an AI system should be deployed. NIST does not present this as Intent Admission, and I do not want to pretend it has endorsed the architecture in this essay. AI RMF 1.0 is also currently being revised, which is exactly why its role here should remain bounded. What it reinforces is the underlying problem: measuring a system well does not relieve us of establishing what the system is for and whether its use makes sense in context. An eval can tell us whether the system executed correctly. An outcome evaluation can tell us whether intended results appeared. Governance can later reconsider whether the objective remains valid. But the objective cannot bootstrap its own legitimacy simply by defining a test and then passing it. That is the eval before the eval. By this point, I do not think it is really another eval. It is admission. Intent Admission Intent Admission should preserve the questions separately rather than collapsing them into one confidence score. 01 Objective + origin What is being requested, what outcome is actually sought, and where did the objective come from? 02 Premises + evidence What has to be true for the objective to make sense, and what evidence supports or contradicts those claims? 03 Authority + governing rules Who gets to decide, what gives them that authority, and what higher-order constraints still apply? 04 Consequences + anti-goals Who or what changes if the objective succeeds, and what must not be sacrificed merely to improve the primary metric? The disposition is operational: EXECUTE · EXECUTE BOUNDED · HOLD · REJECT . New evidence, changed scope, changed authority, expiry, or a legitimate challenge returns the objective through admission rather than bypassing it. The important thing is what the admission mechanism does not do. It does not create authority. It does not make disputed evidence true. It does not erase human discretion. It binds the basis for execution closely enough that the decision can later be reconstructed and challenged. The record behind the decision Once consequential intent is admitted, the system should preserve enough of that decision to reconstruct it later. What objective was approved? Where did it come from? What important assumptions supported it? Which authority applied? What boundaries remained? What would have caused the decision to change? I think of that as an Intent Admission Record . The record does not grant authority. It preserves the basis on which the system believed execution was allowed. Humans should not have to fill out a constitutional form every time an agent does something useful. The system should assemble what it already knows: the originating request, applicable authority, current policy, relevant evidence, affected resources, and existing constraints. It should bring a person only the unresolved part that actually requires judgment. Governance should not turn humans into form-processing middleware. If everything material is already resolved inside an existing delegation, ordinary work should remain ordinary. If one important question remains, that is the question the person should see. The record also gives intent lineage. Objectives change during real work. Someone narrows a request. New evidence changes scope. An agent discovers that fulfilling the original objective requires a materially different action. The system should not quietly mutate the objective and pretend it was always the same task. Intent can evolve while authority does not automatically evolve with it. That distinction will matter enormously as agents become better at decomposing broad goals into increasingly specific sub-goals. The gate can be wrong too None of this makes Intent Admission infallible. Its evidence can be wrong. Its policy can be stale. The institution behind it can be captured. The people defining the rules can design them to protect themselves. The answer cannot be to declare the admission mechanism authoritative merely because it is the admission mechanism. Its rules, sources, owners, challenges, and changes need provenance too. Architecture cannot guarantee good institutions. It can make bad institutional behavior harder to hide behind software. That is a humbler promise than machine morality. It is also one we can actually engineer. A mature system should know what gives it the right to continue For most of computing history, intelligence was scarce. Humans decided what should happen and software executed comparatively narrow instructions. Much of the hard work involved getting enough human reasoning into a form the machine could use. That constraint is disappearing. Reasoning is becoming cheaper. Generation is becoming cheaper. Software creation is becoming cheaper. Execution is becoming cheaper. Autonomous systems can carry an objective much farther before a human needs to intervene. Authority is not becoming cheaper with them. Neither are legitimacy, truth, accountability, or good judgment. Those become more important precisely because capability is becoming abundant. The question used to be whether the machine could do something. Then we learned to ask whether the machine was permitted to do it. There is now another question in front of both: why are we doing this, and what gives the system the right to act on it? Capability cannot answer that. Credentials cannot answer it. An eval designed around the objective cannot answer it. The fact that the instruction came from a human cannot answer it by itself. Intent Admission does not place AI above human intent. It keeps increasingly powerful intelligence from blindly amplifying intent that has never been grounded or bounded. The objective itself becomes part of the governed system. Its origin, evidence, authority, constraints, lifetime, and consequences remain visible rather than disappearing behind the automation that carries it out. When those materially change, the basis for execution changes too. That is not less autonomy. It is a more mature form of it. A powerful system should not merely know how to continue. It should know what gives it the right to continue. Evidence behind the thesis Check the work, not just the conclusion. Public research, authority, lineage, and author testimony are labeled separately. Sources can corroborate, challenge, or bound the argument; they do not replace Tony Malott's judgment. Portable public record Take the complete artifact with you. The deterministic package contains a self-contained offline article, the exact public-route snapshot, canonical public metadata, receipt, source text when available, plain-text context, claim ledger, source records, and a member-hash manifest. Download full artifact package Read plain-text context Inspect package manifest 4 public sources Sources, authority, and lineage Each record states the role it plays. Research support and governance provenance are not treated as interchangeable. External Primary Source Holding Pattern Intellectual catalyst for the affected-people and legitimacy boundary; not authority for Intent Admission or GhostMesh doctrine. Intellectual catalyst for the affected-people and legitimacy boundary; not authority for Intent Admission or GhostMesh doctrine. Open source External Primary Source AI Risk Management Framework Current evidence that AI RMF 1.0 is being revised and background authority for NIST risk-management framing. Current evidence that AI RMF 1.0 is being revised and background authority for NIST risk-management framing. Open source External Primary Source AI RMF Playbook: Map Independent support for intended purpose, assumptions, context, impacts, affected people, and go/no-go reasoning upstream of ordinary measurement. Independent support for intended purpose, assumptions, context, impacts, affected people, and go/no-go reasoning upstream of ordinary measurement. Open source Governing Issue SharePlane Platform Issue #610 Governs semantic lock, Creative Lock, public-safe boundary, implementation, UAT, and publication of this Work. Governs semantic lock, Creative Lock, public-safe boundary, implementation, UAT, and publication of this Work. Open source Claim discipline What is asserted—and how it is bounded Research, author analysis, and personal testimony remain distinct. Supporting links and caveats stay attached to each claim. Author Architectural Analysis claim:eval-before-eval:001 An evaluation can establish performance against a chosen reference frame without, by itself, establishing that the objective defining that frame was legitimate, sufficiently grounded, or appropriate for autonomous execution. Support AI RMF Playbook: Map Boundary NIST supports upstream purpose and context reasoning but does not define or endorse the Intent Admission architecture proposed here. Externally Supported Author Analysis claim:eval-before-eval:002 People who carry the consequences of an AI system can have a legitimate role in how the system is tested and governed even when they do not control the technical machinery. Support Holding Pattern Boundary Tang is used as an intellectual catalyst and not as authority for GhostMesh implementation doctrine. Original Ghostmesh Doctrine claim:eval-before-eval:003 Intent Admission should preserve evidence, authority, governing constraints, consequences, and uncertainty as separate dimensions and produce an operational execution disposition rather than a moral score. Support Author testimony or analysis; no external source is claimed. Boundary This is proposed architecture, not an empirical scientific finding or an external standard. Public boundary. Tang and NIST support parts of the problem framing. Intent Admission, intent laundering, intent provenance, intent hierarchy, constitutional intent, semantic-versus-authority derivation, the four admission dispositions, the Intent Admission Record, and the self-authorization constraint are SharePlane/GhostMesh synthesis and are not attributed to those external sources. 4 sources 3 governed claims 1 portable package Connected work Continue the thinking Each connection explains why the next work belongs here. The graph records the edge; this layer makes it useful to a reader. Foundations Capability and credentials do not create governing authority The Agent Is Not the Security Boundary The security-boundary thesis separates agent capability from enforceable system controls; this Work extends that separation upstream by treating the objective itself as a governed input rather than an unquestioned command. Do not ask whether the agent is trustworthy. Ask whether the system remains safe when the agent is wrong. Systems Essay By Tony Malott 2026-07-25 16 min · Standard long-form article Companions Capability abundance moves the hard work outward The Code Is No Longer the Hard Part The Code Is No Longer the Hard Part argues that AI makes implementation cheaper while intent, authority, context, evidence, and recovery become scarcer engineering work; this Work drills into the intent boundary itself. As AI reduces the relative cost of implementation, durable engineering value moves outward into intent, authority, context, orchestration, evidence, recovery, and institutional integration. Systems Essay By Tony Malott 2026-07-26 12 min · Standard long-form article Explore the complete graph SOURCE REFERENCES Holding Pattern https://au.civic.ai/p/holding-pattern AI Risk Management Framework https://www.nist.gov/itl/ai-risk-management-framework AI RMF Playbook: Map https://airc.nist.gov/airmf-resources/playbook/map/ SharePlane Platform Issue #610 https://github.com/pinklon/shareplane-platform/issues/610