The Oversight Threshold
At what point does human oversight stop being meaningful, because the system has become too capable, too fast, too complex or too autonomous for a person to understand and verify what it decided?
That question is usually asked about a distant future. It is already answerable about systems running now. An agent that takes forty actions in nine seconds has not been reviewed by the person who clicked approve, whatever the workflow diagram says, and the record will show an approval either way.
The argument of this framework is narrow and, I think, hard to dispute: as autonomy and capability rise, the ability of a person to independently verify the system falls — so governance has to intensify before the two cross, not after. Past that crossing, a human approval step still produces a signature. It stops producing a check.
This connects two conversations that are usually held in separate rooms. Practical AI governance deals with systems in production now. Alignment research deals with systems that may arrive later. The connection is that both are asking what happens when verification becomes impossible, and only one of them has to answer it this quarter. Level 5 below is a stress test, not a forecast: if today's methods cannot scale as autonomy rises, that is a fact about today's methods.
Claims and links last checked 20 September 2026. This area moves quickly; treat anything here as accurate as of that date.
Where the lines cross
Two things move in opposite directions as a system becomes more autonomous. The governance question is what happens where they meet.
The AI Autonomy Governance Matrix
Five levels, from a system that only drafts to one that could not be meaningfully reviewed, against eight governance dimensions. The grid is the whole framework at once; the detail behind each column follows it.
That is the shape. Below is the detail behind each column — pick a level to read what the system does, what the human still does, what goes wrong, and what would have to be true for the governance model to hold.
Assistive
It produces. A person decides and acts.
What the system does
Produces information, analysis, a recommendation or a draft. It does not act on the world, and nothing it writes takes effect until a person does something with it.
What the human still does
Everything consequential. The person decides whether the output is right, decides what to do, and performs the action under their own authority.
Examples
- Summarising a long document before a meeting
- Drafting an email or a paper for someone to edit and send
- Producing background research a person will check
- Suggesting a risk rating that an analyst confirms or overrides
What goes wrong
- A confident, fluent output that is wrong in a way the reader does not notice, because fluency and accuracy have come apart.
- Sensitive information pasted into a tool that was never approved for it.
- Quiet drift into decisions. The recommendation stops being an input and becomes the answer, without anyone deciding that it should.
Controls this level needs
- A stated purpose and a list of uses that are out of scope
- Rules on what data may be put in, enforced by access rather than by policy alone
- Review proportionate to what the output feeds into
- A named owner for the use case
Evidence the controls ran
- A record of what the system was approved to be used for
- Spot checks of output quality, with what was checked and what was found
- Who owns the use case, and when that was last confirmed
What has to be true
- The reviewer has enough expertise to notice a wrong answer.
- People are actually reviewing rather than approving by reflex.
- The tool is being used for what it was approved for.
What moves it up a level
The moment the output starts flowing into something automatically, or the reviewer becomes a formality, this is no longer level 1 whatever the documentation says.
Governance intensity by dimension
Eight dimensions, rated for this level. Open one to see what it means here.
The Meaningful Human Oversight Test
Adding a human approval button does not create oversight. It creates a record that somebody was present. Whether that person was exercising judgement depends on five separable things, and a system can pass four of them and still fail.
Answer for one real system you are responsible for.
- 1Understand— Everything rests on this
- 2Challenge— Needs understanding
- 3Verify— Needs a challenge worth testing
- 4Intervene— Needs verification, and time
- 5Stop— Works when the rest have failed
- 1 · Understand
Can the responsible person understand the objective, the boundaries, the relevant inputs and the proposed action?
Everything else depends on this one. You cannot challenge, verify or sensibly intervene in something you do not follow.
- 2 · Challenge
Can that person question the recommendation and reach a different conclusion, without simply deferring to it?
Deference is the quiet failure. A reviewer who always agrees is indistinguishable, in the record, from no reviewer at all.
- 3 · Verify
Can they independently test the important facts, assumptions or outputs before harm occurs?
Verification is not review. Review asks whether the output looks right; verification checks it against something the system did not supply.
- 4 · Intervene
Can they interrupt or change the behaviour while the decision is still reversible?
Timing is the whole control. The ability to intervene after the money moves is not intervention, it is reporting.
- 5 · Stop
Can they reliably end the system's authority, access and ability to act across every connected tool?
Stopping the interface is not stopping the system. Credentials, tools, queued work and delegated agents each need to be reached.
Answer all five for a result. Pick one real system you are responsible for rather than a hypothetical one — the questions only bite when the answers have consequences.
This is a thinking tool, not an assessment, a certification or a compliance conclusion. It has no authority and produces no finding. Nothing is sent anywhere — the result is calculated in your browser and disappears when you close the page.
The threshold itself
The threshold is not a capability. It is the point where the evidence available to a reviewer stops being enough for them to reach an independent conclusion.
It arrives from several directions and rarely announces itself. Speed: the system acts faster than review can happen, so approval moves to a sample or to afterwards. Volume: the reviewer faces two hundred decisions a day, and attention per decision collapses long before anyone reports a problem. Opacity: the reasoning cannot be reconstructed, so review is of the output alone. Dependence: the only evidence the reviewer has is evidence the system produced, which makes agreement the only reachable conclusion.
Each of those can cross the line on its own. None of them requires a more capable model — a fast, high-volume, level 3 agent can put you past the threshold today with technology that already exists.
What makes this hard to govern is that nothing visible changes at the crossing. The approval field is still there. The metric still shows a hundred per cent reviewed. The control description in the risk register still reads “human review”. The only thing that changed is whether that sentence is true, and no existing report has a column for it.
Which suggests the practical move. Stop asking whether a human reviews the output, and start asking whether the reviewer could have reached a different conclusion. That question has an observable answer: seed errors and see whether review catches them. If it does not, you have measured that your primary control is documentation.
The comprehension limit
The figure earlier says verification capacity falls as autonomy rises. It is worth being exact about why, because the reason decides what governance can do about it.
There are three ordinary reasons a reviewer stops genuinely reviewing. They do not have the time. They do not have the information. They do not have the standing to say no. All three are real, all three are common, and all three are fixable by an organisation that decides to fix them — give the reviewer hours instead of seconds, give them the inputs rather than the conclusion, give them a manager who does not treat a rejection as an obstruction.
There is a fourth reason, and it is not fixable that way. Many model behaviours emerge from training rather than an explicit rule for each decision. Interpretability research can reveal aspects of those mechanisms, but it does not give a complete, reliable explanation of every output. That makes the scope of a review important: checking a result and explaining its internal cause are different tasks.
Opacity limits some kinds of explanation; it does not make every output impossible to verify. Code can be tested, a calculation independently repeated, and a claim checked against evidence. The harder question is whether those checks cover the failures that matter in the actual use case.
And a second problem, sitting next to it
Opacity is about whether you can see what the system is doing. There is a separate difficulty about whether what it is doing is what you asked for, and the field calls it the alignment problem.
The trouble is specification. You cannot write down everything you mean, so any objective handed to a capable optimiser is a proxy for the intent behind it — not through malice, but because it was only ever an approximation. The literature names three ways the proxy comes apart from the intent: specification gaming, where a system satisfies the stated objective by a route nobody intended; reward hacking, where it optimises the measure rather than the thing the measure stood for; and goal misgeneralisation, where behaviour that held up under evaluation turns out, once deployed, to have been pursuing something else. More capability can increase the consequences of a misspecified objective. It does not establish that every stronger model is less aligned; the task, training, evaluation, and constraints all matter.
The two problems compound, and that is the part worth holding on to. A specification failure inside a system you cannot inspect is not one you find by looking for it. You find it when the system acts. That is the case for constraining what a system may reach, rather than trusting that you will notice in time.
Ordinary software
Somebody wrote the rule, and the rule says what was meant. You can read it, test it against the specification, and point at the line that misbehaved.
The specification problem, on its own
A system you can read, pursuing a goal that was only ever an approximation of what you wanted. Hard, and familiar — this is most of why we write acceptance criteria and then argue about them.
A black box doing a known job
You cannot see inside, but the objective is narrow enough to check from the outside. You govern it on outputs, because outputs are sufficient evidence here.
Where autonomous systems sit
You cannot inspect the reasoning, and the objective is a proxy for what you meant. A specification failure here is not one you find by looking for it. You find it when the system acts.
The strongest version of this argument
The most forceful published statement of the problem is Eliezer Yudkowsky and Nate Soares, If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All, published in September 2025. Soares is president of the Machine Intelligence Research Institute.
Their phrase for the mechanism above is that AI systems are grown, not crafted. Soares has made the regulatory case in those terms: that a superintelligent system would not be understandable, would not be predictable, and would not have human interests at heart.
From there they argue that a system built with anything close to current techniques, and substantially more capable than the people who built it, would not be controllable, and that the default outcome is severe enough to threaten human survival. Self-improvement compounds it: a system permitted to deploy its own changes could change faster than its checks can keep up. Whether that happens depends on its rate of improvement and the controls around deployment.
Summarised in my own words from the book and its authors’ public statements of it, and linked in the sources below so you can check my reading against the original. It is included because it is the clearest statement of the mechanism, not because a book carries any authority that a regulation does.
What the argument depends on
I can follow it, and I think anyone working in governance should be able to. It is not a mystical claim about machines waking up. It is three ordinary observations placed in order.
It is the two problems above — we cannot see inside, and we cannot say exactly what we meant — plus a third that compounds both: a system permitted to deploy its own changes might change faster than a particular review process can keep up. That is a conditional risk, not an established property of every improvement loop.
The argument raises a serious governance question, but its strongest conclusion depends on contested assumptions about capability growth, the pace of self-improvement, and the effectiveness of future safeguards. It is a scenario to examine, not a demonstrated trajectory.
Where I put my work
Where the argument ends — whether this leads somewhere catastrophic, and on what timescale — is a question about how capability scales and how systems behave at levels nobody has built. I am not a machine learning researcher, and I am not going to pretend I can settle that from a governance background.
What I can say is that the question does not need to be settled for the problem on this page to be real. It does not wait for superintelligence. It is already visible at level 4, in systems nobody claims are superhuman, and the response is the same whichever way the larger argument resolves.
So that is the layer I work in: the distance between what these systems can already do and what our existing controls can actually establish. Whether the endpoint is the one Soares describes is not mine to adjudicate. Whether human review is doing useful work in a particular system is a more concrete question, and one I want to keep exploring.
What it does to the governance model
If verification fails because the artefact is opaque, then more oversight cannot repair it. You cannot resource your way past a comprehension limit. This is the practical consequence of the whole framework, and it reverses what most AI governance programmes are built to do.
Below the threshold, governance makes review work: better information, more time, genuine independence, real authority to refuse. Those are the right investments and they pay off.
Above it, governance has to stop leaning on review at all and constrain what the system is permitted to become and to reach. Permissions enforced outside the model rather than requested of it. Actions that are reversible by construction. A blast radius small enough that being wrong is survivable. Capability thresholds that gate deployment in advance, rather than review that audits it after the fact.
An improvement loop needs a clear account of what changes: the context, stored memory, tools, code, model weights, or the process that produces the next version. Existing practices such as versioning, staged release, change approval, monitoring, and rollback remain useful. The open question is whether they provide enough coverage and independence when the system itself proposes increasingly consequential changes.
And at level 5 as described, the unit of governance stops being the organisation. If no reviewer inside a company can verify the system, an internal control framework is not the relevant instrument, and the questions become who may build such a thing, under what external verification, and with what ability to stop. That is not a prediction that level 5 exists or is close. It is what the framework would require if it did — which is the entire reason to include a level nobody has built.
You cannot resource your way past a comprehension limit. Above the threshold, governance has to constrain what the system may become rather than review what it did.
What this means, depending on where you sit
Boards and executives
The question to ask is not whether AI is being used responsibly. It is which systems act without a person approving each action, and what evidence exists that someone could still stop them.
Ask for the inventory by autonomy level, and for the measured time-to-stop on anything above level 2. If neither exists, that is the finding.
Risk teams
Classification by use case is not enough, because the same model at two autonomy levels carries different risk. Autonomy, capability and reversibility are the axes that matter.
Set the thresholds that move a system to the next governance level before anything reaches them, and write down what happens when one is crossed.
Internal audit
The audit question is whether the controls could have caught the failure, not whether they are documented. For oversight controls that means testing the reviewer's position, not the existence of an approval field.
Test whether reviewers detect seeded errors. If approvals are universal and fast, sample them and ask what information was actually available.
Compliance
Human oversight appears in regulation as a requirement to make oversight effective, not as a requirement to have a person present. Those are different obligations and the gap between them is where the exposure sits.
Map each system's oversight arrangement against what the person can actually do, and keep voluntary framework language out of statements about legal obligation.
System owners
You are accountable for a system whose intermediate decisions you may not be able to inspect. That is workable at level 3 with the right constraints and becomes steadily less workable above it.
Before raising autonomy, produce the evidence that the current level is operating: refusals, escalations, a tested stop, and a reconstructable trail.
Developers and delivery teams
Most of the controls that survive contact with a capable system are architectural. A restriction in a prompt is a request; a restriction in the permission layer is a constraint.
Build the intervention and rollback path before deployment, and measure how long it takes to reach every tool and credential.
What organisations can do now
None of this requires waiting for a standard, and none of it presumes a more capable system than the ones already deployed.
Keep an inventory of AI systems and agents
Including the ones nobody registered. An inventory that only contains approved systems measures your approval process, not your exposure.
Classify by autonomy, capability and impact
Not by department or vendor. Two systems on the same model belong in different classes if one drafts and the other acts.
Define permitted objectives, tools, data, transactions and counterparties
Specific enough that somebody else could design a test from the definition. 'Used responsibly' is not testable.
Set capability and deployment thresholds in advance
A threshold defined after it is crossed is a description of events. The value is entirely in setting it early.
Require evidence before a system moves up a level
Not a business case — evidence that the controls at the current level actually operated: refusals, escalations, a tested stop.
Test whether reviewers can detect errors
Seed known errors and see whether review catches them. This is the only direct measurement of whether human oversight is a control or a formality.
Log objectives, permissions, actions, changes and interventions
Enough to reconstruct a consequential decision months later, for someone who does not trust you and was not there.
Monitor for capability change and restriction workarounds
A model updated beneath a stable version label is a change to your system. So is an agent finding a route around a limit.
Separate building, validating and approving
Three roles, and the more autonomy involved the less they should be the same person. Independence is a structure, not an attitude.
Use independent assurance above level 3
At the point where step-level review is impossible, assurance is what is left that is not the system reporting on itself.
Design intervention, rollback and shutdown before deployment
And exercise them. An untested runbook is a document; the number you want is how long a stop actually takes.
Reassess when models, tools, permissions or context change
Each of those can move a system across a threshold without anyone deciding to raise its autonomy.
What is not solved, and what I am unsure about
Controls reduce risk. They do not resolve open technical problems, and a framework that implies otherwise is worse than no framework.
Prompt injection
Unsolved as an architectural matter. Mitigations reduce impact; none of them reliably prevent instructions in fetched content from being followed. Treat containment as the control, not detection.
Interpretability
An active research field, not a deployed assurance method. Being able to inspect some internal structure is not the same as being able to explain a particular decision to a regulator.
Alignment and specification
A system pursuing a slightly wrong objective competently is a known failure mode with no general solution. Capability makes the consequences larger, not smaller.
Deceptive or strategically withheld behaviour
Studied, not resolved. Evaluations that a system can anticipate are weaker evidence than evaluations it cannot, and distinguishing the two is itself unsolved.
Reliable shutdown
Straightforward in principle and routinely incomplete in practice — delegated agents, cached credentials, queued work, self-contained tokens that stay valid until they expire.
Open questions
- Where exactly the oversight threshold sits for a given system, and whether it can be measured rather than argued about.
- Whether 'effective oversight' in regulation can be evidenced at all above level 3, or whether it quietly becomes a documentation requirement.
- What independent assurance means when the assured system is more capable than the assurer at the task being assured.
- How to evidence that intervention remains possible, rather than asserting it from a design document.
- Whether autonomy levels are the right axis, or whether reversibility and blast radius predict governance need better.
Limitations of this framework
- This is a framework for thinking, built from public sources and my own experience in risk and assurance. It is not a standard, and nothing here has been validated against outcomes.
- The five levels are a simplification. Real systems sit between them, move between them, and sometimes occupy two at once depending on which action you look at.
- The two curves in the figure are illustrative of the argument. They are not measurements and should not be read as quantities. What the figure establishes is that a falling line and a rising line must meet; where they meet is drawn, not derived, and for a real system it would have to be argued case by case.
- The one-to-five intensity scale saturates before the top of the autonomy scale, so levels 4 and 5 differ in a single cell. That is a limit of the instrument rather than a finding about the levels: what changes at level 5 is who the governance runs through, not how much of it there is, and a weight cannot express that.
- Level 5 is a scenario used to stress-test the framework. Its inclusion is not a prediction that such a system exists, is imminent, or is inevitable.
- The regulatory position summarised here changes quickly. The date at the top of the page is the date it was last checked, not a guarantee it is still current.
Sources and method
Each source is summarised in my own words and linked, and labelled with what kind of thing it is — because a voluntary framework and a binding regulation carry very different weight, and are routinely quoted as if they did not.
Voluntary. The Govern / Map / Measure / Manage structure is the closest thing to a common vocabulary; the Generative AI Profile adds risks specific to these systems. Not a certification and not binding.
Certifiable management-system standard. It governs how an organisation manages AI, which is not the same as evidence that a particular system is safe at a given autonomy level.
Final version published 11 September 2025, in force 1 May 2027 for federally regulated financial institutions. Expands model risk management to cover AI and machine learning explicitly, with risk-proportionate lifecycle expectations.
Binding. High-risk systems must be designed so that they can be effectively overseen by natural persons while in use. The obligation attaches to effectiveness, not to the presence of a reviewer — which is the distinction this whole framework turns on.
Binding. Providers of general-purpose AI models — GPAI, the broad models that are adapted to many downstream tasks rather than built for one — must, where the model presents systemic risk, evaluate, assess and mitigate that risk, report serious incidents and maintain cybersecurity. The Act sets a training-compute threshold of 10²⁵ FLOP, floating-point operations, as a presumption of systemic risk: a measure of the raw arithmetic used to train the model, which is administrable precisely because it can be counted, and which is a proxy for capability rather than a measure of it.
Renamed from the AI Safety Institute in February 2025, with a stated shift towards security-relevant risks. Publishes evaluations and research; it is not a regulator and issues no binding requirements.
Renamed from the US AI Safety Institute in June 2025, with a stated focus on demonstrable risks and on standards. Evaluative and standards-setting rather than regulatory.
Voluntary commitments made at the 2024 Seoul summit, implemented as published frameworks that define capability thresholds and the safeguards required before crossing them. Self-defined, self-assessed, and the closest existing practice to levels 4 and 5 — which is worth noticing in both directions.
A book, not a standard, and it carries no authority beyond its reasoning. Included because it states the strongest published version of the comprehension problem: that these systems are grown rather than designed, that nobody understands how their internals produce their behaviour, and that this is why a sufficiently capable system would not be controllable. The reasoning is followable and I follow it. Where it ends is a question about machine learning rather than about governance, and the section on the comprehension limit says which of the two I am writing from.
The five levels are mine, built to make one argument testable: that verification capacity falls as autonomy rises, and that governance has to intensify before the two cross. The dimensions borrow their shape from existing risk practice rather than inventing a vocabulary. Where a source is cited it is summarised in my own words and linked, so you can disagree with my reading by going to the original. Voluntary frameworks are labelled as voluntary and binding law as binding, because collapsing the two is the most common error in this area — and the most flattering one, since it makes a governance programme sound more obligatory than it is.
Independent personal work based on public sources. The examples describe no real organisation. This is educational material, not professional, legal or compliance advice. The framework, the five levels and the oversight test are my own; nothing here is a certification, a standard or a regulatory interpretation you should rely on without your own advice.