Skip to main content
Mubienclarity for what comes next
← Research
Interactive framework

The Oversight Threshold

As AI becomes more autonomous, human approval can remain in the workflow while disappearing in substance.

At what point does human oversight stop being meaningful, because the system has become too capable, too fast, too complex or too autonomous for a person to understand and verify what it decided?

That question is usually asked about a distant future. It is already answerable about systems running now. An agent that takes forty actions in nine seconds has not been reviewed by the person who clicked approve, whatever the workflow diagram says, and the record will show an approval either way.

The argument of this framework is narrow and, I think, hard to dispute: as autonomy and capability rise, the ability of a person to independently verify the system falls — so governance has to intensify before the two cross, not after. Past that crossing, a human approval step still produces a signature. It stops producing a check.

This connects two conversations that are usually held in separate rooms. Practical AI governance deals with systems in production now. Alignment research deals with systems that may arrive later. The connection is that both are asking what happens when verification becomes impossible, and only one of them has to answer it this quarter. Level 5 below is a stress test, not a forecast: if today's methods cannot scale as autonomy rises, that is a fact about today's methods.

Claims and links last checked 20 September 2026. This area moves quickly; treat anything here as accurate as of that date.

Where the lines cross

Two things move in opposite directions as a system becomes more autonomous. The governance question is what happens where they meet.

Oversight thresholdL1AssistiveL2DelegatedL3ConditionalL4HighL5ScenarioHuman verification capacityGovernance intensity required
The lines cross between level 3 and level 4. That crossing is the oversight threshold: the point at which a person's ability to independently verify the system falls below the governance intensity the system now requires. Both lines are illustrative of the argument rather than measured quantities.

The AI Autonomy Governance Matrix

Five levels, from a system that only drafts to one that could not be meaningfully reviewed, against eight governance dimensions. The grid is the whole framework at once; the detail behind each column follows it.

Dimension
L1
L2
L3
L4
L5
Purpose and permitted use
1
2
4
5
5
Authority and permissions
1
3
4
5
5
Human oversight
2
3
3
4
5
Testing and evaluation
1
3
4
5
5
Monitoring and evidence
1
3
4
5
5
Intervention and containment
1
2
4
5
5
Accountability and escalation
2
3
4
5
5
Independent assurance
1
2
3
5
5
Intensity
1Light2Moderate3Substantial4Heavy5Maximal
Oversight threshold
Every dimension climbs and none of them comes back down: there is no level at which some part of governance gets easier. Two things in the grid are worth reading closely. Human oversight is the only dimension not at maximum by level 4, because it is the one that cannot simply be turned up — that is the argument of this page, visible in a single column. And the scale saturates: levels 4 and 5 differ in one cell, which is a limit of a one-to-five weight rather than a claim that the two levels need the same governance. What changes at level 5 is not how much, but who — the reliance moves off individual review and onto external arrangements. The numbers are relative weights for comparing levels against each other, not measurements, not a maturity score, and not something to total up.

That is the shape. Below is the detail behind each column — pick a level to read what the system does, what the human still does, what goes wrong, and what would have to be true for the governance model to hold.

Level 1

Assistive

It produces. A person decides and acts.

What the system does

Produces information, analysis, a recommendation or a draft. It does not act on the world, and nothing it writes takes effect until a person does something with it.

What the human still does

Everything consequential. The person decides whether the output is right, decides what to do, and performs the action under their own authority.

Examples

  • Summarising a long document before a meeting
  • Drafting an email or a paper for someone to edit and send
  • Producing background research a person will check
  • Suggesting a risk rating that an analyst confirms or overrides

What goes wrong

  • A confident, fluent output that is wrong in a way the reader does not notice, because fluency and accuracy have come apart.
  • Sensitive information pasted into a tool that was never approved for it.
  • Quiet drift into decisions. The recommendation stops being an input and becomes the answer, without anyone deciding that it should.

Controls this level needs

  • A stated purpose and a list of uses that are out of scope
  • Rules on what data may be put in, enforced by access rather than by policy alone
  • Review proportionate to what the output feeds into
  • A named owner for the use case

Evidence the controls ran

  • A record of what the system was approved to be used for
  • Spot checks of output quality, with what was checked and what was found
  • Who owns the use case, and when that was last confirmed

What has to be true

  • The reviewer has enough expertise to notice a wrong answer.
  • People are actually reviewing rather than approving by reflex.
  • The tool is being used for what it was approved for.

What moves it up a level

The moment the output starts flowing into something automatically, or the reviewer becomes a formality, this is no longer level 1 whatever the documentation says.

Governance intensity by dimension

Eight dimensions, rated for this level. Open one to see what it means here.

The Meaningful Human Oversight Test

Adding a human approval button does not create oversight. It creates a record that somebody was present. Whether that person was exercising judgement depends on five separable things, and a system can pass four of them and still fail.

Answer for one real system you are responsible for.

  1. 1Understand— Everything rests on this
  2. 2Challenge— Needs understanding
  3. 3Verify— Needs a challenge worth testing
  4. 4Intervene— Needs verification, and time
  5. 5Stop— Works when the rest have failed
They run in order, and each one rests on the one before it. A system can satisfy four and still fail, because the one it misses is the one that mattered. Stopping is set apart because it is the only condition that still holds when the others have gone — which is why it is worth exercising, and why it is the one almost nobody has exercised.
  1. 1 · Understand

    Can the responsible person understand the objective, the boundaries, the relevant inputs and the proposed action?

    Everything else depends on this one. You cannot challenge, verify or sensibly intervene in something you do not follow.

  2. 2 · Challenge

    Can that person question the recommendation and reach a different conclusion, without simply deferring to it?

    Deference is the quiet failure. A reviewer who always agrees is indistinguishable, in the record, from no reviewer at all.

  3. 3 · Verify

    Can they independently test the important facts, assumptions or outputs before harm occurs?

    Verification is not review. Review asks whether the output looks right; verification checks it against something the system did not supply.

  4. 4 · Intervene

    Can they interrupt or change the behaviour while the decision is still reversible?

    Timing is the whole control. The ability to intervene after the money moves is not intervention, it is reporting.

  5. 5 · Stop

    Can they reliably end the system's authority, access and ability to act across every connected tool?

    Stopping the interface is not stopping the system. Credentials, tools, queued work and delegated agents each need to be reached.

Answer all five for a result. Pick one real system you are responsible for rather than a hypothetical one — the questions only bite when the answers have consequences.

This is a thinking tool, not an assessment, a certification or a compliance conclusion. It has no authority and produces no finding. Nothing is sent anywhere — the result is calculated in your browser and disappears when you close the page.

The threshold itself

The threshold is not a capability. It is the point where the evidence available to a reviewer stops being enough for them to reach an independent conclusion.

It arrives from several directions and rarely announces itself. Speed: the system acts faster than review can happen, so approval moves to a sample or to afterwards. Volume: the reviewer faces two hundred decisions a day, and attention per decision collapses long before anyone reports a problem. Opacity: the reasoning cannot be reconstructed, so review is of the output alone. Dependence: the only evidence the reviewer has is evidence the system produced, which makes agreement the only reachable conclusion.

Each of those can cross the line on its own. None of them requires a more capable model — a fast, high-volume, level 3 agent can put you past the threshold today with technology that already exists.

What makes this hard to govern is that nothing visible changes at the crossing. The approval field is still there. The metric still shows a hundred per cent reviewed. The control description in the risk register still reads “human review”. The only thing that changed is whether that sentence is true, and no existing report has a column for it.

Which suggests the practical move. Stop asking whether a human reviews the output, and start asking whether the reviewer could have reached a different conclusion. That question has an observable answer: seed errors and see whether review catches them. If it does not, you have measured that your primary control is documentation.

The comprehension limit

The figure earlier says verification capacity falls as autonomy rises. It is worth being exact about why, because the reason decides what governance can do about it.

There are three ordinary reasons a reviewer stops genuinely reviewing. They do not have the time. They do not have the information. They do not have the standing to say no. All three are real, all three are common, and all three are fixable by an organisation that decides to fix them — give the reviewer hours instead of seconds, give them the inputs rather than the conclusion, give them a manager who does not treat a rejection as an obstruction.

There is a fourth reason, and it is not fixable that way. Many model behaviours emerge from training rather than an explicit rule for each decision. Interpretability research can reveal aspects of those mechanisms, but it does not give a complete, reliable explanation of every output. That makes the scope of a review important: checking a result and explaining its internal cause are different tasks.

Opacity limits some kinds of explanation; it does not make every output impossible to verify. Code can be tested, a calculation independently repeated, and a claim checked against evidence. The harder question is whether those checks cover the failures that matter in the actual use case.

And a second problem, sitting next to it

Opacity is about whether you can see what the system is doing. There is a separate difficulty about whether what it is doing is what you asked for, and the field calls it the alignment problem.

The trouble is specification. You cannot write down everything you mean, so any objective handed to a capable optimiser is a proxy for the intent behind it — not through malice, but because it was only ever an approximation. The literature names three ways the proxy comes apart from the intent: specification gaming, where a system satisfies the stated objective by a route nobody intended; reward hacking, where it optimises the measure rather than the thing the measure stood for; and goal misgeneralisation, where behaviour that held up under evaluation turns out, once deployed, to have been pursuing something else. More capability can increase the consequences of a misspecified objective. It does not establish that every stronger model is less aligned; the task, training, evaluation, and constraints all matter.

The two problems compound, and that is the part worth holding on to. A specification failure inside a system you cannot inspect is not one you find by looking for it. You find it when the system acts. That is the case for constraining what a system may reach, rather than trusting that you will notice in time.

Can inspect · Can specify

Ordinary software

Somebody wrote the rule, and the rule says what was meant. You can read it, test it against the specification, and point at the line that misbehaved.

Can inspect · Cannot specify

The specification problem, on its own

A system you can read, pursuing a goal that was only ever an approximation of what you wanted. Hard, and familiar — this is most of why we write acceptance criteria and then argue about them.

Cannot inspect · Can specify

A black box doing a known job

You cannot see inside, but the objective is narrow enough to check from the outside. You govern it on outputs, because outputs are sufficient evidence here.

Cannot inspect · Cannot specifyWe are here

Where autonomous systems sit

You cannot inspect the reasoning, and the objective is a proxy for what you meant. A specification failure here is not one you find by looking for it. You find it when the system acts.

Three of these four are tractable, and governance has decades of practice in all three. The fourth is not a harder version of the others — it is the one where the usual move, look at what it did and judge whether that was right, stops returning an answer you can rely on. That is the case for constraining what a system may reach, rather than trusting that somebody notices in time.
Argument, not a standard

The strongest version of this argument

The most forceful published statement of the problem is Eliezer Yudkowsky and Nate Soares, If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All, published in September 2025. Soares is president of the Machine Intelligence Research Institute.

Their phrase for the mechanism above is that AI systems are grown, not crafted. Soares has made the regulatory case in those terms: that a superintelligent system would not be understandable, would not be predictable, and would not have human interests at heart.

From there they argue that a system built with anything close to current techniques, and substantially more capable than the people who built it, would not be controllable, and that the default outcome is severe enough to threaten human survival. Self-improvement compounds it: a system permitted to deploy its own changes could change faster than its checks can keep up. Whether that happens depends on its rate of improvement and the controls around deployment.

Summarised in my own words from the book and its authors’ public statements of it, and linked in the sources below so you can check my reading against the original. It is included because it is the clearest statement of the mechanism, not because a book carries any authority that a regulation does.

What the argument depends on

I can follow it, and I think anyone working in governance should be able to. It is not a mystical claim about machines waking up. It is three ordinary observations placed in order.

It is the two problems above — we cannot see inside, and we cannot say exactly what we meant — plus a third that compounds both: a system permitted to deploy its own changes might change faster than a particular review process can keep up. That is a conditional risk, not an established property of every improvement loop.

The argument raises a serious governance question, but its strongest conclusion depends on contested assumptions about capability growth, the pace of self-improvement, and the effectiveness of future safeguards. It is a scenario to examine, not a demonstrated trajectory.

Where I put my work

Where the argument ends — whether this leads somewhere catastrophic, and on what timescale — is a question about how capability scales and how systems behave at levels nobody has built. I am not a machine learning researcher, and I am not going to pretend I can settle that from a governance background.

What I can say is that the question does not need to be settled for the problem on this page to be real. It does not wait for superintelligence. It is already visible at level 4, in systems nobody claims are superhuman, and the response is the same whichever way the larger argument resolves.

So that is the layer I work in: the distance between what these systems can already do and what our existing controls can actually establish. Whether the endpoint is the one Soares describes is not mine to adjudicate. Whether human review is doing useful work in a particular system is a more concrete question, and one I want to keep exploring.

What it does to the governance model

If verification fails because the artefact is opaque, then more oversight cannot repair it. You cannot resource your way past a comprehension limit. This is the practical consequence of the whole framework, and it reverses what most AI governance programmes are built to do.

Below the threshold, governance makes review work: better information, more time, genuine independence, real authority to refuse. Those are the right investments and they pay off.

Above it, governance has to stop leaning on review at all and constrain what the system is permitted to become and to reach. Permissions enforced outside the model rather than requested of it. Actions that are reversible by construction. A blast radius small enough that being wrong is survivable. Capability thresholds that gate deployment in advance, rather than review that audits it after the fact.

An improvement loop needs a clear account of what changes: the context, stored memory, tools, code, model weights, or the process that produces the next version. Existing practices such as versioning, staged release, change approval, monitoring, and rollback remain useful. The open question is whether they provide enough coverage and independence when the system itself proposes increasingly consequential changes.

And at level 5 as described, the unit of governance stops being the organisation. If no reviewer inside a company can verify the system, an internal control framework is not the relevant instrument, and the questions become who may build such a thing, under what external verification, and with what ability to stop. That is not a prediction that level 5 exists or is close. It is what the framework would require if it did — which is the entire reason to include a level nobody has built.

You cannot resource your way past a comprehension limit. Above the threshold, governance has to constrain what the system may become rather than review what it did.

What this means, depending on where you sit

Boards and executives

The question to ask is not whether AI is being used responsibly. It is which systems act without a person approving each action, and what evidence exists that someone could still stop them.

Ask for the inventory by autonomy level, and for the measured time-to-stop on anything above level 2. If neither exists, that is the finding.

Risk teams

Classification by use case is not enough, because the same model at two autonomy levels carries different risk. Autonomy, capability and reversibility are the axes that matter.

Set the thresholds that move a system to the next governance level before anything reaches them, and write down what happens when one is crossed.

Internal audit

The audit question is whether the controls could have caught the failure, not whether they are documented. For oversight controls that means testing the reviewer's position, not the existence of an approval field.

Test whether reviewers detect seeded errors. If approvals are universal and fast, sample them and ask what information was actually available.

Compliance

Human oversight appears in regulation as a requirement to make oversight effective, not as a requirement to have a person present. Those are different obligations and the gap between them is where the exposure sits.

Map each system's oversight arrangement against what the person can actually do, and keep voluntary framework language out of statements about legal obligation.

System owners

You are accountable for a system whose intermediate decisions you may not be able to inspect. That is workable at level 3 with the right constraints and becomes steadily less workable above it.

Before raising autonomy, produce the evidence that the current level is operating: refusals, escalations, a tested stop, and a reconstructable trail.

Developers and delivery teams

Most of the controls that survive contact with a capable system are architectural. A restriction in a prompt is a request; a restriction in the permission layer is a constraint.

Build the intervention and rollback path before deployment, and measure how long it takes to reach every tool and credential.

What organisations can do now

None of this requires waiting for a standard, and none of it presumes a more capable system than the ones already deployed.

1

Keep an inventory of AI systems and agents

Including the ones nobody registered. An inventory that only contains approved systems measures your approval process, not your exposure.

2

Classify by autonomy, capability and impact

Not by department or vendor. Two systems on the same model belong in different classes if one drafts and the other acts.

3

Define permitted objectives, tools, data, transactions and counterparties

Specific enough that somebody else could design a test from the definition. 'Used responsibly' is not testable.

4

Set capability and deployment thresholds in advance

A threshold defined after it is crossed is a description of events. The value is entirely in setting it early.

5

Require evidence before a system moves up a level

Not a business case — evidence that the controls at the current level actually operated: refusals, escalations, a tested stop.

6

Test whether reviewers can detect errors

Seed known errors and see whether review catches them. This is the only direct measurement of whether human oversight is a control or a formality.

7

Log objectives, permissions, actions, changes and interventions

Enough to reconstruct a consequential decision months later, for someone who does not trust you and was not there.

8

Monitor for capability change and restriction workarounds

A model updated beneath a stable version label is a change to your system. So is an agent finding a route around a limit.

9

Separate building, validating and approving

Three roles, and the more autonomy involved the less they should be the same person. Independence is a structure, not an attitude.

10

Use independent assurance above level 3

At the point where step-level review is impossible, assurance is what is left that is not the system reporting on itself.

11

Design intervention, rollback and shutdown before deployment

And exercise them. An untested runbook is a document; the number you want is how long a stop actually takes.

12

Reassess when models, tools, permissions or context change

Each of those can move a system across a threshold without anyone deciding to raise its autonomy.

What is not solved, and what I am unsure about

Controls reduce risk. They do not resolve open technical problems, and a framework that implies otherwise is worse than no framework.

Prompt injection

Unsolved as an architectural matter. Mitigations reduce impact; none of them reliably prevent instructions in fetched content from being followed. Treat containment as the control, not detection.

Interpretability

An active research field, not a deployed assurance method. Being able to inspect some internal structure is not the same as being able to explain a particular decision to a regulator.

Alignment and specification

A system pursuing a slightly wrong objective competently is a known failure mode with no general solution. Capability makes the consequences larger, not smaller.

Deceptive or strategically withheld behaviour

Studied, not resolved. Evaluations that a system can anticipate are weaker evidence than evaluations it cannot, and distinguishing the two is itself unsolved.

Reliable shutdown

Straightforward in principle and routinely incomplete in practice — delegated agents, cached credentials, queued work, self-contained tokens that stay valid until they expire.

Open questions

  • Where exactly the oversight threshold sits for a given system, and whether it can be measured rather than argued about.
  • Whether 'effective oversight' in regulation can be evidenced at all above level 3, or whether it quietly becomes a documentation requirement.
  • What independent assurance means when the assured system is more capable than the assurer at the task being assured.
  • How to evidence that intervention remains possible, rather than asserting it from a design document.
  • Whether autonomy levels are the right axis, or whether reversibility and blast radius predict governance need better.

Limitations of this framework

  • This is a framework for thinking, built from public sources and my own experience in risk and assurance. It is not a standard, and nothing here has been validated against outcomes.
  • The five levels are a simplification. Real systems sit between them, move between them, and sometimes occupy two at once depending on which action you look at.
  • The two curves in the figure are illustrative of the argument. They are not measurements and should not be read as quantities. What the figure establishes is that a falling line and a rising line must meet; where they meet is drawn, not derived, and for a real system it would have to be argued case by case.
  • The one-to-five intensity scale saturates before the top of the autonomy scale, so levels 4 and 5 differ in a single cell. That is a limit of the instrument rather than a finding about the levels: what changes at level 5 is who the governance runs through, not how much of it there is, and a weight cannot express that.
  • Level 5 is a scenario used to stress-test the framework. Its inclusion is not a prediction that such a system exists, is imminent, or is inevitable.
  • The regulatory position summarised here changes quickly. The date at the top of the page is the date it was last checked, not a guarantee it is still current.

Sources and method

Each source is summarised in my own words and linked, and labelled with what kind of thing it is — because a voluntary framework and a binding regulation carry very different weight, and are routinely quoted as if they did not.

FrameworkNIST (US National Institute of Standards and Technology)
AI Risk Management Framework (AI RMF 1.0) and the Generative AI Profile

Voluntary. The Govern / Map / Measure / Manage structure is the closest thing to a common vocabulary; the Generative AI Profile adds risks specific to these systems. Not a certification and not binding.

StandardISO/IEC (International Organization for Standardization / International Electrotechnical Commission)
ISO/IEC 42001 — AI management systems

Certifiable management-system standard. It governs how an organisation manages AI, which is not the same as evidence that a particular system is safe at a given autonomy level.

Supervisory guidanceOSFI — Office of the Superintendent of Financial Institutions (Canada)
Guideline E-23 — Model Risk Management

Final version published 11 September 2025, in force 1 May 2027 for federally regulated financial institutions. Expands model risk management to cover AI and machine learning explicitly, with risk-proportionate lifecycle expectations.

RegulationEuropean Union
Artificial Intelligence Act — Article 14, human oversight

Binding. High-risk systems must be designed so that they can be effectively overseen by natural persons while in use. The obligation attaches to effectiveness, not to the presence of a reviewer — which is the distinction this whole framework turns on.

RegulationEuropean Union
Artificial Intelligence Act — general-purpose AI with systemic risk

Binding. Providers of general-purpose AI models — GPAI, the broad models that are adapted to many downstream tasks rather than built for one — must, where the model presents systemic risk, evaluate, assess and mitigate that risk, report serious incidents and maintain cybersecurity. The Act sets a training-compute threshold of 10²⁵ FLOP, floating-point operations, as a presumption of systemic risk: a measure of the raw arithmetic used to train the model, which is administrable precisely because it can be counted, and which is a proxy for capability rather than a measure of it.

FrameworkUnited Kingdom
AI Security Institute

Renamed from the AI Safety Institute in February 2025, with a stated shift towards security-relevant risks. Publishes evaluations and research; it is not a regulator and issues no binding requirements.

FrameworkNIST / US Department of Commerce
Center for AI Standards and Innovation (CAISI)

Renamed from the US AI Safety Institute in June 2025, with a stated focus on demonstrable risks and on standards. Evaluative and standards-setting rather than regulatory.

Industry commitmentFrontier AI developers
Frontier safety frameworks

Voluntary commitments made at the 2024 Seoul summit, implemented as published frameworks that define capability thresholds and the safeguards required before crossing them. Self-defined, self-assessed, and the closest existing practice to levels 4 and 5 — which is worth noticing in both directions.

ArgumentEliezer Yudkowsky and Nate Soares (MIRI), September 2025
If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All

A book, not a standard, and it carries no authority beyond its reasoning. Included because it states the strongest published version of the comprehension problem: that these systems are grown rather than designed, that nobody understands how their internals produce their behaviour, and that this is why a sufficiently capable system would not be controllable. The reasoning is followable and I follow it. Where it ends is a question about machine learning rather than about governance, and the section on the comprehension limit says which of the two I am writing from.

The five levels are mine, built to make one argument testable: that verification capacity falls as autonomy rises, and that governance has to intensify before the two cross. The dimensions borrow their shape from existing risk practice rather than inventing a vocabulary. Where a source is cited it is summarised in my own words and linked, so you can disagree with my reading by going to the original. Voluntary frameworks are labelled as voluntary and binding law as binding, because collapsing the two is the most common error in this area — and the most flattering one, since it makes a governance programme sound more obligatory than it is.

Independent personal work based on public sources. The examples describe no real organisation. This is educational material, not professional, legal or compliance advice. The framework, the five levels and the oversight test are my own; nothing here is a certification, a standard or a regulatory interpretation you should rely on without your own advice.