Verification and stop conditions keep the builder from being the judge and keep stops as hard limits the model cannot talk past. The judge needs independent evidence, not a self-report, and time, iteration, cost, and scope are enforced in the loop, including no external send without approval.1

Overview

Agreement is not a check. A frozen copy of the same model can verify a pool of rollouts against the environment, and unsettled claims stay marked.2 The numbers reported for this approach are the authors’ own.

Mechanism

  1. The builder produces the artifact.1
  2. A separate judge reads the brief and the work and returns a verdict with evidence.1
  3. Time, iteration, cost, and scope are enforced in the loop, including no external send without approval.1
  4. A cheap completion check runs on the goal,3 and the work escalates when confidence is low or the stake is high.4
  5. The stop rules are proven with failure tests, and the suite is rerun if any test fails.5
  6. A role is written in five fields before the day’s assignment: Owns, Inputs, May, Must ask before, Done when.6
  7. Before any tool is used, the role is trialled in text: what it would do, in order, naming every tool and every unsure judgment.6
  8. Done is a checkable list, not “useful”. When unsure, the agent stops and asks.6
  9. Probation is three runs: observe, correct on a comparable task without hand-reminding, then release.6
  10. The rule is repaired, not the artifact. A transient tool failure is retried twice, malformed output is repaired once, conflicting evidence means stop and ask, three failed corrections mean escalation, and the cost ceiling means stop.6
  11. Autonomy is a ladder that can be demoted: 0 observe, 1 prepare, 2 execute with approval, 3 scheduled or triggered, 4 coordinate. Promotion comes only after five consecutive clean runs, verification every time, zero unresolved side effects, rollback tested once, and the approval park actually seen. Demotion comes if quality drops, an integration changes, or the agent needs correcting two weeks in a row.6
  12. A weekly receipt is kept (runs, passed, human repairs, runtime, repeated failure, keep-or-drop), one artifact is spot-checked, and a routine nobody would miss is deleted.6
  13. A frozen pool of rollouts is checked in the task environment, not by voting. A disagreement resolver tests competing claims against the environment and drops what the evidence contradicts. A consensus challenger tests how a shared claim or an omission could be wrong. Unsettled claims stay marked, and the writer still does not grade itself.2
  14. The reviewer comes from a different model family than the builder, because models from one family tend to miss the same mistakes. Long runs declare their phases and gates up front, the runtime enforces them, and negative results are carried forward with the positive ones.7

Applications

The approach applies to bounded unattended or parallel work that must not grade itself.

Limitations

Stops placed only in the prompt can be talked past, and a writer that verifies itself is not a check.1 Agreement between models is not a check either.2

Worked example

The article quoted in the source post describes unattended parallel work.1

  1. Only independent tasks run unattended, each agent in its own isolated workspace.
  2. The builder never grades its own work. A separate judge checks independent evidence, such as tests, the requirements, and real execution, not a self-report. High-stakes work gets a second pass in a fresh context with no memory of how the work was produced.
  3. Hard limits are set in logic, not in the prompt: maximum time, maximum iterations, and a cost ceiling the system cannot reason past.
  4. An explicit scope boundary is set: no deploy, delete, spend, or external send without approval.
  5. The evidence is reviewed afterwards, not the success flags. The work starts small and scales only after the checks and the stops have been tested.

See also

Further reading

References

Footnotes

  1. https://x.com/cyrilxbt/status/2087057424373182860 ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  2. https://arxiv.org/abs/2610.00972 ↩ ↩2 ↩3

  3. https://x.com/omarsar0/status/2101443311454036477 ↩

  4. https://x.com/i/article/2101413250797572096 ↩

  5. https://x.com/kocer_eth/status/2103133414333218876 ↩

  6. https://x.com/i/article/2105262478535847936 ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7

  7. https://arxiv.org/abs/2609.39551 ↩