The harness runtime is how one harness handles durable instructions, gates tools, compacts context, and runs hooks.1 Read-only tools may run together, while mutating tools run one at a time.1
Overview
Permissions live in config, deny first, with more than one layer.1 Compaction is a ladder applied before the context window runs out, and hooks can refuse a tool call.1
Mechanism
- Read-only tools are batched, writes are serialized, and oversized results are moved out of the window.1
- A permission cascade is applied.1 As defense in depth, the loaded rules are verified, because a deny rule can fail to match.2
- Context is compacted in stages before each model call.1
- Lifecycle hooks serve as gates.1 Editorial note: work is called done only against a task contract plus evidence.
- Gate thresholds sit in one testable policy function with three verdicts: run, ask a person, block. What each part does when its classifier cannot be reached is decided in advance: a tool gate fails closed, and a model router falls back to the stronger model. The gate judges the tool call itself, and anything code can check for certain, such as a path outside the project, stays in code.3
- Each step is checkpointed before the loop moves on, and a tool call’s intent is stored before it runs. After a crash, only tools marked safe to replay are rerun; for any other tool, the model is told the call was interrupted and decides what to do. Each submission has an id, so a retry cannot run it twice.45
- Where staged compaction loses working state, the model can be given its live context as a file it may edit; each edit becomes the next turn’s context.6
Applications
The runtime rules apply when a thin loop still needs concrete rules for tools, permissions, and context.
Limitations
Bypass is not the default permission mode.1 A deny list alone is not enough.2
Worked example
The source paper’s running example is the request “Fix the failing test in auth.test.ts”.1
- The request enters one simple loop: call the model, run the tools it asks for, repeat.
- Before every model call, a staged compaction pipeline shapes the context, cheaper steps first and a full summary last.
- The model asks to run the test command. The request passes a deny-first permission check: deny rules beat ask rules, ask rules beat allow rules, and an unmatched risky action asks the user.
- Read-only tool calls may run in parallel, while calls that change state run one at a time. Results go back into the loop in request order.
- If the classifier or a deny rule blocks a call, the model receives the denial reason, revises its approach, and tries a safer alternative on the next loop iteration.
- The loop stops when the model answers with text only, or earlier on a turn limit, a context overflow, a hook that halts it, or an abort.
See also
- Hermes, Block Buzz – tool instances
- Thin harness, fat skills – the thin loop and fat skills philosophy
- Agentic foundations – files and git persistence
- Verification and stop conditions – hooks and automatic permission
- Multi-agent teams – fork, teammate, and worktree shapes
- ML Intern – a dual runtime with a doom-loop check
- Agentic harness engineering – evolving the harness
- Loop engineering – the concept layer above the runtime
Further reading
- https://x.com/0xcodez/status/2062926176117469477
- https://x.com/iiiichigo_chan/status/2099482995849666786
- https://arxiv.org/abs/2609.37725