Agentic harness engineering treats the files around a model as the object of change. A harness edit is a file-level action that can be reverted, and every edit declares a prediction that the next round checks against outcomes.1
Overview
Large trajectories are distilled into evidence that the evolving process can use.1 Some behavior belongs in an inner skill rather than in more orchestration glue.2 A workflow that wins in one harness can lose in another, even with the same prompt.3
With the backbone model fixed, papers are grouped into mechanism families per harness module. One module is mutated at a time, and a combination is kept only when the validation gain holds. Later cycles can be driven by new papers rather than only by observed failures.4
A separate line of work trains a proposer, with optional teacher supervised fine-tuning and then GRPO on harness scores. With both models frozen, test-time execution feedback revises the solver’s harness code, and that revision skill transfers to unseen tasks.5 This differs from paper-sourced module mutation with a fixed backbone, and from an untrained propose, validate, and revert loop: the proposer is trained, and at test time the weights do not move.54
Mechanism
- Each editable harness piece is represented as a file.1
- A change is edited in, the tasks are run, the prediction is checked, and the change is kept or reverted.1
- Tool, middleware, and memory changes are preferred before the system prompt is enlarged.1
- Editorial note: unbounded generation is separated from bounded decisions (see Verification and stop conditions).
- A mechanism is kept only if it survives selection across more than one environment.6
- When a harness learns from a stronger model’s trace, only the failing turn is rewritten.7
- With the backbone fixed, papers are grouped into mechanism families per harness module and one module is mutated at a time; a combination is kept only when the validation gain holds, and a later cycle can be driven by new papers.4
- A trained proposer revises the solver’s harness code from test-time execution feedback while both models stay frozen.5
Applications
The approach applies when the files around a model need to change and the change must be checkable.
Limitations
The pattern does not license a self-evolving harness without a stop condition. A stronger model’s full trace is not copied into a harness built for a weaker planner; only the failing turn is rewritten.7 A score measured in a different harness is not evidence that a change works in the target harness.3
Worked example
SoL-Pi illustrates the approach.6
- Instead of one harness being hand-tuned, auto-research loops run at the harness layer.
- Those loops are scaled across many and varied environments, so a change is tested on more than the setting it was found in.
- Only the mechanisms that survive selection are kept. In the paper, four survive: one for action execution, one for context compaction, one for observation handling, and one for delegated reading.
- Because the search was not tied to one environment, the paper reports that the kept mechanisms carry over beyond their development setting.
See also
- Verification and stop conditions – generation versus bounded decisions, and the decision-density audit
- Autoresearch loop – a metric loop on task code, as distinct from harness meta-evolution
- Harness runtime – the runtime of a single harness
- Thin harness, fat skills – the thin-loop philosophy
- Loop engineering – designing loops that prompt agents, above the harness runtime
Further reading
- https://x.com/omarsar0/status/2102485592420606054
- https://x.com/omarsar0/status/2097958286146605446
- https://x.com/omarsar0/status/2049492169887748365
- https://x.com/omarsar0/status/2051678102934454330
- https://x.com/omarsar0/status/2106619582463312315
- https://x.com/omarsar0/status/2106529236068819394