The autoresearch loop divides the work between a human, who iterates the prompt, and an agent, which iterates one allowed code file. Each run has a fixed budget and is kept or discarded by a named metric, and accepted experiments are recorded as commits.12

Overview

The same loop can target any metric that is efficient to evaluate.3 An outer loop may edit how the inner loop searches, and that change is import-validated or reverted.45

Mechanism

  1. The agent modifies the allowed file, runs under a fixed budget, and keeps or discards the change by the metric.1
  2. Additive keeps are stacked, and promising ideas are promoted from a smaller scale to a larger one.36
  3. An outer loop proposes a new search mechanism, and a mechanism that is not import-validated is reverted.45
  4. Harness edits are regularized with a shrinking budget, untried directions, a critic, and a pruner.7
  5. Long runs are parallelized, share memory across failures, and keep a playbook of failed attempts.8
  6. The next environment is hardened from the objective on a schedule, not only from the agent’s own misses.9
  7. When one search path stalls, it is split into branches. Each branch keeps the cases it solves better than the others, drops cases every branch already solves, and revises its own way of proposing edits. Which branch’s result to use is chosen per input from development data only.10

Module search from published papers, and trained test-time harness revision, are covered on Agentic harness engineering.

Applications

The loop applies when progress is a measured keep-or-discard loop rather than a wiki compile.

Limitations

A metric keep is not permission to ship, and an overnight run needs hard stops (see Verification and stop conditions). A gain on the search tasks is not proof that the harness improved on real tasks.

Worked example

The loop in the source post and its repository runs in four steps.12

  1. The human edits the agent prompt, program.md. The agent may edit one file, the training code train.py. The data preparation file stays fixed.
  2. The agent changes the code and trains for a fixed five-minute wall-clock budget.
  3. The run is kept or discarded by one metric, validation bits per byte, where lower is better.
  4. A kept change is committed on a feature branch, so the accepted experiments are the git history.

The post notes that research progress can then be compared across prompts and across agents.1

See also

Further reading

References

Footnotes

  1. https://x.com/karpathy/status/2030371219518931079 ↩ ↩2 ↩3 ↩4

  2. https://github.com/karpathy/autoresearch ↩ ↩2

  3. https://x.com/karpathy/status/2031135152349524125 ↩ ↩2

  4. https://github.com/EdwardOptimization/Bilevel-Autoresearch ↩ ↩2

  5. https://arxiv.org/abs/2603.23420 ↩ ↩2

  6. https://github.com/karpathy/nanochat/commit/6ed7d1d82cee16c2e26f45d559ad3338447a6c1b ↩

  7. https://arxiv.org/abs/2609.24972 ↩

  8. https://arxiv.org/abs/2609.10922 ↩

  9. https://arxiv.org/abs/2609.04128 ↩

  10. https://arxiv.org/abs/2609.37834 ↩