Harness notes is a collection of reference pages on the design of agent harnesses: the loop that runs a language model, the skills the loop loads, the checks placed on the loop’s output, and the memory it reads. Its recurring themes are that the loop stays thin and routes, skills hold the procedure, checks stop the builder from grading itself, and memory is a wiki compiled once instead of a search on every question.

Key lessons

  1. A wiki is compiled once from immutable sources and then queried, instead of the raw sources being re-read on every question. Karpathy LLM wiki foundation
  2. The harness stays thin: it loops the model, reads and writes files, manages context, and enforces safety. The procedure lives in skills that load when the task matches. Thin harness, fat skills
  3. The builder is not the judge. A separate check needs independent evidence, not a self-report, and time, iteration, cost, and scope limits belong in the loop, not only in the prompt. Verification and stop conditions
  4. A skill is a folder, not one file loaded every time. SKILL.md states when to load the skill and its constraints, and heavy docs stay in references until they are needed. Skillify authoring
  5. A loop needs feedback and a halt. Missing validation and loops without an exit are gaps, and stops such as max iterations and max cost belong in hard logic, not in a prompt the model can talk past. Loop engineering Verification and stop conditions
  6. A harness edit is a file change that can be reverted. Each edit declares a prediction, and the next round checks it against outcomes. Agentic harness engineering

Memory

A wiki is compiled from immutable sources and then queried, unlike RAG. Karpathy LLM wiki foundation

Avid memory stack Shared brain Ephemeral wiki compile-lint Design docs as source

Loop

The harness stays thin and routes; skills stay fat. Thin harness, fat skills

Loop engineering Harness runtime Multi-agent teams Autoresearch loop

Checks

The builder is not the judge. Verification and stop conditions

Output oversight

Skills

A skill is a folder, not one always-loaded file. Skillify authoring

Editorial note: a failure adds a gotcha to the skill, and a skill is rewritten only from real runs. A model may draft a skill or its evals; a human edits and signs it, and nothing ships unedited (see Contradictions).

Tool notes

Tool notes describe third-party products. The conventions for them are on Conventions.

Suggested reading order

  1. Karpathy LLM wiki foundation
  2. Thin harness, fat skills
  3. Verification and stop conditions
  4. Skillify authoring
  5. Ephemeral wiki compile-lint
  6. Multi-agent teams

Notes for automated readers

The pages are arranged so that an automated reader can start at the index, open one page, and cite that page’s path without inventing a source. The references on a page are evidence, not policy. Where two pages disagree, both are cited and the disagreement is reported as a conflict. A claim that is not on a page is reported as absent.

See also