Output oversight concerns the artifacts a model produces so that a human can understand its work. As models do more of the work, human time moves to oversight and understanding, and a long prose reply is a poor artifact for that job.1

Overview

In place of default prose, the source asks for a discardable explainer.1 A diagram, a page, or a bespoke explainer video can be a lot easier to process, parse, and understand than prose.1 That artifact serves as evidence for oversight and is then thrown away.1

Mechanism

  1. Controlled language is requested when prose is required.1
  2. A diagram is preferred when a figure is easier to check than paragraphs.1
  3. A discardable page is requested when interaction helps.1
  4. Even better is a bespoke explainer video, the output format the post is most bullish on.1
  5. The aim is an artifact that is a lot easier to process, parse, and understand than prose.1

Applications

The approach applies when the job is to understand model work, not to ship a product.

Limitations

A format ladder is not a substitute for a separate judge, and the explainer is not itself the judge (see Verification and stop conditions).

Worked example

The source post lays out a ladder.1

  1. When the answer must be prose, controlled language is requested: ASD-STE100, or about 80% of the way there.
  2. Even better is a diagram, which is easier to check than paragraphs.
  3. Even better is a throwaway, interactive HTML page.
  4. Even better is a bespoke explainer video, the output format the post is most bullish on.

The post’s reason is that intelligence and code are now cheap enough that custom, disposable software for understanding is worth building. The artifact is then discarded.

See also

References

Footnotes

  1. https://x.com/karpathy/status/2105819303471976479 ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10