Skip to content

Bootstrap a corpus with fill

Ten minutes to write a good eval, times two hundred pages, is not a project anyone starts. manni docevals fill asks your configured LLM to propose evals for each page from its content, each with a self-reported 0–1 confidence.

Always --dry-run first:

Terminal window
npx @hawkeyexl/manni docevals fill --dry-run docs/
Terminal window
proposed docs/tests/overview.mdx +3 evals (all-test-statuses-defined 0.90, test-hierarchy-explained 0.85, skip-conditions-enumerated 0.75) — dry run, not written
proposed docs/tests/writing.mdx +2 evals (steps-are-numbered 0.88, examples-are-runnable 0.71) — dry run, not written
Threshold: 0.7 · inference calls: 2

Two reasons, and the second is the one people miss. You see every proposal before it touches the repo, and the dry run is the pass. One page costs one inference call, and its raw proposals are cached, so dropping --dry-run afterwards replays them and makes no further calls. Looking first is free.

Raw proposals are cached before the confidence gate. Re-running at a different --confidence costs nothing:

Terminal window
npx @hawkeyexl/manni docevals fill --dry-run --confidence 0.85 docs/

Knowing this changes how you work. Without it, the threshold feels like a one-shot irreversible decision, so people agonise, pick wrong, and pay twice. Sweep it instead, across 0.6, 0.7, and 0.85, and look at what each admits.

Terminal window
npx @hawkeyexl/manni docevals fill --max-turns 40 docs/

--max-turns (and fill.maxTurns) counts uncached inference calls, exactly one per page filled. Pages already in the cache do not count against it, and a page the budget cannot cover is reported as skipped with (turn budget exhausted) rather than silently dropped. Set it even when you are confident; the run you regret is the one you did not bound. See Caching and turn budgets.

Drop --dry-run:

Terminal window
npx @hawkeyexl/manni docevals fill docs/

Only proposals at or above the threshold (default 0.7, config fill.confidenceThreshold) are appended to the page’s frontmatter. The rest are reported so you can see what was nearly good enough.

What fill will not do:

  • Modify an existing eval. Ever. It only appends.
  • Create duplicates. A proposal whose name collides with the page’s resolved plan is dropped, whether that plan is inline, referenced, or suite-expanded.
  • Touch a skipped page. eval-skip: true is respected.

fill.maxEvalsPerPage (default 3) bounds how many it proposes per page.

fill proposes; it does not decide. This step is the job, not a formality.

Proposals are always ai-graded with explicit examples. That is the right conservative default for a machine-written assertion, and it also means an unreviewed corpus is now an expensive corpus. Read them as assertions:

  • Is it judgeable, or is it a feeling? See Write good assertions.
  • Does it encode the page’s current state as the standard? A proposal derived from the content will happily assert what the page already does, which measures nothing.
  • Should it be code? Most “the page contains X” proposals should be, and promote finds them.

Batch by directory so the review stays reviewable. Nobody can approve a 3,000-page pull request.

SituationPage
The corpus has never been measured and an honest run would be all redRetrofit a legacy corpus
Everything is an ai eval and the bill scales with the corpusPromote to deterministic
You want the generated scripts reviewed properlyReview generated scripts