Skip to content

Retrofit a legacy corpus

Point an honest quality bar at a corpus that has never been measured and nearly everything fails at once. The result is accurate and useless: unmergeable, untriageable, and it teaches your team that the tool is wrong rather than that the docs are.

The instinct that follows is to soften the assertions until the build goes green. That is the failure this page exists to prevent, because it is irreversible in practice. An assertion weakened to accommodate the corpus permanently encodes the corpus as the standard, and nobody ever tightens it back.

Turn the check on at error today. Record what is already broken. Fail only on what is new.

A baseline is a committed file listing the findings your corpus has right now. A run subtracts it before deciding anything, so pre-existing findings report as accepted and a fresh one fails the build. The standard tightens this Monday instead of after a cleanup nobody staffed.

docevals:
baseline: .manni-docevals-baseline.json

Do this before recording. baseline: is what makes an ordinary run read the file; without it a recording run writes a file nothing will ever load, and says so:

Terminal window
warn .manni-docevals-baseline.json Recorded .manni-docevals-baseline.json, but `baseline:` in
/repo/manni.config.yaml is not set. An ordinary run will not read this file — point the key at it,
or pass --baseline .manni-docevals-baseline.json.
Terminal window
npx @hawkeyexl/manni docevals run --write-baseline

A bare --write-baseline records into the configured path, not the default. That matters if you put the file somewhere other than the repo root. Without that rule, a repo pointing baseline: at config/docs-baseline.json would record into .manni-docevals-baseline.json. Nothing would read it, and the ratchet would do nothing while looking like it worked.

Terminal window
Baseline .manni-docevals-baseline.json: recorded 412 finding(s) (+412, -0).

The recording run applies the baseline it just wrote, so it exits 0. Recording a finding is declaring it accepted. You do not have to wrap the command in || true, which is the reliable way to lose the exit code that matters on the next run.

Terminal window
git add .manni-docevals-baseline.json

It is team state, not a cache. Sorted keys and sorted fingerprints mean it diffs and merges legibly, and every later change to it shows up in review, which is the point. Shape and re-record mechanics are in Files and state.

Nothing more to configure. An ordinary run now reports what it forgave and fails only on the rest:

Terminal window
Baseline .manni-docevals-baseline.json: 412 finding(s) suppressed of 412 recorded.

Add a page with a fresh violation and that run exits 1, with only the new finding in the report. Use --no-baseline to see the full backlog again for one run without touching the file.

What a baseline covers, and what it does not

Section titled “What a baseline covers, and what it does not”

Read this before you rely on it. The gate a baseline gives you is real but weaker than it looks.

Identity is per rule per file, not per occurrence

Section titled “Identity is per rule per file, not per occurrence”

A finding’s identity is (file, eval name, rule id). The line number and the message text are deliberately excluded. A fingerprint that moved with the line would present a pure reordering as a wall of new findings. One that included the message would be invalidated for every consuming repo the next time a grader rewords a string.

The consequence: a file baselined for a no-todo-markers finding will not fail when a second TODO appears. That is not an oversight waiting on a fix. Per-occurrence identity needs a stable anchor, and prose has none. A text snippet churns on every edit, and an ordinal renumbers on every insertion.

Plan around it. A baseline stops a new kind of problem entering a file; it does not stop more of a problem that file already has. Where you need the stricter guarantee, fix the file and drop its entry.

It covers findings, which means deterministic graders

Section titled “It covers findings, which means deterministic graders”

command and tool:regex graders produce findings, so those are what a baseline records. An ai-graded eval’s verdict carries no rule identity to fingerprint. Freezing a judge verdict in a backlog file is not what you want anyway. Human review is the mechanism for those.

Two more things a baseline does not touch:

  • An eval that errored. error is “the check could not run”, not “the page is bad”, and there is no finding to record. It still exits 1.
  • A suite below its target-pass-rate. Suppressed findings do raise the pass rate, because suppressing the last error-severity finding makes the eval pass. The target itself is a separate lever. See Set capability targets near reality below.

Renaming an eval invalidates its fingerprints

Section titled “Renaming an eval invalidates its fingerprints”

The eval name is part of the identity, so the tool cannot tell no-todo and no-todo-markers are the same check. Rename one and the next run reports that eval’s whole backlog as fresh. The remedy is a re-record; the surprise is worth knowing before it happens in CI.

Every re-record reports both halves of the diff:

Terminal window
Baseline .manni-docevals-baseline.json: recorded 1 finding(s) (+0, -1).
1 previously recorded finding(s) are no longer in the baseline. If this run covered less of the
corpus than the last one, they have just been forgiven.

removed is the load-bearing number. A baseline’s failure mode is over-forgiveness, and it is silent by construction. The whole job of the file is to make a red run green. An accidental --write-baseline over a narrowed glob, a mistyped collection exclude, a run with the wrong --collection: each forgives everything it did not see. The fingerprints are opaque hashes, so a diff dropping 200 of them reads as noise unless a number names it.

Re-record deliberately, over the same scope as your gate, and read that line.

A hand-broken baseline is rejected rather than silently ignored, for the same reason. A typed fingerprint can never match a real finding, so the symptom would otherwise be “that finding came back” with nothing to explain it:

Terminal window
manni: Baseline ".manni-docevals-baseline.json": entries["docs/install.md"] contains
"not-a-fingerprint", which is not a fingerprint (16 lowercase hex characters).

Exit 2, the pipeline owner’s problem, not the author’s.

The baseline is what makes day one mergeable. These steps are what make the following quarter converge.

5. Decide what should not be evaluated at all

Section titled “5. Decide what should not be evaluated at all”

Triage before you record. Deprecated sections, generated reference, and archives should be excluded, not baselined. A baseline entry is a promise to come back, and nobody is coming back for docs/archive/.

Exclude them from the collection the run reads. node_modules is never read, so it needs no entry:

collections:
- name: site
paths: ["docs/**/*.{md,mdx}"]
exclude:
- "docs/archive/**"
- "docs/api/generated/**"

Per page:

---
title: Deprecated integration guide
eval-skip: true
---

This is a legitimate first step, not giving up. A reader who feels that excluding is cheating will instead spend a week fixing content nobody reads.

Do it first, because narrowing scope after a baseline exists is exactly the shape that produces a scary removed count on the next re-record.

Terminal window
npx @hawkeyexl/manni docevals fill --dry-run --max-turns 40 docs/how-to/

Batching keeps the review reviewable and the work bounded. See Bootstrap a corpus.

A baseline forgives findings. Suites are the judged analog, and their lever is the target pass rate:

docevals:
suites:
how-to-adopting:
target-pass-rate: 0.6 # where the corpus is today
evals: [explains-why-before-how, no-future-promises]

Set the target near the current rate, then raise it as pages improve. “60% of how-to pages explain why before how, and last quarter it was 45%” is a real report. See Regression vs capability.

Pick the section that matters most, usually the one with the most traffic, not the most findings. Fix its pages, then re-record so the fixed findings leave the file. Fix a failing eval is the page to hand contributors.

Findings you fix without re-recording still pass; they just linger in the file as entries that no longer occur, and a run says so:

Terminal window
Baseline .manni-docevals-baseline.json: 340 finding(s) suppressed of 412 recorded, 72 no longer occur.

That number falling is the burndown. Re-record when it is worth the diff.

A retrofitted corpus is disproportionately ai-graded, because fill proposes ai evals by construction. Before the recurring cost becomes the story, push them down the hierarchy with Promote to deterministic.

When severity inversion is still the answer

Section titled “When severity inversion is still the answer”

Dropping an eval to severity: warning reports everywhere and blocks nowhere. It is the narrower tool, for one specific case:

A finding class you do not intend to fix, on pages you do not intend to exclude.

A banned word inside quoted error messages, or a TODO in vendored reference, is worth seeing and not worth gating. Say that with severity, which is honest about the class, rather than with a baseline, which claims you will come back.

SituationReach for
Existing findings you intend to fix eventuallyBaseline
A whole finding class you will never gate onseverity: warning
Content that should never have been in scopeA collection’s exclude / eval-skip
A judged eval the corpus does not meet yettarget-pass-rate below 1.0

Severity still ratchets per page and per section while you work, which is the fine-grained companion to a corpus-wide baseline:

evals:
- use: no-todo-markers
severity: error # this section is done; hold the line

What good looks like at the end of a quarter

Section titled “What good looks like at the end of a quarter”
  • Every eval on at error, gating new work from day one.
  • A committed baseline whose recorded count is visibly falling.
  • Capability targets set near reality and rising.
  • No assertion weakened to get there.
TemptingDo instead
Loosen the assertion until the build is greenKeep it; record a baseline and gate on new findings
Turn every eval down to warning to adopt itReserve warning for classes you will never gate on
Re-record the baseline whenever CI goes redRead the removed count; a red run is usually a real new finding
Set every suite target to what today’s corpus scoresSet it slightly above, and raise it
Run fill over everything and merge in one PRBatch by directory
Delete the check because it “doesn’t fit our docs”Ask whether the check is right and the docs are wrong

The first column produces a green build immediately and a meaningless standard permanently.