Skip to content

How docevals works

One concept, one preference order, one kind of verdict. Everything else in these docs is a consequence of the three ideas on this page.

An eval is a named, testable assertion about a page, plus a grader that decides pass or fail.

- id: install-command-present
assertion: The page shows the command that installs Doc Detective.
grader: tool:regex
options:
pattern: "npm i -g doc-detective"

Matching a pattern, running a generated Node script, asking a model whether a page over-promises, and routing something to a human are all evals. What differs is the grader.

That matters practically. One config, one report, one exit code, and one place to look when something goes red.

Four graders sit in a preference order, cheapest and most explicable first.

1. Code, via command and tool:regex. Deterministic, free, fast, and arguable in a pull request. tool:regex matches a pattern in the page. command runs a script, either one you wrote or one manni docevals generated from the assertion, and passes on its exit code. Any CLI check your corpus needs fits in a command eval.

Checks with a home elsewhere in manni run there. Frontmatter is validated against its schema by manni meta validate, and page structure is checked by manni lint. Each reports on the same severity scale with the same exit codes, so a pipeline runs them side by side.

2. The AI judge, ai. For assertions that need interpretation: “does this page explain why before how?” No linter expresses that. The judge does, with safeguards described below.

3. A human, via human and the review zone. For what neither code nor judge should decide alone.

The rule to carry is simple. If you can express the criterion as code, do. An ai eval is the choice of last resort, not the default. Left alone, everything drifts upward into the expensive tier, which is why promote and generate exist.

Deterministic graders produce findings with a severity. An eval fails only on an error-severity finding; warning and notice report and pass. That is the mechanism behind entering a new check at warning and ratcheting it up.

The one exception is a finding that says the grader never reached a verdict. A command eval has no script yet, its command could not start, or it timed out. Those are marked diagnostic and fail the eval at any severity. Passing is a claim about the page, and a command that never ran makes no claim at all. Severity still controls how the finding is displayed.

The judge runs the eval N times independently (default 3), each a separate request with no shared context so the runs cannot reinforce one another. Consensus aggregates them, then confidence decides the zone:

ZoneConditionOutcome
auto-passpassing consensus, mean confidence ≥ judge.zones.autoPasspass
auto-failfailing consensus, mean confidence ≥ judge.zones.autoFailfail
human reviewanything betweenneeds-review

Two asymmetries are deliberate. A partial verdict counts as a fail, and an errored run counts against consensus. Errors can push an eval toward human review; they can never produce a silent pass.

Verdicts are cached by content, so an unchanged page with an unchanged assertion never re-judges.

Every eval is pass or fail. That looks crude for prose until you see where the nuance lives: the suite pass rate.

  • A regression suite guards behavior that must keep working, at target 1.0.
  • A capability suite measures reach, at target ~0.7.

So “70% of our how-to pages explain why before how” is expressible. No individual eval has to return a mushy 0.7. type defaults to regression, because most evals guard something. See Regression vs capability.

discover pages
→ resolve frontmatter against config (suites, named evals, inline evals)
→ generation pass (write scripts for command evals that lack one)
→ deterministic graders (cheap first)
→ AI judge
→ apply persisted human reviews
→ aggregate per suite → exit code

Deterministic graders run before the judge on purpose: they are free, and a page that fails a cheap check rarely needs an expensive opinion.