How docevals works
One concept, one preference order, one kind of verdict. Everything else in these docs is a consequence of the three ideas on this page.
One concept, the eval
Section titled “One concept, the eval”An eval is a named, testable assertion about a page, plus a grader that decides pass or fail.
- id: install-command-present assertion: The page shows the command that installs Doc Detective. grader: tool:regex options: pattern: "npm i -g doc-detective"Matching a pattern, running a generated Node script, asking a model whether a page over-promises, and routing something to a human are all evals. What differs is the grader.
That matters practically. One config, one report, one exit code, and one place to look when something goes red.
The grader hierarchy
Section titled “The grader hierarchy”Four graders sit in a preference order, cheapest and most explicable first.
1. Code, via command and tool:regex. Deterministic, free, fast, and arguable in a pull
request. tool:regex matches a pattern in the page. command runs a script, either one you wrote
or one manni docevals generated from the assertion, and passes on its exit code. Any CLI check
your corpus needs fits in a command eval.
Checks with a home elsewhere in manni run there. Frontmatter is validated against its schema by
manni meta validate, and page structure is checked by
manni lint. Each reports on the same severity scale with the same exit codes, so
a pipeline runs them side by side.
2. The AI judge, ai. For assertions that need interpretation: “does this page explain why
before how?” No linter expresses that. The judge does, with safeguards described below.
3. A human, via human and the review zone. For what neither code nor judge should decide alone.
The rule to carry is simple. If you can express the criterion as code, do. An ai eval is the
choice of last resort, not the default. Left alone, everything drifts upward into the expensive
tier, which is why promote and generate
exist.
How a verdict is reached
Section titled “How a verdict is reached”Deterministic graders produce findings with a
severity. An eval fails only on an error-severity finding; warning and notice report and pass.
That is the mechanism behind entering a new check at warning and
ratcheting it up.
The one exception is a finding that says the grader never reached a verdict. A command eval has
no script yet, its command could not start, or it timed out. Those are marked diagnostic and fail
the eval at any severity. Passing is a claim about the page, and a command that never ran
makes no claim at all. Severity still controls how the finding is displayed.
The judge runs the eval N times independently (default 3), each a separate request with no shared context so the runs cannot reinforce one another. Consensus aggregates them, then confidence decides the zone:
| Zone | Condition | Outcome |
|---|---|---|
| auto-pass | passing consensus, mean confidence ≥ judge.zones.autoPass | pass |
| auto-fail | failing consensus, mean confidence ≥ judge.zones.autoFail | fail |
| human review | anything between | needs-review |
Two asymmetries are deliberate. A partial verdict counts as a fail, and an errored run counts
against consensus. Errors can push an eval toward human review; they can never produce a silent
pass.
Verdicts are cached by content, so an unchanged page with an unchanged assertion never re-judges.
Verdicts are binary; pass rates are not
Section titled “Verdicts are binary; pass rates are not”Every eval is pass or fail. That looks crude for prose until you see where the nuance lives: the suite pass rate.
- A regression suite guards behavior that must keep working, at target 1.0.
- A capability suite measures reach, at target ~0.7.
So “70% of our how-to pages explain why before how” is expressible. No individual eval has to
return a mushy 0.7. type defaults to regression, because most evals guard something. See
Regression vs capability.
The pipeline
Section titled “The pipeline”discover pages → resolve frontmatter against config (suites, named evals, inline evals) → generation pass (write scripts for command evals that lack one) → deterministic graders (cheap first) → AI judge → apply persisted human reviews → aggregate per suite → exit codeDeterministic graders run before the judge on purpose: they are free, and a page that fails a cheap check rarely needs an expensive opinion.
- Declare evals in frontmatter, the contract
- Write good assertions, the skill that decides whether any of this works
- Trust the judge, the mechanics above, in full, plus how to prove them