Glossary
Assertion : The claim an eval makes about a page, true or false, not an aspiration. “The page states its prerequisites before the first command” is an assertion; “the page is well-written” is not. See Write good assertions.
Capability suite
: A suite that measures reach rather than correctness, so its target-pass-rate sits below 1.0
(~0.7 is typical). Contrast regression suite.
Confidence zone
: Which of three bands a judged eval lands in. At or above judge.zones.autoPass a passing
consensus auto-passes; at or above judge.zones.autoFail a failing one auto-fails; anything
between goes to human review. The judge is being asked to know when it is unsure, not to be right
every time.
Consensus
: How N independent judge runs become one verdict. A partial verdict counts as a fail, and an
errored run counts against consensus. Errors can only push an eval toward human review,
never toward a silent pass.
Ensemble
: The N independent judge runs themselves (judge.ensembleRuns, default 3). Each is a separate
request with no shared context, so they do not reinforce one another.
Eval : The unit. A named, testable assertion plus a grader that decides pass or fail. There is one concept, not a split between “runners” and “evals”. What differs between checks is the grader.
Evidence
: The evidence field, naming what part of the page the judge should look at. Scoping it stops the judge
reasoning from sections the assertion was never about.
Finding
: A normalized result from a deterministic grader: eval name, file, message, severity, and usually a
rule id and line. Only error-severity findings fail the build.
Golden set
: 20–50 human-verified cases in .manni/docevals/golden/*.yaml that
calibration scores the judge against. This is the artifact you hand a
skeptic.
Grader
: What decides an eval. command and tool:regex are code, ai is the judge, human is a person.
They sit in that preference order. See Graders.
Grader hierarchy
: The rule that code beats judge beats human. If you can express the criterion as code, do. It is
cheaper, faster, and explicable in a review. promote and generate are how you keep moving evals
down it.
Human-review zone : The band between the two confidence thresholds, where verdicts route to a person. It is not an admission of failure. It is what makes a binary verdict acceptable.
Pass rate
: The weighted share of a suite’s graded evals (pass, fail, and error) that passed. Each eval or
criterion counts by its weight. Verdicts are binary; the nuance lives here, which is why
a suite carries a target and an individual eval does not.
Regression suite
: A suite guarding behavior that must keep working, so its target-pass-rate is 1.0. The default
eval type is regression because most evals guard something.
Severity
: error, warning, or notice on a deterministic eval. Only error fails the build, which is
what lets you enter a new check at warning and ratchet it up. See
Retrofit a legacy corpus.
Suite : A named group of evals with a target pass rate, applied to a page by name.