Severity and findings
Deterministic graders produce findings. Severity decides which of them fail the build. It is the mechanism behind adopting a new check without breaking everyone.
Only error fails
Section titled “Only error fails”| Severity | Reported | Fails the eval |
|---|---|---|
error | Yes | Yes |
warning | Yes | No |
notice | Yes | No |
An eval fails only if it produced at least one error-severity finding. warning and notice show up
in the report and in PR annotations, and pass.
docevals: evals: names-an-action: assertion: The page documents at least one Doc Detective action. grader: tool:regex options: { pattern: "(goTo|find|click|checkLink|httpRequest)" } severity: warning # reports, does not block no-todo-markers: assertion: The page carries no TODO, TBD or FIXME markers. grader: tool:regex options: { pattern: "\\b(TODO|TBD|FIXME)\\b", match: not-contains } severity: error # blocksseverity defaults to error.
The ratchet
Section titled “The ratchet”This is what severity is for. Adding an honest check to an existing corpus fails a lot of pages at once, which is accurate and useless, since nobody can merge it.
The move is not to weaken the assertion. It is to lower the severity:
- Add the check at
severity: warning. It reports on every page; nothing blocks. - Burn down the findings, section by section.
- Flip to
severity: error. Now it holds.
Weakening the assertion instead permanently encodes the corpus’s current state as your standard, and nobody ever tightens it back.
You can ratchet per page while you work, using an override:
evals: - use: names-an-action severity: error # this section is done; hold the line hereSeverity is not weight
Section titled “Severity is not weight”They answer different questions and it is worth keeping them apart.
severitydecides whether a failure fails the run. Onlyerrordoes.weightdecides how much an outcome moves its suite’s pass rate. It never changes the eval’s own pass/fail.
A check lifted from a style rule is usually both. Use severity: warning so it does not break the
build, and a weight below 1 so it does not swing the rate either. Reaching for severity alone
to say “this matters less” overstates the claim. It says the failure does not count at all.
What a finding carries
Section titled “What a finding carries” FAIL no-todo-markers error:14 [regex/found] Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absentseverity:line [ruleId] message. In JSON:
{ "evalName": "no-todo-markers", "file": "docs/actions/goTo.mdx", "ruleId": "regex/found", "message": "Pattern /\\b(TODO|TBD|FIXME)\\b/ found in body, expected absent", "severity": "error", "line": 14}ruleId, line, and col are present when the grader can supply them. tool:regex names the
line in the file, which is what makes --format github able to annotate the exact line.
The same eval at warning reports and passes. docs/get-started/concepts.md overrides
no-todo-markers to severity: warning, so its TODO shows as warning:21 [regex/found] under a
pass.
Severity also maps onto annotation levels: error → ::error, warning → ::warning, notice →
::notice. See Output and exit codes.
Severity is not the judge’s
Section titled “Severity is not the judge’s”Severity applies to deterministic graders. A judged eval’s outcome comes from consensus and confidence zones instead. See How judging works. The capability suite target is the judged analog of the ratchet; see Regression vs capability.