Output and exit codes
Verified against src/docevals/reporters/ and src/docevals/types.ts, with samples captured from real runs against
the fixture corpus in test/docevals/fixtures/pages, run from that directory.
Exit codes
Section titled “Exit codes”| Code | Meaning | Who acts |
|---|---|---|
0 | Everything passed. | Nobody. |
1 | Failures, errors, or a suite below its target pass rate. | The page author. |
2 | Operational or usage error, such as bad config, a missing provider, a malformed flag, or a run that would check nothing. | Whoever owns the pipeline. |
1 and 2 route to different people, which is the whole reason they are distinct. A CI recipe
that treats any non-zero exit as “the docs are bad” blames authors for infrastructure problems. It
also earns the check a reputation for flakiness it does not deserve.
--fail-on-review additionally makes exit 1 cover evals sitting in the human-review zone.
A run narrowed by --eval or --suite cannot exit 1 on a suite target: its suite summaries
are marked partial and are not enforced. Exit 1 then comes only from the evals that
actually ran, or from an error-level resolution problem.
A run that would check nothing is exit 2
Section titled “A run that would check nothing is exit 2”If no page the run would check resolves a single eval, run stops with exit 2 instead of
reporting success over an empty report. Nothing is wired up: no defaults.suite in the config, and
no eval-suite or evals key on any page. The message names how many pages were affected, the
keys that would attach an eval, and manni docevals list, which prints the resolved plan.
This is the same answer --eval no-such-eval already gave. Reaching “nothing would have been
checked” through a config file rather than a flag does not make it a passing build.
Two empty runs are not errors, because both are choices rather than misconfigurations:
- Every page carrying
eval-skip. Pages you skipped are not counted. - A
--sincescope that selects no pages (see Scoping a run to what changed).
Neither passes in silence. Whenever no eval reaches a verdict, the report carries a warning saying the run graded nothing and established nothing about the corpus. A green exit code is never the only thing you have to go on.
A manifest the tool cannot write to is exit 2
Section titled “A manifest the tool cannot write to is exit 2”A collection may keep evals, eval-suite and eval-skip in its
external-metadata manifest. Two
declarations of that manifest leave a writer nowhere to write, so both stop the run before any page
is read:
| Refusal | stderr | Exit |
|---|---|---|
| A URL manifest owns an eval key | manni: manni.config.yaml: collection site: evals cannot come from a URL manifest, because docevals writes them. | 2 |
| Two of a page’s collections own one | manni: docs/install.md is in collections site and guides, and both keep evals in a manifest. | 2 |
A page that carries a key its manifest owns is the third case, and it is exit 1. That one is a
page problem rather than a configuration error, because the fix is an edit to the page:
error docs/install.md:4 "evals" is owned by manifest site.metadata.yaml (collection site); remove it from the documentWhich formats each command accepts
Section titled “Which formats each command accepts”-f/--format is not one shared set. run has seven reporters; list and fill render only two.
| Command | Accepted formats |
|---|---|
run | pretty, json, markdown, github, sarif, junit, html |
list | pretty, json |
fill | pretty, json |
Anything else is a usage error, giving exit 2, with the allowed set in the message. Matching
is exact, so JSON and Json are rejected rather than folded, and a typo surfaces instead of being
guessed at.
Passing run’s markdown to list is rejected for the same reason.
$ manni docevals list docs/ --format xmlmanni: --format must be one of pretty | json, got "xml"$ echo $?2That matters most in CI, where --format json is piped into a parser. Exit 2 stops the pipeline;
silently falling back to the pretty report would hand jq ANSI colour codes and still exit 0.
The writers say where a value went
Section titled “The writers say where a value went”fill, generate and promote --write put each key where its
location says, which is the page or a
manifest. Each reports which, so a reviewer knows what to open.
A fill result carries wroteTo: "page", or the manifest’s path as the run spells it. It is
absent when nothing was written, and present under --dry-run, where it names the file the write
would go to:
{ "file": "docs/install.md", "status": "filled", "wroteTo": "site.metadata.yaml"}The pretty report prints the same thing under the page’s line, and prints nothing when the evals stayed on the page:
filled docs/install.md +1 evals (states-the-problem 0.90) evals → site.metadata.yamlmetaProvenance carries the destination of its own entry, since a manifest may own
meta-provenance without owning the eval keys. destination is the manifest that holds the entry,
and absent means the page:
| Shape | Meaning |
|---|---|
{"written": true, "entry": …, "destination": "site.metadata.yaml"} | The entry went into that manifest. |
{"written": false, "skipReason": "schema-mismatch"} | The page’s meta-provenance is not a list. |
{"written": false, "skipReason": "manifest-owned", "manifest": "…"} | A URL manifest owns the key, and nothing fetched can be written. |
{"written": false, "skipReason": "unwritable"} | The owning manifest joins on a field this page does not carry. |
A page a joined manifest has no entry for is refused rather than written somewhere else. fill
reports it as an error result with no wroteTo, and the run exits 1:
docs/no-id.md carries no doc-id, which site.metadata.yaml joins on, so its evals has no entry there.generate carries refusals, one per eval whose command reference had nowhere to go, each with its
file, evalName and the sentence above. It is decided before the model is asked, so no script is
written for a refused eval. A script nothing can reference helps nobody, and paying for one helps
less. The refused target still counts, so the run exits 1 for generating fewer scripts than it had
targets.
Where no manifest owns the key, the page keeps it. On a terminal fill offers the relocation that
would give it a manifest, once per collection. Off one it writes the page and says so once for the
whole run:
manni: wrote evals to 3 pages; the schema prefers external metadata. Run manni meta relocate to give it a manifest.Under --dry-run the same line reads would write and nothing is offered.
--format pretty
Section titled “--format pretty”The default. Per page, one line per eval; then per-suite aggregates.
$ manni docevals run docs/actions/goTo.mdx --deterministic-only --no-generatedocs/actions/goTo.mdx skip no-future-promises judge skipped (--deterministic-only) pass names-an-action FAIL no-todo-markers error:14 [regex/found] Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absent
Suites reference: 1/2 passed — 0% vs target 100% below target (1 skipped)The finding line is severity:line [ruleId] message. Colour is applied only on a TTY, and never
under NO_COLOR or --no-color. The other formats carry no colour.
A finding may also carry "diagnostic": true, which appears in --format json and sarif. It
means a command eval could not reach a verdict. It had no command yet, its command failed to
start, or it timed out. The eval fails whatever its severity says. Consumers that gate on
severity alone will miss these; read the eval’s outcome instead.
Note the suite line. The skipped eval is excluded, so the count is 1/2 rather than 1/3. The rate is
weighted and the count is not. names-an-action and no-todo-markers form a failing criterion of
weight 2, scored in place of its members. So 0 of 2 weight passed, which is 0%.
--format json
Section titled “--format json”One object, for programmatic consumers.
{ "pages": 1, "evalResults": [ { "evalName": "no-todo-markers", "type": "regression", "grader": "tool:regex", "file": "docs/actions/goTo.mdx", "outcome": "fail", "findings": [ { "evalName": "no-todo-markers", "file": "docs/actions/goTo.mdx", "ruleId": "regex/found", "message": "Pattern /\\b(TODO|TBD|FIXME)\\b/ found in body, expected absent", "severity": "error", "line": 14 } ], "durationMs": 0, "suite": "reference", "weight": 1 } ], "suites": [ { "suite": "reference", "total": 3, "passed": 1, "failed": 1, "needsReview": 0, "skipped": 1, "errored": 0, "passRate": 0, "targetPassRate": 1, "meetsTarget": false, "criteria": { "total": 1, "passed": 0, "failed": 1, "suspended": 0 } } ], "usage": { "totalTokens": 0, "cachedEvals": 0, "judgedEvals": 0 }, "generated": [], "exitCode": 1, "problems": []}The sample shows one of the three evalResults. The other two are the skipped ai eval and the
passing names-an-action.
evalResults[]
Section titled “evalResults[]”| Field | Notes |
|---|---|
outcome | pass, fail, needs-review, skipped, or error. |
consensus | Present for ai-graded evals only. |
findings | Present for command/tool evals that produced findings. |
generated | True when this run generated the eval’s check script. |
via | "human-review" when a persisted review resolved the outcome. |
selfPreference | {axis, model} when the model that judged the eval also produced what it graded (axis: "content") or proposed the eval (axis: "criterion"). Absent otherwise. See Self-preference. |
suite, weight | The suite the result reports under, and its weight in that suite’s rate. |
baselined | Count of findings a baseline suppressed on this eval. Present only when a baseline was in effect, so its absence means “no baseline”, not “nothing was forgiven”. |
skipReason, durationMs |
suites[]
Section titled “suites[]”passRate is the weighted share of graded evals (pass, fail, and error) that passed, and 1 when
nothing was graded. A scored criterion counts once, by its own weight, in place of its members.
meetsTarget compares it to the suite’s target-pass-rate from the config. criteria is present
when the suite lists criteria.
partial is present, and true, only when --eval or --suite filtered the run, or --since
left pages out. It says the run
measured part of the suite, so the summary carries numbers but no verdict:
{ "suite": "reference", "total": 5, "passed": 5, "failed": 0, "needsReview": 0, "skipped": 0, "errored": 0, "passRate": 1, "targetPassRate": 1, "meetsTarget": false, "partial": true, "criteria": { "total": 5, "passed": 0, "failed": 0, "suspended": 5 }}That is the reference suite from manni docevals run --deterministic-only --no-generate --eval names-an-action.
Each criterion is suspended, because the filter left one of its two members ungraded.
meetsTargetisfalseon every partial summary, including this one, wherepassRateclears the target. It is not a claim that the suite failed.- A partial suite neither passes nor fails the run, so it contributes nothing to the exit code.
Read partial first. A consumer that reads meetsTarget on its own will read a filtered run as
a failing suite. An unfiltered run omits the key entirely, so partial !== true is the test for
“this verdict is real”. See
Selecting evals suspends suite enforcement.
problems[]
Section titled “problems[]”Resolution problems, one entry per page problem, reported rather than thrown so one bad page does not stop the run. The pretty, markdown, github and html reporters all print them.
| Field | Notes |
|---|---|
file | The file to edit. Usually the page, and the manifest when a manifest supplied the value. |
line | 1-based, in file. Absent when the problem has no line. |
level | error or warning. An error drops the page from grading and makes the run exit 1. |
message | What is wrong. |
Read file rather than assuming the page. A corpus that keeps its eval keys in a
manifest has nothing at that pointer in
the page. An annotation placed there points at the wrong line of the wrong file:
{ "file": "site.metadata.yaml", "line": 2, "level": "error", "message": "Unknown suite \"reference\" (not defined in manni.config.yaml)"}baseline
Section titled “baseline”Present only when a findings baseline was read or written. Absent means no baseline applied.
{ "path": ".manni-docevals-baseline.json", "recorded": 412, "suppressed": 340, "stale": 72}| Field | Notes |
|---|---|
path | The baseline as the user spelled it, relative to the config file. |
recorded | Fingerprints the baseline holds for the files this run checked, not for the whole file. |
suppressed | Findings this run produced that the baseline already held. |
stale | Recorded fingerprints for checked files that no longer occur. This is the burndown number. |
written | Present only on --write-baseline: { added, removed, total } against the previous file. |
recorded and stale are scoped to the files the run touched on purpose. Counted over the whole
file, a single-page run would announce that hundreds of entries no longer occur. Acting on that
advice would destroy them.
written.removed is the field to alert on. A re-record rewrites the file from what that run
covered, so a narrowed scope silently forgives everything it did not see. See
A baselined finding does not affect the exit
code.
--format markdown
Section titled “--format markdown”The suite table plus per-page findings, suitable for a PR comment or a job summary.
--format github
Section titled “--format github”Workflow commands for inline annotations, followed by the markdown summary. The same output
both annotates the diff and works as $GITHUB_STEP_SUMMARY.
::error file=docs/actions/goTo.mdx,line=14,title=manni docevals%3A no-todo-markers::Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absent
## manni docevals results
| Suite | Passed | Failed | Review | Pass rate | Target | ||---|---|---|---|---|---|---|| reference | 1 | 1 | 0 | 0% | 100% | ❌ |
### Findings
- ❌ **no-todo-markers** — `docs/actions/goTo.mdx` - error:14: Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absentSeverity maps to the annotation level: error → ::error, warning → ::warning, notice →
::notice. Values are escaped per GitHub’s rules. That covers %, CR, and LF in data, plus : and
, in properties, which is why the title reads manni docevals%3A.
A failing ai eval annotates the file with the mean confidence and the judge’s reasoning, since it has no line to point at:
::error file=...,title=manni docevals%3A no-future-promises::AI judge: fail (confidence 0.93). <reasoning>--format html
Section titled “--format html”One self-contained HTML file, with no CDN, no external stylesheet, no web font, and no script. It is the format for a person rather than a machine. It shows which pages failed, what the judge actually quoted, and whether each suite met its target.
npx @hawkeyexl/manni docevals run --format html > docevals.htmlSelf-containment is the point. The file is meant to survive being attached to a pull request,
mailed, or opened from a CI artifact directory. All of those strip or block external requests, and
any of them would otherwise render it unstyled. It follows the reader’s prefers-color-scheme.
It carries the judge’s observed quotation alongside each verdict, because a verdict without the
text it rests on is not reviewable. It also carries findings grouped by page, the weighted suite
summary including criteria, and baseline and review state. A self-preference marker appears when the
judging model also produced what it graded.
--format sarif
Section titled “--format sarif”A SARIF 2.1.0 log, for a code-scanning dashboard.
github annotates one pull request and scrolls away with it. SARIF is ingested and kept: which
finding, on which line, first seen when, still open or not. Upload it with
github/codeql-action/upload-sarif.
npx @hawkeyexl/manni docevals run --format sarif > docevals.sarifTwo properties are load-bearing, and both are asserted in CI:
- File URIs are repo-relative and forward-slashed. An absolute or backslashed path uploads successfully and then matches no file, so every finding lands on nothing.
- Every reported rule is declared in
tool.driver.rules, so the dashboard can name it instead of showing a bare id.
A failing ai-graded eval appears as a result too, with the judge’s confidence and reasoning. It
has no Finding, and omitting it would read as “the AI evals all passed”.
--format junit
Section titled “--format junit”JUnit XML, which every CI system already knows how to render as a test report.
One <testsuite> per eval suite, one <testcase> per (page, eval) pair, with classname as the
page and name as the eval. That is the grouping JUnit viewers offer, so you expand a file and see
its checks.
npx @hawkeyexl/manni docevals run --format junit > docevals.junit.xml| Outcome | Element |
|---|---|
pass | (none) |
fail | <failure> with the findings, or the judge’s reasoning |
error | <error> |
skipped | <skipped> with the skip reason |
needs-review | <skipped message="needs human review"> |
needs-review is a skip rather than a failure on purpose. JUnit has no third state, and reporting
a human-review queue as a broken build is the wrong signal for the person reading it.