Skip to content

Output and exit codes

Verified against src/docevals/reporters/ and src/docevals/types.ts, with samples captured from real runs against the fixture corpus in test/docevals/fixtures/pages, run from that directory.

CodeMeaningWho acts
0Everything passed.Nobody.
1Failures, errors, or a suite below its target pass rate.The page author.
2Operational or usage error, such as bad config, a missing provider, a malformed flag, or a run that would check nothing.Whoever owns the pipeline.

1 and 2 route to different people, which is the whole reason they are distinct. A CI recipe that treats any non-zero exit as “the docs are bad” blames authors for infrastructure problems. It also earns the check a reputation for flakiness it does not deserve.

--fail-on-review additionally makes exit 1 cover evals sitting in the human-review zone.

A run narrowed by --eval or --suite cannot exit 1 on a suite target: its suite summaries are marked partial and are not enforced. Exit 1 then comes only from the evals that actually ran, or from an error-level resolution problem.

If no page the run would check resolves a single eval, run stops with exit 2 instead of reporting success over an empty report. Nothing is wired up: no defaults.suite in the config, and no eval-suite or evals key on any page. The message names how many pages were affected, the keys that would attach an eval, and manni docevals list, which prints the resolved plan.

This is the same answer --eval no-such-eval already gave. Reaching “nothing would have been checked” through a config file rather than a flag does not make it a passing build.

Two empty runs are not errors, because both are choices rather than misconfigurations:

Neither passes in silence. Whenever no eval reaches a verdict, the report carries a warning saying the run graded nothing and established nothing about the corpus. A green exit code is never the only thing you have to go on.

A manifest the tool cannot write to is exit 2

Section titled “A manifest the tool cannot write to is exit 2”

A collection may keep evals, eval-suite and eval-skip in its external-metadata manifest. Two declarations of that manifest leave a writer nowhere to write, so both stop the run before any page is read:

RefusalstderrExit
A URL manifest owns an eval keymanni: manni.config.yaml: collection site: evals cannot come from a URL manifest, because docevals writes them.2
Two of a page’s collections own onemanni: docs/install.md is in collections site and guides, and both keep evals in a manifest.2

A page that carries a key its manifest owns is the third case, and it is exit 1. That one is a page problem rather than a configuration error, because the fix is an edit to the page:

error docs/install.md:4 "evals" is owned by manifest site.metadata.yaml (collection site); remove it from the document

-f/--format is not one shared set. run has seven reporters; list and fill render only two.

CommandAccepted formats
runpretty, json, markdown, github, sarif, junit, html
listpretty, json
fillpretty, json

Anything else is a usage error, giving exit 2, with the allowed set in the message. Matching is exact, so JSON and Json are rejected rather than folded, and a typo surfaces instead of being guessed at. Passing run’s markdown to list is rejected for the same reason.

Terminal window
$ manni docevals list docs/ --format xml
manni: --format must be one of pretty | json, got "xml"
$ echo $?
2

That matters most in CI, where --format json is piped into a parser. Exit 2 stops the pipeline; silently falling back to the pretty report would hand jq ANSI colour codes and still exit 0.

fill, generate and promote --write put each key where its location says, which is the page or a manifest. Each reports which, so a reviewer knows what to open.

A fill result carries wroteTo: "page", or the manifest’s path as the run spells it. It is absent when nothing was written, and present under --dry-run, where it names the file the write would go to:

{
"file": "docs/install.md",
"status": "filled",
"wroteTo": "site.metadata.yaml"
}

The pretty report prints the same thing under the page’s line, and prints nothing when the evals stayed on the page:

filled docs/install.md +1 evals (states-the-problem 0.90)
evals → site.metadata.yaml

metaProvenance carries the destination of its own entry, since a manifest may own meta-provenance without owning the eval keys. destination is the manifest that holds the entry, and absent means the page:

ShapeMeaning
{"written": true, "entry": …, "destination": "site.metadata.yaml"}The entry went into that manifest.
{"written": false, "skipReason": "schema-mismatch"}The page’s meta-provenance is not a list.
{"written": false, "skipReason": "manifest-owned", "manifest": "…"}A URL manifest owns the key, and nothing fetched can be written.
{"written": false, "skipReason": "unwritable"}The owning manifest joins on a field this page does not carry.

A page a joined manifest has no entry for is refused rather than written somewhere else. fill reports it as an error result with no wroteTo, and the run exits 1:

docs/no-id.md carries no doc-id, which site.metadata.yaml joins on, so its evals has no entry there.

generate carries refusals, one per eval whose command reference had nowhere to go, each with its file, evalName and the sentence above. It is decided before the model is asked, so no script is written for a refused eval. A script nothing can reference helps nobody, and paying for one helps less. The refused target still counts, so the run exits 1 for generating fewer scripts than it had targets.

Where no manifest owns the key, the page keeps it. On a terminal fill offers the relocation that would give it a manifest, once per collection. Off one it writes the page and says so once for the whole run:

manni: wrote evals to 3 pages; the schema prefers external metadata. Run manni meta relocate to give it a manifest.

Under --dry-run the same line reads would write and nothing is offered.

The default. Per page, one line per eval; then per-suite aggregates.

Terminal window
$ manni docevals run docs/actions/goTo.mdx --deterministic-only --no-generate
docs/actions/goTo.mdx
skip no-future-promises
judge skipped (--deterministic-only)
pass names-an-action
FAIL no-todo-markers
error:14 [regex/found] Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absent
Suites
reference: 1/2 passed — 0% vs target 100% below target (1 skipped)

The finding line is severity:line [ruleId] message. Colour is applied only on a TTY, and never under NO_COLOR or --no-color. The other formats carry no colour.

A finding may also carry "diagnostic": true, which appears in --format json and sarif. It means a command eval could not reach a verdict. It had no command yet, its command failed to start, or it timed out. The eval fails whatever its severity says. Consumers that gate on severity alone will miss these; read the eval’s outcome instead.

Note the suite line. The skipped eval is excluded, so the count is 1/2 rather than 1/3. The rate is weighted and the count is not. names-an-action and no-todo-markers form a failing criterion of weight 2, scored in place of its members. So 0 of 2 weight passed, which is 0%.

One object, for programmatic consumers.

{
"pages": 1,
"evalResults": [
{
"evalName": "no-todo-markers",
"type": "regression",
"grader": "tool:regex",
"file": "docs/actions/goTo.mdx",
"outcome": "fail",
"findings": [
{
"evalName": "no-todo-markers",
"file": "docs/actions/goTo.mdx",
"ruleId": "regex/found",
"message": "Pattern /\\b(TODO|TBD|FIXME)\\b/ found in body, expected absent",
"severity": "error",
"line": 14
}
],
"durationMs": 0,
"suite": "reference",
"weight": 1
}
],
"suites": [
{
"suite": "reference",
"total": 3,
"passed": 1,
"failed": 1,
"needsReview": 0,
"skipped": 1,
"errored": 0,
"passRate": 0,
"targetPassRate": 1,
"meetsTarget": false,
"criteria": { "total": 1, "passed": 0, "failed": 1, "suspended": 0 }
}
],
"usage": { "totalTokens": 0, "cachedEvals": 0, "judgedEvals": 0 },
"generated": [],
"exitCode": 1,
"problems": []
}

The sample shows one of the three evalResults. The other two are the skipped ai eval and the passing names-an-action.

FieldNotes
outcomepass, fail, needs-review, skipped, or error.
consensusPresent for ai-graded evals only.
findingsPresent for command/tool evals that produced findings.
generatedTrue when this run generated the eval’s check script.
via"human-review" when a persisted review resolved the outcome.
selfPreference{axis, model} when the model that judged the eval also produced what it graded (axis: "content") or proposed the eval (axis: "criterion"). Absent otherwise. See Self-preference.
suite, weightThe suite the result reports under, and its weight in that suite’s rate.
baselinedCount of findings a baseline suppressed on this eval. Present only when a baseline was in effect, so its absence means “no baseline”, not “nothing was forgiven”.
skipReason, durationMs

passRate is the weighted share of graded evals (pass, fail, and error) that passed, and 1 when nothing was graded. A scored criterion counts once, by its own weight, in place of its members. meetsTarget compares it to the suite’s target-pass-rate from the config. criteria is present when the suite lists criteria.

partial is present, and true, only when --eval or --suite filtered the run, or --since left pages out. It says the run measured part of the suite, so the summary carries numbers but no verdict:

{
"suite": "reference",
"total": 5,
"passed": 5,
"failed": 0,
"needsReview": 0,
"skipped": 0,
"errored": 0,
"passRate": 1,
"targetPassRate": 1,
"meetsTarget": false,
"partial": true,
"criteria": { "total": 5, "passed": 0, "failed": 0, "suspended": 5 }
}

That is the reference suite from manni docevals run --deterministic-only --no-generate --eval names-an-action. Each criterion is suspended, because the filter left one of its two members ungraded.

  • meetsTarget is false on every partial summary, including this one, where passRate clears the target. It is not a claim that the suite failed.
  • A partial suite neither passes nor fails the run, so it contributes nothing to the exit code.

Read partial first. A consumer that reads meetsTarget on its own will read a filtered run as a failing suite. An unfiltered run omits the key entirely, so partial !== true is the test for “this verdict is real”. See Selecting evals suspends suite enforcement.

Resolution problems, one entry per page problem, reported rather than thrown so one bad page does not stop the run. The pretty, markdown, github and html reporters all print them.

FieldNotes
fileThe file to edit. Usually the page, and the manifest when a manifest supplied the value.
line1-based, in file. Absent when the problem has no line.
levelerror or warning. An error drops the page from grading and makes the run exit 1.
messageWhat is wrong.

Read file rather than assuming the page. A corpus that keeps its eval keys in a manifest has nothing at that pointer in the page. An annotation placed there points at the wrong line of the wrong file:

{
"file": "site.metadata.yaml",
"line": 2,
"level": "error",
"message": "Unknown suite \"reference\" (not defined in manni.config.yaml)"
}

Present only when a findings baseline was read or written. Absent means no baseline applied.

{
"path": ".manni-docevals-baseline.json",
"recorded": 412,
"suppressed": 340,
"stale": 72
}
FieldNotes
pathThe baseline as the user spelled it, relative to the config file.
recordedFingerprints the baseline holds for the files this run checked, not for the whole file.
suppressedFindings this run produced that the baseline already held.
staleRecorded fingerprints for checked files that no longer occur. This is the burndown number.
writtenPresent only on --write-baseline: { added, removed, total } against the previous file.

recorded and stale are scoped to the files the run touched on purpose. Counted over the whole file, a single-page run would announce that hundreds of entries no longer occur. Acting on that advice would destroy them.

written.removed is the field to alert on. A re-record rewrites the file from what that run covered, so a narrowed scope silently forgives everything it did not see. See A baselined finding does not affect the exit code.

The suite table plus per-page findings, suitable for a PR comment or a job summary.

Workflow commands for inline annotations, followed by the markdown summary. The same output both annotates the diff and works as $GITHUB_STEP_SUMMARY.

::error file=docs/actions/goTo.mdx,line=14,title=manni docevals%3A no-todo-markers::Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absent
## manni docevals results
| Suite | Passed | Failed | Review | Pass rate | Target | |
|---|---|---|---|---|---|---|
| reference | 1 | 1 | 0 | 0% | 100% | ❌ |
### Findings
- ❌ **no-todo-markers** — `docs/actions/goTo.mdx`
- error:14: Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absent

Severity maps to the annotation level: error → ::error, warning → ::warning, notice → ::notice. Values are escaped per GitHub’s rules. That covers %, CR, and LF in data, plus : and , in properties, which is why the title reads manni docevals%3A.

A failing ai eval annotates the file with the mean confidence and the judge’s reasoning, since it has no line to point at:

::error file=...,title=manni docevals%3A no-future-promises::AI judge: fail (confidence 0.93). <reasoning>

One self-contained HTML file, with no CDN, no external stylesheet, no web font, and no script. It is the format for a person rather than a machine. It shows which pages failed, what the judge actually quoted, and whether each suite met its target.

Terminal window
npx @hawkeyexl/manni docevals run --format html > docevals.html

Self-containment is the point. The file is meant to survive being attached to a pull request, mailed, or opened from a CI artifact directory. All of those strip or block external requests, and any of them would otherwise render it unstyled. It follows the reader’s prefers-color-scheme.

It carries the judge’s observed quotation alongside each verdict, because a verdict without the text it rests on is not reviewable. It also carries findings grouped by page, the weighted suite summary including criteria, and baseline and review state. A self-preference marker appears when the judging model also produced what it graded.

A SARIF 2.1.0 log, for a code-scanning dashboard.

github annotates one pull request and scrolls away with it. SARIF is ingested and kept: which finding, on which line, first seen when, still open or not. Upload it with github/codeql-action/upload-sarif.

Terminal window
npx @hawkeyexl/manni docevals run --format sarif > docevals.sarif

Two properties are load-bearing, and both are asserted in CI:

  • File URIs are repo-relative and forward-slashed. An absolute or backslashed path uploads successfully and then matches no file, so every finding lands on nothing.
  • Every reported rule is declared in tool.driver.rules, so the dashboard can name it instead of showing a bare id.

A failing ai-graded eval appears as a result too, with the judge’s confidence and reasoning. It has no Finding, and omitting it would read as “the AI evals all passed”.

JUnit XML, which every CI system already knows how to render as a test report.

One <testsuite> per eval suite, one <testcase> per (page, eval) pair, with classname as the page and name as the eval. That is the grouping JUnit viewers offer, so you expand a file and see its checks.

Terminal window
npx @hawkeyexl/manni docevals run --format junit > docevals.junit.xml
OutcomeElement
pass(none)
fail<failure> with the findings, or the judge’s reasoning
error<error>
skipped<skipped> with the skip reason
needs-review<skipped message="needs human review">

needs-review is a skip rather than a failure on purpose. JUnit has no third state, and reporting a human-review queue as a broken build is the wrong signal for the person reading it.