Exit codes and annotations
Three exit codes, and the distinction between two of them is the difference between a check people trust and one they route around.
The contract
Section titled “The contract”| Code | Meaning | Whose problem |
|---|---|---|
0 | Everything passed | Nobody |
1 | Findings, errors, or a suite below its target pass rate | The page author |
2 | Operational or usage error, such as bad config, a missing provider, or a malformed flag | The pipeline owner |
Never collapse 1 and 2
Section titled “Never collapse 1 and 2”A recipe like this is wrong:
- run: npx @hawkeyexl/manni docevals run || echo "docs check failed"So is anything that reports every non-zero exit as “the docs are bad”. When the provider key is
missing or the config has a typo, exit 2 fires. Blaming the author for it is precisely how a
check earns a reputation for flakiness it does not deserve. Two or three of those and someone removes
the step.
Route them separately:
- name: manni docevals id: docevals continue-on-error: true run: | set +e npx @hawkeyexl/manni docevals run --format github code=$? echo "exit-code=$code" >> "$GITHUB_OUTPUT" exit $code
- name: Fail the PR on findings if: steps.docevals.outcome == 'failure' run: | code=${{ steps.docevals.outputs.exit-code }} if [ "$code" = "2" ]; then echo "::error::manni docevals could not run — this is a pipeline problem, not a docs problem." fi exit 1The simplest correct version is to let the step fail naturally. Make sure whoever triages CI knows
2 means “look at the config, not the page”.
A baselined finding does not affect the exit code
Section titled “A baselined finding does not affect the exit code”If the repo has a findings baseline,
a run subtracts it before deciding anything. A recorded finding is suppressed, the eval’s outcome is
recomputed without it, and the suite pass rate rises accordingly. A corpus with a full backlog on
file therefore exits 0, and only a new finding turns the job red.
Two lines in the log carry that, and they are dim, so they are easy to scroll past:
Baseline .manni-docevals-baseline.json: 412 finding(s) suppressed of 412 recorded.Baseline .manni-docevals-baseline.json: recorded 411 finding(s) (+0, -1).The second form is the one to watch, and removed is the number. A re-record rewrites the whole
file from what that run covered. A narrowed glob, a mistyped collection exclude, or a job that
passed the wrong --collection silently forgives everything it did not see. Nothing else in the log says
so: the fingerprints are opaque hashes, and a diff that drops 200 of them reads as noise. A
removed above zero prints a warning beneath the line; treat it as a review gate on the pull
request, not a log entry.
The corollary for a CI recipe: --write-baseline does not belong in the gating job. Record from
a deliberate, human-run invocation over the full corpus and commit the result. A job that re-records
on every run cannot fail.
Two failures a baseline never suppresses, because neither produces a finding to record. One is an
eval that errored (the check could not run). The other is a suite below its target pass
rate. Both still exit 1.
A malformed baseline file is exit 2, the pipeline owner’s problem, with the offending entry named.
--fail-on-review
Section titled “--fail-on-review”By default, evals in the human-review zone do not fail the build. With --fail-on-review, they exit
1.
This is a genuine policy fork, not a best practice:
| Aspect | Blocks on review | Doesn’t block |
|---|---|---|
| The queue | Never ignored | Can rot |
| Pull requests | Sometimes wait on a person | Keep moving |
| Needs | Someone who owns the queue | Nobody |
The deciding question is whether anyone owns the review queue. If nobody does, blocking on it just converts a rotting queue into blocked contributors. See Human review.
Inline annotations
Section titled “Inline annotations”--format github emits workflow commands before the markdown summary, so one invocation both
annotates the diff and produces a job summary.
::error file=docs/actions/goTo.mdx,line=14,title=manni docevals%3A no-todo-markers::Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absentSeverity maps to the annotation level:
| Finding severity | Annotation |
|---|---|
error | ::error |
warning | ::warning |
notice | ::notice |
Values are escaped per GitHub’s rules. That covers %, CR and LF in data, plus : and , in
properties, which is why the title renders as manni docevals%3A no-todo-markers.
Judged failures annotate the file
Section titled “Judged failures annotate the file”An ai eval has no line to point at, so it annotates the file with the mean confidence and the judge’s reasoning:
::error file=docs/tutorial.mdx,title=manni docevals%3A no-future-promises::AI judge: fail (confidence 0.93). <reasoning>That reasoning is what a contributor acts on. It is only as good as the assertion behind it. See Write good assertions.
Suites below target fail the run
Section titled “Suites below target fail the run”Suites reference: 1/2 passed — 0% vs target 100% below target (1 skipped)Exit 1, even if no individual eval errored. A capability suite left at the default target of 1.0
is the most common cause of a surprising red build. See
Regression vs capability.
A filtered run is not a gate
Section titled “A filtered run is not a gate”--eval and --suite narrow a run to named evals. They are local debugging flags, and they are not
a CI invocation: while either is active, suite targets are reported but not enforced.
Suites reference: 1/1 passed — 100% vs target 100% partial — filtered run, target not evaluatedThat is deliberate. A run narrowed to one passing eval out of a suite of twelve would otherwise
compute 1/1 = 100%. That clears a target of 1.0 and exits 0 having checked almost nothing, a
green check indistinguishable from a real one. So a job built around manni docevals run --eval …
does not get a faster gate. It gives the gate up, on every suite line of its own report.
A filter that matches nothing exits 2 rather than 0, which is the right routing. An eval renamed
out from under a pipeline is the pipeline owner’s problem, and it stops the job instead of passing
it.
Full behavior in Selecting evals suspends suite enforcement.
Job summary
Section titled “Job summary” - run: npx @hawkeyexl/manni docevals run --format github | tee "$GITHUB_STEP_SUMMARY"