Skip to content

Exit codes and annotations

Three exit codes, and the distinction between two of them is the difference between a check people trust and one they route around.

CodeMeaningWhose problem
0Everything passedNobody
1Findings, errors, or a suite below its target pass rateThe page author
2Operational or usage error, such as bad config, a missing provider, or a malformed flagThe pipeline owner

A recipe like this is wrong:

- run: npx @hawkeyexl/manni docevals run || echo "docs check failed"

So is anything that reports every non-zero exit as “the docs are bad”. When the provider key is missing or the config has a typo, exit 2 fires. Blaming the author for it is precisely how a check earns a reputation for flakiness it does not deserve. Two or three of those and someone removes the step.

Route them separately:

- name: manni docevals
id: docevals
continue-on-error: true
run: |
set +e
npx @hawkeyexl/manni docevals run --format github
code=$?
echo "exit-code=$code" >> "$GITHUB_OUTPUT"
exit $code
- name: Fail the PR on findings
if: steps.docevals.outcome == 'failure'
run: |
code=${{ steps.docevals.outputs.exit-code }}
if [ "$code" = "2" ]; then
echo "::error::manni docevals could not run — this is a pipeline problem, not a docs problem."
fi
exit 1

The simplest correct version is to let the step fail naturally. Make sure whoever triages CI knows 2 means “look at the config, not the page”.

A baselined finding does not affect the exit code

Section titled “A baselined finding does not affect the exit code”

If the repo has a findings baseline, a run subtracts it before deciding anything. A recorded finding is suppressed, the eval’s outcome is recomputed without it, and the suite pass rate rises accordingly. A corpus with a full backlog on file therefore exits 0, and only a new finding turns the job red.

Two lines in the log carry that, and they are dim, so they are easy to scroll past:

Terminal window
Baseline .manni-docevals-baseline.json: 412 finding(s) suppressed of 412 recorded.
Terminal window
Baseline .manni-docevals-baseline.json: recorded 411 finding(s) (+0, -1).

The second form is the one to watch, and removed is the number. A re-record rewrites the whole file from what that run covered. A narrowed glob, a mistyped collection exclude, or a job that passed the wrong --collection silently forgives everything it did not see. Nothing else in the log says so: the fingerprints are opaque hashes, and a diff that drops 200 of them reads as noise. A removed above zero prints a warning beneath the line; treat it as a review gate on the pull request, not a log entry.

The corollary for a CI recipe: --write-baseline does not belong in the gating job. Record from a deliberate, human-run invocation over the full corpus and commit the result. A job that re-records on every run cannot fail.

Two failures a baseline never suppresses, because neither produces a finding to record. One is an eval that errored (the check could not run). The other is a suite below its target pass rate. Both still exit 1. A malformed baseline file is exit 2, the pipeline owner’s problem, with the offending entry named.

By default, evals in the human-review zone do not fail the build. With --fail-on-review, they exit 1.

This is a genuine policy fork, not a best practice:

AspectBlocks on reviewDoesn’t block
The queueNever ignoredCan rot
Pull requestsSometimes wait on a personKeep moving
NeedsSomeone who owns the queueNobody

The deciding question is whether anyone owns the review queue. If nobody does, blocking on it just converts a rotting queue into blocked contributors. See Human review.

--format github emits workflow commands before the markdown summary, so one invocation both annotates the diff and produces a job summary.

::error file=docs/actions/goTo.mdx,line=14,title=manni docevals%3A no-todo-markers::Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absent

Severity maps to the annotation level:

Finding severityAnnotation
error::error
warning::warning
notice::notice

Values are escaped per GitHub’s rules. That covers %, CR and LF in data, plus : and , in properties, which is why the title renders as manni docevals%3A no-todo-markers.

An ai eval has no line to point at, so it annotates the file with the mean confidence and the judge’s reasoning:

::error file=docs/tutorial.mdx,title=manni docevals%3A no-future-promises::AI judge: fail (confidence 0.93). <reasoning>

That reasoning is what a contributor acts on. It is only as good as the assertion behind it. See Write good assertions.

Terminal window
Suites
reference: 1/2 passed — 0% vs target 100% below target (1 skipped)

Exit 1, even if no individual eval errored. A capability suite left at the default target of 1.0 is the most common cause of a surprising red build. See Regression vs capability.

--eval and --suite narrow a run to named evals. They are local debugging flags, and they are not a CI invocation: while either is active, suite targets are reported but not enforced.

Terminal window
Suites
reference: 1/1 passed — 100% vs target 100% partial — filtered run, target not evaluated

That is deliberate. A run narrowed to one passing eval out of a suite of twelve would otherwise compute 1/1 = 100%. That clears a target of 1.0 and exits 0 having checked almost nothing, a green check indistinguishable from a real one. So a job built around manni docevals run --eval … does not get a faster gate. It gives the gate up, on every suite line of its own report.

A filter that matches nothing exits 2 rather than 0, which is the right routing. An eval renamed out from under a pipeline is the pipeline owner’s problem, and it stops the job instead of passing it.

Full behavior in Selecting evals suspends suite enforcement.

- run: npx @hawkeyexl/manni docevals run --format github | tee "$GITHUB_STEP_SUMMARY"