Skip to content

Consume results programmatically

Sometimes the report needs to feed something other than a human: a dashboard, a PR bot, a quality scorecard. Use --format json or the library directly.

One object per run. Full field reference in Output and exit codes.

Terminal window
npx @hawkeyexl/manni docevals run --format json > docevals.json
{
"pages": 1,
"evalResults": [ { "evalName": "no-todo-markers", "outcome": "fail", "findings": [ … ] } ],
"suites": [ { "suite": "reference", "passRate": 0, "targetPassRate": 1, "meetsTarget": false } ],
"usage": { "totalTokens": 0, "cachedEvals": 0, "judgedEvals": 0 },
"generated": [],
"exitCode": 1
}

Primary output goes to stdout, so redirect it. Diagnostics go to stderr.

Findings, flattened for a table:

Terminal window
npx @hawkeyexl/manni docevals run --format json \
| jq -r '.evalResults[] | select(.findings) | .findings[]
| [.file, .line // "-", .severity, .evalName, .message] | @tsv'

Suites missing their target:

Terminal window
npx @hawkeyexl/manni docevals run --format json \
| jq -r '.suites[] | select(.meetsTarget | not)
| "\(.suite): \(.passRate * 100 | floor)% vs \(.targetPassRate * 100)%"'

Cache effectiveness, the number that tells you whether your CI cache is working:

Terminal window
npx @hawkeyexl/manni docevals run --format json | jq '.usage | {cachedEvals, judgedEvals, totalTokens}'

Evals awaiting a person:

Terminal window
npx @hawkeyexl/manni docevals run --format json \
| jq -r '.evalResults[] | select(.outcome == "needs-review") | "\(.file) \(.evalName)"'

suites[].passRate is the metric worth storing. Push it wherever your team already looks:

Terminal window
npx @hawkeyexl/manni docevals run --format json \
| jq -c '{ts: now, suites: [.suites[] | {suite, passRate, meetsTarget}]}' \
>> metrics/docs-quality.jsonl

Capability suites are the interesting series. A regression suite should sit flat at 1.0, while a capability suite’s rate is the actual measure of a quality goal spreading. See Regression vs capability.

The package exports the page vocabulary for consumers that validate pages themselves:

import { docevals } from "@hawkeyexl/manni";
const { frontmatterSchema, FRONTMATTER_SCHEMA_ID } = docevals;

frontmatterSchema is the parsed schema object, the evals vocabulary manni docevals run validates every page against. FRONTMATTER_SCHEMA_ID is its $id, manni:evals:1.0.0. No file copy ships; see Validating the frontmatter itself.

--format json is for scripts. When the audience is a human who wants to see what the judge actually said, use --format html. It writes one self-contained file, with no CDN and no external stylesheet. That file survives being attached to a pull request or opened from an artifact directory:

Terminal window
npx @hawkeyexl/manni docevals run --format html > docevals.html

--format pretty is for people, and its layout is not a contract. Colour depends on TTY detection, and the wording can change in any release. Parse json; it is the shape that is versioned.

The JSON carries exitCode, and the process exits with it. Scripts should branch on the process exit code rather than re-deriving pass/fail from the report. Keep 1 and 2 distinct. See Exit codes and annotations.