manni docevals
Every quality check on a documentation page is an eval: a named, testable assertion with a grader that decides pass or fail. Graders run in preference order. Code first, an AI judge second, a human last.
The 60-second version
Section titled “The 60-second version”Name the evals a page must meet in its frontmatter. no-todo-markers is one that
manni docevals init writes into manni.config.yaml, a pattern check for TODO, TBD and
FIXME.
---title: goTodescription: Navigate to a specified URL.eval-suite: referenceevals: - use: no-todo-markers severity: error---
The `goTo` action navigates the browser to a specified URL.
TBD: say which page `goTo` waits for before the next step runs.Run it:
$ npx @hawkeyexl/manni docevals run docs/actions/goTo.mdx --deterministic-onlydocs/actions/goTo.mdx skip no-future-promises judge skipped (--deterministic-only) pass names-an-action FAIL no-todo-markers error:14 [regex/found] Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absent
Suites reference: 1/2 passed — 0% vs target 100% below target (1 skipped)Exit 1. That’s a docs regression, caught the way a test would catch a code one. The finding
names line 14 of the page, and this run needed no API key. The reference suite added the other
two evals.
What are you trying to do?
Section titled “What are you trying to do?”| Goal | Start here |
|---|---|
| Stand up a first eval gate on my repo | Get started |
| Understand the model before I commit to it | How manni docevals works |
| Write evals my team can maintain | Write evals |
| Cover a corpus I can’t annotate by hand | Adopt at scale |
| Run the gate in CI safely and cheaply | Run it in CI |
| Convince myself the judge is trustworthy | Trust the judge |
| Unblock a pull request that just went red | Fix a failing eval |
| Look up a flag, key, or grader | Reference |
What makes it different
Section titled “What makes it different”It grades what no other check can. Four graders do the work. ai judges an assertion
against the page, command runs a script you can review, human records a person’s verdict, and
tool:regex matches a pattern. Checks with a home elsewhere in manni run there. Frontmatter is
validated by manni meta validate, and page structure by
manni lint. Every domain reports on one severity scale with one set of exit
codes, so the steps compose in one pipeline.
The judge is the last resort, not the pitch. If a criterion can be expressed as code, it should be.
manni docevals promote finds ai evals that could be deterministic, and manni docevals generate writes the
scripts as reviewable files beside your docs, never inline in frontmatter.
The judge has safeguards, and you can audit them. Temperature 0, a model you pin in config, and a 3-run ensemble.
Confidence zones route the uncertain cases to a person instead of guessing. manni docevals calibrate scores the judge against your own human-verified set, so “is this trustworthy?” has a
number rather than an opinion.
Adoption is incremental. manni docevals fill proposes evals for a whole corpus with a confidence
gate. Severity levels and capability suites let you land an honest standard on a legacy corpus
without a day-one wall of red.
Requirements
Section titled “Requirements”Node.js 24+. No hosted service, no account. It runs in your existing pipeline.
npm i -D @hawkeyexl/manninpx @hawkeyexl/manni docevals init