Skip to content

manni docevals

Deterministic and LLM-as-judge evals for documentation pages, driven by frontmatter.

Every quality check on a documentation page is an eval: a named, testable assertion with a grader that decides pass or fail. Graders run in preference order. Code first, an AI judge second, a human last.

Name the evals a page must meet in its frontmatter. no-todo-markers is one that manni docevals init writes into manni.config.yaml, a pattern check for TODO, TBD and FIXME.

---
title: goTo
description: Navigate to a specified URL.
eval-suite: reference
evals:
- use: no-todo-markers
severity: error
---
The `goTo` action navigates the browser to a specified URL.
TBD: say which page `goTo` waits for before the next step runs.

Run it:

Terminal window
$ npx @hawkeyexl/manni docevals run docs/actions/goTo.mdx --deterministic-only
docs/actions/goTo.mdx
skip no-future-promises
judge skipped (--deterministic-only)
pass names-an-action
FAIL no-todo-markers
error:14 [regex/found] Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absent
Suites
reference: 1/2 passed — 0% vs target 100% below target (1 skipped)

Exit 1. That’s a docs regression, caught the way a test would catch a code one. The finding names line 14 of the page, and this run needed no API key. The reference suite added the other two evals.

GoalStart here
Stand up a first eval gate on my repoGet started
Understand the model before I commit to itHow manni docevals works
Write evals my team can maintainWrite evals
Cover a corpus I can’t annotate by handAdopt at scale
Run the gate in CI safely and cheaplyRun it in CI
Convince myself the judge is trustworthyTrust the judge
Unblock a pull request that just went redFix a failing eval
Look up a flag, key, or graderReference

It grades what no other check can. Four graders do the work. ai judges an assertion against the page, command runs a script you can review, human records a person’s verdict, and tool:regex matches a pattern. Checks with a home elsewhere in manni run there. Frontmatter is validated by manni meta validate, and page structure by manni lint. Every domain reports on one severity scale with one set of exit codes, so the steps compose in one pipeline.

The judge is the last resort, not the pitch. If a criterion can be expressed as code, it should be. manni docevals promote finds ai evals that could be deterministic, and manni docevals generate writes the scripts as reviewable files beside your docs, never inline in frontmatter.

The judge has safeguards, and you can audit them. Temperature 0, a model you pin in config, and a 3-run ensemble. Confidence zones route the uncertain cases to a person instead of guessing. manni docevals calibrate scores the judge against your own human-verified set, so “is this trustworthy?” has a number rather than an opinion.

Adoption is incremental. manni docevals fill proposes evals for a whole corpus with a confidence gate. Severity levels and capability suites let you land an honest standard on a legacy corpus without a day-one wall of red.

Node.js 24+. No hosted service, no account. It runs in your existing pipeline.

Terminal window
npm i -D @hawkeyexl/manni
npx @hawkeyexl/manni docevals init