Skip to content

Deterministic checks

Deterministic graders are free, fast, and arguable in a pull request. Prefer them. Reach for the judge only when no pattern or script can express the assertion. This page covers the two deterministic graders, tool:regex and command, including the scripts manni docevals writes for you.

Some checks already have a home in manni, and they run there rather than as evals.

Any other CLI tool you run becomes a command eval, described below.

tool:regex is the cheapest rung below the judge. An assertion like “the page carries no TODO markers” is a verifiable fact. Matching it is faster and more trustworthy than asking a model three times.

docevals:
evals:
no-todo-markers:
assertion: The page carries no TODO, TBD or FIXME markers.
grader: tool:regex
options:
pattern: "\\b(TODO|TBD|FIXME)\\b"
match: not-contains
severity: error
OptionDefaultWhat it does
patternnoneRequired. A JavaScript regular expression.
flags""JS RegExp flags (d g i m s u v y).
matchcontainscontains, not-contains, or count:N for exactly n matches.

count:N catches a heading that a bad merge duplicated. A pattern that does not compile is a configuration error, not a page failure.

The eval’s target picks the bytes the pattern runs against. The default is the page body. Set raw for the whole file, frontmatter for the metadata, or a companion file.

evals:
- id: has-an-owner
grader: tool:regex
target: frontmatter
options:
pattern: "^owner:"

A finding names its rule and the line in the file:

docs/actions/goTo.mdx
skip no-future-promises
judge skipped (--deterministic-only)
pass names-an-action
FAIL no-todo-markers
error:14 [regex/found] Pattern /\b(TODO|TBD|FIXME)\b/ found in body, expected absent

The rule ids are regex/not-found, regex/found and regex/count. Full option tables are in Graders.

Any executable. {file} expands to the page’s absolute path.

- id: no-absolute-internal-links
assertion: The page uses relative links for internal docs.
grader: command
command: [node, scripts/check-links.mjs, "{file}"]
success-exit-codes: [0]

Exit code decides. 0, or anything in success-exit-codes, passes. Anything else fails with the output tail as the message. Timeouts are reported distinctly.

Commands declared in page frontmatter need a grant

Section titled “Commands declared in page frontmatter need a grant”

A command eval defined in manni.config.yaml runs without a grant, because the config is yours. A page that declares one in its own frontmatter is content asking to run code. That command runs only under execution.allow: [frontmatter-commands] or --allow-execution frontmatter-commands. Without the grant it reports as skipped, and the run goes on.

docevals:
execution:
allow: [frontmatter-commands]

Grant it only to a corpus whose editors you trust. Untrusted pull requests covers the risk.

Plain-language checks manni docevals writes for you

Section titled “Plain-language checks manni docevals writes for you”

A command eval with an assertion and no command is a deterministic check you have described but not implemented. manni docevals run, or manni docevals generate, has your configured LLM write a small Node script for it. It saves the script beside the doc and records the command back into frontmatter:

- id: install-command-present
assertion: The page contains a bash code block with `npm i -g doc-detective`.
grader: command
command: [node, manni-docevals/installation.install-command-present.mjs, "{file}"]
generated-assertion-hash: aefaa89e…

The model is used once, at generation time. After that the check is fully deterministic, with no model in the loop and no per-run cost. The recorded command lives in page frontmatter, so it runs under the frontmatter-commands grant like any other.

Three consequences worth internalising:

  • The script is ordinary source. It lands in the repo beside the page, shows up in pull requests, and you can edit it by hand. Review it. A generated script that passes for the wrong reason is worse than the ai eval it replaced, because it is now silent and cheap.
  • generated-assertion-hash keeps them honest. Edit the assertion and the hash stops matching, so the script regenerates rather than quietly checking the old thing.
  • --no-generate turns it off, which is what you want in CI where generation should not happen implicitly.

Already have ai evals that should have been code? promote finds them.