Skip to content

Named evals and suites

The gate worked on one page, so it got copied. Now the same assertion exists in twelve pages with three slightly different wordings. Nobody can say what a how-to page is actually checked for.

This is where manni docevals stops being a linter and becomes a standard.

Move it into manni.config.yaml under evals:

docevals:
evals:
no-future-promises:
type: regression
assertion: The page makes no claims about unreleased or future functionality.
grader: ai
evidence: All prose sections
examples:
pass: Describes only shipped behavior.
fail: Says "coming soon" or references an unreleased version.

Pages reference it by name:

evals:
- use: no-future-promises

Change the wording in one place and every page that uses it changes. Names match ^[a-z0-9][a-z0-9-]*$.

A suite is a named group plus a target pass rate, the checks that apply to a kind of page.

docevals:
suites:
reference:
target-pass-rate: 1.0
evals: [no-future-promises, names-an-action, no-todo-markers]
tutorial:
target-pass-rate: 0.7
evals: [explains-why-before-how, no-future-promises]
how-to:
target-pass-rate: 1.0
evals: [no-future-promises, no-todo-markers]
---
title: Installation
eval-suite: how-to
---

Set defaults.suite to apply one to every page that names none, and pages only carry an evals key when they need something specific.

This is the piece people miss, and missing it makes manni docevals look far too strict.

Individual verdicts are binary. The nuance lives in the rate across a suite:

  • 1.0 for a regression suite. These checks must all hold.
  • ~0.7 for a capability suite. This measures how far a quality goal has spread, and a corpus partway there is the expected state, not a failure.
Terminal window
Suites
reference: 1/2 passed — 0% vs target 100% below target (1 skipped)

A suite below target fails the run. Leaving a capability suite at the default 1.0 turns every capability finding into a build break. That is the most common way to conclude the tool is unusable.

Skipped evals are excluded from the rate entirely, and so are evals awaiting human review. The rate is computed over the graded set, which is pass, fail, and error. The rate is weighted and the count is not, which is why 1/2 above reads 0%. Both graded evals belong to one failing criterion, scored once. Weights and criteria come next.

By default each eval contributes 1. weight changes how much an outcome moves the rate, and nothing else. The eval still passes or fails on its own terms, and the counts stay counts.

docevals:
evals:
names-an-action:
assertion: The page documents at least one Doc Detective action.
grader: tool:regex
options:
pattern: "(goTo|find|click|checkLink|httpRequest)"
severity: warning
weight: 0.5 # secondary: it reports, it does not dominate

Zero is rejected. A weightless eval is a silent disable, and eval-skip already means that, loudly.

Sometimes two checks only mean something together. A criterion groups them and scores the group once, so writing three checks as a group cannot outvote three standalone evals:

docevals:
criteria:
action-page-is-trustworthy:
evals: [names-an-action, no-todo-markers]
combine: all # or `any`
weight: 2
suites:
reference:
evals: [no-future-promises, names-an-action, no-todo-markers]
criteria: [action-page-is-trustworthy]

Members keep their own results, so a report still says which one failed. A criterion whose members were not all graded, through a filtered run or a skipped member, is suspended rather than failed. A group measured in part has numbers but no verdict.

--suite <name> narrows run and list to the evals that report under a suite:

Terminal window
npx @hawkeyexl/manni docevals list --suite reference
npx @hawkeyexl/manni docevals run --suite reference --deterministic-only

It selects the evals that report under that suite. That is the same suite you see in list output and in the summary. A page gets it from its own eval-suite or from defaults.suite. So --suite reference runs everything on the reference pages and nothing on a how-to page, even where the two suites are composed of overlapping evals.

Two things it does not do. It does not redefine anything: target-pass-rate is untouched, and a page still reports under its own suite. And it does not measure a target. The run saw only part of the corpus, so the summary comes back partial, printing its numbers and withholding the verdict:

Terminal window
Suites
reference: 9/10 passed — 79% vs target 100% partial — filtered run, target not evaluated (6 skipped)

That run still exits 1, because one of the ten genuinely failed. Suspending the target does not suspend the findings. A filtered run stops claiming that the suite as a whole is in good standing. It does not stop reporting that something it looked at was wrong.

Naming a suite the config does not define is exit 2, with the defined names in the message. Full behavior, including the per-eval --eval filter, in Selecting evals suspends suite enforcement.

Two mistakes fail immediately with exit 2 rather than surfacing later as a confusing per-page problem:

  • A suite referencing an eval that is not defined.
  • defaults.suite naming a suite that does not exist.

When one page needs a variant, override at the reference rather than defining a near-duplicate:

evals:
- use: no-todo-markers
severity: warning

options merges over the base; type, severity, and skip replace. A second eval with the id no-todo-markers-but-softer is how a library rots.

Suites, references, and inline definitions can all contribute the same name, and page entries win on collision. Do not guess:

Terminal window
npx @hawkeyexl/manni docevals list docs/actions/find.mdx
Terminal window
docs/actions/find.mdx (suite: reference)
- no-future-promises [ai, regression, config]
- names-an-action [tool:regex, regression, config]
- no-todo-markers [tool:regex, regression, config]
- has-examples-heading [command, regression, page]
1 pages, 4 evals resolved

source, the third item, is config or page, which is exactly what you need when one page behaves unlike its neighbours.