Frontmatter reference
Eval declarations live in page frontmatter under three keys, validated against
manni:evals:1.0.0, the built-in vocabulary manni meta
ships for per-page quality assertions. manni docevals validates every page against that schema,
bundled into the CLI, and implements the graders behind it. Which machines wrote the page and
proposed its evals is recorded in two keys from
manni:ai-context:1.0.0, described
below. Verified against those schemas,
src/docevals/core/resolve.ts and src/docevals/core/external.ts.
The three page keys
Section titled “The three page keys”Settings are page-level facts, not an enclosing object. The page root stays open, so a sibling
tool’s keys pass untouched. But the eval- prefix is reserved: an unrecognized eval-* key is
an error, not something silently ignored.
| Key | Type | Meaning |
|---|---|---|
evals | string, or non-empty list | One assertion, or the list of entries |
eval-suite | string | Named suite from the config whose evals apply. An unknown name is a page error. |
eval-skip | boolean (false) | Skip this page entirely; it reports as skipped. |
Any other eval-* key is a page error, exit 1, and the message lists the two settings that
exist:
error docs/limits.md:3 frontmatter/eval-provenance: unknown key "eval-provenance". The "eval-" prefix is reserved, and the only settings under it are eval-suite, eval-skip — so a typo is an error here rather than a key nothing reads.---title: Installationevals: - use: no-future-promises - id: install-command-present assertion: The page contains a bash code block with `npm i -g doc-detective`. grader: command------title: Conceptseval-suite: how-toeval-skip: falseevals: - use: no-future-promises---Where evals live
Section titled “Where evals live”The vocabulary marks all three keys x-manni-location: external. The mark says they serve maintainers
and CI rather than whoever fetches the page, so their home is the collection’s
external-metadata manifest. A page keeps them until a
manifest claims them, and manni meta relocate is the move.
A collection names the keys its manifest owns:
collections: - name: site paths: ["docs/**/*.md"] externalMetadata: - file: ./site.metadata.yaml keys: [evals, eval-suite, eval-skip]docs/install.md: eval-suite: reference evals: - use: no-todo-markersThe owning manifest wins. Every command reads a page’s metadata as its frontmatter plus the keys
its manifests own. The resolved plan is therefore the same before and after the move. Membership comes from
every collection the config declares, whatever --collection or the positional paths selected.
--no-config, a config with no collections, and a page in no collection each mean frontmatter alone.
Another vocabulary marks keys this tool reads. provenance and meta-provenance are
external in ai-context. A manifest may hold those without holding the eval keys.
Writers follow the same rule. fill, generate and promote --write put each key where its
location says, so the page and the manifest never both hold it. See
Move an annotated corpus into a manifest.
The refusals
Section titled “The refusals”| When | Reported as | Exit |
|---|---|---|
| A page carries a key its manifest owns | page error | 1 |
A URL manifest owns evals, eval-suite or eval-skip | operational error | 2 |
| Two of a page’s collections both keep one of those keys in a manifest | operational error | 2 |
Two declarations of what to check have no tiebreak, so the first is an error rather than a preference:
error docs/install.md:4 "evals" is owned by manifest site.metadata.yaml (collection site); remove it from the documentThe other two stop the run before any page is graded. docevals writes eval keys where manni meta only reads them, so a home it cannot write to is refused rather than worked around:
$ manni docevals runmanni: manni.config.yaml: collection site: evals cannot come from a URL manifest, because docevals writes them.$ echo $?2$ manni docevals runmanni: docs/install.md is in collections site and guides, and both keep evals in a manifest.$ echo $?2Entry forms
Section titled “Entry forms”An entry in the evals list takes one of three forms.
A bare string is an assertion, judged by AI at error severity. It is the one form with no id of its own, which is the trade for being one line.
evals: - The documented install command matches the current package name.A single assertion needs no list at all:
evals: The documented install command matches the current package name.A reference with overrides narrows a named eval for this page only:
evals: - use: no-todo-markers severity: warning options: { flags: i }| Override | Effect |
|---|---|
type | Replaces the eval’s type. |
severity | Replaces the severity. |
options | Merged over the base options, not replaced. |
skip | Skips this one eval on this page. |
target selects which bytes are graded
Section titled “target selects which bytes are graded”evidence tells the judge where to look within what it is given. target decides what it is
given in the first place, which is why deterministic graders honour it too. A regex grader has no
focus, it has a target.
| Value | What the grader receives |
|---|---|
body (default) | The page after its frontmatter block, byte for byte. MDX imports and components stay in. |
raw | The file verbatim, frontmatter included. |
frontmatter | The page’s metadata alone, written back out as YAML. That is the frontmatter block plus every key a manifest supplies. The fences and any comments are not included, so use raw for the literal text. |
{source: file, path: <rel>} | A companion file, resolved against the page’s directory. |
evals: - id: has-an-owner grader: tool:regex target: frontmatter options: pattern: "^owner:"The two differ on a corpus that keeps its evals in a manifest. frontmatter is what the tool read,
manifest values included, so a grader sees the same metadata wherever it lives. raw is the bytes on
disk, so a key the manifest owns is absent from it.
A target that cannot be read is an error, not a quiet fall back to the page body. A verdict about the wrong bytes is worse than no verdict. A companion path may not be absolute and may not climb out of the page’s directory.
An inline definition declares an eval that exists only on this page. It requires an id, kebab-case and unique on the page, and takes every field below. Ids are required rather than derived
from position, because a position-derived name orphans every cached verdict as soon as entries move.
Eval fields
Section titled “Eval fields”| Field | Type | Default | Notes |
|---|---|---|---|
assertion | string | none | Required when grader is set. The claim about the page that is true or false. |
type | regression | capability | regression | Most evals guard behavior that must keep working. |
grader | string | ai | ai, command, human, or tool:<name>. It is ai, not llm, because the judge may be an agent rather than a bare model. |
provider | string | docevals.provider, then providers.provider | Which provider judges this ai eval: auto, anthropic, openai, claude-cli, or llama-cpp. A CLI --provider still wins, and --local replaces it with llama-cpp. An unknown name makes this eval an error. |
model | string | the model of whichever level chose the provider | Which model judges this ai eval, within provider. A CLI --model still wins. It needs a named provider, from this eval, --provider, docevals.provider or providers.provider. A model under auto makes this eval an error. Use it to name a stronger judge for an eval that matters, or a judge other than the model that wrote the page. |
runs | integer 1–50 | judge.ensembleRuns | Ensemble runs for this eval alone. A CLI --runs still wins. Capped because it multiplies judge cost directly. |
evidence | string | none | Scopes what the judge looks at within what is graded. A hint, not a selector; see target. |
target | see below | body | Which bytes the grader receives. |
examples.pass / examples.fail | string or string[] | none | Pins the boundary; one anchor or several. An inline ai eval without them raises a warning. |
command | string[] | none | For command graders. {file} expands to the page’s absolute path. |
success-exit-codes | integer[] | [0] | Exit codes treated as a pass. |
timeout-ms | integer ≥ 1 | scripts.timeoutMs | For command graders. |
generated-assertion-hash | string | none | Written by script generation, always beside command; a mismatch means the script is stale. |
options | object | {} | Grader-specific. See Graders. |
severity | error | warning | notice | error | Only error findings fail the build. |
weight | number > 0 | 1 | How much this eval moves its suite’s pass rate. It does not change the eval’s own pass/fail. Zero is rejected, because a weightless eval is a silent disable, and skip already means that, loudly. |
Resolution order
Section titled “Resolution order”For each page, the plan is built in this order:
- Suite evals from the config, via the page’s
eval-suiteordefaults.suite. - Page entries, in document order. A reference merges onto the suite’s version of that eval; an inline definition replaces it.
Page entries win on id collision. Local override beats global default, which is the right
behavior and invisible until it bites. Check the resolved plan with manni docevals list rather than
by running.
Every resolved eval records its source as config or page, which list prints.
Page problems
Section titled “Page problems”These are reported per page rather than thrown, so one bad page does not stop the run:
| Problem | Level |
|---|---|
| Frontmatter fails schema validation | error |
| Unknown suite name | error |
| Unknown eval reference | error |
| Duplicate eval id on one page | error |
Unrecognized eval-* page key, eval-provenance included | error |
| A page key an owning manifest also holds | error |
| String shorthand matching a config eval id | warning |
Inline ai eval with no examples | warning |
An inline eval may carry severity-map, because the shared vocabulary allows it. No registered
grader reads it, so manni docevals warns on stderr once per eval and grades as if it were absent:
manni: <page>: eval "<name>" sets severity-map, which no registered grader reads; it has no effect.Errors carry the source line, resolved through manni meta’s JSON-Pointer line map. A problem about a value a manifest supplied names that manifest and its line there, because the page has nothing at that pointer to fix.
provenance and meta-provenance
Section titled “provenance and meta-provenance”Two page-level keys say what machines did to a page. manni docevals reads both, and fill writes
the second. Neither is an eval-* key, so the prefix reservation does not apply to them.
| Key | Records | Written by | Read by |
|---|---|---|---|
provenance | The body lines each machine wrote, as generated-by, lines and an integrity pin | manni meta derive | the self-preference check, for target: body and raw |
meta-provenance | The metadata each machine proposed: fields as JSON Pointers, evals as ids, and a confidence for each | manni meta fill for fields, manni docevals fill for evals | the self-preference check: fields for target: frontmatter and raw, evals for the eval itself |
A human deletes a meta-provenance entry once its fields and evals are reviewed, so a surviving
entry means unreviewed machine-proposed metadata.
---title: Install the operatormeta-provenance: - generated-by: claude-fable-5 fields: [/description] evals: [install-verified, eks-coverage] confidence: /description: 0.9 install-verified: 0.88 eks-coverage: 0.74evals: - id: install-verified assertion: The Helm install steps produce a Ready operator pod.---| Field | Type | Meaning |
|---|---|---|
generated-by | string | Required. The model that proposed them. |
fields | string[] | JSON Pointers of the fields it proposed. One of fields or evals is required. |
evals | string[] | Ids of the evals it proposed. |
confidence | object | 0–1 per field or eval, keyed by the pointer or the id. A pointer starts with / and an id cannot, so the two never collide. |
manni docevals fill merges into the first entry whose generated-by is the run’s model. It appends
the ids it wrote after the entry’s own and sets their confidence. Everything else in that entry, and
every other entry, is kept. Without an entry for the model, it appends one. That is the merge
manni meta fill does for fields, so a page has one entry per model whichever command proposed
what. The field definitions themselves are
ai-context’s.
Both keys are marked external in that vocabulary, so the entry goes wherever
its location puts it. A manifest that owns meta-provenance holds the entry,
and the page then carries none.
Validating the frontmatter itself
Section titled “Validating the frontmatter itself”manni docevals run validates every page against
manni:evals:1.0.0 before it grades anything, and
reports a declaration the schema rejects as a page error.
The schema is built into manni meta too, so manni meta validate checks the same declarations
without grading them:
manni meta validate --schema manni:evals:1.0.0 docs/Programmatic consumers can import the schema as an object:
import { docevals } from "@hawkeyexl/manni";const { frontmatterSchema, FRONTMATTER_SCHEMA_ID } = docevals;// FRONTMATTER_SCHEMA_ID is "manni:evals:1.0.0"