Skip to content

manni evals vocabulary

Built-in id: manni:evals:1.0.0

Published at https://hawkeyexl.github.io/manni/schemas/evals/1.0.0.json. A $schema that names this URL resolves to the bundled copy, with no network call.

The question it answers: what must be true of this page? Evals are per-page quality assertions, such as “the install command matches the current package name.” The vocabulary describes the data. Any grader, judge, or CI tool can read it, and other schemas can compose on top of it. A docs team adopts it when it wants those checks written beside the page they protect. manni meta checks that the declarations are well formed. It does not run them.

Evals is in the default set, so a run checks it on every page with no config line. A config’s own schemas add to the default set. An override replaces the set, so an entry that should keep evals sets defaults: true:

meta:
overrides:
- collection: guides
defaults: true
schemas:
- ./schemas/guide.json

Three page-level keys. Two carry the reserved eval- prefix, and the third is evals itself.

FieldTypeLocationNotes
evalsstring, or non-empty list of entriesexternalOne AI-judged assertion as a string, or the list. A page with no evals omits the key, because an empty list fails
eval-suitestring, non-emptyexternalA named suite from the grading tool’s config. Its evals apply beside the page’s own, and a page entry wins on an id collision
eval-skipboolean, default falseexternalSkip the page’s evals. A tool reports the skip rather than ignoring the page

The Location column is the x-manni-location mark each field carries. page means the value belongs in the document’s own metadata and reaches delivered output. external means it belongs in the collection’s external-metadata manifest. A field is page when an agent fetching the page acts on it. All three keys are external, because evals belong to CI, and an agent reading the page does not act on them. The keys inside an entry carry no mark of their own. See field location for how the marks are used.

An entry takes one of three forms.

  • A string. An AI-judged assertion at error severity, and the only form with no id.
  • A use: reference. It joins an eval the grading tool’s config defines, by name.
  • An inline definition. An object with a required id.

Both object forms are closed, so an unknown key inside one fails.

KeyTypeFormNotes
usestring, non-emptyreferenceThe name of an eval the config defines. Its presence makes the object a reference
idkebab-case stringinline, requiredUnique within the page. Two entries with one id is an error, which a tool enforces because the schema cannot
assertionstring, non-emptyinlineThe testable claim, in plain language. Required for ai and human graders, and when grader is absent
graderai, command, human, or tool:<kebab>inlineHow the assertion is checked. Default ai
typecapability or regressionbothcapability probes a boundary and is expected to fail sometimes. regression, the default, protects behavior that works
severityerror, warning, or noticebothOnly an error-severity failure fails a run. Default error. The same scale every manni tool uses
weightnumber above 0bothHow much this eval’s outcome moves an aggregate score. Default 1
optionsobjectbothGrader-specific options, which the grader validates
skipbooleanbothSkip this one eval
targetbody, raw, frontmatter, or a file objectinlineWhat the grader receives. See Target
evidencestring, non-emptyinlineA hint scoping where the judge looks within what it receives
examplesobject with pass and failinlineAnchor examples for the judge. Each member is one string or a non-empty list of unique strings
providerstring, non-emptyinlineThe inference provider or agent runner that judges an ai eval. Omit it for the config default
modelstring, non-emptyinlineThe model that judges an ai eval, within provider
runsinteger, 1 to 50inlineEnsemble runs for an ai eval, overriding the configured default
commandnon-empty list of stringsinlineThe argv of a command eval. {file} stands for the page path. Exit 0 passes
success-exit-codesnon-empty list of integersinlineExit codes that count as a pass. Default [0]
timeout-msinteger, 1 or moreinlineThe command’s wall-clock budget, in milliseconds
generated-assertion-hashstringinlineThe hash of the assertion a generated command was built from
severity-mapobject of severitiesinlineMaps a tool:* grader’s own severities onto the three above

ai is the default. It is a model or an agent, chosen per eval with provider. command runs an executable. human puts the entry in a review queue. tool:<kebab> names an integration, and that namespace is open. The first three are closed, because the schema’s conditionals branch on them.

The conditionals hold these rules:

  • ai and human entries, and entries with no grader, require assertion.
  • A command entry requires assertion or command.
  • An entry with command, success-exit-codes, or timeout-ms must say grader: command.
  • provider, model, and runs are legal only on an ai entry.
  • generated-assertion-hash is never legal without command.

A command entry with an assertion and no command is the generation contract. A tool generates a check script from the assertion, then writes command and generated-assertion-hash back. A command with a hash that differs from the current assertion’s is a signal to regenerate it.

target selects the bytes the grader receives.

ValueThe grader receives
bodyThe page body with the frontmatter stripped. The default
rawThe file verbatim, frontmatter included
frontmatterThe parsed frontmatter alone
{ source: file, path: … }A companion file, at a path relative to the page’s directory

Allowed at the page root, with one exception. The root is open, so a generator’s keys and the other vocabularies’ keys pass beside these three. The eval- prefix is reserved, and any eval-* key other than eval-suite and eval-skip fails. Entry objects are closed.

The minimum is one assertion:

---
title: Install the operator on Kubernetes
description: Deploy the operator with Helm and verify the rollout.
evals: The documented install command matches the current package name.
---

A suite plus four entries, one per grader shape:

---
title: Install the operator on Kubernetes
description: Deploy the operator with Helm and verify the rollout.
eval-suite: how-to
evals:
- id: install-command-current
assertion: The install command matches the current package name.
- id: links-resolve
grader: command
command: [npx, linkinator, "{file}"]
- id: install-verified
assertion: The Helm install steps produce a Ready operator pod.
grader: human
severity: warning
- id: description-says-what-for
assertion: The description says what the reader can do after reading.
target: frontmatter
weight: 2
runs: 3
---

A misspelled settings key fails, even though the root is open.

---
title: Install the operator on Kubernetes
description: Deploy the operator with Helm and verify the rollout.
eval-sute: how-to
---
$ manni meta validate page.md --no-config -s manni:core:1.0.0 -s manni:evals:1.0.0
✗ page.md
/eval-sute boolean schema is false (line 4) [manni:evals:1.0.0]
1 file checked, 0 passed, 1 failed, 1 error

The exit code is 1. “Boolean schema is false” is how the reserved prefix reports a key it does not know.

  • Flat, like the rest of the family. There is no settings object. The list is the value, and the settings are page-level eval-* keys. The root stays open, so the schema rejects any unknown eval-* key. That gives the root the property a closed container has, meaning a misspelled key fails instead of being ignored.
  • Ids are required on object entries. The string shorthand is the only id-less form. Names derived from position orphan cached verdicts whenever entries move.
  • Errors name the actual fault. Entries branch with if/then rather than oneOf. A misspelled key inside an object is reported against that key, instead of collapsing to “must be string.”
  • A reference overrides only what the page claims. A use: entry may set skip, type, severity, options, and weight. It may not set provider, model, runs, or target. Those say how the tool executes, and a named eval reads the same bytes on every page that uses it.
  • weight scores. It does not judge. It changes how much an outcome moves a suite or run total, and never the eval’s own pass or fail. SARIF, JUnit and findings baselines consume that binary outcome. Zero is excluded, because a weightless eval is a silent disable, and skip already says that openly.
  • target selects the bytes, and evidence hints where to look. A deterministic grader has nothing to focus on, but it can still read the frontmatter or a companion file. So the selector is structural, and every grader honors it. That is why it is named target and not focus.
  • runs and model belong to ai evals. runs is capped at 50 because runs multiply cost directly. model lets an eval name a judge other than the machines that wrote the page. That makes a self-preference check possible, against the machines that ai-context records.
  • Machine-proposed evals go in the family’s one trail. Their ids sit under evals in a page-level meta-provenance entry, which ai-context defines. A human deletes the entry after review, so a surviving entry means unreviewed machine-proposed evals.
  • Artifacts share the entry shape. Skills and agent definitions carry evals under metadata, in artifact-evals. The envelope, the target values and the grader list differ there, each for a stated reason.

Strict overlay id: manni:evals-strict:1.0.0

Published at https://hawkeyexl.github.io/manni/schemas/evals-strict/1.0.0.json. The overlay holds only what strict adds. It narrows the form of a value that is present and requires no key. A string entry passes it as the open vocabulary reads it.

FieldStrict adds
eval-suite, useKebab case, meaning lowercase letters, digits and hyphens that start with a letter or digit
providerKebab case, such as anthropic or llama-cpp
modelA model id with no spaces, of letters, digits, ., _, :, /, \, @ and -. A Windows path to a local model passes
success-exit-codesUnique, each from 0 to 255, the range a process returns
target.pathA relative path, with no leading /, \ or ~, no drive letter and no URL scheme
generated-assertion-hashsha256- and sixty-four lowercase hex digits

strict: true stacks it right after the vocabulary, with every other default’s overlay beside its own base:

meta:
strict: true

To adopt this overlay alone, list its id. The default set already carries the vocabulary:

meta:
schemas:
- manni:evals-strict:1.0.0

Each schema is checked on its own, and a finding names the one that produced it. A strict-only failure reads as one.

---
title: Install the operator on Kubernetes
description: Deploy the operator with Helm and verify the rollout.
eval-suite: How-To
evals:
- id: links-resolve
grader: command
command: [npx, linkinator, "{file}"]
success-exit-codes: [0, 256]
---
$ manni meta validate page.md --no-config -s manni:core:1.0.0 -s manni:evals:1.0.0 -s manni:evals-strict:1.0.0
✗ page.md
/evals/0/success-exit-codes/1 must be <= 255 (line 9) [manni:evals-strict:1.0.0]
/eval-suite must match pattern "^[a-z0-9][a-z0-9-]*$" (line 4) [manni:evals-strict:1.0.0]
/eval-suite warning "eval-suite" is stored in the page; manni:evals:1.0.0 prefers external metadata, and --no-config leaves it no manifest. (line 4) [location:external]
/evals warning "evals" is stored in the page; manni:evals:1.0.0 prefers external metadata, and --no-config leaves it no manifest. (line 5) [location:external]
1 file checked, 0 passed, 1 failed, 2 errors, 2 warnings

The exit code is 1. The two warnings are the external location marks at work, and they never fail a run. The open vocabulary accepts both values, so the overlay alone fails the page.