manni evals vocabulary
Built-in id: manni:evals:1.0.0
Published at https://hawkeyexl.github.io/manni/schemas/evals/1.0.0.json. A
$schema that names this URL resolves to the bundled copy, with no network
call.
The question it answers: what must be true of this page? Evals are per-page quality assertions, such as “the install command matches the current package name.” The vocabulary describes the data. Any grader, judge, or CI tool can read it, and other schemas can compose on top of it. A docs team adopts it when it wants those checks written beside the page they protect. manni meta checks that the declarations are well formed. It does not run them.
Use it
Section titled “Use it”Evals is in the
default set, so a
run checks it on every page with no config line. A config’s own schemas add
to the default set. An override replaces the set, so an entry that should keep
evals sets defaults: true:
meta: overrides: - collection: guides defaults: true schemas: - ./schemas/guide.jsonFields
Section titled “Fields”Three page-level keys. Two carry the reserved eval- prefix, and the third is
evals itself.
| Field | Type | Location | Notes |
|---|---|---|---|
evals | string, or non-empty list of entries | external | One AI-judged assertion as a string, or the list. A page with no evals omits the key, because an empty list fails |
eval-suite | string, non-empty | external | A named suite from the grading tool’s config. Its evals apply beside the page’s own, and a page entry wins on an id collision |
eval-skip | boolean, default false | external | Skip the page’s evals. A tool reports the skip rather than ignoring the page |
The Location column is the x-manni-location mark each field carries. page
means the value belongs in the document’s own metadata and reaches delivered
output. external means it belongs in the collection’s external-metadata
manifest. A field is page when an agent fetching the page acts on it. All
three keys are external, because evals belong to CI, and an agent reading
the page does not act on them. The keys inside an entry carry no mark of their
own. See field location for how the
marks are used.
Entries
Section titled “Entries”An entry takes one of three forms.
- A string. An AI-judged assertion at error severity, and the only form with no id.
- A
use:reference. It joins an eval the grading tool’s config defines, by name. - An inline definition. An object with a required
id.
Both object forms are closed, so an unknown key inside one fails.
| Key | Type | Form | Notes |
|---|---|---|---|
use | string, non-empty | reference | The name of an eval the config defines. Its presence makes the object a reference |
id | kebab-case string | inline, required | Unique within the page. Two entries with one id is an error, which a tool enforces because the schema cannot |
assertion | string, non-empty | inline | The testable claim, in plain language. Required for ai and human graders, and when grader is absent |
grader | ai, command, human, or tool:<kebab> | inline | How the assertion is checked. Default ai |
type | capability or regression | both | capability probes a boundary and is expected to fail sometimes. regression, the default, protects behavior that works |
severity | error, warning, or notice | both | Only an error-severity failure fails a run. Default error. The same scale every manni tool uses |
weight | number above 0 | both | How much this eval’s outcome moves an aggregate score. Default 1 |
options | object | both | Grader-specific options, which the grader validates |
skip | boolean | both | Skip this one eval |
target | body, raw, frontmatter, or a file object | inline | What the grader receives. See Target |
evidence | string, non-empty | inline | A hint scoping where the judge looks within what it receives |
examples | object with pass and fail | inline | Anchor examples for the judge. Each member is one string or a non-empty list of unique strings |
provider | string, non-empty | inline | The inference provider or agent runner that judges an ai eval. Omit it for the config default |
model | string, non-empty | inline | The model that judges an ai eval, within provider |
runs | integer, 1 to 50 | inline | Ensemble runs for an ai eval, overriding the configured default |
command | non-empty list of strings | inline | The argv of a command eval. {file} stands for the page path. Exit 0 passes |
success-exit-codes | non-empty list of integers | inline | Exit codes that count as a pass. Default [0] |
timeout-ms | integer, 1 or more | inline | The command’s wall-clock budget, in milliseconds |
generated-assertion-hash | string | inline | The hash of the assertion a generated command was built from |
severity-map | object of severities | inline | Maps a tool:* grader’s own severities onto the three above |
Graders
Section titled “Graders”ai is the default. It is a model or an agent, chosen per eval with
provider. command runs an executable. human puts the entry in a review
queue. tool:<kebab> names an integration, and that namespace is open. The
first three are closed, because the schema’s conditionals branch on them.
The conditionals hold these rules:
aiandhumanentries, and entries with nograder, requireassertion.- A
commandentry requiresassertionorcommand. - An entry with
command,success-exit-codes, ortimeout-msmust saygrader: command. provider,model, andrunsare legal only on anaientry.generated-assertion-hashis never legal withoutcommand.
A command entry with an assertion and no command is the generation contract.
A tool generates a check script from the assertion, then writes command and
generated-assertion-hash back. A command with a hash that differs from the
current assertion’s is a signal to regenerate it.
Target
Section titled “Target”target selects the bytes the grader receives.
| Value | The grader receives |
|---|---|
body | The page body with the frontmatter stripped. The default |
raw | The file verbatim, frontmatter included |
frontmatter | The parsed frontmatter alone |
{ source: file, path: … } | A companion file, at a path relative to the page’s directory |
Additional properties
Section titled “Additional properties”Allowed at the page root, with one exception. The root is open, so a
generator’s keys and the other vocabularies’ keys pass beside these three. The
eval- prefix is reserved, and any eval-* key other than eval-suite and
eval-skip fails. Entry objects are closed.
Example
Section titled “Example”The minimum is one assertion:
---title: Install the operator on Kubernetesdescription: Deploy the operator with Helm and verify the rollout.evals: The documented install command matches the current package name.---A suite plus four entries, one per grader shape:
---title: Install the operator on Kubernetesdescription: Deploy the operator with Helm and verify the rollout.eval-suite: how-toevals: - id: install-command-current assertion: The install command matches the current package name. - id: links-resolve grader: command command: [npx, linkinator, "{file}"] - id: install-verified assertion: The Helm install steps produce a Ready operator pod. grader: human severity: warning - id: description-says-what-for assertion: The description says what the reader can do after reading. target: frontmatter weight: 2 runs: 3---A common mistake
Section titled “A common mistake”A misspelled settings key fails, even though the root is open.
---title: Install the operator on Kubernetesdescription: Deploy the operator with Helm and verify the rollout.eval-sute: how-to---$ manni meta validate page.md --no-config -s manni:core:1.0.0 -s manni:evals:1.0.0✗ page.md /eval-sute boolean schema is false (line 4) [manni:evals:1.0.0]
1 file checked, 0 passed, 1 failed, 1 errorThe exit code is 1. “Boolean schema is false” is how the reserved prefix reports a key it does not know.
Design decisions
Section titled “Design decisions”- Flat, like the rest of the family. There is no settings object. The list
is the value, and the settings are page-level
eval-*keys. The root stays open, so the schema rejects any unknowneval-*key. That gives the root the property a closed container has, meaning a misspelled key fails instead of being ignored. - Ids are required on object entries. The string shorthand is the only id-less form. Names derived from position orphan cached verdicts whenever entries move.
- Errors name the actual fault. Entries branch with
if/thenrather thanoneOf. A misspelled key inside an object is reported against that key, instead of collapsing to “must be string.” - A reference overrides only what the page claims. A
use:entry may setskip,type,severity,options, andweight. It may not setprovider,model,runs, ortarget. Those say how the tool executes, and a named eval reads the same bytes on every page that uses it. weightscores. It does not judge. It changes how much an outcome moves a suite or run total, and never the eval’s own pass or fail. SARIF, JUnit and findings baselines consume that binary outcome. Zero is excluded, because a weightless eval is a silent disable, andskipalready says that openly.targetselects the bytes, andevidencehints where to look. A deterministic grader has nothing to focus on, but it can still read the frontmatter or a companion file. So the selector is structural, and every grader honors it. That is why it is namedtargetand notfocus.runsandmodelbelong toaievals.runsis capped at 50 because runs multiply cost directly.modellets an eval name a judge other than the machines that wrote the page. That makes a self-preference check possible, against the machines that ai-context records.- Machine-proposed evals go in the family’s one trail. Their ids sit under
evalsin a page-levelmeta-provenanceentry, which ai-context defines. A human deletes the entry after review, so a surviving entry means unreviewed machine-proposed evals. - Artifacts share the entry shape. Skills and agent definitions carry
evals under
metadata, in artifact-evals. The envelope, thetargetvalues and the grader list differ there, each for a stated reason.
Strict overlay
Section titled “Strict overlay”Strict overlay id: manni:evals-strict:1.0.0
Published at https://hawkeyexl.github.io/manni/schemas/evals-strict/1.0.0.json.
The overlay holds only what strict adds. It narrows the form of a value that is
present and requires no key. A string entry passes it as the open vocabulary
reads it.
| Field | Strict adds |
|---|---|
eval-suite, use | Kebab case, meaning lowercase letters, digits and hyphens that start with a letter or digit |
provider | Kebab case, such as anthropic or llama-cpp |
model | A model id with no spaces, of letters, digits, ., _, :, /, \, @ and -. A Windows path to a local model passes |
success-exit-codes | Unique, each from 0 to 255, the range a process returns |
target.path | A relative path, with no leading /, \ or ~, no drive letter and no URL scheme |
generated-assertion-hash | sha256- and sixty-four lowercase hex digits |
strict: true stacks it right after the vocabulary, with every other
default’s overlay beside its own base:
meta: strict: trueTo adopt this overlay alone, list its id. The default set already carries the vocabulary:
meta: schemas: - manni:evals-strict:1.0.0Each schema is checked on its own, and a finding names the one that produced it. A strict-only failure reads as one.
---title: Install the operator on Kubernetesdescription: Deploy the operator with Helm and verify the rollout.eval-suite: How-Toevals: - id: links-resolve grader: command command: [npx, linkinator, "{file}"] success-exit-codes: [0, 256]---$ manni meta validate page.md --no-config -s manni:core:1.0.0 -s manni:evals:1.0.0 -s manni:evals-strict:1.0.0✗ page.md /evals/0/success-exit-codes/1 must be <= 255 (line 9) [manni:evals-strict:1.0.0] /eval-suite must match pattern "^[a-z0-9][a-z0-9-]*$" (line 4) [manni:evals-strict:1.0.0] /eval-suite warning "eval-suite" is stored in the page; manni:evals:1.0.0 prefers external metadata, and --no-config leaves it no manifest. (line 4) [location:external] /evals warning "evals" is stored in the page; manni:evals:1.0.0 prefers external metadata, and --no-config leaves it no manifest. (line 5) [location:external]
1 file checked, 0 passed, 1 failed, 2 errors, 2 warningsThe exit code is 1. The two warnings are the external location marks at
work, and they never fail a run. The open vocabulary accepts both values, so
the overlay alone fails the page.