Configuration reference
Every key, verified against src/docevals/core/config-schema.json and the defaults applied in
src/docevals/core/config.ts.
The docevals namespace
Section titled “The docevals namespace”manni.config.yaml is shared by every tool in the manni family. Each tool reads one top-level key,
so every setting on this page lives under docevals::
docevals: judge: ensembleRuns: 3The key names in the tables below are written relative to that namespace: judge.ensembleRuns
means docevals.judge.ensembleRuns in the file.
The section’s own keys are camelCase, as every manni tool’s are. The entries under evals,
criteria, and suites keep the kebab-case spelling of the
frontmatter vocabulary, because they are the same entries a
page carries: judge.ensembleRuns beside target-pass-rate. A kebab-case spelling of a section key
is an unknown key, and the message names the key it meant:
manni: Invalid config in /repo/manni.config.yaml: /docevals/judge: unknown key "ensemble-runs"; did you mean "ensembleRuns"?Root keys manni docevals does not recognize belong to sibling tools and are ignored. Another
tool’s settings can sit beside yours in the same file. Inside docevals:, unknown keys are rejected
at every level, so a typo there is a startup error rather than a silently ignored setting.
One consequence is worth knowing. Because the root is permissive, misspelling the namespace itself
(docevlas:) is not an error. manni docevals reads it as another tool’s config and runs on pure
defaults, with no named evals and no suites. Pages that carry their own evals still run. A corpus
that relied on the config’s suites stops with exit 2 and No evals resolved. manni docevals init
writes the key correctly; if a run reports far fewer evals than you expect, check that spelling
first.
With no config file present, every default below applies and there are no named evals or suites.
The family file is the only one read. A docevals.config.yaml is not a config file, however it is
spelled, so move its settings under docevals: in manni.config.yaml.
Which pages a run reads
Section titled “Which pages a run reads”The pages are not a docevals: key. They are declared once for every manni tool, in the top-level
collections: list beside docevals:. The
configuration reference for collections lists
every key a collection takes:
collections: - name: site paths: ["docs/**/*.{md,mdx}"] exclude: ["docs/drafts/**"]
docevals:- With no positional paths,
run,list,generate,fill, andpromoteread every declared collection, or the ones--collection <name>names. A collection’spathsandexcluderesolve from the config file’s directory, so a run from a subdirectory reads the same pages. - Positional paths replace the collections for that run and resolve from the working directory.
node_modulesand.gitare never read, whichever of the two supplied the pages.- No paths and no collections is exit 2. There is no default set, so a run never evaluates everything beneath the working directory by accident.
A files: key under docevals: is refused by name with exit 2, its message pointing to the
collections reference, rather than reported as an unknown key. Move its include globs to a collection’s paths and its exclude
globs to that collection’s exclude.
Where the eval keys may live
Section titled “Where the eval keys may live”A collection may also declare an
external-metadata manifest, and that manifest may own the
three page keys evals, eval-suite and eval-skip. It is the same externalMetadata: entry
every manni tool reads, and it is not a docevals: key either:
collections: - name: site paths: ["docs/**/*.{md,mdx}"] externalMetadata: - file: ./site.metadata.yaml keys: [evals, eval-suite, eval-skip]
docevals:A page’s metadata is then its frontmatter plus the keys its manifests own, for reading and for writing alike. Three rules follow, and the frontmatter reference has the refusals:
- Membership comes from every declared collection, whatever
--collectionor the positional paths selected. A page named by path is still a member of whatever contains it. - Only a manifest owning a key this tool reads is loaded. Besides the three keys above, that is
provenanceandmeta-provenance. A manifest owning a sibling tool’s keys alone costs a run nothing. --no-configmeans frontmatter alone. With no config there are no collections, so there is no manifest to read.
file: may be a URL for the other three keys. For an eval key it is exit 2, because docevals
writes those and a fetched file is not somewhere to write.
baseline
Section titled “baseline”| Key | Type | Default |
|---|---|---|
baseline | string | null | null |
Path to the findings baseline, relative to the config file rather than to the working
directory. A run from a subdirectory therefore finds the same file. null means no baseline.
docevals: baseline: .manni-docevals-baseline.jsonSetting it makes every ordinary run subtract the file: pre-existing findings are suppressed and only
new ones fail. --baseline [path] overrides it, --no-baseline suppresses it for one run, and a
bare --write-baseline records into this path. A repo that points it somewhere custom cannot
accidentally record into a file nothing reads.
Leaving this unset while running --write-baseline is reported as a warning: the file is written and
then ignored by every later run. Set the key first.
The ratchet and its limits are on Retrofit a legacy corpus; the file’s shape is in Files and state.
defaults
Section titled “defaults”| Key | Type | Default | Notes |
|---|---|---|---|
defaults.suite | string | null | null | Suite applied to pages that name none. Must be a defined suite or startup fails. |
defaults.failFast | boolean | false | |
defaults.concurrency | integer ≥ 1 | 4 | Bounded concurrency across targets. |
provider and model
Section titled “provider and model”docevals: provider: auto # auto | anthropic | openai | claude-cli | llama-cpp model: <model-id> # needs a named provider; unset takes the provider's default| Key | Type | Default | Notes |
|---|---|---|---|
provider | auto | anthropic | openai | claude-cli | llama-cpp | the family’s providers.provider, else auto | Which provider judges, generates, fills and promotes. auto detects one exactly as manni meta fill does: an Anthropic key, then an OpenAI key, then the Claude CLI, then a local model. --provider wins over it, and so does --local. An unknown name is exit 2. |
model | string | none | The model within provider. manni sets no default: unset, the inference library’s default for the provider is used. It needs a provider named here, by --provider or in providers.provider, because a model name does not say which provider owns it. --model wins over it. |
The connection settings are not docevals: keys. An endpoint, a key’s
environment variable, the Claude CLI command and the local model’s settings
live in the family’s top-level providers: map. manni meta fill reads the
same map. The configuration reference for
providers lists every key, the
full precedence and the limits of detection:
providers: provider: anthropic # what every tool falls back to openai: baseUrl: http://localhost:11434/v1
docevals: provider: openai # wins over providers.provider for docevalsAn ai eval’s own provider: and model: sit between the flags and these
keys, and follow the same rules. A model belongs to the provider its own level
names. So an eval that names a different provider takes that provider’s
default, not the config’s model. An eval naming an unknown provider, or a
model with no provider to own it, reports as an error for that eval. The run
then exits 1.
--local runs inference on this machine with llama-cpp, over each of these
levels. It names each provider it replaced once on stderr. See
Keep pages on this machine.
Which provider to pick, and what each one implies about where your pages go, is covered in Choose a provider.
| Key | Type | Default | Notes |
|---|---|---|---|
judge.ensembleRuns | integer ≥ 1 | 3 | Independent runs aggregated by consensus. |
judge.temperature | number ≥ 0 | 0 | |
judge.zones.autoPass | number 0–1 | 0.8 | Confidence at or above which a passing consensus auto-passes. |
judge.zones.autoFail | number 0–1 | 0.8 | Confidence at or above which a failing consensus auto-fails. |
judge.falsePositiveAlert | number 0–1 | 0.15 | Calibration alerts above this false-positive rate. |
judge.cacheDir | string | .manni/docevals/cache | |
judge.concurrency | integer ≥ 1 | defaults.concurrency | Pages judged in parallel. Separate from defaults.concurrency because a local llama-cpp model runs in-process against one GPU and serves one context at a time. Deterministic graders are happy at 4. |
judge.maxTurns | integer ≥ 1 | null | null | Stop judging after this many uncached ensemble runs. One ai eval spends judge.ensembleRuns turns; a cached ensemble spends none. null is unbounded. |
judge.chunkChars | integer ≥ 1 | 12000 | Characters of page per judge call. A longer page is read in parts. Each part contributes the passages bearing on the assertion, and one judge then answers against the collected evidence. Content that fits is judged in a single call, exactly as before. |
Anything between the two zone thresholds lands in human review. See How judging works.
scripts
Section titled “scripts”| Key | Type | Default | Notes |
|---|---|---|---|
scripts.dir | string | {docDir}/manni-docevals | Where generated scripts go. {docDir} expands to the page’s directory. |
scripts.configDir | string | manni-docevals-scripts | Used when scripts.dir has no {docDir}; resolved relative to the config file. |
scripts.timeoutMs | integer ≥ 1 | 30000 |
execution controls what content is allowed to run
Section titled “execution controls what content is allowed to run”One thing in a content file can reach a shell, and that is a command eval declared in page
frontmatter. Its argv is code that whoever edits the page controls, so it is denied until you
grant it. A page eval that is denied reports as skipped.
| Key | Type | Default | Notes |
|---|---|---|---|
execution.allow | list | [] | frontmatter-commands is the one value. It lets command evals declared in page frontmatter run. Any other value is exit 2, unknown execution grant. |
docevals: execution: allow: [frontmatter-commands]A command eval defined in this file is the operator’s own code and needs no grant.
--allow-execution frontmatter-commands adds the grant for one run; --no-execution clears every
grant for one run. Any other value is a usage error, exit 2:
manni: --allow-execution must be one of frontmatter-commands, got "shell-steps"| Key | Type | Default | Notes |
|---|---|---|---|
fill.confidenceThreshold | number 0–1 | 0.7 | |
fill.maxEvalsPerPage | integer ≥ 1 | 3 | |
fill.temperature | number ≥ 0 | 0 | |
fill.cacheDir | string | .manni/docevals/cache/fill | |
fill.maxTurns | integer ≥ 1 | null | null | Stop after this many uncached inference calls, one per page filled. |
fill.chunkChars | integer ≥ 1 | 12000 | Characters of page per fill call. A longer page is proposed against in parts and the proposals merged by id, keeping the highest confidence. |
A map of eval id to definition. Ids match ^[a-z0-9][a-z0-9-]*$.
Every field an eval can carry is documented in Frontmatter reference. The same shape applies whether the eval is defined here or inline on a page.
docevals: evals: no-future-promises: type: regression assertion: The page makes no claims about unreleased or future functionality. grader: ai evidence: All prose sections examples: pass: Describes only shipped behavior. fail: Says "coming soon" or references an unreleased version.suites
Section titled “suites”A map of suite name to definition.
| Key | Type | Default |
|---|---|---|
target-pass-rate | number 0–1 | 1.0 |
evals | string[] | [] |
criteria | string[] | [] |
docevals: suites: reference: target-pass-rate: 1.0 evals: [no-future-promises, names-an-action, no-todo-markers] tutorial: target-pass-rate: 0.7 evals: [explains-why-before-how, no-future-promises]criteria
Section titled “criteria”Several evals scored as one unit. A criterion contributes a single weighted outcome to its suite, and its members contribute none. Writing three checks as a group therefore cannot outvote three standalone evals just for being a group. Members keep their own results, so a report still says which one failed.
| Key | Type | Default | Notes |
|---|---|---|---|
evals | string[] | none | Required. The eval ids this criterion is made of. |
combine | all | any | all | all: every member must pass. any: one passing member is enough. |
weight | number > 0 | 1 | The criterion’s contribution to its suite’s pass rate. |
docevals: criteria: install-path-is-complete: evals: [has-prereqs, has-verify-step] combine: all weight: 2 suites: reference: evals: [no-todo-markers] criteria: [install-path-is-complete]A criterion whose members were not all graded in a run is suspended rather than failed. That
covers a run where --eval selected only some of them, or where a member was skipped. A group
measured in part has numbers but no verdict.
Referential integrity
Section titled “Referential integrity”Four checks run at load time and fail the process with exit 2 rather than surfacing later as a confusing per-page problem:
- A suite may only reference evals defined under
evals. - A criterion may only reference evals defined under
evals. - A suite may only reference criteria defined under
criteria. defaults.suitemust name a defined suite.