Skip to content

Configuration reference

Every key, verified against src/docevals/core/config-schema.json and the defaults applied in src/docevals/core/config.ts.

manni.config.yaml is shared by every tool in the manni family. Each tool reads one top-level key, so every setting on this page lives under docevals::

docevals:
judge:
ensembleRuns: 3

The key names in the tables below are written relative to that namespace: judge.ensembleRuns means docevals.judge.ensembleRuns in the file.

The section’s own keys are camelCase, as every manni tool’s are. The entries under evals, criteria, and suites keep the kebab-case spelling of the frontmatter vocabulary, because they are the same entries a page carries: judge.ensembleRuns beside target-pass-rate. A kebab-case spelling of a section key is an unknown key, and the message names the key it meant:

Terminal window
manni: Invalid config in /repo/manni.config.yaml:
/docevals/judge: unknown key "ensemble-runs"; did you mean "ensembleRuns"?

Root keys manni docevals does not recognize belong to sibling tools and are ignored. Another tool’s settings can sit beside yours in the same file. Inside docevals:, unknown keys are rejected at every level, so a typo there is a startup error rather than a silently ignored setting.

One consequence is worth knowing. Because the root is permissive, misspelling the namespace itself (docevlas:) is not an error. manni docevals reads it as another tool’s config and runs on pure defaults, with no named evals and no suites. Pages that carry their own evals still run. A corpus that relied on the config’s suites stops with exit 2 and No evals resolved. manni docevals init writes the key correctly; if a run reports far fewer evals than you expect, check that spelling first.

With no config file present, every default below applies and there are no named evals or suites. The family file is the only one read. A docevals.config.yaml is not a config file, however it is spelled, so move its settings under docevals: in manni.config.yaml.

The pages are not a docevals: key. They are declared once for every manni tool, in the top-level collections: list beside docevals:. The configuration reference for collections lists every key a collection takes:

collections:
- name: site
paths: ["docs/**/*.{md,mdx}"]
exclude: ["docs/drafts/**"]
docevals:
  • With no positional paths, run, list, generate, fill, and promote read every declared collection, or the ones --collection <name> names. A collection’s paths and exclude resolve from the config file’s directory, so a run from a subdirectory reads the same pages.
  • Positional paths replace the collections for that run and resolve from the working directory.
  • node_modules and .git are never read, whichever of the two supplied the pages.
  • No paths and no collections is exit 2. There is no default set, so a run never evaluates everything beneath the working directory by accident.

A files: key under docevals: is refused by name with exit 2, its message pointing to the collections reference, rather than reported as an unknown key. Move its include globs to a collection’s paths and its exclude globs to that collection’s exclude.

A collection may also declare an external-metadata manifest, and that manifest may own the three page keys evals, eval-suite and eval-skip. It is the same externalMetadata: entry every manni tool reads, and it is not a docevals: key either:

collections:
- name: site
paths: ["docs/**/*.{md,mdx}"]
externalMetadata:
- file: ./site.metadata.yaml
keys: [evals, eval-suite, eval-skip]
docevals:

A page’s metadata is then its frontmatter plus the keys its manifests own, for reading and for writing alike. Three rules follow, and the frontmatter reference has the refusals:

  • Membership comes from every declared collection, whatever --collection or the positional paths selected. A page named by path is still a member of whatever contains it.
  • Only a manifest owning a key this tool reads is loaded. Besides the three keys above, that is provenance and meta-provenance. A manifest owning a sibling tool’s keys alone costs a run nothing.
  • --no-config means frontmatter alone. With no config there are no collections, so there is no manifest to read.

file: may be a URL for the other three keys. For an eval key it is exit 2, because docevals writes those and a fetched file is not somewhere to write.

KeyTypeDefault
baselinestring | nullnull

Path to the findings baseline, relative to the config file rather than to the working directory. A run from a subdirectory therefore finds the same file. null means no baseline.

docevals:
baseline: .manni-docevals-baseline.json

Setting it makes every ordinary run subtract the file: pre-existing findings are suppressed and only new ones fail. --baseline [path] overrides it, --no-baseline suppresses it for one run, and a bare --write-baseline records into this path. A repo that points it somewhere custom cannot accidentally record into a file nothing reads.

Leaving this unset while running --write-baseline is reported as a warning: the file is written and then ignored by every later run. Set the key first.

The ratchet and its limits are on Retrofit a legacy corpus; the file’s shape is in Files and state.

KeyTypeDefaultNotes
defaults.suitestring | nullnullSuite applied to pages that name none. Must be a defined suite or startup fails.
defaults.failFastbooleanfalse
defaults.concurrencyinteger ≥ 14Bounded concurrency across targets.
docevals:
provider: auto # auto | anthropic | openai | claude-cli | llama-cpp
model: <model-id> # needs a named provider; unset takes the provider's default
KeyTypeDefaultNotes
providerauto | anthropic | openai | claude-cli | llama-cppthe family’s providers.provider, else autoWhich provider judges, generates, fills and promotes. auto detects one exactly as manni meta fill does: an Anthropic key, then an OpenAI key, then the Claude CLI, then a local model. --provider wins over it, and so does --local. An unknown name is exit 2.
modelstringnoneThe model within provider. manni sets no default: unset, the inference library’s default for the provider is used. It needs a provider named here, by --provider or in providers.provider, because a model name does not say which provider owns it. --model wins over it.

The connection settings are not docevals: keys. An endpoint, a key’s environment variable, the Claude CLI command and the local model’s settings live in the family’s top-level providers: map. manni meta fill reads the same map. The configuration reference for providers lists every key, the full precedence and the limits of detection:

providers:
provider: anthropic # what every tool falls back to
openai:
baseUrl: http://localhost:11434/v1
docevals:
provider: openai # wins over providers.provider for docevals

An ai eval’s own provider: and model: sit between the flags and these keys, and follow the same rules. A model belongs to the provider its own level names. So an eval that names a different provider takes that provider’s default, not the config’s model. An eval naming an unknown provider, or a model with no provider to own it, reports as an error for that eval. The run then exits 1.

--local runs inference on this machine with llama-cpp, over each of these levels. It names each provider it replaced once on stderr. See Keep pages on this machine.

Which provider to pick, and what each one implies about where your pages go, is covered in Choose a provider.

KeyTypeDefaultNotes
judge.ensembleRunsinteger ≥ 13Independent runs aggregated by consensus.
judge.temperaturenumber ≥ 00
judge.zones.autoPassnumber 0–10.8Confidence at or above which a passing consensus auto-passes.
judge.zones.autoFailnumber 0–10.8Confidence at or above which a failing consensus auto-fails.
judge.falsePositiveAlertnumber 0–10.15Calibration alerts above this false-positive rate.
judge.cacheDirstring.manni/docevals/cache
judge.concurrencyinteger ≥ 1defaults.concurrencyPages judged in parallel. Separate from defaults.concurrency because a local llama-cpp model runs in-process against one GPU and serves one context at a time. Deterministic graders are happy at 4.
judge.maxTurnsinteger ≥ 1 | nullnullStop judging after this many uncached ensemble runs. One ai eval spends judge.ensembleRuns turns; a cached ensemble spends none. null is unbounded.
judge.chunkCharsinteger ≥ 112000Characters of page per judge call. A longer page is read in parts. Each part contributes the passages bearing on the assertion, and one judge then answers against the collected evidence. Content that fits is judged in a single call, exactly as before.

Anything between the two zone thresholds lands in human review. See How judging works.

KeyTypeDefaultNotes
scripts.dirstring{docDir}/manni-docevalsWhere generated scripts go. {docDir} expands to the page’s directory.
scripts.configDirstringmanni-docevals-scriptsUsed when scripts.dir has no {docDir}; resolved relative to the config file.
scripts.timeoutMsinteger ≥ 130000

execution controls what content is allowed to run

Section titled “execution controls what content is allowed to run”

One thing in a content file can reach a shell, and that is a command eval declared in page frontmatter. Its argv is code that whoever edits the page controls, so it is denied until you grant it. A page eval that is denied reports as skipped.

KeyTypeDefaultNotes
execution.allowlist[]frontmatter-commands is the one value. It lets command evals declared in page frontmatter run. Any other value is exit 2, unknown execution grant.
docevals:
execution:
allow: [frontmatter-commands]

A command eval defined in this file is the operator’s own code and needs no grant.

--allow-execution frontmatter-commands adds the grant for one run; --no-execution clears every grant for one run. Any other value is a usage error, exit 2:

manni: --allow-execution must be one of frontmatter-commands, got "shell-steps"
KeyTypeDefaultNotes
fill.confidenceThresholdnumber 0–10.7
fill.maxEvalsPerPageinteger ≥ 13
fill.temperaturenumber ≥ 00
fill.cacheDirstring.manni/docevals/cache/fill
fill.maxTurnsinteger ≥ 1 | nullnullStop after this many uncached inference calls, one per page filled.
fill.chunkCharsinteger ≥ 112000Characters of page per fill call. A longer page is proposed against in parts and the proposals merged by id, keeping the highest confidence.

A map of eval id to definition. Ids match ^[a-z0-9][a-z0-9-]*$.

Every field an eval can carry is documented in Frontmatter reference. The same shape applies whether the eval is defined here or inline on a page.

docevals:
evals:
no-future-promises:
type: regression
assertion: The page makes no claims about unreleased or future functionality.
grader: ai
evidence: All prose sections
examples:
pass: Describes only shipped behavior.
fail: Says "coming soon" or references an unreleased version.

A map of suite name to definition.

KeyTypeDefault
target-pass-ratenumber 0–11.0
evalsstring[][]
criteriastring[][]
docevals:
suites:
reference:
target-pass-rate: 1.0
evals: [no-future-promises, names-an-action, no-todo-markers]
tutorial:
target-pass-rate: 0.7
evals: [explains-why-before-how, no-future-promises]

Several evals scored as one unit. A criterion contributes a single weighted outcome to its suite, and its members contribute none. Writing three checks as a group therefore cannot outvote three standalone evals just for being a group. Members keep their own results, so a report still says which one failed.

KeyTypeDefaultNotes
evalsstring[]noneRequired. The eval ids this criterion is made of.
combineall | anyallall: every member must pass. any: one passing member is enough.
weightnumber > 01The criterion’s contribution to its suite’s pass rate.
docevals:
criteria:
install-path-is-complete:
evals: [has-prereqs, has-verify-step]
combine: all
weight: 2
suites:
reference:
evals: [no-todo-markers]
criteria: [install-path-is-complete]

A criterion whose members were not all graded in a run is suspended rather than failed. That covers a run where --eval selected only some of them, or where a member was skipped. A group measured in part has numbers but no verdict.

Four checks run at load time and fail the process with exit 2 rather than surfacing later as a confusing per-page problem:

  • A suite may only reference evals defined under evals.
  • A criterion may only reference evals defined under evals.
  • A suite may only reference criteria defined under criteria.
  • defaults.suite must name a defined suite.