Skip to content

CLI reference

Eight commands under one subcommand. Every flag below is verified against src/docevals/cli.ts, and the exit codes are exercised by the inline tests at the foot of this page.

Terminal window
manni docevals [command] [command options] [arguments]

The commands that read pages, run, list, generate, fill, and promote, take positional paths: files, directories, and globs, relative to the working directory. With none given, they read every collection the config declares, or the ones --collection names. A collection’s paths resolve from the config file’s directory, wherever the command runs. No paths and no collections is exit 2. node_modules and .git are never read.

FlagCommandsMeaning
-c, --config <path>all except init and reviewPath to manni.config.yaml. Without it, manni docevals looks for that file from the working directory up to the repository root. It reads the file’s docevals: key and its top-level collections: and providers: keys. With no file at all, built-in defaults apply and there are no named evals, suites, or collections.
--no-configrun, list, generate, fill, promoteIgnore any discovered config file and run on built-in defaults. -c and --no-config set the same option, so the one written later wins.
--collection <name>same as aboveRead only this configured collection. Repeatable, one name per occurrence, never comma-separated. Cannot be combined with positional paths, and needs a config file to select from. Each of those is exit 2, as is a name the config does not declare.
--exclude <glob>same as aboveGlob to exclude. Repeatable. Applies to typed paths and collections alike. A collection’s own exclude: applies only when the run reads that collection.
-f, --format <format>list, run, fillOutput format. run accepts pretty, json, markdown, github, sarif, junit, html; list and fill accept pretty and json. Anything else is a usage error: exit 2, allowed set in the message. That includes run’s markdown passed to list. See Output and exit codes.
--provider <name>run, generate, fill, promote, calibrateauto, anthropic, openai, claude-cli, or llama-cpp. Overrides docevals.provider, the family’s providers.provider, and an eval’s own provider:. auto, the default, detects a provider exactly as manni meta fill does. An unknown name is exit 2.
--model <model>same as aboveThe model within the provider. Overrides docevals.model, and an eval’s own model:. It needs a named provider, from --provider, docevals.provider or providers.provider, because a model name does not say which provider owns it: a model under auto is exit 2. Unset, the provider’s own default model is used.

Flags override config; they do not bypass it. An unset flag falls through to the resolved config value, so a config file and the CLI reach the same code paths.

Two of those flags do less than they look like they do to a corpus whose eval keys live in an external-metadata manifest. --collection and positional paths narrow which pages are read, never which manifests supply their keys. A page named by path is still a member of whatever collection contains it. --no-config does remove them, since there are then no collections and so no manifests. A run under it reads frontmatter alone.

Each command’s table below lists every flag it accepts, the shared ones included. A table is the whole surface of its command.

Global options are accepted before the subcommand.

OptionDescription
-V, --versionPrint the manni version and exit.
-h, --helpPrint help for the program or a subcommand.

The evals tool. Every command on this page lives under it. docevals accepts its own -V, --version, which reports the manni version, and -h, --help.

Terminal window
manni docevals [options] [command] [command options] [arguments]
OptionDescription
--no-colorDisable colored output. Color applies to pretty output only, on a TTY, and never under NO_COLOR; the same rules as the metadata tool.

Runs every resolved eval. Deterministic graders come first (cheapest first), then the AI judge, then human-review resolution.

Terminal window
manni docevals run [paths...] [options]
ArgumentDescription
[paths...]Files, directories, or globs to evaluate, relative to the working directory. Optional; without them, the configured collections.
OptionArgumentDefaultDescription
-c, --config<path>discoveredPath to manni.config.yaml. See Common conventions.
--no-confign/aoffIgnore any discovered config file. See Common conventions.
--collection<name>every collectionRead only this configured collection. Repeatable. See Common conventions.
--exclude<glob>n/aGlob to exclude. Repeatable.
-f, --format<pretty|json|markdown|github|sarif|junit|html>prettyOutput format. An unknown value is a usage error (exit 2).
--deterministic-onlyn/aoffRun only command and tool:regex graders. AI-graded evals report as skipped. No provider needed. This is the offline path.
--ai-onlyn/aoffRun only AI-judged evals; skip deterministic graders.
--allow-execution<kind>n/aGrant content-authored execution: frontmatter-commands. That lets command evals declared in page frontmatter run. Adds to execution.allow for this run. Any other kind is a usage error (exit 2).
--no-executionn/aoffClear every execution grant for this run, whatever the config says.
--no-generaten/aoffDo not generate scripts for command evals that lack a command.
--no-cachen/aoffBypass the judge response cache.
--fail-on-reviewn/aoffExit 1 when any eval lands in the human-review zone.
--provider<name>config provider, else autoJudge provider: auto, anthropic, openai, claude-cli, or llama-cpp.
--model<model>config model, else the provider’s defaultJudge model. Needs a named provider.
--localn/aoffRun inference on this machine with llama-cpp, over docevals.provider, the family’s providers.provider and an eval’s own provider:. A --provider other than llama-cpp or auto is a usage error (exit 2).
--runs<n>config judge.ensembleRunsEnsemble runs per eval. Overrides the config default of 3, and an eval’s own runs.
--chunk-chars<n>config judge.chunkCharsCharacters of page per judge call. A longer page is read in parts and judged once against the collected evidence. The config default is 12000.
--eval<name>n/aRun only this eval. Repeat the flag for more than one (--eval a --eval b). Names are not space-separated, because a variadic flag would swallow the following glob. Suite targets are not evaluated while a filter is active; see below. A name that matches nothing is a usage error: exit 2.
--suite<name>n/aRun only the evals that report under the named suite. Same suspension of suite targets. A suite the config does not define is exit 2, with the defined names in the message.
--since<ref>n/aEvaluate only pages whose file or eval manifest changed between this git ref and HEAD. Suite targets are suspended when the scope leaves pages out. Nothing changed is exit 0, not an error; see below.
--max-turns<n>config judge.maxTurnsStop judging after this many uncached ensemble runs. One ai eval spends judge.ensembleRuns turns; a cached ensemble spends none. A run may make a second provider call when the first response fails schema validation. This is therefore a floor on API calls rather than an exact count. Unbounded by default.
--baseline[path]n/aFail only on findings a recorded baseline does not already hold. With no path, the configured baseline:, then .manni-docevals-baseline.json.
--no-baselinen/aoffIgnore the configured baseline: for this run, so the full backlog reports again.
--write-baseline[path]n/aRecord this run’s findings as the baseline. With no path, the configured one; see below.

A baseline records today’s findings so a run fails only on new ones. It covers findings, which means command and tool:regex graders. It does not cover judged verdicts, an eval that errored, or a suite target. Full behavior in Retrofit a legacy corpus; the file itself in Files and state.

A bare --write-baseline records into the configured path, not the default. Get that backwards and a repo pointing baseline: somewhere custom records into a file nothing reads. Every run still fails, an unreferenced file grows in the diff, and the ratchet does nothing while looking like it works. Recording with no baseline: set at all is reported as a warning naming the config file.

A recording run applies the baseline it just wrote, so it exits 0 on findings. Recording a finding is declaring it accepted. A --write-baseline you have to wrap in || true loses the exit code that matters on the next run. A suite below target or an errored eval still exits 1.

Every re-record reports the diff against the previous file:

Terminal window
Baseline .manni-docevals-baseline.json: recorded 1 finding(s) (+0, -1).
1 previously recorded finding(s) are no longer in the baseline. If this run covered less of the
corpus than the last one, they have just been forgiven.

removed is the number to watch. It is the only thing in a CI log that makes an over-forgiving re-record visible. A baseline whose JSON, version, or fingerprints are malformed is exit 2, with the offending entry named.

Selecting evals suspends suite enforcement

Section titled “Selecting evals suspends suite enforcement”

--eval and --suite narrow the run after resolution and before grading. Resolution still happens in full, so a bad eval-* key or an unknown grader on a page you filtered out still reports.

A suite target is a claim about a body of checks. A run narrowed to one passing eval out of a suite of twelve would compute 1/1 = 100%. That clears a target of 1.0 and exits 0 having checked almost nothing. So while either flag is active, suite targets are reported but not enforced. Each summary carries the numbers it measured and withholds the verdict:

Terminal window
$ manni docevals run --deterministic-only --no-generate --eval no-todo-markers
...
Suites
default: 1/1 passed — 100% vs target 100% partial — filtered run, target not evaluated
how-to: 2/2 passed — 100% vs target 100% partial — filtered run, target not evaluated
reference: 3/4 passed — 75% vs target 100% partial — filtered run, target not evaluated (1 skipped)

That run still exits 1, because no-todo-markers fails on one page. The reference suite at 75% does not add a second reason.

What still holds:

  • A fail or an error among the evals that ran exits 1.
  • An error-level resolution problem exits 1.
  • A filter matching no evals exits 2. A green run over zero evals is never the answer.
  • --suite filters, it does not redefine. It selects the evals that report under the named suite, which a page takes from its eval-suite or defaults.suite. target-pass-rate is untouched.

This is why --eval is a local debugging flag and not a CI invocation. See Exit codes and annotations and the partial field in Output and exit codes.

--since <ref> grades only the pages that differ between <ref> and HEAD, so a job on a pull request does not re-judge a corpus nobody touched.

A page differs when its own file changed, or when a manifest that supplies its evals, eval-suite or eval-skip changed. That covers a page’s own sidecar, such as install.meta.yaml, and a collection’s shared manifest. An edit to one eval’s assertion therefore grades the page it belongs to. Any change to such a manifest counts, including one to a key that is not an eval key. A changed collection manifest selects every page it supplies eval keys to.

Terminal window
$ manni docevals run --since origin/main --format github

The comparison is <ref>...HEAD, with three dots, so the diff is against the merge base. That is what a pull request means by “changed”: commits that landed on the base branch after yours forked are not your changes. It also means uncommitted working-tree edits are not included, which is right for CI and surprising locally. Edit a page, run --since main, and that page is not graded until the edit is committed.

The scope does not narrow resolution. Every page is still read and resolved. A bad eval-* key on a page outside the scope still reports, and still exits 1. An unknown grader anywhere is still exit 2.

Three things it does change:

  • Nothing changed is exit 0, the opposite of an --eval that matches nothing. A branch that touched only source is a correct answer, not a typo. The run says so rather than going quiet:

    Terminal window
    No pages changed since origin/main — nothing was evaluated.
  • Suite targets are reported but not enforced when the scope left pages out. A run that measured a sample of the corpus withholds the verdict, exactly as under --eval. The condition is coverage actually lost, not the flag’s presence. A pull request that touched every page in the corpus is a full run, and its suite targets are enforced normally. Deriving it from the flag would have turned the aggregate gate off permanently for anyone who adopted --since in CI. That is what ADR 01040 intends it for.

  • --write-baseline is refused: exit 2, before git is even run. A re-record rebuilds the file from this run’s findings, so recording from a scoped run would forgive every finding the scope excluded. Reading a baseline is unaffected, and pages outside the scope are not reported as no longer occurring.

In GitHub Actions, actions/checkout clones with fetch-depth: 1, so origin/main is usually not present locally and --since origin/main is exit 2 with a message naming the ref. Fetch the full history:

- uses: actions/checkout@v7
with:
fetch-depth: 0

--since has no config key, for the same reason --eval and --suite have none. It selects work for one invocation rather than setting policy, and the ref you compare against belongs to the job, not to the repository. list takes no --since.

--since and the judge cache are complements. --since avoids dispatching work on a page that did not change. The cache avoids paying twice for a page that did change but whose prompt did not.

Shows the resolved eval plan per page and runs nothing. This is the dry run for named evals and suites. Use it when a page’s frontmatter, a referenced eval, and a suite could each contribute the same name.

Terminal window
manni docevals list [paths...] [options]
Terminal window
$ manni docevals list docs/actions/goTo.mdx
docs/actions/goTo.mdx (suite: reference)
- no-future-promises [ai, regression, config]
- names-an-action [tool:regex, regression, config]
- no-todo-markers [tool:regex, regression, config]
1 pages, 3 evals resolved

Each line is name [grader, type, source], where source is config or page.

ArgumentDescription
[paths...]Files, directories, or globs to list, relative to the working directory. Optional; without them, the configured collections.
OptionArgumentDefaultDescription
-c, --config<path>discoveredPath to manni.config.yaml. See Common conventions.
--no-confign/aoffIgnore any discovered config file. See Common conventions.
--collection<name>every collectionRead only this configured collection. Repeatable. See Common conventions.
--exclude<glob>n/aGlob to exclude. Repeatable.
-f, --format<pretty|json>prettyOutput format. run’s wider set is rejected here rather than falling back to pretty (exit 2).
--eval<name>n/aShow only this eval. Repeatable, like the run flag.
--suite<name>n/aShow only the evals that report under the named suite.

These are the same two filters run takes, and they fail the same way: a name that matches nothing is exit 2. Reach for them to answer “what would --eval X actually have run” without spending a grading pass.

Writes a check script for every command eval that has a plain-language assertion but no command, then records the command where the eval lives. See Deterministic checks.

Terminal window
manni docevals generate [paths...] [options]

Exits 1 if any target failed to generate.

A page-defined eval whose owning manifest has no entry for the page has nowhere to hold the command. That eval is refused before the model is asked, and no script is written for it. Each refusal prints under the generated paths and counts against the exit code.

ArgumentDescription
[paths...]Files, directories, or globs to read, relative to the working directory. Optional; without them, the configured collections.
OptionArgumentDefaultDescription
-c, --config<path>discoveredPath to manni.config.yaml. See Common conventions.
--no-confign/aoffIgnore any discovered config file. See Common conventions.
--collection<name>every collectionRead only this configured collection. Repeatable. See Common conventions.
--exclude<glob>n/aGlob to exclude. Repeatable.
--provider<name>config provider, else autoProvider that writes the scripts: auto, anthropic, openai, claude-cli, or llama-cpp.
--model<model>config model, else the provider’s defaultModel. Needs a named provider.
--localn/aoffRun inference on this machine with llama-cpp, over docevals.provider, the family’s providers.provider and an eval’s own provider:. A --provider other than llama-cpp or auto is a usage error (exit 2).

Proposes evals for each page from its content, gated on a self-reported confidence.

Terminal window
manni docevals fill [paths...] [options]
ArgumentDescription
[paths...]Files, directories, or globs to fill, relative to the working directory. Optional; without them, the configured collections.
OptionArgumentDefaultDescription
-c, --config<path>discoveredPath to manni.config.yaml. See Common conventions.
--no-confign/aoffIgnore any discovered config file. See Common conventions.
--collection<name>every collectionRead only this configured collection. Repeatable. See Common conventions.
--exclude<glob>n/aGlob to exclude. Repeatable.
-f, --format<pretty|json>prettyOutput format. An unknown value is rejected before any work is done (exit 2).
--dry-runn/aoffReport proposals without writing frontmatter. Run this first. You see every proposal, and how many inference calls the pass made, before anything touches the repo.
--confidence<n>config fill.confidenceThresholdMinimum confidence to write, 0 to 1. The config default is 0.7. Values above 1 are a usage error.
--max-turns<n>config fill.maxTurnsStop after this many uncached inference calls, exactly one per page filled. Unbounded by default.
--no-cachen/aoffBypass the proposal cache.
--chunk-chars<n>config fill.chunkCharsCharacters of page per fill call. A longer page is proposed in parts.
--provider<name>config provider, else autoProvider: auto, anthropic, openai, claude-cli, or llama-cpp.
--model<model>config model, else the provider’s defaultModel. Needs a named provider.
--localn/aoffRun inference on this machine with llama-cpp, over docevals.provider, the family’s providers.provider and an eval’s own provider:. A --provider other than llama-cpp or auto is a usage error (exit 2).

Raw proposals are cached before the confidence gate, so re-running at a different --confidence spends no turns.

Each key goes where its location puts it. A local manifest that owns evals takes them, spliced into the page’s entry with every other byte of the file left alone. The result’s wroteTo names it, "page" or the manifest’s path, and the pretty report prints it under the page’s line:

filled docs/install.md +1 evals (states-the-problem 0.90)
evals → site.metadata.yaml

Where no manifest owns the key, the page keeps it. On a terminal fill offers the relocation that would give it one, asked once per collection before the first model request. Off a terminal, and on a declined offer, the evals go to the page and one line per collection says so.

A manifest that joins on a field the page lacks has no entry to hold the value. That file is refused rather than written elsewhere. It reports as an error result with no wroteTo, the rest of the run is filled, and the exit code is 1. See The writers say where a value went.

The evals fill writes are recorded in meta-provenance, in the same edit, merged into the entry for the run’s model. That key follows its own location, so the entry lands on the page or in the manifest that owns it. See provenance and meta-provenance. The pretty report prints the entry under the page’s line, with every id it now holds:

filled docs/limits.md +1 evals (links-resolve 0.80)
meta-provenance claude-fable-5: retry-named, links-resolve

A page whose meta-provenance is not a list still gets its evals, and the entry is not written:

filled docs/limits.md +1 evals (links-resolve 0.80)
meta-provenance not written: this page's schemas do not allow it

With -f json, each result carries metaProvenance in manni meta fill’s shape. It is reported under --dry-run too, and absent on a page where nothing was written:

{
"file": "docs/limits.md",
"status": "proposed",
"metaProvenance": {
"written": true,
"entry": {
"generated-by": "claude-fable-5",
"fields": ["/description"],
"evals": ["retry-named", "links-resolve"],
"confidence": { "/description": 0.9, "retry-named": 0.75, "links-resolve": 0.8 }
}
}
}

A written entry carries destination when a manifest holds it. An entry that was not written reads {"written": false, "skipReason": …}, with one of three reasons; the report reference lists them. fill never sends a page’s frontmatter to the model, so neither record can be proposed.

Reviews ai-graded evals and reports which are expressible as deterministic checks. Report-only by default; --write applies the conversions. See Promote to deterministic.

Terminal window
manni docevals promote [paths...] [options]

--write rewrites each eval where it lives, so a manifest that owns the page’s evals is the file that changes. A promotable eval with nowhere to go carries an error saying why, applied stays false, and the pretty report prints the sentence under the proposal.

ArgumentDescription
[paths...]Files, directories, or globs to read, relative to the working directory. Optional; without them, the configured collections.
OptionArgumentDefaultDescription
-c, --config<path>discoveredPath to manni.config.yaml. See Common conventions.
--no-confign/aoffIgnore any discovered config file. See Common conventions.
--collection<name>every collectionRead only this configured collection. Repeatable. See Common conventions.
--exclude<glob>n/aGlob to exclude. Repeatable.
--writen/aoffApply promotions. It writes the scripts and rewrites the evals.
--provider<name>config provider, else autoProvider: auto, anthropic, openai, claude-cli, or llama-cpp.
--model<model>config model, else the provider’s defaultModel. Needs a named provider.
--localn/aoffRun inference on this machine with llama-cpp, over docevals.provider, the family’s providers.provider and an eval’s own provider:. A --provider other than llama-cpp or auto is a usage error (exit 2).

Measures judge agreement against a human-verified golden set.

Terminal window
manni docevals calibrate [options]
OptionArgumentDefaultDescription
-c, --config<path>discoveredPath to manni.config.yaml. See Common conventions.
--golden<dir>.manni/docevals/goldenGolden set directory.
--seedn/aoffWrite golden candidates from recorded reviews into <golden>/from-reviews.yaml, then exit. Judges nothing and constructs no provider, so it runs with no API key. Idempotent on (file, eval): re-running updates rather than duplicates, and never un-reviews a confirmed case. Every new case lands reviewed: false.
--provider<name>config provider, else autoProvider: auto, anthropic, openai, claude-cli, or llama-cpp.
--model<model>config model, else the provider’s defaultModel. Needs a named provider.
--localn/aoffRun inference on this machine with llama-cpp, over docevals.provider, the family’s providers.provider and an eval’s own provider:. A --provider other than llama-cpp or auto is a usage error (exit 2).
--runs<n>config judge.ensembleRunsEnsemble runs per case.
--max-turns<n>config judge.maxTurnsStop judging after this many uncached ensemble runs, the same unit as run, since calibration goes through the same judge. One case spends judge.ensembleRuns turns. Run-wide, not per case: every case is judged in one batched call. A cached case spends none.
--no-cachen/aoffBypass the judge response cache.

Exits 1 when agreement is below the threshold. Unreviewed and stale cases are judged and counted toward that rate, flagged [unreviewed] / [stale] per case with a summary line naming the counts. See Calibrate.

Writes a starter manni.config.yaml in the working directory, with the settings under a docevals: key. Takes no options.

Terminal window
manni docevals init

With no arguments, lists the evals awaiting a human verdict. With all three, records one.

Terminal window
manni docevals review [file] [eval] [verdict] [options]
ArgumentDescription
[file]Page path.
[eval]Eval name.
[verdict]pass or fail.

Supplying a file without an eval and verdict is a usage error. See Human review.

OptionArgumentDefaultDescription
--reviewer<name>n/aReviewer recorded with the verdict.
--note<text>n/aOptional note.
CodeMeaning
0Everything passed.
1Findings, errors, or a suite below its target pass rate. The author’s problem.
2Operational or usage error, such as bad config, a missing provider, or a malformed flag. Not the author’s problem.

Full detail in Output and exit codes.