Skip to content

FAQ

Short answers for people who just want their pull request green. The triage table is on Fix a failing eval.

Why did this fail when I only changed one sentence?

Section titled “Why did this fail when I only changed one sentence?”

An AI verdict is cached against the page body. Editing any part of the page invalidates that cache, so an eval that was passing gets judged again. A page that was borderline can land differently.

If the verdict looks wrong rather than newly-true, say so in the pull request. A borderline case flipping is useful signal about the assertion.

The check passed yesterday and fails today, with no change from me

Section titled “The check passed yesterday and fails today, with no change from me”

Usually one of:

  • Someone edited the assertion, so every page using it was re-judged.
  • The model was changed, which invalidates every cached verdict.
  • A command eval checks something besides the page, such as a file or a tool it calls. That can change while nobody touches the page.

The last one is not a bug. It is the check doing its job.

You can, and you should ask first. A page-level skip:

---
title: My page
eval-skip: true
---

Skipping is right for generated reference, archives, and deprecated content. It is not right for “the check is inconvenient”, which is a conversation with whoever owns the eval.

AI verdicts are about the page, not a character position. Read the rationale alongside the eval’s assertion and examples.fail. See An AI verdict.

What does needs-review mean and how do I clear it?

Section titled “What does needs-review mean and how do I clear it?”

The judge was not confident enough either way. You cannot clear it. Someone with standing runs manni docevals review <file> <eval> pass|fail. Ask whoever owns docs quality in your repo.

I ran it locally and got a different result

Section titled “I ran it locally and got a different result”

Most likely a different judge graded it. Under provider: auto, the default, a missing API key does not skip the AI evals. Detection moves on to the Claude CLI if it is installed. Otherwise it picks a local llama-cpp model, which installs its runtime and downloads its weights on first use. Your verdicts then come from a different model than CI’s. The run names the provider it picked on stderr, inference: provider "auto" — auto-selected "claude-cli", so compare that line with CI’s log.

AI evals report skipped only when no provider can be had at all. The run then warns provider unavailable — …. Running deterministic evals only. That happens when the config names anthropic or openai and its key is not set. It also happens under auto when there is no key, no Claude CLI, and the local model is ruled out, for example with INFERENCE_NO_AUTO_INSTALL set. See Choose a provider.

--deterministic-only is the local check for everything except judged evals. Also check you are running against the same config. CI may pass -c with a specific path.

No. The config names anthropic or openai, and that provider’s key is not set, so the judge was skipped and the deterministic evals ran. To say so on purpose, and skip the judge without the warning, use:

Terminal window
npx @hawkeyexl/manni docevals run docs/your-page.mdx --deterministic-only

That needs no credential. To get AI verdicts without a key, name claude-cli or llama-cpp instead.

Look at the Suites section:

Terminal window
Suites
tutorial: 7/10 passed — 70% vs target 100% below target

A suite below its target pass rate fails the run even when no individual eval errored. That usually means a capability suite is set to an unrealistic target, which is a config problem rather than a page problem.

What is the difference between warning and error?

Section titled “What is the difference between warning and error?”

Only error fails the build. warning and notice report and pass. If your finding says warning, something else turned the build red.

  • 1 means findings. A docs problem; that is yours.
  • 2 means operational. Bad config, missing key, malformed flag. Not yours, so tell whoever owns the pipeline.

It might be. Generated scripts are ordinary committed source and can be edited by hand. If it checks something other than what the assertion says, that is worth raising. See Review generated scripts.

Whoever added it. manni docevals list <your file> shows each eval’s source. config means it is defined centrally in manni.config.yaml, and page means it is inline in the page’s frontmatter. git blame the definition from there.