Judge & consensus
Run the same question past a model N times and turn the answers into a decision you can defend.
This page covers the ensemble, the consensus math, and the zone routing. Caching and budgets build on it.
The promise first
Section titled “The promise first”Everything on this page exists to deliver one guarantee:
Only a unanimous, high-confidence ensemble auto-resolves. Anything split, low-confidence, or containing an errored run routes to
human-review.
An errored run can push a result toward review. It can never produce a silent pass. If you are building a product promise on top of this — and you probably are, because “the model said so” is not an answer your users will accept — that is the sentence to build it on.
Run one
Section titled “Run one”judge() does the whole thing: N runs, consensus, and zone routing in one call.
// An LLM-as-judge ensemble, and proof that an errored run can never pass silently.// Runs with no API key: MockProvider stands in for a real provider.
import { MockProvider, judge, mockVerdict } from "@hawkeyexl/inference";
const system = "You evaluate whether a page satisfies an assertion.";const user = "# Assertion\nThe page documents authentication.\n\n# Page\nUse a bearer token.";
// Three confident agreeing runs.const unanimous = new MockProvider([ mockVerdict("pass", 0.95), mockVerdict("pass", 0.93), mockVerdict("pass", 0.97),]);
const clean = await judge({ provider: unanimous, system, user, runs: 3 });
console.log("verdict:", clean.verdict);console.log("zone:", clean.zone);console.log("votes:", clean.votes);console.log("agreement:", clean.agreement.toFixed(2));console.log("meanConfidence:", clean.meanConfidence.toFixed(2));
// The same ensemble, with a provider failure scripted in.//// Each run retries once before giving up, so a single scripted error would be// absorbed by the retry and never reach the consensus. Two consecutive errors// exhaust one run's attempts and produce a genuinely errored run.const flaky = new MockProvider([ mockVerdict("pass", 0.95), { error: "429 rate limited" }, { error: "429 rate limited" }, mockVerdict("pass", 0.97),]);
const degraded = await judge({ provider: flaky, system, user, runs: 3 });
console.log("degraded.verdict:", degraded.verdict);console.log("degraded.zone:", degraded.zone);console.log("degraded.votes:", degraded.votes);
// An errored run counts against consensus. It can push a result toward review;// it can never produce a silent pass.console.log("errored run forced review:", degraded.zone === "human-review");verdict: passzone: auto-passvotes: { pass: 3, fail: 0, partial: 0, error: 0 }agreement: 1.00meanConfidence: 0.95degraded.verdict: passdegraded.zone: human-reviewdegraded.votes: { pass: 2, fail: 0, partial: 0, error: 1 }errored run forced review: trueThe second half is the guarantee, demonstrated. Two runs still passed confidently — and the result did not auto-resolve, because one run errored.
Read the result
Section titled “Read the result”ConsensusResult is what you act on:
| Field | Meaning |
|---|---|
zone |
auto-pass, auto-fail, or human-review — the field you branch on |
verdict |
pass or fail — the binary answer |
votes |
{ pass, fail, partial, error } — the full breakdown |
agreement |
0–1 across non-errored runs |
meanConfidence |
mean over non-errored runs |
runs |
every JudgeRun, for logging and cost accounting |
Branch on zone. Use agreement and meanConfidence afterwards, to explain the decision to
whoever asks.
The exact boundaries
Section titled “The exact boundaries”These are guarantees, so they are stated precisely rather than approximately.
partialcounts toward fail for the binary verdict, but stays visible invotes.- A tie is not a pass.
auto-passrequirespass > 0, and zerofail,partial, anderrorvotes. auto-failrequires zero passes and zero errors, with at least onefailorpartial.- Any errored run forces
human-review, regardless of how confident the others were. agreementandmeanConfidencecover non-errored runs only. All-errored yields0agreement.- Both zones also require
meanConfidenceto clear their threshold.
Tune the thresholds
Section titled “Tune the thresholds”const consensus = await judge({ provider, system, user, runs: 3, zones: { autoPass: 0.9, autoFail: 0.75 },});DEFAULT_ZONES is { autoPass: 0.8, autoFail: 0.8 }. Raising a threshold sends more results to
review; lowering it sends fewer.
What thresholds cannot do is override the vote rules. A split ensemble goes to review at any threshold, and so does one containing an error. Confidence is a second gate, not a substitute for agreement.
When you need the runs
Section titled “When you need the runs”judge() is runEnsemble + computeConsensus + zoneFor. Call them separately when you need the
individual runs in between — to cache them, to price them, or to log each verdict.
import { computeConsensus, runEnsemble, zoneFor } from "@hawkeyexl/inference";
const runs = await runEnsemble({ provider, system, user, runs: 3 });const base = computeConsensus(runs);const zone = zoneFor(base, { autoPass: 0.9, autoFail: 0.8 });Temperature
Section titled “Temperature”Defaults to 0. Raising it warns once per process, because independent samples at temperature zero
are already independent: the provider is called separately each time, with no shared state. If you
want diversity, prefer more runs over more randomness.
- Stop paying for unchanged subjects → Caching
- Keep spend inside a ceiling → Cost and budgets
- Word the verdict in your own domain → Custom verdict schema
- Test all of it offline → Testing your integration