Skip to content

Judge & consensus

Run the same question past a model N times and turn the answers into a decision you can defend.

This page covers the ensemble, the consensus math, and the zone routing. Caching and budgets build on it.

Everything on this page exists to deliver one guarantee:

Only a unanimous, high-confidence ensemble auto-resolves. Anything split, low-confidence, or containing an errored run routes to human-review.

An errored run can push a result toward review. It can never produce a silent pass. If you are building a product promise on top of this — and you probably are, because “the model said so” is not an answer your users will accept — that is the sentence to build it on.

judge() does the whole thing: N runs, consensus, and zone routing in one call.

examples/judge-ensemble.mjs
// An LLM-as-judge ensemble, and proof that an errored run can never pass silently.
// Runs with no API key: MockProvider stands in for a real provider.
import { MockProvider, judge, mockVerdict } from "@hawkeyexl/inference";
const system = "You evaluate whether a page satisfies an assertion.";
const user = "# Assertion\nThe page documents authentication.\n\n# Page\nUse a bearer token.";
// Three confident agreeing runs.
const unanimous = new MockProvider([
mockVerdict("pass", 0.95),
mockVerdict("pass", 0.93),
mockVerdict("pass", 0.97),
]);
const clean = await judge({ provider: unanimous, system, user, runs: 3 });
console.log("verdict:", clean.verdict);
console.log("zone:", clean.zone);
console.log("votes:", clean.votes);
console.log("agreement:", clean.agreement.toFixed(2));
console.log("meanConfidence:", clean.meanConfidence.toFixed(2));
// The same ensemble, with a provider failure scripted in.
//
// Each run retries once before giving up, so a single scripted error would be
// absorbed by the retry and never reach the consensus. Two consecutive errors
// exhaust one run's attempts and produce a genuinely errored run.
const flaky = new MockProvider([
mockVerdict("pass", 0.95),
{ error: "429 rate limited" },
{ error: "429 rate limited" },
mockVerdict("pass", 0.97),
]);
const degraded = await judge({ provider: flaky, system, user, runs: 3 });
console.log("degraded.verdict:", degraded.verdict);
console.log("degraded.zone:", degraded.zone);
console.log("degraded.votes:", degraded.votes);
// An errored run counts against consensus. It can push a result toward review;
// it can never produce a silent pass.
console.log("errored run forced review:", degraded.zone === "human-review");
verdict: pass
zone: auto-pass
votes: { pass: 3, fail: 0, partial: 0, error: 0 }
agreement: 1.00
meanConfidence: 0.95
degraded.verdict: pass
degraded.zone: human-review
degraded.votes: { pass: 2, fail: 0, partial: 0, error: 1 }
errored run forced review: true

The second half is the guarantee, demonstrated. Two runs still passed confidently — and the result did not auto-resolve, because one run errored.

ConsensusResult is what you act on:

Field Meaning
zone auto-pass, auto-fail, or human-review — the field you branch on
verdict pass or fail — the binary answer
votes { pass, fail, partial, error } — the full breakdown
agreement 0–1 across non-errored runs
meanConfidence mean over non-errored runs
runs every JudgeRun, for logging and cost accounting

Branch on zone. Use agreement and meanConfidence afterwards, to explain the decision to whoever asks.

These are guarantees, so they are stated precisely rather than approximately.

  • partial counts toward fail for the binary verdict, but stays visible in votes.
  • A tie is not a pass. auto-pass requires pass > 0, and zero fail, partial, and error votes.
  • auto-fail requires zero passes and zero errors, with at least one fail or partial.
  • Any errored run forces human-review, regardless of how confident the others were.
  • agreement and meanConfidence cover non-errored runs only. All-errored yields 0 agreement.
  • Both zones also require meanConfidence to clear their threshold.
const consensus = await judge({
provider, system, user, runs: 3,
zones: { autoPass: 0.9, autoFail: 0.75 },
});

DEFAULT_ZONES is { autoPass: 0.8, autoFail: 0.8 }. Raising a threshold sends more results to review; lowering it sends fewer.

What thresholds cannot do is override the vote rules. A split ensemble goes to review at any threshold, and so does one containing an error. Confidence is a second gate, not a substitute for agreement.

judge() is runEnsemble + computeConsensus + zoneFor. Call them separately when you need the individual runs in between — to cache them, to price them, or to log each verdict.

import { computeConsensus, runEnsemble, zoneFor } from "@hawkeyexl/inference";
const runs = await runEnsemble({ provider, system, user, runs: 3 });
const base = computeConsensus(runs);
const zone = zoneFor(base, { autoPass: 0.9, autoFail: 0.8 });

Defaults to 0. Raising it warns once per process, because independent samples at temperature zero are already independent: the provider is called separately each time, with no shared state. If you want diversity, prefer more runs over more randomness.