Running at scale
The ensemble works for one subject. Now it has to run over hundreds.
Everything here lives above the library. That is deliberate, and this page explains both how to build it and the two ways it changes guarantees stated elsewhere.
Where concurrency belongs
Section titled “Where concurrency belongs”Runs inside an ensemble are sequential and stay that way. runs: 5 is five calls one after
another. Concurrency belongs one level up — across subjects — where you control the pool size, the
ordering, and the budget.
That is a real design choice, not an omission. Parallelising inside an ensemble would multiply the
request rate by runs for no benefit: the samples are already independent.
A bounded pool
Section titled “A bounded pool”// Running an ensemble across many subjects, and what concurrency does to a budget gate.// Runs with no API key: MockProvider stands in for a real provider.
import { MockProvider, costOfRuns, mockVerdict, pricingFor, runEnsemble } from "@hawkeyexl/inference";
const subjects = Array.from({ length: 12 }, (_, i) => `page-${String(i + 1).padStart(2, "0")}.md`);const system = "You evaluate whether a page satisfies an assertion.";
function makeProviderFor() { // One provider per worker: MockProvider records requests, and sharing one across // workers would interleave them. A real provider is stateless and can be shared. return new MockProvider([mockVerdict("pass", 0.95)], "claude-sonnet-4-5");}
const pricing = pricingFor("claude-sonnet-4-5");
// ---------------------------------------------------------------------------// A bounded worker pool over a shared index cursor.//// Not Promise.all(subjects.map(...)) — that runs all 12 at once. The pool caps// how many are in flight, and results are assigned BY INDEX so the output keeps// the input order that a pool would otherwise lose.// ---------------------------------------------------------------------------
async function runAll({ concurrency, maxCostUsd }) { const results = new Array(subjects.length); let cursor = 0; let spent = 0; let skipped = 0;
const worker = async () => { const provider = makeProviderFor(); while (cursor < subjects.length) { const i = cursor++;
// The gate. Read-then-act on a shared counter: with K workers, up to K // subjects can pass this check before any of them adds to `spent`. if (pricing !== undefined && spent >= maxCostUsd) { results[i] = { subject: subjects[i], status: "skipped-budget" }; skipped += 1; continue; }
const runs = await runEnsemble({ provider, system, user: `# Page\n${subjects[i]}`, runs: 3, }); spent += costOfRuns(runs, pricing); results[i] = { subject: subjects[i], status: "judged" }; } };
await Promise.all( Array.from({ length: Math.min(concurrency, subjects.length) }, () => worker()), );
return { results, spent, skipped };}
const ceiling = 0.045; // five ensembles at $0.009 each
const sequential = await runAll({ concurrency: 1, maxCostUsd: ceiling });const parallel = await runAll({ concurrency: 4, maxCostUsd: ceiling });
console.log("ceiling: $" + ceiling.toFixed(3));console.log("sequential: spent $" + sequential.spent.toFixed(3), "| skipped", sequential.skipped);console.log("concurrency 4: spent $" + parallel.spent.toFixed(3), "| skipped", parallel.skipped);
// The overshoot is bounded by the number of workers, not by one call.const overshoot = (parallel.spent - sequential.spent) / 0.009;console.log("extra ensembles under concurrency:", Math.round(overshoot));
// Order survives the pool, because results are assigned by index.console.log("order preserved:", parallel.results.every((r, i) => r.subject === subjects[i]));console.log("first three:", parallel.results.slice(0, 3).map((r) => r.subject).join(" "));ceiling: $0.045sequential: spent $0.045 | skipped 7concurrency 4: spent $0.072 | skipped 4extra ensembles under concurrency: 3order preserved: truefirst three: page-01.md page-02.md page-03.mdThree details in that code are easy to get wrong:
- Not
Promise.all(subjects.map(...)). That launches every subject at once. The pool caps how many are in flight by having a fixed number of workers pull from a shared cursor. - Assign results by index. A pool finishes out of order. Writing
results[i] = …keeps the output aligned with the input, whichPromise.allwould have given you for free and a pool takes away. - One provider per worker if you are using
MockProvider. It records every request, and sharing one across workers interleaves them. A real provider is stateless and can be shared.
The budget gate under concurrency
Section titled “The budget gate under concurrency”The gate is a read-then-act on a shared counter:
if (pricing !== undefined && spent >= maxCostUsd) { /* skip */ }// ... await the call ...spent += costOfRuns(runs, pricing);With K workers, up to K subjects can pass the check before any of them adds to spent. The sample
output measures it: the same ceiling, the same subjects, $0.045 sequential against $0.072 at
concurrency 4 — three extra ensembles, one per additional worker.
So the rule generalises to: the overshoot is bounded by your pool size. Budget for
maxCostUsd + (K - 1) × cost_per_subject, or gate on a reservation rather than on spend if you need
a hard ceiling.
What a 429 actually becomes
Section titled “What a 429 actually becomes”This is the consequence that surprises people, and it follows from two rules already documented separately.
- A rate limit is recorded on
run.errorlike any other provider failure. It does not throw. - Any errored run in an ensemble
forces
human-review.
Therefore: a rate limit does not surface as a rate limit. It surfaces as a growing human-review queue.
That reads like the model becoming less certain about your subjects, which invites tuning thresholds — exactly the wrong response. If your review pile grows right after you raise concurrency, lower the concurrency before you touch anything else.
There is no backoff in the library, and none of the existing consumers implement one. If you need it, it goes in your worker loop.
Choosing runs: N
Section titled “Choosing runs: N”runs multiplies both cost and latency linearly, and buys confidence:
runs |
Buys | Costs |
|---|---|---|
| 1 | a single opinion; agreement is always 1.0 and meaningless |
1× |
| 3 | a majority, and disagreement becomes visible | 3× |
| 5 | finer-grained agreement ratios | 5× |
3 is the common default — both production eval consumers arrived at it independently. Below 3 the consensus math has nothing to work with; above it you are paying linearly for a slowly improving signal.
If you want diversity, prefer more runs to a higher temperature. Independent samples at temperature 0 are already independent, and raising temperature warns for a reason.
Estimating a run
Section titled “Estimating a run”The library reports durationMs per run but makes no latency promises. Because ensembles are
sequential, the arithmetic is at least predictable:
wall clock ≈ (subjects / concurrency) × runs × per_call_latencycost ≈ subjects × runs × per_call_cost (before cache hits)Run one subject first and read durationMs and usage off the result. That is a far better
estimate than any number this page could print, and it costs one call.
- Caching — the cheapest call is the one you do not make
- Budgets and errors — the sequential statement of the gate
- Wire it into a CLI — statuses and exit codes around this loop