Skip to content

Running at scale

The ensemble works for one subject. Now it has to run over hundreds.

Everything here lives above the library. That is deliberate, and this page explains both how to build it and the two ways it changes guarantees stated elsewhere.

Runs inside an ensemble are sequential and stay that way. runs: 5 is five calls one after another. Concurrency belongs one level up — across subjects — where you control the pool size, the ordering, and the budget.

That is a real design choice, not an omission. Parallelising inside an ensemble would multiply the request rate by runs for no benefit: the samples are already independent.

examples/orchestrate-concurrency.mjs
// Running an ensemble across many subjects, and what concurrency does to a budget gate.
// Runs with no API key: MockProvider stands in for a real provider.
import { MockProvider, costOfRuns, mockVerdict, pricingFor, runEnsemble } from "@hawkeyexl/inference";
const subjects = Array.from({ length: 12 }, (_, i) => `page-${String(i + 1).padStart(2, "0")}.md`);
const system = "You evaluate whether a page satisfies an assertion.";
function makeProviderFor() {
// One provider per worker: MockProvider records requests, and sharing one across
// workers would interleave them. A real provider is stateless and can be shared.
return new MockProvider([mockVerdict("pass", 0.95)], "claude-sonnet-4-5");
}
const pricing = pricingFor("claude-sonnet-4-5");
// ---------------------------------------------------------------------------
// A bounded worker pool over a shared index cursor.
//
// Not Promise.all(subjects.map(...)) — that runs all 12 at once. The pool caps
// how many are in flight, and results are assigned BY INDEX so the output keeps
// the input order that a pool would otherwise lose.
// ---------------------------------------------------------------------------
async function runAll({ concurrency, maxCostUsd }) {
const results = new Array(subjects.length);
let cursor = 0;
let spent = 0;
let skipped = 0;
const worker = async () => {
const provider = makeProviderFor();
while (cursor < subjects.length) {
const i = cursor++;
// The gate. Read-then-act on a shared counter: with K workers, up to K
// subjects can pass this check before any of them adds to `spent`.
if (pricing !== undefined && spent >= maxCostUsd) {
results[i] = { subject: subjects[i], status: "skipped-budget" };
skipped += 1;
continue;
}
const runs = await runEnsemble({
provider,
system,
user: `# Page\n${subjects[i]}`,
runs: 3,
});
spent += costOfRuns(runs, pricing);
results[i] = { subject: subjects[i], status: "judged" };
}
};
await Promise.all(
Array.from({ length: Math.min(concurrency, subjects.length) }, () => worker()),
);
return { results, spent, skipped };
}
const ceiling = 0.045; // five ensembles at $0.009 each
const sequential = await runAll({ concurrency: 1, maxCostUsd: ceiling });
const parallel = await runAll({ concurrency: 4, maxCostUsd: ceiling });
console.log("ceiling: $" + ceiling.toFixed(3));
console.log("sequential: spent $" + sequential.spent.toFixed(3), "| skipped", sequential.skipped);
console.log("concurrency 4: spent $" + parallel.spent.toFixed(3), "| skipped", parallel.skipped);
// The overshoot is bounded by the number of workers, not by one call.
const overshoot = (parallel.spent - sequential.spent) / 0.009;
console.log("extra ensembles under concurrency:", Math.round(overshoot));
// Order survives the pool, because results are assigned by index.
console.log("order preserved:", parallel.results.every((r, i) => r.subject === subjects[i]));
console.log("first three:", parallel.results.slice(0, 3).map((r) => r.subject).join(" "));
ceiling: $0.045
sequential: spent $0.045 | skipped 7
concurrency 4: spent $0.072 | skipped 4
extra ensembles under concurrency: 3
order preserved: true
first three: page-01.md page-02.md page-03.md

Three details in that code are easy to get wrong:

  1. Not Promise.all(subjects.map(...)). That launches every subject at once. The pool caps how many are in flight by having a fixed number of workers pull from a shared cursor.
  2. Assign results by index. A pool finishes out of order. Writing results[i] = … keeps the output aligned with the input, which Promise.all would have given you for free and a pool takes away.
  3. One provider per worker if you are using MockProvider. It records every request, and sharing one across workers interleaves them. A real provider is stateless and can be shared.

The gate is a read-then-act on a shared counter:

if (pricing !== undefined && spent >= maxCostUsd) { /* skip */ }
// ... await the call ...
spent += costOfRuns(runs, pricing);

With K workers, up to K subjects can pass the check before any of them adds to spent. The sample output measures it: the same ceiling, the same subjects, $0.045 sequential against $0.072 at concurrency 4 — three extra ensembles, one per additional worker.

So the rule generalises to: the overshoot is bounded by your pool size. Budget for maxCostUsd + (K - 1) × cost_per_subject, or gate on a reservation rather than on spend if you need a hard ceiling.

This is the consequence that surprises people, and it follows from two rules already documented separately.

  1. A rate limit is recorded on run.error like any other provider failure. It does not throw.
  2. Any errored run in an ensemble forces human-review.

Therefore: a rate limit does not surface as a rate limit. It surfaces as a growing human-review queue.

That reads like the model becoming less certain about your subjects, which invites tuning thresholds — exactly the wrong response. If your review pile grows right after you raise concurrency, lower the concurrency before you touch anything else.

There is no backoff in the library, and none of the existing consumers implement one. If you need it, it goes in your worker loop.

runs multiplies both cost and latency linearly, and buys confidence:

runs Buys Costs
1 a single opinion; agreement is always 1.0 and meaningless 1×
3 a majority, and disagreement becomes visible 3×
5 finer-grained agreement ratios 5×

3 is the common default — both production eval consumers arrived at it independently. Below 3 the consensus math has nothing to work with; above it you are paying linearly for a slowly improving signal.

If you want diversity, prefer more runs to a higher temperature. Independent samples at temperature 0 are already independent, and raising temperature warns for a reason.

The library reports durationMs per run but makes no latency promises. Because ensembles are sequential, the arithmetic is at least predictable:

wall clock ≈ (subjects / concurrency) × runs × per_call_latency
cost ≈ subjects × runs × per_call_cost (before cache hits)

Run one subject first and read durationMs and usage off the result. That is a far better estimate than any number this page could print, and it costs one call.