Wiring it into a CLI
Turning a working call into a subcommand somebody else runs.
Every consumer of this package built the same scaffolding around it — the same statuses, the same exit codes, the same dry-run semantics. That convergence is why this page exists: it is not incidental glue, it is the shape the library implies.
Give every subject a status
Section titled “Give every subject a status”A per-subject result is not a boolean. Four outcomes, and each needs to be distinguishable:
| Status | Meaning | Called the model? |
|---|---|---|
filled |
proposed and written | yes |
complete |
nothing was missing | no |
low-confidence |
proposed, but below your threshold | yes |
skipped-budget |
the ceiling stopped it first | no |
error |
the call failed, or never validated | yes |
skipped-budget is the one people collapse. Folding it into error makes a deliberate cost
decision look like a failure; folding it into complete hides work that never happened.
The whole loop
Section titled “The whole loop”// Turning a call into a subcommand: per-subject statuses, an exit-code contract,// a dry run that makes the real run free, and a confidence gate.//// Runs with no API key: MockProvider stands in for a real provider.
import { mkdtempSync, rmSync } from "node:fs";import { tmpdir } from "node:os";import { join } from "node:path";
import { JsonCache, MockProvider, buildCacheKey, completeValidatedJSON, costOfUsage, pricingFor, sha256,} from "@hawkeyexl/inference";
const cacheDir = mkdtempSync(join(tmpdir(), "inference-cli-"));const PROMPT_VERSION = 1;const MIN_CONFIDENCE = 0.7;const MAX_COST_USD = 0.009; // three calls, at $0.003 each
const schema = { type: "object", required: ["title", "confidence"], properties: { title: { type: "string" }, // Ask the model to score its own answer. A schema-valid answer can still be a bad one. confidence: { type: "number", minimum: 0, maximum: 1 }, }, additionalProperties: false,};
const documents = [ { path: "auth.md", missing: ["title"] }, { path: "limits.md", missing: ["title"] }, { path: "webhooks.md", missing: [] }, // nothing to do — must not cost a call { path: "errors.md", missing: ["title"] }, { path: "retries.md", missing: ["title"] },];
const provider = new MockProvider( [ { json: { title: "Authentication", confidence: 0.94 } }, { json: { title: "Rate limits", confidence: 0.41 } }, // below the gate { json: { title: "Errors", confidence: 0.88 } }, ], "claude-sonnet-4-5",);
const pricing = pricingFor(provider.modelName());const cache = new JsonCache(cacheDir, true, "my-tool");
// The dry-run flag is deliberately NOT part of the key. That is what lets the real// run replay what the dry run already paid for.const keyFor = (doc) => buildCacheKey([ provider.provider(), provider.modelName(), `v${PROMPT_VERSION}`, // The existing state belongs in the key: a re-run after a partial fill then asks // for what is still missing rather than replaying a proposal already applied. sha256([...doc.missing].sort().join(",")), sha256(doc.path), ]);
async function run({ dryRun }) { const statuses = []; let spent = 0; let calls = 0; let replays = 0;
for (const doc of documents) { // Skip before you spend: nothing missing means no call at all. if (doc.missing.length === 0) { statuses.push([doc.path, "complete"]); continue; } if (pricing !== undefined && spent >= MAX_COST_USD) { statuses.push([doc.path, "skipped-budget"]); continue; }
const key = keyFor(doc); let proposal = cache.get(key); if (proposal !== undefined) replays += 1;
if (proposal === undefined) { const result = await completeValidatedJSON({ provider, system: "You propose a title for a documentation page.", user: `Path: ${doc.path}\nMissing: ${doc.missing.join(", ")}`, schema, }); calls += 1; spent += costOfUsage(result.usage, pricing); if (result.error !== undefined) { statuses.push([doc.path, "error"]); continue; } // Cache BEFORE gating, so re-tuning the threshold costs nothing. proposal = result.result; cache.set(key, proposal); }
if (proposal.confidence < MIN_CONFIDENCE) { statuses.push([doc.path, "low-confidence"]); continue; } // The dry run does everything except write. if (!dryRun) applyToDisk(doc, proposal); statuses.push([doc.path, "filled"]); }
return { statuses, spent, calls, replays };}
function applyToDisk() { /* the only thing --dry-run skips */}
const dry = await run({ dryRun: true });console.log("--- dry run ---");for (const [path, status] of dry.statuses) console.log(" ", status.padEnd(15), path);console.log(" calls:", dry.calls, "| replays:", dry.replays, "| spent $" + dry.spent.toFixed(6));
const real = await run({ dryRun: false });console.log("--- real run ---");for (const [path, status] of real.statuses) console.log(" ", status.padEnd(15), path);console.log(" calls:", real.calls, "| replays:", real.replays, "| spent $" + real.spent.toFixed(6));
// Everything the dry run reached is replayed for free. And because replays cost// nothing, the budget now stretches further — retries.md gets judged this time,// where the dry run's ceiling had already stopped it.console.log(" replayed from the dry run:", real.replays);console.log(" reached further on the re-run:", dry.statuses.at(-1)[1], "->", real.statuses.at(-1)[1]);
// Exit codes: 0 clean, 1 findings, 2 operational. An operational failure would have// thrown before any of this and is mapped separately.const hadError = real.statuses.some(([, s]) => s === "error");const hadFindings = real.statuses.some(([, s]) => s === "low-confidence" || s === "skipped-budget");console.log("exit code:", hadError ? 1 : hadFindings ? 1 : 0);
rmSync(cacheDir, { recursive: true, force: true });--- dry run --- filled auth.md low-confidence limits.md complete webhooks.md filled errors.md skipped-budget retries.md calls: 3 | replays: 0 | spent $0.009000--- real run --- filled auth.md low-confidence limits.md complete webhooks.md filled errors.md filled retries.md calls: 1 | replays: 3 | spent $0.003000 replayed from the dry run: 3 reached further on the re-run: skipped-budget -> filledexit code: 1A dry run is not free
Section titled “A dry run is not free”The natural reading of --dry-run is “skip the expensive part.” For an LLM-backed command the
expensive part has already happened by the time you decide not to write — the call was made and
paid for. Only the write is skipped.
What makes the pattern work is the cache:
In the sample output, the real run makes one call instead of four: three subjects replay from what the dry run already bought.
Cache before you gate
Section titled “Cache before you gate”Notice the ordering:
proposal = result.result;cache.set(key, proposal); // cache the raw proposal...if (proposal.confidence < MIN_CONFIDENCE) { /* ...then gate */ }Caching the pre-gated proposal means sweeping the threshold costs nothing — re-run with a
different --confidence and every subject re-gates from cache. Gate before caching and every sweep
is a fresh bill.
Skip before you spend
Section titled “Skip before you spend”Two things belong in the key that are easy to leave out:
buildCacheKey([ provider.provider(), provider.modelName(), `v${PROMPT_VERSION}`, sha256([...doc.missing].sort().join(",")), // the existing state sha256(doc.path),]);Putting what is still missing in the key makes a re-run additive: after a partial fill, the
second run asks the model for what remains rather than replaying a proposal already applied. And a
subject with nothing missing short-circuits to complete before any of this — no key, no call.
Sort the field set before hashing, or two orderings of the same fields become two keys.
Exit codes
Section titled “Exit codes”0 clean1 findings — something needs a human2 operational — the tool could not run2 is where a translated InferenceError lands, which is why
translating at the boundary
matters: an untranslated one escapes as a crash with the wrong code.
low-confidence and skipped-budget are findings, not failures — the tool worked, the result needs
attention.
A cached re-run gets further
Section titled “A cached re-run gets further”The last line of the sample output is worth reading twice:
reached further on the re-run: skipped-budget -> filledretries.md was cut off by the ceiling during the dry run. On the real run the first three subjects
replayed for free, so the budget was still intact when the loop reached it.
That is the intended interaction, and it is a good reason to make re-running cheap and habitual — but it also means a budget-limited run is not deterministic across invocations. If you need reproducibility, cache first and gate on a fixed subject list rather than on remaining budget.
- Running at scale — concurrency around this loop
- Caching — key composition and where the cache lives
- Budgets and errors — the two error classes