Diagnosing a failure
Something failed. Start here.
Every other page in this documentation is organised by what you are building. This one is organised by what you are seeing, because when a run fails you do not yet know which layer you are in.
The one question
Section titled “The one question”Did it throw, or did it come back on
run.error?
That single split does most of the diagnosis:
| Threw | Came back on run.error |
|
|---|---|---|
| Type | InferenceError |
a string on the run |
| When | at construction, before any call | after the call, after both attempts |
| Means | your configuration is wrong | the model or the provider failed |
| Costs | nothing — no call was made | tokens, usually |
| Fix | change the spec, the environment, or the call | change the schema, the prompt, or your handling |
Everything else follows. If you have a message in hand, go straight to the error reference and use your browser’s find.
Both classes, side by side
Section titled “Both classes, side by side”// Every failure class, provoked for real, so the messages on the error reference// cannot drift away from the ones the library actually produces.//// Runs with no API key and no network.
import { InferenceError, MockProvider, completeValidatedJSON, extractJson, judge, makeProvider, makeProviderAsync, mockVerdict,} from "@hawkeyexl/inference";
// ---------------------------------------------------------------------------// Class 1 — operational. Thrown as InferenceError, at construction.// ---------------------------------------------------------------------------
const thrown = (fn) => { try { fn(); return "DID NOT THROW"; } catch (error) { return `${error instanceof InferenceError ? "InferenceError" : error.constructor.name}: ${error.message.split("\n")[0]}`; }};
console.log("[operational]");console.log(" 1.", thrown(() => makeProvider({ provider: "anthropic", apiKeyEnv: "NOT_SET_XYZ" })));console.log(" 2.", thrown(() => makeProvider({})));console.log(" 3.", thrown(() => makeProvider({ provider: "llama-cpp", model: "auto" })));console.log(" 4.", thrown(() => makeProvider({ provider: "nope" })));console.log(" 5.", thrown(() => makeProvider({ provider: "llama-cpp", model: "not-a-model" })));console.log(" 6.", thrown(() => new MockProvider([])));
// ---------------------------------------------------------------------------// Class 2 — model failure. Never thrown; recorded on run.error.// ---------------------------------------------------------------------------
const schema = { type: "object", required: ["ok"], properties: { ok: { type: "boolean" } }, additionalProperties: false,};
console.log("\n[model failure — recorded, never thrown]");
// Validation exhausted after both attempts.const invalid = await completeValidatedJSON({ provider: new MockProvider([{ json: { wrong: 1 } }]), system: "s", user: "u", schema,});console.log(" 7. run.error:", invalid.error.split(";")[0]);console.log(" run.result:", invalid.result);
// A provider that rejects — here a scripted 429, in production a real one.const rejected = await completeValidatedJSON({ provider: new MockProvider([{ error: "429 rate limited" }]), system: "s", user: "u", schema,});console.log(" 8. run.error:", rejected.error);
// A response with no JSON in it at all. Shared by openai, claude-cli and llama-cpp.console.log(" 9.", thrown(() => extractJson("I'm afraid I can't do that.")));
// ---------------------------------------------------------------------------// The Claude CLI subprocess failures, through an injected exec seam.// ---------------------------------------------------------------------------
const cliError = async (result) => { const provider = makeProvider({ provider: "claude-cli", exec: async () => result }); const run = await completeValidatedJSON({ provider, system: "s", user: "u", schema }); return run.error;};
console.log("\n[claude-cli subprocess]");console.log("10.", await cliError({ code: null, stdout: "", stderr: "", timedOut: false, spawnError: "ENOENT" }));console.log("11.", await cliError({ code: null, stdout: "", stderr: "", timedOut: true }));console.log("12.", await cliError({ code: 1, stdout: "", stderr: "not logged in", timedOut: false }));console.log("13.", await cliError({ code: 0, stdout: "{}", stderr: "", timedOut: false }));console.log("14.", await cliError({ code: 0, stdout: "Welcome to Claude Code!\n\nRun /login to continue.", stderr: "", timedOut: false }));
// ---------------------------------------------------------------------------// What a failure does to a verdict — the consequence readers miss.// ---------------------------------------------------------------------------
const flaky = new MockProvider([ mockVerdict("pass", 0.95), { error: "429 rate limited" }, { error: "429 rate limited" }, mockVerdict("pass", 0.97),]);const consensus = await judge({ provider: flaky, system: "s", user: "u", runs: 3 });
console.log("\n[what it costs you]");console.log("15. two runs passed confidently, one errored");console.log(" votes:", JSON.stringify(consensus.votes));console.log(" zone:", consensus.zone, "— a rate limit becomes a review queue, not an error");[operational] 1. InferenceError: Anthropic provider needs NOT_SET_XYZ set (or choose another provider) 2. InferenceError: No provider specified. Detecting one probes the environment, the Claude CLI and the local model runtime, which cannot be done synchronously — use resolveProviderIdentityAsync/makeProviderAsync, or name a provider (anthropic, openai, claude-cli, mock, llama-cpp). 3. InferenceError: llama-cpp model "auto" is a selector and cannot be resolved synchronously — picking a tier probes GPU memory. Use resolveProviderIdentityAsync/makeProviderAsync, or name a concrete model (e.g. "qwen3.5-4b"). 4. InferenceError: Unknown provider "nope". Available: anthropic, openai, claude-cli, mock, llama-cpp. 5. InferenceError: Unknown llama-cpp model "not-a-model". Use a selector (auto, fast, balanced, quality), a curated alias (granite-4.1-3b-q2, qwen3.5-4b, qwen3.5-9b, gemma-4-e2b, gemma-4-e4b, gemma-4-12b, gemma-4-26b-a4b, gemma-4-e2b-q2), an hf: URI, or a path to a .gguf file. 6. Error: MockProvider needs at least one scripted response
[model failure — recorded, never thrown] 7. run.error: Response failed schema validation: must have required property 'ok' run.result: undefined 8. run.error: 429 rate limited 9. Error: Response contained no parseable JSON object
[claude-cli subprocess]10. Failed to run claude: ENOENT (is the Claude CLI installed?)11. Claude CLI timed out12. Claude CLI exited 1: not logged in13. Claude CLI returned no result field14. Claude CLI printed non-JSON output (is it logged in?): Welcome to Claude Code! Run /login to continue.
[what it costs you]15. two runs passed confidently, one errored votes: {"pass":2,"fail":0,"partial":0,"error":1} zone: human-review — a rate limit becomes a review queue, not an errorWork through it
Section titled “Work through it”-
Read the last line of the message, not the first. Nearly every operational error names the fix at the end —
use makeProviderAsync,npm i node-llama-cpp,Available: …. -
If it threw, nothing was spent. Operational failures happen at construction. You can retry freely.
-
If it came back on the run, check
run.error— not atry/catch. A model failure never throws. If you are wrapping the call intry/catchand seeing nothing, this is why. -
Look for a warning you missed. The library writes eight of them to
console.warn, and five fire exactly once per process. See the warnings reference. -
Check whether the failure was silent. An errored run inside an ensemble does not surface as an error — it forces
human-review. A growing review queue is a symptom worth reading as infrastructure trouble rather than model uncertainty.
The one that is easy to misread
Section titled “The one that is easy to misread”A 429 is recorded like any other provider failure, and any errored run in an ensemble
forces human-review.
So a rate limit does not appear as a rate limit. It appears as more results needing human attention — which reads like the model getting less certain, and invites exactly the wrong response. If your review queue grows after you raise concurrency, read running at scale before you tune anything else.
- Common failures — the ones worth prose
- Error reference — every message, verbatim
- Warnings reference — the eight
console.warnpaths - Budgets and errors — how to handle each class in code