Skip to content

Diagnosing a failure

Something failed. Start here.

Every other page in this documentation is organised by what you are building. This one is organised by what you are seeing, because when a run fails you do not yet know which layer you are in.

Did it throw, or did it come back on run.error?

That single split does most of the diagnosis:

Threw Came back on run.error
Type InferenceError a string on the run
When at construction, before any call after the call, after both attempts
Means your configuration is wrong the model or the provider failed
Costs nothing — no call was made tokens, usually
Fix change the spec, the environment, or the call change the schema, the prompt, or your handling

Everything else follows. If you have a message in hand, go straight to the error reference and use your browser’s find.

examples/diagnose-errors.mjs
// Every failure class, provoked for real, so the messages on the error reference
// cannot drift away from the ones the library actually produces.
//
// Runs with no API key and no network.
import {
InferenceError,
MockProvider,
completeValidatedJSON,
extractJson,
judge,
makeProvider,
makeProviderAsync,
mockVerdict,
} from "@hawkeyexl/inference";
// ---------------------------------------------------------------------------
// Class 1 — operational. Thrown as InferenceError, at construction.
// ---------------------------------------------------------------------------
const thrown = (fn) => {
try {
fn();
return "DID NOT THROW";
} catch (error) {
return `${error instanceof InferenceError ? "InferenceError" : error.constructor.name}: ${error.message.split("\n")[0]}`;
}
};
console.log("[operational]");
console.log(" 1.", thrown(() => makeProvider({ provider: "anthropic", apiKeyEnv: "NOT_SET_XYZ" })));
console.log(" 2.", thrown(() => makeProvider({})));
console.log(" 3.", thrown(() => makeProvider({ provider: "llama-cpp", model: "auto" })));
console.log(" 4.", thrown(() => makeProvider({ provider: "nope" })));
console.log(" 5.", thrown(() => makeProvider({ provider: "llama-cpp", model: "not-a-model" })));
console.log(" 6.", thrown(() => new MockProvider([])));
// ---------------------------------------------------------------------------
// Class 2 — model failure. Never thrown; recorded on run.error.
// ---------------------------------------------------------------------------
const schema = {
type: "object",
required: ["ok"],
properties: { ok: { type: "boolean" } },
additionalProperties: false,
};
console.log("\n[model failure — recorded, never thrown]");
// Validation exhausted after both attempts.
const invalid = await completeValidatedJSON({
provider: new MockProvider([{ json: { wrong: 1 } }]),
system: "s",
user: "u",
schema,
});
console.log(" 7. run.error:", invalid.error.split(";")[0]);
console.log(" run.result:", invalid.result);
// A provider that rejects — here a scripted 429, in production a real one.
const rejected = await completeValidatedJSON({
provider: new MockProvider([{ error: "429 rate limited" }]),
system: "s",
user: "u",
schema,
});
console.log(" 8. run.error:", rejected.error);
// A response with no JSON in it at all. Shared by openai, claude-cli and llama-cpp.
console.log(" 9.", thrown(() => extractJson("I'm afraid I can't do that.")));
// ---------------------------------------------------------------------------
// The Claude CLI subprocess failures, through an injected exec seam.
// ---------------------------------------------------------------------------
const cliError = async (result) => {
const provider = makeProvider({ provider: "claude-cli", exec: async () => result });
const run = await completeValidatedJSON({ provider, system: "s", user: "u", schema });
return run.error;
};
console.log("\n[claude-cli subprocess]");
console.log("10.", await cliError({ code: null, stdout: "", stderr: "", timedOut: false, spawnError: "ENOENT" }));
console.log("11.", await cliError({ code: null, stdout: "", stderr: "", timedOut: true }));
console.log("12.", await cliError({ code: 1, stdout: "", stderr: "not logged in", timedOut: false }));
console.log("13.", await cliError({ code: 0, stdout: "{}", stderr: "", timedOut: false }));
console.log("14.", await cliError({ code: 0, stdout: "Welcome to Claude Code!\n\nRun /login to continue.", stderr: "", timedOut: false }));
// ---------------------------------------------------------------------------
// What a failure does to a verdict — the consequence readers miss.
// ---------------------------------------------------------------------------
const flaky = new MockProvider([
mockVerdict("pass", 0.95),
{ error: "429 rate limited" },
{ error: "429 rate limited" },
mockVerdict("pass", 0.97),
]);
const consensus = await judge({ provider: flaky, system: "s", user: "u", runs: 3 });
console.log("\n[what it costs you]");
console.log("15. two runs passed confidently, one errored");
console.log(" votes:", JSON.stringify(consensus.votes));
console.log(" zone:", consensus.zone, "— a rate limit becomes a review queue, not an error");
[operational]
1. InferenceError: Anthropic provider needs NOT_SET_XYZ set (or choose another provider)
2. InferenceError: No provider specified. Detecting one probes the environment, the Claude CLI and the local model runtime, which cannot be done synchronously — use resolveProviderIdentityAsync/makeProviderAsync, or name a provider (anthropic, openai, claude-cli, mock, llama-cpp).
3. InferenceError: llama-cpp model "auto" is a selector and cannot be resolved synchronously — picking a tier probes GPU memory. Use resolveProviderIdentityAsync/makeProviderAsync, or name a concrete model (e.g. "qwen3.5-4b").
4. InferenceError: Unknown provider "nope". Available: anthropic, openai, claude-cli, mock, llama-cpp.
5. InferenceError: Unknown llama-cpp model "not-a-model". Use a selector (auto, fast, balanced, quality), a curated alias (granite-4.1-3b-q2, qwen3.5-4b, qwen3.5-9b, gemma-4-e2b, gemma-4-e4b, gemma-4-12b, gemma-4-26b-a4b, gemma-4-e2b-q2), an hf: URI, or a path to a .gguf file.
6. Error: MockProvider needs at least one scripted response
[model failure — recorded, never thrown]
7. run.error: Response failed schema validation: must have required property 'ok'
run.result: undefined
8. run.error: 429 rate limited
9. Error: Response contained no parseable JSON object
[claude-cli subprocess]
10. Failed to run claude: ENOENT (is the Claude CLI installed?)
11. Claude CLI timed out
12. Claude CLI exited 1: not logged in
13. Claude CLI returned no result field
14. Claude CLI printed non-JSON output (is it logged in?): Welcome to Claude Code! Run /login to continue.
[what it costs you]
15. two runs passed confidently, one errored
votes: {"pass":2,"fail":0,"partial":0,"error":1}
zone: human-review — a rate limit becomes a review queue, not an error
  1. Read the last line of the message, not the first. Nearly every operational error names the fix at the end — use makeProviderAsync, npm i node-llama-cpp, Available: ….

  2. If it threw, nothing was spent. Operational failures happen at construction. You can retry freely.

  3. If it came back on the run, check run.error — not a try/catch. A model failure never throws. If you are wrapping the call in try/catch and seeing nothing, this is why.

  4. Look for a warning you missed. The library writes eight of them to console.warn, and five fire exactly once per process. See the warnings reference.

  5. Check whether the failure was silent. An errored run inside an ensemble does not surface as an error — it forces human-review. A growing review queue is a symptom worth reading as infrastructure trouble rather than model uncertainty.

A 429 is recorded like any other provider failure, and any errored run in an ensemble forces human-review.

So a rate limit does not appear as a rate limit. It appears as more results needing human attention — which reads like the model getting less certain, and invites exactly the wrong response. If your review queue grows after you raise concurrency, read running at scale before you tune anything else.