Common failures
Nine failures that account for most of the trouble, each with more context than the error reference can give in a row.
No provider is available
Section titled “No provider is available”No inference provider is available. Tried: anthropic: ANTHROPIC_API_KEY is not set openai: OPENAI_API_KEY is not set and no baseUrl was given claude-cli: could not run `claude` (is the Claude CLI installed?) llama-cpp: node-llama-cpp is not installed (npm i node-llama-cpp)This is the most likely first failure, because the zero-config path is the one we recommend. You
called makeProviderAsync({}) on a machine with no credentials and no local runtime, so every probe
failed.
Any one of these fixes it:
| Fix | Cost |
|---|---|
export ANTHROPIC_API_KEY=… or OPENAI_API_KEY=… |
an account |
Point baseUrl at a local OpenAI-compatible server |
a running server, no key |
| Install and log into the Claude CLI | an existing Claude subscription |
npm install node-llama-cpp |
a few GB of disk, no account |
The last row is the one people miss: it needs no credential at all. See run models locally.
A selector reached a synchronous factory
Section titled “A selector reached a synchronous factory”llama-cpp model "auto" is a selector and cannot be resolved synchronously — picking a tier probesGPU memory. Use resolveProviderIdentityAsync/makeProviderAsync, or name a concrete model.makeProvider({ provider: "llama-cpp" }) throws, because the default model is auto and
resolving it reads GPU memory.
// throwsconst provider = makeProvider({ provider: "llama-cpp" });
// worksconst provider = await makeProviderAsync({ provider: "llama-cpp" });The refusal is deliberate rather than a limitation: recording the literal "auto" as cache-key
material would let a 2.62 GB and a 6.72 GB model share cached results, and make one key mean
different things on two machines.
makeProviderAsync delegates to makeProvider for every other provider, so switching wholesale is
safe.
The Claude CLI is there but not logged in
Section titled “The Claude CLI is there but not logged in”Claude CLI exited 1: <the last 300 characters of stderr>The stderr tail is included because it carries the real reason, and “not logged in” is by far the most common one.
Check it directly:
claude --version # is it installed and on PATH?claude -p "hi" # are you authenticated?If the first fails you will see Failed to run claude: ENOENT instead — a different error with a
different fix (install it, or set command to its path).
A model that answers in prose
Section titled “A model that answers in prose”Response contained no parseable JSON objectThe provider got a reply with no JSON in it — no bare object, no fenced block, nothing embedded in
prose. Shared by openai, claude-cli, and llama-cpp, since all three parse text rather than
receiving native structured output.
Usually the system prompt. Ask for the object directly and give the model nothing else to do. If it
persists, anthropic constrains output with a forced tool call and cannot answer in prose at all.
A schema the model cannot satisfy
Section titled “A schema the model cannot satisfy”Response failed schema validation: must have required property 'summary'; …Two attempts, both rejected by Ajv. Before blaming the model, check the schema is expressible:
deeply nested objects, long enum lists, and tight pattern constraints are all common causes.
A 429 that does not look like one
Section titled “A 429 that does not look like one”A rate limit arrives as an ordinary errored run — 429 rate limited, or whatever your provider
returns — and an errored run
forces human-review.
So the symptom is a growing human-review queue, not an error. It reads like the model becoming less certain, which invites tuning thresholds instead of backing off the request rate.
If your review pile grew right after you raised concurrency, that is this. See running at scale.
A budget that is not being enforced
Section titled “A budget that is not being enforced”Your maxCostUsd never trips and spend keeps climbing.
The gate is almost certainly inert rather than satisfied: pricingFor returned undefined, so
every cost is 0 and every comparison passes.
const pricing = pricingFor(provider.modelName(), config.pricing);if (pricing === undefined) warn(`no price for ${provider.modelName()} — the budget gate is inert`);Three roads lead here: a hosted model absent from the price table, claude-cli (which reports no
token usage at all), and any local model. Full detail in
cost and budgets.
A local model run dies with a CUDA error
Section titled “A local model run dies with a CUDA error”D:ode-llama-cppode-llama-cpp\llama\llama.cpp\ggml\src\ggml-cuda\ggml-cuda.cu:106: CUDA errorOn some Windows machines with an NVIDIA GPU, node-llama-cpp’s prebuilt CUDA backend aborts partway
through generation. It is a native abort, so there is no exception to catch. Before version 0.4.0
the abort ended your process, along with every result it had computed. Environment switches such as
GGML_CUDA_DISABLE_GRAPHS and CUDA_LAUNCH_BLOCKING do not stop it.
Since 0.4.0 the model runs in a worker process, so the abort ends only the worker. What happens next depends on whether you chose the backend:
| Backend | What you see |
|---|---|
"auto" (the default) |
one warning, CUDA backend crashed … Retrying on Vulkan. The call succeeds, and the rest of the process runs on Vulkan. |
pinned with llamaCpp.gpu or NODE_LLAMA_CPP_GPU |
the call fails with an InferenceError that names the crash, and an ensemble records it on run.error |
Fix: on such a machine, start on Vulkan with NODE_LLAMA_CPP_GPU=vulkan or
llamaCpp.gpu: "vulkan". That skips the crash and the retry it costs. On an RTX 4090, Vulkan judged
a call in about 12 seconds, where the CPU took over ten minutes.
Something ran twice that should have been cached
Section titled “Something ran twice that should have been cached”Cache misses that should be hits usually come down to one of three things:
- Both
cacheandcacheKeyare required. The cache is consulted only when you pass both. - A fresh schema object per call defeats
validatorFor’s identity-keyed cache — a different problem with the same smell. See the performance trap. - A resolved selector changed the model name. Weights cached under
autoon a small machine will not be hit on a larger one, because the key names the concrete model that ran. That is the async factory working as intended.
Still stuck
Section titled “Still stuck”- Error reference — every message, verbatim
- Warnings reference — what the library logs, and when
- Open an issue — include the exact message and which provider you were on