Skip to content

Common failures

Nine failures that account for most of the trouble, each with more context than the error reference can give in a row.

No inference provider is available. Tried:
anthropic: ANTHROPIC_API_KEY is not set
openai: OPENAI_API_KEY is not set and no baseUrl was given
claude-cli: could not run `claude` (is the Claude CLI installed?)
llama-cpp: node-llama-cpp is not installed (npm i node-llama-cpp)

This is the most likely first failure, because the zero-config path is the one we recommend. You called makeProviderAsync({}) on a machine with no credentials and no local runtime, so every probe failed.

Any one of these fixes it:

Fix Cost
export ANTHROPIC_API_KEY=… or OPENAI_API_KEY=… an account
Point baseUrl at a local OpenAI-compatible server a running server, no key
Install and log into the Claude CLI an existing Claude subscription
npm install node-llama-cpp a few GB of disk, no account

The last row is the one people miss: it needs no credential at all. See run models locally.

llama-cpp model "auto" is a selector and cannot be resolved synchronously — picking a tier probes
GPU memory. Use resolveProviderIdentityAsync/makeProviderAsync, or name a concrete model.

makeProvider({ provider: "llama-cpp" }) throws, because the default model is auto and resolving it reads GPU memory.

// throws
const provider = makeProvider({ provider: "llama-cpp" });
// works
const provider = await makeProviderAsync({ provider: "llama-cpp" });

The refusal is deliberate rather than a limitation: recording the literal "auto" as cache-key material would let a 2.62 GB and a 6.72 GB model share cached results, and make one key mean different things on two machines.

makeProviderAsync delegates to makeProvider for every other provider, so switching wholesale is safe.

Claude CLI exited 1: <the last 300 characters of stderr>

The stderr tail is included because it carries the real reason, and “not logged in” is by far the most common one.

Check it directly:

Terminal window
claude --version # is it installed and on PATH?
claude -p "hi" # are you authenticated?

If the first fails you will see Failed to run claude: ENOENT instead — a different error with a different fix (install it, or set command to its path).

Response contained no parseable JSON object

The provider got a reply with no JSON in it — no bare object, no fenced block, nothing embedded in prose. Shared by openai, claude-cli, and llama-cpp, since all three parse text rather than receiving native structured output.

Usually the system prompt. Ask for the object directly and give the model nothing else to do. If it persists, anthropic constrains output with a forced tool call and cannot answer in prose at all.

Response failed schema validation: must have required property 'summary'; …

Two attempts, both rejected by Ajv. Before blaming the model, check the schema is expressible: deeply nested objects, long enum lists, and tight pattern constraints are all common causes.

A rate limit arrives as an ordinary errored run — 429 rate limited, or whatever your provider returns — and an errored run forces human-review.

So the symptom is a growing human-review queue, not an error. It reads like the model becoming less certain, which invites tuning thresholds instead of backing off the request rate.

If your review pile grew right after you raised concurrency, that is this. See running at scale.

Your maxCostUsd never trips and spend keeps climbing.

The gate is almost certainly inert rather than satisfied: pricingFor returned undefined, so every cost is 0 and every comparison passes.

const pricing = pricingFor(provider.modelName(), config.pricing);
if (pricing === undefined) warn(`no price for ${provider.modelName()} — the budget gate is inert`);

Three roads lead here: a hosted model absent from the price table, claude-cli (which reports no token usage at all), and any local model. Full detail in cost and budgets.

D:
ode-llama-cpp
ode-llama-cpp\llama\llama.cpp\ggml\src\ggml-cuda\ggml-cuda.cu:106: CUDA error

On some Windows machines with an NVIDIA GPU, node-llama-cpp’s prebuilt CUDA backend aborts partway through generation. It is a native abort, so there is no exception to catch. Before version 0.4.0 the abort ended your process, along with every result it had computed. Environment switches such as GGML_CUDA_DISABLE_GRAPHS and CUDA_LAUNCH_BLOCKING do not stop it.

Since 0.4.0 the model runs in a worker process, so the abort ends only the worker. What happens next depends on whether you chose the backend:

Backend What you see
"auto" (the default) one warning, CUDA backend crashed … Retrying on Vulkan. The call succeeds, and the rest of the process runs on Vulkan.
pinned with llamaCpp.gpu or NODE_LLAMA_CPP_GPU the call fails with an InferenceError that names the crash, and an ensemble records it on run.error

Fix: on such a machine, start on Vulkan with NODE_LLAMA_CPP_GPU=vulkan or llamaCpp.gpu: "vulkan". That skips the crash and the retry it costs. On an RTX 4090, Vulkan judged a call in about 12 seconds, where the CPU took over ten minutes.

Something ran twice that should have been cached

Section titled “Something ran twice that should have been cached”

Cache misses that should be hits usually come down to one of three things:

  1. Both cache and cacheKey are required. The cache is consulted only when you pass both.
  2. A fresh schema object per call defeats validatorFor’s identity-keyed cache — a different problem with the same smell. See the performance trap.
  3. A resolved selector changed the model name. Weights cached under auto on a small machine will not be hit on a larger one, because the key names the concrete model that ran. That is the async factory working as intended.