Error reference
Every message the library can produce, verbatim, with its trigger and its fix.
If you arrived with a message in hand, use your browser’s find. If you are not sure which class of failure you have, start at diagnosing a failure.
Operational — thrown as InferenceError
Section titled “Operational — thrown as InferenceError”Anthropic provider needs <VAR> set (or choose another provider)
Section titled “Anthropic provider needs <VAR> set (or choose another provider)”The Anthropic provider was constructed and process.env[apiKeyEnv] was empty. <VAR> is
ANTHROPIC_API_KEY unless you set apiKeyEnv.
Fix: export the key, name a different provider, or omit provider entirely and let
detection pick one that works.
OpenAI provider needs <VAR> set (or point baseUrl at a local server)
Section titled “OpenAI provider needs <VAR> set (or point baseUrl at a local server)”Same, for the OpenAI-compatible provider. A key is required only when baseUrl contains
api.openai.com; a keyless local server is allowed.
Fix: export the key, or set baseUrl to your local endpoint.
No inference provider is available. Tried:
Section titled “No inference provider is available. Tried:”Detection ran every probe and none succeeded. The message continues with one line per provider
explaining why each was unavailable, then Pass an explicit provider, set one of the keys above, or install node-llama-cpp.
The reason lines are:
| Reason | Meaning |
|---|---|
ANTHROPIC_API_KEY is not set |
no Anthropic key in the environment |
OPENAI_API_KEY is not set and no baseUrl was given |
no OpenAI key and no local endpoint |
could not run `claude` (is the Claude CLI installed?) |
the CLI is absent or not on PATH |
node-llama-cpp is not installed (npm i node-llama-cpp) |
the optional peer dependency is missing |
node-llama-cpp could not start (…) |
installed, but the native binding failed to load |
node-llama-cpp is installed but failed to load (…) |
resolved and then threw — ABI, system library, or Node version. Installing cannot fix it, so it is not attempted |
node-llama-cpp is not installed and INFERENCE_NO_AUTO_INSTALL is set |
absent, and the automatic install was refused |
not auto-selectable |
only ever shown for mock, which detection never picks |
Fix: whichever line is closest to your situation. This is the most common first failure on a machine with no credentials — see common failures.
No provider specified. Detecting one probes the environment, …
Section titled “No provider specified. Detecting one probes the environment, …”No provider specified. Detecting one probes the environment, the Claude CLI and the localmodel runtime, which cannot be done synchronously — useresolveProviderIdentityAsync/makeProviderAsync, or name a provider (anthropic, openai,claude-cli, mock, llama-cpp).You called the synchronous makeProvider or resolveProviderIdentity with provider omitted
or set to "auto".
Fix: await makeProviderAsync(spec), or name a provider.
Model "<M>" was given without a provider, and a model name does not say which provider owns it
Section titled “Model "<M>" was given without a provider, and a model name does not say which provider owns it”Model "gpt-4o-mini" was given without a provider, and a model name does not say whichprovider owns it. Name the provider too (anthropic, openai, claude-cli, mock, llama-cpp),or drop the model to take the detected provider's default.You set model but left provider omitted or "auto". A model name belongs to exactly one
provider, so it cannot be carried into whichever provider detection happens to pick — before this
check, { model: "gpt-4o-mini" } on a machine with an Anthropic key selected anthropic and then
404’d at call time, after you had already paid for detection.
model: null is still “use the provider’s default” — only a real name is ambiguous.
Fix: name the provider alongside the model, or drop model.
node-llama-cpp is installed but failed to load (<ERR>)
Section titled “node-llama-cpp is installed but failed to load (<ERR>)”node-llama-cpp is installed but failed to load (<ERR>). This is the copy resolved fromyour own node_modules, so reinstalling it here will not help — check the Node version andthe platform build.The package resolved and then failed: an ABI mismatch, a missing system library, or a Node version the native build does not support.
This is deliberately not treated as “not installed”. Only ERR_MODULE_NOT_FOUND triggers the
automatic install; anything else would fetch the same broken package again and bury the real cause
under a download. Detection reports the same thing as
node-llama-cpp is installed but failed to load (…) in its reason lines.
Fix: whatever the inner error says — usually rebuilding for your Node version, or a platform
with no prebuilt binary and no toolchain. INFERENCE_NO_AUTO_INSTALL is irrelevant here; nothing
was going to be installed.
The llama-cpp provider needs node-llama-cpp, and INFERENCE_NO_AUTO_INSTALL is set
Section titled “The llama-cpp provider needs node-llama-cpp, and INFERENCE_NO_AUTO_INSTALL is set”The local runtime is absent and you have refused the automatic install.
Fix: npm i node-llama-cpp@^3.19.0 yourself, unset INFERENCE_NO_AUTO_INSTALL, or name a
provider that does not need it.
Installing node-llama-cpp into <DIR> failed (exit <N>)
Section titled “Installing node-llama-cpp into <DIR> failed (exit <N>)”npm ran and returned non-zero. The message carries the last lines of npm’s output, then repeats the manual command. Nothing is left behind: the prefix is marked ready only after a clean exit, so the next call retries rather than importing a half-installed directory.
Fix: whatever npm’s output says — usually network, registry auth, or a missing build toolchain on a platform with no prebuilt binary.
Could not run npm to install node-llama-cpp (<ERR>)
Section titled “Could not run npm to install node-llama-cpp (<ERR>)”npm itself could not be spawned — common on images that ship node without npm, or where only pnpm
or yarn is on PATH.
Fix: install the package with whatever package manager you do have, or name a different provider.
Installing node-llama-cpp into <DIR> timed out
Section titled “Installing node-llama-cpp into <DIR> timed out”The install exceeded 15 minutes. On a platform without a prebuilt binary this is a source build, and a slow machine can genuinely exceed it.
Fix: retry, raise timeoutMs, or install it yourself.
Timed out waiting for another process to install node-llama-cpp into <DIR>
Section titled “Timed out waiting for another process to install node-llama-cpp into <DIR>”Another process holds the prefix lock. Concurrent runs serialise so two npm processes never write
one node_modules.
Fix: wait for the other run. If nothing else is running, the lock was orphaned by a killed
process — remove the named .install.lock and retry.
llama-cpp model "<M>" is a selector and cannot be resolved synchronously — …
Section titled “llama-cpp model "<M>" is a selector and cannot be resolved synchronously — …”A selector (auto, fast, balanced, quality) reached a synchronous factory. Resolving it reads
GPU memory.
Fix: await makeProviderAsync(spec), or name a concrete model. The refusal is deliberate:
recording "auto" as cache-key material would let two differently-sized models share results.
llama-cpp model "<M>" is a selector and needs a hardware probe to resolve.
Section titled “llama-cpp model "<M>" is a selector and needs a hardware probe to resolve.”The same cause reaching a different entry point — resolveLlamaModelRef, and through it
blobNameFor, isModelDownloaded, or clearLlamaModels({ models }).
Fix: resolve the selector first, or pass a concrete alias to those helpers.
llama-cpp model "<M>" is a selector. Constructing a provider directly needs a concrete model …
Section titled “llama-cpp model "<M>" is a selector. Constructing a provider directly needs a concrete model …”new LlamaCppProvider("auto") — the class constructor rather than a factory.
Fix: await makeProviderAsync({ provider: "llama-cpp" }).
llamaCpp.contextSize must be a positive integer number of tokens, got <X>.
Section titled “llamaCpp.contextSize must be a positive integer number of tokens, got <X>.”llamaCpp.contextSize was zero, negative, fractional, or not a number. Thrown when the provider is
constructed.
Fix: pass a whole number of tokens, such as 8192, or leave it unset so each context is sized to
its prompt.
llamaCpp.gpu must be "auto", "cuda", "vulkan", "metal" or false, got <X>.
Section titled “llamaCpp.gpu must be "auto", "cuda", "vulkan", "metal" or false, got <X>.”llamaCpp.gpu named a backend node-llama-cpp does not have, such as "rocm", or was not a string
or false. Thrown when the provider is constructed.
Fix: use one of the listed values, or leave it unset so NODE_LLAMA_CPP_GPU decides, and
failing that "auto". The CPU is false, not "cpu".
Unknown provider "<P>". Available: anthropic, openai, claude-cli, mock, llama-cpp.
Section titled “Unknown provider "<P>". Available: anthropic, openai, claude-cli, mock, llama-cpp.”A provider value outside the five. Usually a typo, or free text arriving from a CLI flag.
Fix: use one of the listed names. If the value came from user input, validate it before building the spec — every existing consumer does exactly this.
Unknown llama-cpp model "<M>". Use a selector …, a curated alias …, an hf: URI, or a path to a .gguf file.
Section titled “Unknown llama-cpp model "<M>". Use a selector …, a curated alias …, an hf: URI, or a path to a .gguf file.”The model reference matched no accepted form.
The llama-cpp provider needs the optional peer dependency node-llama-cpp. Install it with: npm i node-llama-cpp
Section titled “The llama-cpp provider needs the optional peer dependency node-llama-cpp. Install it with: npm i node-llama-cpp”The dynamic import failed. The original error is appended in parentheses.
Fix: npm install node-llama-cpp. It is an optional peer dependency and is never installed for
you.
Model failures — recorded on run.error
Section titled “Model failures — recorded on run.error”None of these throw. Each arrives with result (or verdict) absent.
Response failed schema validation: <path> <message>; …
Section titled “Response failed schema validation: <path> <message>; …”Ajv rejected the response on the final attempt. One clause per validation error, joined with ; .
Fix: usually the schema, not the model. Check the shape is expressible — and if you are on a
local model, that you are not relying on required or numeric bounds, which
a grammar does not enforce.
Response contained no parseable JSON object
Section titled “Response contained no parseable JSON object”extractJson found nothing usable — no bare object, no fenced block, no object embedded in prose.
Shared by openai, claude-cli, and llama-cpp.
Fix: most often a model that answered in prose. Tighten the system prompt, or use a provider with native structured output.
<provider message> — or HTTP <status>
Section titled “<provider message> — or HTTP <status>”The OpenAI-compatible endpoint returned a non-2xx. The provider’s own error.message is passed
through verbatim; when the body has none you get HTTP 401, HTTP 429, HTTP 500.
Empty completion response
Section titled “Empty completion response”The endpoint returned 2xx but choices[0].message.content was missing or empty.
Anthropic response hit max_tokens (<N>) before completing the tool call — raise the anthropic.maxTokens option.
Section titled “Anthropic response hit max_tokens (<N>) before completing the tool call — raise the anthropic.maxTokens option.”The forced tool call was truncated. Guarded early and reported as itself rather than being allowed to fail validation, which would have burned the retry on a misleading error.
Fix: raise anthropic.maxTokens (default 1024).
Anthropic response contained no tool_use block
Section titled “Anthropic response contained no tool_use block”The response came back without the forced tool call. Rare; usually a refusal.
llama-cpp generation hit the token limit before completing the JSON (maxTokens: <N>) — raise llamaCpp.maxTokens, or shorten the prompt if the context is full.
Section titled “llama-cpp generation hit the token limit before completing the JSON (maxTokens: <N>) — raise llamaCpp.maxTokens, or shorten the prompt if the context is full.”The local-model equivalent of the above.
Fix: raise llamaCpp.maxTokens, or shorten the prompt.
llama-cpp generation filled the <S>-token context before completing the JSON — raise llamaCpp.contextSize, or set llamaCpp.maxTokens to bound the response.
Section titled “llama-cpp generation filled the <S>-token context before completing the JSON — raise llamaCpp.contextSize, or set llamaCpp.maxTokens to bound the response.”With maxTokens unset, the response may use whatever room the context has left after the prompt,
and no more. The answer ran past that room. Uncapped, llama.cpp would have shifted the prompt out of
the context to keep generating, and the rest of the answer would follow a prompt you did not send.
Fix: raise llamaCpp.contextSize so the context holds a longer answer, or set
llamaCpp.maxTokens to bound the response and size the context around it.
llama-cpp prompt needs <N> tokens of context, more than this model's training context of <T> tokens. Counted: …
Section titled “llama-cpp prompt needs <N> tokens of context, more than this model's training context of <T> tokens. Counted: …”The prompt does not fit the largest context the model supports. <N> is the system prompt with its
restated schema, plus the user prompt, plus 512 tokens of chat-template overhead, plus room for the
response: maxTokens, or 2048 when it is unset, plus thoughtTokens. The message lists each count.
Nothing is generated. Sending the prompt anyway would make llama.cpp shift its start out of the
context, and the model would answer a prompt you did not send. Called directly, completeJSON
throws this as an InferenceError.
Fix: shorten the prompt, lower llamaCpp.maxTokens if you set it, or use a model trained on a
longer context.
llama-cpp prompt needs <N> tokens of context, more than llamaCpp.contextSize (<S>). Counted: …
Section titled “llama-cpp prompt needs <N> tokens of context, more than llamaCpp.contextSize (<S>). Counted: …”The same count, against the fixed size you set with llamaCpp.contextSize.
Fix: raise llamaCpp.contextSize, or set llamaCpp.maxTokens to shrink the response reserve,
or leave contextSize unset so each context is sized to its prompt. A fixed size below about 2600
tokens holds no prompt at all until maxTokens is set, because the default reserve is 2048.
llama.cpp's <B> backend crashed the local-model worker (<exit>). <B> was chosen explicitly (<setting>), so the library did not switch backends. …
Section titled “llama.cpp's <B> backend crashed the local-model worker (<exit>). <B> was chosen explicitly (<setting>), so the library did not switch backends. …”llama.cpp aborted the local-model worker
on backend <B>, for example with exit code 127: …ggml-cuda.cu:106: CUDA error. You pinned that
backend with llamaCpp.gpu or NODE_LLAMA_CPP_GPU, so the call was not moved to another one. Your
process is unaffected. Only the worker died, and the next call starts a fresh one on the same
backend.
Fix: pin a different backend (NODE_LLAMA_CPP_GPU=vulkan, or false for the CPU), or unset it
so the library falls back on its own. For CUDA, the ggml-cuda.cu line in the message is the one to
report upstream.
llama.cpp crashed the local-model worker on every backend this machine offers — <B> (<exit>), …, CPU (<exit>). The request was not answered.
Section titled “llama.cpp crashed the local-model worker on every backend this machine offers — <B> (<exit>), …, CPU (<exit>). The request was not answered.”With gpu on "auto", the worker crashed on each GPU backend in turn and then on the CPU, the last
resort. When every backend fails, the request itself is the likelier cause. Later calls still try
the CPU.
Fix: read the per-backend exits in the message, which hold llama.cpp’s own last line. Try a smaller prompt or a different model, and report the abort to node-llama-cpp.
llama.cpp's local-model worker crashed before it chose a backend (<exit>). The request was not answered.
Section titled “llama.cpp's local-model worker crashed before it chose a backend (<exit>). The request was not answered.”The worker died before it named a backend, so there was nothing to fall back from. That usually means node-llama-cpp itself could not load in the worker process.
Fix: run npx node-llama-cpp inspect gpu to check the binding on this machine, and read the
exit in the message.
The local-model worker was shut down by disposeLlamaModels while a call was still running.
Section titled “The local-model worker was shut down by disposeLlamaModels while a call was still running.”disposeLlamaModels() ran while a call was in flight, and stopped the worker serving it.
Fix: await your calls before disposing.
Could not start the local-model worker (<ERR>).
Section titled “Could not start the local-model worker (<ERR>).”Node could not fork the worker process: the OS refused a new process, or the Node executable could not be run again.
Fix: read <ERR>. This is a machine-level failure, not a model one.
The local-model worker was not initialised. — or The local-model worker has no model <N>., The local-model worker has no session <N>.
Section titled “The local-model worker was not initialised. — or The local-model worker has no model <N>., The local-model worker has no session <N>.”An internal consistency check in the worker failed. These should never appear.
Fix: report it, with the call that produced it.
Failed to run <cmd>: <spawnError> (is the Claude CLI installed?)
Section titled “Failed to run <cmd>: <spawnError> (is the Claude CLI installed?)”The Claude CLI process could not start.
Fix: install the CLI, or set command to its path. ENOENT here means “not on PATH”.
Claude CLI timed out
Section titled “Claude CLI timed out”Exceeded timeoutMs (default 180000).
Claude CLI exited <code>: <last 300 chars of stderr>
Section titled “Claude CLI exited <code>: <last 300 chars of stderr>”Non-zero exit. The stderr tail is included because it usually holds the real reason — most often that you are not logged in.
Fix: run claude yourself once to confirm it works and that you are authenticated.
Claude CLI returned no result field
Section titled “Claude CLI returned no result field”The CLI’s JSON envelope parsed, but had no string result.
Claude CLI printed non-JSON output (is it logged in?): <first 200 chars>
Section titled “Claude CLI printed non-JSON output (is it logged in?): <first 200 chars>”The CLI exited 0 and printed something that is not the --output-format json envelope at all —
an update banner, a login prompt, an auth error, a proxy interception page. The excerpt is what it
actually printed, whitespace collapsed to one line and capped at 200 characters so a page of HTML
cannot become the error message. (no output) means it printed nothing.
Fix: read the excerpt. A login prompt means claude is not authenticated — run it yourself
once. A banner or an HTML page means something is intercepting the command; try claude -p --output-format json by hand and see what comes back.
MockProvider needs at least one scripted response
Section titled “MockProvider needs at least one scripted response”new MockProvider([]). Test-only.
Empty command
Section titled “Empty command”Not an error but a field: realExec([]) returns spawnError: "Empty command". See the
exec reference.
See it for yourself
Section titled “See it for yourself”Every message above is provoked for real by
examples/diagnose-errors.mjs,
which runs in CI. scripts/check-error-coverage.mjs fails the build if a throw in src/ has no
entry on this page — so this reference cannot quietly fall behind the code.