Skip to content

Error reference

Every message the library can produce, verbatim, with its trigger and its fix.

If you arrived with a message in hand, use your browser’s find. If you are not sure which class of failure you have, start at diagnosing a failure.

Anthropic provider needs <VAR> set (or choose another provider)

Section titled “Anthropic provider needs <VAR> set (or choose another provider)”

The Anthropic provider was constructed and process.env[apiKeyEnv] was empty. <VAR> is ANTHROPIC_API_KEY unless you set apiKeyEnv.

Fix: export the key, name a different provider, or omit provider entirely and let detection pick one that works.

OpenAI provider needs <VAR> set (or point baseUrl at a local server)

Section titled “OpenAI provider needs <VAR> set (or point baseUrl at a local server)”

Same, for the OpenAI-compatible provider. A key is required only when baseUrl contains api.openai.com; a keyless local server is allowed.

Fix: export the key, or set baseUrl to your local endpoint.

No inference provider is available. Tried:

Section titled “No inference provider is available. Tried:”

Detection ran every probe and none succeeded. The message continues with one line per provider explaining why each was unavailable, then Pass an explicit provider, set one of the keys above, or install node-llama-cpp.

The reason lines are:

Reason Meaning
ANTHROPIC_API_KEY is not set no Anthropic key in the environment
OPENAI_API_KEY is not set and no baseUrl was given no OpenAI key and no local endpoint
could not run `claude` (is the Claude CLI installed?) the CLI is absent or not on PATH
node-llama-cpp is not installed (npm i node-llama-cpp) the optional peer dependency is missing
node-llama-cpp could not start (…) installed, but the native binding failed to load
node-llama-cpp is installed but failed to load (…) resolved and then threw — ABI, system library, or Node version. Installing cannot fix it, so it is not attempted
node-llama-cpp is not installed and INFERENCE_NO_AUTO_INSTALL is set absent, and the automatic install was refused
not auto-selectable only ever shown for mock, which detection never picks

Fix: whichever line is closest to your situation. This is the most common first failure on a machine with no credentials — see common failures.

No provider specified. Detecting one probes the environment, …

Section titled “No provider specified. Detecting one probes the environment, …”
No provider specified. Detecting one probes the environment, the Claude CLI and the local
model runtime, which cannot be done synchronously — use
resolveProviderIdentityAsync/makeProviderAsync, or name a provider (anthropic, openai,
claude-cli, mock, llama-cpp).

You called the synchronous makeProvider or resolveProviderIdentity with provider omitted or set to "auto".

Fix: await makeProviderAsync(spec), or name a provider.

Model "<M>" was given without a provider, and a model name does not say which provider owns it

Section titled “Model "<M>" was given without a provider, and a model name does not say which provider owns it”
Model "gpt-4o-mini" was given without a provider, and a model name does not say which
provider owns it. Name the provider too (anthropic, openai, claude-cli, mock, llama-cpp),
or drop the model to take the detected provider's default.

You set model but left provider omitted or "auto". A model name belongs to exactly one provider, so it cannot be carried into whichever provider detection happens to pick — before this check, { model: "gpt-4o-mini" } on a machine with an Anthropic key selected anthropic and then 404’d at call time, after you had already paid for detection.

model: null is still “use the provider’s default” — only a real name is ambiguous.

Fix: name the provider alongside the model, or drop model.

node-llama-cpp is installed but failed to load (<ERR>)

Section titled “node-llama-cpp is installed but failed to load (<ERR>)”
node-llama-cpp is installed but failed to load (<ERR>). This is the copy resolved from
your own node_modules, so reinstalling it here will not help — check the Node version and
the platform build.

The package resolved and then failed: an ABI mismatch, a missing system library, or a Node version the native build does not support.

This is deliberately not treated as “not installed”. Only ERR_MODULE_NOT_FOUND triggers the automatic install; anything else would fetch the same broken package again and bury the real cause under a download. Detection reports the same thing as node-llama-cpp is installed but failed to load (…) in its reason lines.

Fix: whatever the inner error says — usually rebuilding for your Node version, or a platform with no prebuilt binary and no toolchain. INFERENCE_NO_AUTO_INSTALL is irrelevant here; nothing was going to be installed.

The llama-cpp provider needs node-llama-cpp, and INFERENCE_NO_AUTO_INSTALL is set

Section titled “The llama-cpp provider needs node-llama-cpp, and INFERENCE_NO_AUTO_INSTALL is set”

The local runtime is absent and you have refused the automatic install.

Fix: npm i node-llama-cpp@^3.19.0 yourself, unset INFERENCE_NO_AUTO_INSTALL, or name a provider that does not need it.

Installing node-llama-cpp into <DIR> failed (exit <N>)

Section titled “Installing node-llama-cpp into <DIR> failed (exit <N>)”

npm ran and returned non-zero. The message carries the last lines of npm’s output, then repeats the manual command. Nothing is left behind: the prefix is marked ready only after a clean exit, so the next call retries rather than importing a half-installed directory.

Fix: whatever npm’s output says — usually network, registry auth, or a missing build toolchain on a platform with no prebuilt binary.

Could not run npm to install node-llama-cpp (<ERR>)

Section titled “Could not run npm to install node-llama-cpp (<ERR>)”

npm itself could not be spawned — common on images that ship node without npm, or where only pnpm or yarn is on PATH.

Fix: install the package with whatever package manager you do have, or name a different provider.

Installing node-llama-cpp into <DIR> timed out

Section titled “Installing node-llama-cpp into <DIR> timed out”

The install exceeded 15 minutes. On a platform without a prebuilt binary this is a source build, and a slow machine can genuinely exceed it.

Fix: retry, raise timeoutMs, or install it yourself.

Timed out waiting for another process to install node-llama-cpp into <DIR>

Section titled “Timed out waiting for another process to install node-llama-cpp into <DIR>”

Another process holds the prefix lock. Concurrent runs serialise so two npm processes never write one node_modules.

Fix: wait for the other run. If nothing else is running, the lock was orphaned by a killed process — remove the named .install.lock and retry.

llama-cpp model "<M>" is a selector and cannot be resolved synchronously — …

Section titled “llama-cpp model "<M>" is a selector and cannot be resolved synchronously — …”

A selector (auto, fast, balanced, quality) reached a synchronous factory. Resolving it reads GPU memory.

Fix: await makeProviderAsync(spec), or name a concrete model. The refusal is deliberate: recording "auto" as cache-key material would let two differently-sized models share results.

llama-cpp model "<M>" is a selector and needs a hardware probe to resolve.

Section titled “llama-cpp model "<M>" is a selector and needs a hardware probe to resolve.”

The same cause reaching a different entry point — resolveLlamaModelRef, and through it blobNameFor, isModelDownloaded, or clearLlamaModels({ models }).

Fix: resolve the selector first, or pass a concrete alias to those helpers.

llama-cpp model "<M>" is a selector. Constructing a provider directly needs a concrete model …

Section titled “llama-cpp model "<M>" is a selector. Constructing a provider directly needs a concrete model …”

new LlamaCppProvider("auto") — the class constructor rather than a factory.

Fix: await makeProviderAsync({ provider: "llama-cpp" }).

llamaCpp.contextSize must be a positive integer number of tokens, got <X>.

Section titled “llamaCpp.contextSize must be a positive integer number of tokens, got <X>.”

llamaCpp.contextSize was zero, negative, fractional, or not a number. Thrown when the provider is constructed.

Fix: pass a whole number of tokens, such as 8192, or leave it unset so each context is sized to its prompt.

llamaCpp.gpu must be "auto", "cuda", "vulkan", "metal" or false, got <X>.

Section titled “llamaCpp.gpu must be "auto", "cuda", "vulkan", "metal" or false, got <X>.”

llamaCpp.gpu named a backend node-llama-cpp does not have, such as "rocm", or was not a string or false. Thrown when the provider is constructed.

Fix: use one of the listed values, or leave it unset so NODE_LLAMA_CPP_GPU decides, and failing that "auto". The CPU is false, not "cpu".

Unknown provider "<P>". Available: anthropic, openai, claude-cli, mock, llama-cpp.

Section titled “Unknown provider "<P>". Available: anthropic, openai, claude-cli, mock, llama-cpp.”

A provider value outside the five. Usually a typo, or free text arriving from a CLI flag.

Fix: use one of the listed names. If the value came from user input, validate it before building the spec — every existing consumer does exactly this.

Unknown llama-cpp model "<M>". Use a selector …, a curated alias …, an hf: URI, or a path to a .gguf file.

Section titled “Unknown llama-cpp model "<M>". Use a selector …, a curated alias …, an hf: URI, or a path to a .gguf file.”

The model reference matched no accepted form.

The llama-cpp provider needs the optional peer dependency node-llama-cpp. Install it with: npm i node-llama-cpp

Section titled “The llama-cpp provider needs the optional peer dependency node-llama-cpp. Install it with: npm i node-llama-cpp”

The dynamic import failed. The original error is appended in parentheses.

Fix: npm install node-llama-cpp. It is an optional peer dependency and is never installed for you.

None of these throw. Each arrives with result (or verdict) absent.

Response failed schema validation: <path> <message>; …

Section titled “Response failed schema validation: <path> <message>; …”

Ajv rejected the response on the final attempt. One clause per validation error, joined with ; .

Fix: usually the schema, not the model. Check the shape is expressible — and if you are on a local model, that you are not relying on required or numeric bounds, which a grammar does not enforce.

Response contained no parseable JSON object

Section titled “Response contained no parseable JSON object”

extractJson found nothing usable — no bare object, no fenced block, no object embedded in prose. Shared by openai, claude-cli, and llama-cpp.

Fix: most often a model that answered in prose. Tighten the system prompt, or use a provider with native structured output.

The OpenAI-compatible endpoint returned a non-2xx. The provider’s own error.message is passed through verbatim; when the body has none you get HTTP 401, HTTP 429, HTTP 500.

The endpoint returned 2xx but choices[0].message.content was missing or empty.

Anthropic response hit max_tokens (<N>) before completing the tool call — raise the anthropic.maxTokens option.

Section titled “Anthropic response hit max_tokens (<N>) before completing the tool call — raise the anthropic.maxTokens option.”

The forced tool call was truncated. Guarded early and reported as itself rather than being allowed to fail validation, which would have burned the retry on a misleading error.

Fix: raise anthropic.maxTokens (default 1024).

Anthropic response contained no tool_use block

Section titled “Anthropic response contained no tool_use block”

The response came back without the forced tool call. Rare; usually a refusal.

llama-cpp generation hit the token limit before completing the JSON (maxTokens: <N>) — raise llamaCpp.maxTokens, or shorten the prompt if the context is full.

Section titled “llama-cpp generation hit the token limit before completing the JSON (maxTokens: <N>) — raise llamaCpp.maxTokens, or shorten the prompt if the context is full.”

The local-model equivalent of the above.

Fix: raise llamaCpp.maxTokens, or shorten the prompt.

llama-cpp generation filled the <S>-token context before completing the JSON — raise llamaCpp.contextSize, or set llamaCpp.maxTokens to bound the response.

Section titled “llama-cpp generation filled the <S>-token context before completing the JSON — raise llamaCpp.contextSize, or set llamaCpp.maxTokens to bound the response.”

With maxTokens unset, the response may use whatever room the context has left after the prompt, and no more. The answer ran past that room. Uncapped, llama.cpp would have shifted the prompt out of the context to keep generating, and the rest of the answer would follow a prompt you did not send.

Fix: raise llamaCpp.contextSize so the context holds a longer answer, or set llamaCpp.maxTokens to bound the response and size the context around it.

llama-cpp prompt needs <N> tokens of context, more than this model's training context of <T> tokens. Counted: …

Section titled “llama-cpp prompt needs <N> tokens of context, more than this model's training context of <T> tokens. Counted: …”

The prompt does not fit the largest context the model supports. <N> is the system prompt with its restated schema, plus the user prompt, plus 512 tokens of chat-template overhead, plus room for the response: maxTokens, or 2048 when it is unset, plus thoughtTokens. The message lists each count.

Nothing is generated. Sending the prompt anyway would make llama.cpp shift its start out of the context, and the model would answer a prompt you did not send. Called directly, completeJSON throws this as an InferenceError.

Fix: shorten the prompt, lower llamaCpp.maxTokens if you set it, or use a model trained on a longer context.

llama-cpp prompt needs <N> tokens of context, more than llamaCpp.contextSize (<S>). Counted: …

Section titled “llama-cpp prompt needs <N> tokens of context, more than llamaCpp.contextSize (<S>). Counted: …”

The same count, against the fixed size you set with llamaCpp.contextSize.

Fix: raise llamaCpp.contextSize, or set llamaCpp.maxTokens to shrink the response reserve, or leave contextSize unset so each context is sized to its prompt. A fixed size below about 2600 tokens holds no prompt at all until maxTokens is set, because the default reserve is 2048.

llama.cpp's <B> backend crashed the local-model worker (<exit>). <B> was chosen explicitly (<setting>), so the library did not switch backends. …

Section titled “llama.cpp's <B> backend crashed the local-model worker (<exit>). <B> was chosen explicitly (<setting>), so the library did not switch backends. …”

llama.cpp aborted the local-model worker on backend <B>, for example with exit code 127: …ggml-cuda.cu:106: CUDA error. You pinned that backend with llamaCpp.gpu or NODE_LLAMA_CPP_GPU, so the call was not moved to another one. Your process is unaffected. Only the worker died, and the next call starts a fresh one on the same backend.

Fix: pin a different backend (NODE_LLAMA_CPP_GPU=vulkan, or false for the CPU), or unset it so the library falls back on its own. For CUDA, the ggml-cuda.cu line in the message is the one to report upstream.

llama.cpp crashed the local-model worker on every backend this machine offers — <B> (<exit>), …, CPU (<exit>). The request was not answered.

Section titled “llama.cpp crashed the local-model worker on every backend this machine offers — <B> (<exit>), …, CPU (<exit>). The request was not answered.”

With gpu on "auto", the worker crashed on each GPU backend in turn and then on the CPU, the last resort. When every backend fails, the request itself is the likelier cause. Later calls still try the CPU.

Fix: read the per-backend exits in the message, which hold llama.cpp’s own last line. Try a smaller prompt or a different model, and report the abort to node-llama-cpp.

llama.cpp's local-model worker crashed before it chose a backend (<exit>). The request was not answered.

Section titled “llama.cpp's local-model worker crashed before it chose a backend (<exit>). The request was not answered.”

The worker died before it named a backend, so there was nothing to fall back from. That usually means node-llama-cpp itself could not load in the worker process.

Fix: run npx node-llama-cpp inspect gpu to check the binding on this machine, and read the exit in the message.

The local-model worker was shut down by disposeLlamaModels while a call was still running.

Section titled “The local-model worker was shut down by disposeLlamaModels while a call was still running.”

disposeLlamaModels() ran while a call was in flight, and stopped the worker serving it.

Fix: await your calls before disposing.

Could not start the local-model worker (<ERR>).

Section titled “Could not start the local-model worker (<ERR>).”

Node could not fork the worker process: the OS refused a new process, or the Node executable could not be run again.

Fix: read <ERR>. This is a machine-level failure, not a model one.

The local-model worker was not initialised. — or The local-model worker has no model <N>., The local-model worker has no session <N>.

Section titled “The local-model worker was not initialised. — or The local-model worker has no model <N>., The local-model worker has no session <N>.”

An internal consistency check in the worker failed. These should never appear.

Fix: report it, with the call that produced it.

Failed to run <cmd>: <spawnError> (is the Claude CLI installed?)

Section titled “Failed to run <cmd>: <spawnError> (is the Claude CLI installed?)”

The Claude CLI process could not start.

Fix: install the CLI, or set command to its path. ENOENT here means “not on PATH”.

Exceeded timeoutMs (default 180000).

Claude CLI exited <code>: <last 300 chars of stderr>

Section titled “Claude CLI exited <code>: <last 300 chars of stderr>”

Non-zero exit. The stderr tail is included because it usually holds the real reason — most often that you are not logged in.

Fix: run claude yourself once to confirm it works and that you are authenticated.

The CLI’s JSON envelope parsed, but had no string result.

Claude CLI printed non-JSON output (is it logged in?): <first 200 chars>

Section titled “Claude CLI printed non-JSON output (is it logged in?): <first 200 chars>”

The CLI exited 0 and printed something that is not the --output-format json envelope at all — an update banner, a login prompt, an auth error, a proxy interception page. The excerpt is what it actually printed, whitespace collapsed to one line and capped at 200 characters so a page of HTML cannot become the error message. (no output) means it printed nothing.

Fix: read the excerpt. A login prompt means claude is not authenticated — run it yourself once. A banner or an HTML page means something is intercepting the command; try claude -p --output-format json by hand and see what comes back.

MockProvider needs at least one scripted response

Section titled “MockProvider needs at least one scripted response”

new MockProvider([]). Test-only.

Not an error but a field: realExec([]) returns spawnError: "Empty command". See the exec reference.

Every message above is provoked for real by examples/diagnose-errors.mjs, which runs in CI. scripts/check-error-coverage.mjs fails the build if a throw in src/ has no entry on this page — so this reference cannot quietly fall behind the code.