Skip to content

Warnings reference

The library writes to console.warn in exactly eight places. Nothing else in the docset says so, which matters if your tool owns its own output formatting.

Each warning fires at most once, but the scope of “once” differs per warning, and that catches people out in tests.

Warning Fires once per Reset with
Unsupported Node version process resetNodeVersionWarning()
Nonzero judge temperature module resetTemperatureWarning()
Cache write failure cache instance construct a new JsonCache
Provider auto-selected process resetProviderDetectionWarning()
Local weights not downloaded process resetProviderDetectionWarning()
Local runtime being installed process resetRuntimeInstall()
Local backend crashed, switching backend switch none — a crashed backend stays skipped for the process
Local-model worker missing process none

inference: running on Node <version>, older than the Node 24 this package requires. npm only warns about that at install time (EBADENGINE), so nothing has stopped you yet — upgrade Node, or expect failures this library cannot explain.

Section titled “inference: running on Node <version>, older than the Node 24 this package requires. npm only warns about that at install time (EBADENGINE), so nothing has stopped you yet — upgrade Node, or expect failures this library cannot explain.”

Fires on first use — whichever of makeProvider or completeValidatedJSON you reach first — when process.versions.node reports a major version below the >=24 in engines.

It is a warning and not a throw because Node 24 is this package’s support floor, not a technical one: nothing in src/ uses a Node-24-only API, and the strictest dependency floor is node-llama-cpp at >=20. On Node 22 the library will very likely work — but it is untested there, and a failure that surfaces later would never mention your Node version. Refusing to run would break setups that currently do.

Section titled “<label>: judge temperature is <N> — nonzero temperature adds noise to verdicts; 0 is strongly recommended.”

Fires when runEnsemble or judge gets temperature > 0. <label> is EnsembleOptions.label, default inference.

Independent samples at temperature 0 are already independent — the provider is called separately each time with no shared state. If you want diversity, more runs beats more randomness.

<label>: could not write the cache at <dir> (<message>). Continuing without caching.

Section titled “<label>: could not write the cache at <dir> (<message>). Continuing without caching.”

Fires when a cache write fails — a read-only workspace, a full disk, a permissions problem. The run continues, because the cache is an optimization, never a dependency.

inference: <reason> — auto-selected "<provider>". Pass an explicit provider to pin it.

Section titled “inference: <reason> — auto-selected "<provider>". Pass an explicit provider to pin it.”

Fires when detection picks a provider for you. <reason> is either provider "auto" or no provider specified.

It exists because an environment variable appearing or disappearing between runs otherwise silently changes what an eval measured. Pin the provider and the warning goes away.

inference: "<model>" is not downloaded yet — the first run will fetch ~<N> GB. Pre-fetch it, or pass an explicit provider to avoid the local model entirely.

Section titled “inference: "<model>" is not downloaded yet — the first run will fetch ~<N> GB. Pre-fetch it, or pass an explicit provider to avoid the local model entirely.”

Fires when detection lands on llama-cpp and the weights are absent — before the download starts, so you can still stop it.

inference: node-llama-cpp is not installed — fetching it into <DIR> so the local model can run. …

Section titled “inference: node-llama-cpp is not installed — fetching it into <DIR> so the local model can run. …”

Fires once, before the library installs node-llama-cpp into its own runtime directory, because the binding was not found in your node_modules. The install is a native module and can take a while. Set INFERENCE_NO_AUTO_INSTALL=1 to refuse it.

inference: llama.cpp's <B> backend crashed the local-model worker (<exit>). Retrying on <next>, and staying on <next> for the rest of this process. …

Section titled “inference: llama.cpp's <B> backend crashed the local-model worker (<exit>). Retrying on <next>, and staying on <next> for the rest of this process. …”

Fires when llama.cpp aborts the local-model worker on backend <B> while gpu is "auto". The call that crashed is retried on <next> and so is every later call, so you hear this once per switch. On a Windows machine with an NVIDIA GPU it most often reads CUDA backend crashed … ggml-cuda.cu:106: CUDA error … Retrying on Vulkan.

When the next backend is the CPU, the message instead says Retrying on the CPU, which is much slower — expect minutes per call. Pin a backend with NODE_LLAMA_CPP_GPU or llamaCpp.gpu to start on the one that works, or to fail fast instead of waiting.

inference: the local-model worker (llama-worker.js) is missing beside this library — was it bundled? …

Section titled “inference: the local-model worker (llama-worker.js) is missing beside this library — was it bundled? …”

Fires once when the package’s llama-worker.js is not next to its index.js, which happens when a bundler has inlined this library into a file of its own. Local models then run in your process, as they did before the worker existed, and a native crash in llama.cpp will end it.

Fix: mark @hawkeyexl/inference as external in your bundler.

There is no log-level option. If your tool formats its own output, wrap console.warn for the duration of the call:

const warnings = [];
const original = console.warn;
console.warn = (...args) => warnings.push(args.join(" "));
try {
await judge({ provider, system, user, runs: 3 });
} finally {
console.warn = original;
}

Everything the library emits is prefixed — with inference: for detection and the Node version notice, and with your label for the temperature and cache warnings — so filtering by prefix is reliable.

import {
resetNodeVersionWarning,
resetProviderDetectionWarning,
resetTemperatureWarning,
} from "@hawkeyexl/inference";
beforeEach(() => {
resetNodeVersionWarning();
resetTemperatureWarning();
resetProviderDetectionWarning(); // resets both detection warnings
});

Without this, a test that runs second sees no warning and a “warns once” assertion passes for the wrong reason. resetClaudeCliProbe() is a third seam — it forgets the memoised claude --version probe rather than a warning.

src/runtime.ts, src/judge/ensemble.ts, src/cache.ts, src/providers/detect.ts, src/providers/llama-install.ts, and src/providers/llama-host.ts.