Warnings reference
The library writes to console.warn in exactly eight places. Nothing else in the docset says so,
which matters if your tool owns its own output formatting.
Each warning fires at most once, but the scope of “once” differs per warning, and that catches people out in tests.
The eight
Section titled “The eight”| Warning | Fires once per | Reset with |
|---|---|---|
| Unsupported Node version | process | resetNodeVersionWarning() |
| Nonzero judge temperature | module | resetTemperatureWarning() |
| Cache write failure | cache instance | construct a new JsonCache |
| Provider auto-selected | process | resetProviderDetectionWarning() |
| Local weights not downloaded | process | resetProviderDetectionWarning() |
| Local runtime being installed | process | resetRuntimeInstall() |
| Local backend crashed, switching | backend switch | none — a crashed backend stays skipped for the process |
| Local-model worker missing | process | none |
inference: running on Node <version>, older than the Node 24 this package requires. npm only warns about that at install time (EBADENGINE), so nothing has stopped you yet — upgrade Node, or expect failures this library cannot explain.
Section titled “inference: running on Node <version>, older than the Node 24 this package requires. npm only warns about that at install time (EBADENGINE), so nothing has stopped you yet — upgrade Node, or expect failures this library cannot explain.”Fires on first use — whichever of makeProvider or completeValidatedJSON you reach first — when
process.versions.node reports a major version below the >=24 in engines.
It is a warning and not a throw because Node 24 is this package’s support floor, not a technical
one: nothing in src/ uses a Node-24-only API, and the strictest dependency floor is
node-llama-cpp at >=20. On Node 22 the library will very likely work — but it is untested
there, and a failure that surfaces later would never mention your Node version. Refusing to run
would break setups that currently do.
<label>: judge temperature is <N> — nonzero temperature adds noise to verdicts; 0 is strongly recommended.
Section titled “<label>: judge temperature is <N> — nonzero temperature adds noise to verdicts; 0 is strongly recommended.”Fires when runEnsemble or judge gets temperature > 0. <label> is EnsembleOptions.label,
default inference.
Independent samples at temperature 0 are already independent — the provider is called separately each time with no shared state. If you want diversity, more runs beats more randomness.
<label>: could not write the cache at <dir> (<message>). Continuing without caching.
Section titled “<label>: could not write the cache at <dir> (<message>). Continuing without caching.”Fires when a cache write fails — a read-only workspace, a full disk, a permissions problem. The run continues, because the cache is an optimization, never a dependency.
inference: <reason> — auto-selected "<provider>". Pass an explicit provider to pin it.
Section titled “inference: <reason> — auto-selected "<provider>". Pass an explicit provider to pin it.”Fires when detection picks a provider for you. <reason> is either provider "auto" or
no provider specified.
It exists because an environment variable appearing or disappearing between runs otherwise silently changes what an eval measured. Pin the provider and the warning goes away.
inference: "<model>" is not downloaded yet — the first run will fetch ~<N> GB. Pre-fetch it, or pass an explicit provider to avoid the local model entirely.
Section titled “inference: "<model>" is not downloaded yet — the first run will fetch ~<N> GB. Pre-fetch it, or pass an explicit provider to avoid the local model entirely.”Fires when detection lands on llama-cpp and the weights are absent — before the download starts,
so you can still stop it.
inference: node-llama-cpp is not installed — fetching it into <DIR> so the local model can run. …
Section titled “inference: node-llama-cpp is not installed — fetching it into <DIR> so the local model can run. …”Fires once, before the library installs node-llama-cpp into its own runtime directory, because
the binding was not found in your node_modules. The install is a native module and can take a
while. Set INFERENCE_NO_AUTO_INSTALL=1 to refuse it.
inference: llama.cpp's <B> backend crashed the local-model worker (<exit>). Retrying on <next>, and staying on <next> for the rest of this process. …
Section titled “inference: llama.cpp's <B> backend crashed the local-model worker (<exit>). Retrying on <next>, and staying on <next> for the rest of this process. …”Fires when llama.cpp aborts the local-model worker
on backend <B> while gpu is "auto". The call that crashed is retried on <next> and so is
every later call, so you hear this once per switch. On a Windows machine with an NVIDIA GPU it most
often reads CUDA backend crashed … ggml-cuda.cu:106: CUDA error … Retrying on Vulkan.
When the next backend is the CPU, the message instead says Retrying on the CPU, which is much slower — expect minutes per call. Pin a backend with NODE_LLAMA_CPP_GPU or llamaCpp.gpu to
start on the one that works, or to fail fast instead of waiting.
inference: the local-model worker (llama-worker.js) is missing beside this library — was it bundled? …
Section titled “inference: the local-model worker (llama-worker.js) is missing beside this library — was it bundled? …”Fires once when the package’s llama-worker.js is not next to its index.js, which happens when a
bundler has inlined this library into a file of its own. Local models then run in your process, as
they did before the worker existed, and a native crash in llama.cpp will end it.
Fix: mark @hawkeyexl/inference as external in your bundler.
Capturing or silencing them
Section titled “Capturing or silencing them”There is no log-level option. If your tool formats its own output, wrap console.warn for the
duration of the call:
const warnings = [];const original = console.warn;console.warn = (...args) => warnings.push(args.join(" "));try { await judge({ provider, system, user, runs: 3 });} finally { console.warn = original;}Everything the library emits is prefixed — with inference: for detection and the Node version
notice, and with your label for the temperature and cache warnings — so filtering by prefix is
reliable.
Resetting them in tests
Section titled “Resetting them in tests”import { resetNodeVersionWarning, resetProviderDetectionWarning, resetTemperatureWarning,} from "@hawkeyexl/inference";
beforeEach(() => { resetNodeVersionWarning(); resetTemperatureWarning(); resetProviderDetectionWarning(); // resets both detection warnings});Without this, a test that runs second sees no warning and a “warns once” assertion passes for the
wrong reason. resetClaudeCliProbe() is a third seam — it forgets the memoised claude --version
probe rather than a warning.
Source of truth
Section titled “Source of truth”src/runtime.ts, src/judge/ensemble.ts, src/cache.ts, src/providers/detect.ts,
src/providers/llama-install.ts, and src/providers/llama-host.ts.