Run models locally
Remove the credential dependency without changing what your tool does.
The llama-cpp provider runs GGUF weights locally via
node-llama-cpp. No daemon, no API key, no per-token cost.
The model runs in a worker process the library starts and stops for you, so a crash in llama.cpp’s
native code costs a retry rather than your process. See
the worker process and the backend.
What stays the same
Section titled “What stays the same”Everything above the provider contract. LlamaCppProvider satisfies the same
InferenceProvider interface, so your completion code, your judge code, your caching, and your
schemas are all untouched.
What you must do differently
Section titled “What you must do differently”-
Install the optional peer dependency.
Terminal window npm install node-llama-cppIt is not installed with this package unless you ask for it. The
^3.19.0floor is load-bearing — Gemma 4 support landed there, and earlier versions cannot load anything in the catalog. -
Use the async factory.
const provider = await makeProviderAsync({ provider: "llama-cpp" });Not
makeProvider. See below.
What silently behaves differently
Section titled “What silently behaves differently”Two things, and neither announces itself:
- Your budget gate goes inert. Covered below.
- Four schema behaviors change under grammar-constrained decoding. Covered in choosing a model.
Why selectors need the async factory
Section titled “Why selectors need the async factory”The default model is the selector auto, and resolving it reads GPU memory — which needs an
await. So the synchronous forms throw rather than guess.
// Running the llama-cpp provider, and why a selector needs the async factory.//// Runs with no API key, no GPU, and no GGUF weights: LlamaRuntime is the injection// seam the library uses in its own tests. In real use you omit `llamaRuntime`// entirely and the default runtime downloads and loads real weights.
import { InferenceError, judge, makeProvider, makeProviderAsync, resolveProviderIdentityAsync,} from "@hawkeyexl/inference";
// A stand-in for node-llama-cpp. Reports a 16 GB budget and answers with a fixed verdict.const fakeRuntime = { async getMemoryBudgetBytes() { return 16 * 1024 ** 3; }, async resolveModelFile(uri, directory) { return `${directory}/${uri.split("/").pop()}`; }, async loadModel() { return { async createSession() { return { async prompt() { return { text: JSON.stringify({ claim: "The page documents authentication.", observed: "The page describes bearer tokens and refresh.", match: "pass", confidence: 0.91, reasoning: "Both the token type and the refresh interval are stated.", }), usage: { inputTokens: 420, outputTokens: 78 }, stopReason: "stop", }; }, async dispose() {}, }; }, async dispose() {}, }; },};
// The synchronous factory refuses an unresolved selector rather than recording the// literal "auto" as cache-key material — which would let a 2.6 GB and a 6.7 GB model// share cached results, and make one key mean different things on two machines.try { makeProvider({ provider: "llama-cpp", model: "auto", llamaRuntime: fakeRuntime });} catch (error) { console.log("sync factory refused:", error instanceof InferenceError); console.log("message:", error.message);}
// The async forms resolve the selector by measuring the machine, and return the// concrete model. They delegate to the sync forms for every other provider.const identity = await resolveProviderIdentityAsync({ provider: "llama-cpp", model: "auto", llamaRuntime: fakeRuntime,});console.log("auto resolved to:", identity);
const provider = await makeProviderAsync({ provider: "llama-cpp", model: "auto", llamaRuntime: fakeRuntime,});
// The provider satisfies the same contract, so judge and completion code is unchanged.const consensus = await judge({ provider, system: "You evaluate whether a page satisfies an assertion.", user: "# Assertion\nThe page documents authentication.\n\n# Page\nUse a bearer token.", runs: 3,});
console.log("model:", provider.modelName());console.log("verdict:", consensus.verdict, "zone:", consensus.zone);console.log("usage reported:", consensus.runs[0].usage);sync factory refused: truemessage: llama-cpp model "auto" is a selector and cannot be resolved synchronously — picking a tier probes GPU memory. Use resolveProviderIdentityAsync/makeProviderAsync, or name a concrete model (e.g. "qwen3.5-4b").auto resolved to: { provider: 'llama-cpp', model: 'qwen3.5-4b' }model: qwen3.5-4bverdict: pass zone: auto-passusage reported: { inputTokens: 420, outputTokens: 78 }makeProviderAsync and resolveProviderIdentityAsync return the model the selector actually
resolved to. Both delegate to the synchronous forms for every other provider, so you can switch
your whole codebase over rather than branching on provider name. A concrete model — an alias, a URI,
or a path — works with either form.
The budget gate goes inert
Section titled “The budget gate goes inert”Nothing is broken — the run genuinely costs nothing per token. But your gate is inert, not satisfied, and that matters beyond the local run: if you later switch one stage back to a hosted provider whose model is not in the price table, the gate is still inert.
Guard on the price being known, not on the number:
const pricing = pricingFor(provider.modelName());if (pricing !== undefined && spent >= maxCostUsd) break;See cost and budgets.
Testing without weights
Section titled “Testing without weights”LlamaRuntime is the injection seam — the same one this library’s own tests use. Pass a fake and
the entire local path runs with no download, no GPU, and no weights:
await makeProviderAsync({ provider: "llama-cpp", llamaRuntime: myFakeRuntime });Omit llamaRuntime in real use and the default runtime downloads and loads real weights. The sample
above uses a fake, which is why it runs in CI. See
testing your integration.
Long-lived processes
Section titled “Long-lived processes”Weights load once per process and are shared across every provider naming the same model. In a
long-running process, disposeLlamaModels() frees them.
- Choose a model — what
autopicks, and what changes under a grammar - Manage model files — where weights live and how to reclaim the disk