Skip to content

Run models locally

Remove the credential dependency without changing what your tool does.

The llama-cpp provider runs GGUF weights locally via node-llama-cpp. No daemon, no API key, no per-token cost. The model runs in a worker process the library starts and stops for you, so a crash in llama.cpp’s native code costs a retry rather than your process. See the worker process and the backend.

Everything above the provider contract. LlamaCppProvider satisfies the same InferenceProvider interface, so your completion code, your judge code, your caching, and your schemas are all untouched.

  1. Install the optional peer dependency.

    Terminal window
    npm install node-llama-cpp

    It is not installed with this package unless you ask for it. The ^3.19.0 floor is load-bearing — Gemma 4 support landed there, and earlier versions cannot load anything in the catalog.

  2. Use the async factory.

    const provider = await makeProviderAsync({ provider: "llama-cpp" });

    Not makeProvider. See below.

Two things, and neither announces itself:

  • Your budget gate goes inert. Covered below.
  • Four schema behaviors change under grammar-constrained decoding. Covered in choosing a model.

The default model is the selector auto, and resolving it reads GPU memory — which needs an await. So the synchronous forms throw rather than guess.

examples/local-provider.mjs
// Running the llama-cpp provider, and why a selector needs the async factory.
//
// Runs with no API key, no GPU, and no GGUF weights: LlamaRuntime is the injection
// seam the library uses in its own tests. In real use you omit `llamaRuntime`
// entirely and the default runtime downloads and loads real weights.
import {
InferenceError,
judge,
makeProvider,
makeProviderAsync,
resolveProviderIdentityAsync,
} from "@hawkeyexl/inference";
// A stand-in for node-llama-cpp. Reports a 16 GB budget and answers with a fixed verdict.
const fakeRuntime = {
async getMemoryBudgetBytes() {
return 16 * 1024 ** 3;
},
async resolveModelFile(uri, directory) {
return `${directory}/${uri.split("/").pop()}`;
},
async loadModel() {
return {
async createSession() {
return {
async prompt() {
return {
text: JSON.stringify({
claim: "The page documents authentication.",
observed: "The page describes bearer tokens and refresh.",
match: "pass",
confidence: 0.91,
reasoning: "Both the token type and the refresh interval are stated.",
}),
usage: { inputTokens: 420, outputTokens: 78 },
stopReason: "stop",
};
},
async dispose() {},
};
},
async dispose() {},
};
},
};
// The synchronous factory refuses an unresolved selector rather than recording the
// literal "auto" as cache-key material — which would let a 2.6 GB and a 6.7 GB model
// share cached results, and make one key mean different things on two machines.
try {
makeProvider({ provider: "llama-cpp", model: "auto", llamaRuntime: fakeRuntime });
} catch (error) {
console.log("sync factory refused:", error instanceof InferenceError);
console.log("message:", error.message);
}
// The async forms resolve the selector by measuring the machine, and return the
// concrete model. They delegate to the sync forms for every other provider.
const identity = await resolveProviderIdentityAsync({
provider: "llama-cpp",
model: "auto",
llamaRuntime: fakeRuntime,
});
console.log("auto resolved to:", identity);
const provider = await makeProviderAsync({
provider: "llama-cpp",
model: "auto",
llamaRuntime: fakeRuntime,
});
// The provider satisfies the same contract, so judge and completion code is unchanged.
const consensus = await judge({
provider,
system: "You evaluate whether a page satisfies an assertion.",
user: "# Assertion\nThe page documents authentication.\n\n# Page\nUse a bearer token.",
runs: 3,
});
console.log("model:", provider.modelName());
console.log("verdict:", consensus.verdict, "zone:", consensus.zone);
console.log("usage reported:", consensus.runs[0].usage);
sync factory refused: true
message: llama-cpp model "auto" is a selector and cannot be resolved synchronously — picking a tier probes GPU memory. Use resolveProviderIdentityAsync/makeProviderAsync, or name a concrete model (e.g. "qwen3.5-4b").
auto resolved to: { provider: 'llama-cpp', model: 'qwen3.5-4b' }
model: qwen3.5-4b
verdict: pass zone: auto-pass
usage reported: { inputTokens: 420, outputTokens: 78 }

makeProviderAsync and resolveProviderIdentityAsync return the model the selector actually resolved to. Both delegate to the synchronous forms for every other provider, so you can switch your whole codebase over rather than branching on provider name. A concrete model — an alias, a URI, or a path — works with either form.

Nothing is broken — the run genuinely costs nothing per token. But your gate is inert, not satisfied, and that matters beyond the local run: if you later switch one stage back to a hosted provider whose model is not in the price table, the gate is still inert.

Guard on the price being known, not on the number:

const pricing = pricingFor(provider.modelName());
if (pricing !== undefined && spent >= maxCostUsd) break;

See cost and budgets.

LlamaRuntime is the injection seam — the same one this library’s own tests use. Pass a fake and the entire local path runs with no download, no GPU, and no weights:

await makeProviderAsync({ provider: "llama-cpp", llamaRuntime: myFakeRuntime });

Omit llamaRuntime in real use and the default runtime downloads and loads real weights. The sample above uses a fake, which is why it runs in CI. See testing your integration.

Weights load once per process and are shared across every provider naming the same model. In a long-running process, disposeLlamaModels() frees them.