Skip to content

Local models reference

Everything exported for the llama-cpp provider.

For using it, see run models locally.

const LLAMA_MODELS: Readonly<Record<string, LlamaModelEntry>>;
interface LlamaModelEntry {
uri: string; // an exact blob path, never a :QUANT tag
sizeBytes: number;
license: string;
tier?: LlamaTier; // absent for entries outside the tier ladder
notes: string;
}
Alias Size Tier License
granite-4.1-3b-q2 1.41 GB fast Apache-2.0
qwen3.5-4b 2.91 GB balanced Apache-2.0
qwen3.5-9b 5.97 GB quality Apache-2.0
gemma-4-e2b 2.62 GB — Apache-2.0
gemma-4-e4b 4.22 GB — Apache-2.0
gemma-4-12b 6.72 GB — Apache-2.0
gemma-4-26b-a4b 14.25 GB — Apache-2.0
gemma-4-e2b-q2 2.19 GB — Apache-2.0

All entries are unsloth GGUF builds, ungated. Each is deep-frozen.

The untiered Gemma entries backed the three tiers until ADR 01009 retiered the catalog on measured results. They still resolve by name, so an existing pin keeps working; nothing selects them automatically.

gemma-4-e2b-q2 is kept only so existing pins resolve. Do not choose it. It does not reliably terminate: on a 12-page benchmark it left 6 pages unfinished at 120 seconds, with one still running at 400 seconds.

The catalog is a vetted default set, not a whitelist gate — any Hugging Face GGUF reference or local .gguf path also works.

const LLAMA_TIERS: readonly ["fast", "balanced", "quality"];
const LLAMA_SELECTORS: readonly ["auto", "fast", "balanced", "quality"];
type LlamaTier = (typeof LLAMA_TIERS)[number];
type LlamaSelector = (typeof LLAMA_SELECTORS)[number];
function isLlamaSelector(model: string): model is LlamaSelector;
function tierForBudget(budgetBytes: number): LlamaTier;
function aliasForTier(tier: LlamaTier): string;
function uriForTier(tier: LlamaTier): string;
function resolveLlamaModelRef(model: string): string;

tierForBudget requires sizeBytes * 3.5 <= budget, walks tiers smallest to largest, and floors at fast.

resolveLlamaModelRef expands an alias to its pinned URI and passes through hf:, huggingface:, hf.co/, http(s)://, and anything ending .gguf. It throws for a selector, and rejects a bare user/repo as a mistyped alias rather than guessing.

function defaultLlamaModelsDirectory(): string;
function blobNameFor(model: string): string;
function isModelDownloaded(model: string, directory: string): boolean;
function clearLlamaModels(options?: ClearLlamaModelsOptions): Promise<ClearLlamaModelsResult>;
function disposeLlamaModels(): Promise<void>;
interface ClearLlamaModelsOptions {
directory?: string;
models?: string[]; // aliases or URIs; omit to clear everything
dryRun?: boolean;
}
interface ClearLlamaModelsResult {
files: { path: string; sizeBytes: number }[];
freedBytes: number;
directory: string;
dryRun: boolean;
}

defaultLlamaModelsDirectory() is INFERENCE_MODELS_DIR or ~/.hawkeyexl-inference/models — this library’s own directory, deliberately not node-llama-cpp’s shared ~/.node-llama-cpp/models.

Selective clearing matches by suffix, because downloads are prefixed hf_<user>_, and removes every part of a split model. An unknown name is rejected rather than silently matching nothing.

disposeLlamaModels() frees the process-wide weight cache and stops the local-model workers. Weights load once per process and are shared across providers naming the same model.

interface LlamaRuntime {
resolveModelFile(uri: string, directory: string): Promise<string>;
loadModel(path: string): Promise<LlamaLoadedModel>;
getMemoryBudgetBytes(): Promise<number>;
}
interface LlamaLoadedModel {
createSession(systemPrompt: string, contextSize?: number): Promise<LlamaSession>;
dispose(): Promise<void>;
readonly trainContextSize?: number; // the model's training context, in tokens
countTokens?(text: string): number | Promise<number>; // the model's own tokenizer
}
interface LlamaSession {
prompt(text: string, options: LlamaPromptOptions): Promise<LlamaPromptResult>;
dispose(): Promise<void>;
readonly contextSize?: number; // tokens of context actually created
}
interface LlamaPromptOptions {
schema: Record<string, unknown>;
temperature: number;
thoughtTokens: number;
maxTokens?: number;
}
interface LlamaPromptResult {
text: string;
usage?: TokenUsage;
stopReason?: string;
}
function defaultLlamaRuntime(options?: { gpu?: LlamaGpu }): LlamaRuntime;
type LlamaGpu = "auto" | "cuda" | "vulkan" | "metal" | false;

defaultLlamaRuntime() is lazy: constructing it starts no process and loads no native code. Its first use finds node-llama-cpp (a missing module raises an InferenceError naming npm i node-llama-cpp) and starts a worker process to run it. The built dist/index.js contains only a dynamic import, so consumers without the package still typecheck.

Inject a fake through ProviderSpec.llamaRuntime to test the whole local path with no weights and no GPU. See testing your integration.

getMemoryBudgetBytes returns max(free VRAM, totalmem() / 2), falling back to totalmem() / 2 when the GPU probe fails.

The provider passes createSession the context size it computed for the prompt. The real runtime creates exactly that many tokens of context, rounded up to a multiple of 256 by llama.cpp, and 8192 when no size is passed. trainContextSize and countTokens are optional, so a fake written without them still satisfies the interface. Without countTokens, the provider uses the 8192-token default and does not check that the prompt fits. countTokens may return a promise; the real runtime’s does, because its tokenizer lives in the worker.

class LlamaCppProvider implements InferenceProvider {
constructor(model: string, options?: LlamaCppProviderOptions);
}
interface LlamaCppProviderOptions {
runtime?: LlamaRuntime;
thoughtTokens?: number; // default 0
maxTokens?: number;
contextSize?: number; // default: sized to each prompt, 8192 tokens at least
modelsDirectory?: string;
gpu?: LlamaGpu; // default: NODE_LLAMA_CPP_GPU, else "auto"
}

A fresh session per call — the contract is single-shot. The schema is compiled to a GBNF grammar and restated in the system prompt, because the grammar hides descriptions from the model. Truncation at maxTokens is reported as an error naming the limit rather than a validation failure.

Each call gets a fresh context, sized once the prompt is known. The context is most of what a call costs in memory beyond the weights, so this is the number that decides whether a run fits a machine.

The provider counts the tokens the call needs with the model’s own tokenizer:

Part Tokens
System prompt, with the schema restated counted
User prompt counted
Chat-template overhead 512
Response maxTokens, or 2048 when unset, plus thoughtTokens
contextSize Context created When the call does not fit
unset (default) 8192 tokens, or the count when it is larger, up to the model’s training context error naming the training context
a number exactly that many tokens error naming llamaCpp.contextSize

A call that does not fit is never sent, because llama.cpp would shift the start of the prompt out of the context. It fails with an InferenceError, which completeValidatedJSON and judge record on run.error. The messages are in the error reference.

Earlier releases let node-llama-cpp choose the size, and it chose the largest context free memory allowed. For granite-4.1-3b-q2, trained on 131072 tokens, a one-field prompt of about 4000 tokens took 12.9 GB on CPU. With the 8192-token default the same call takes 2.85 GB. ADR 01011 records the measurements and the options weighed.

The real runtime does not run llama.cpp in your process. It forks a worker, llama-worker.js from beside the package’s index.js, and sends it each load, session and prompt over IPC. Your process never initialises a GPU backend.

The reason is that llama.cpp reports a fatal GPU error by aborting the process (GGML_ABORT), and no try can catch that. On some Windows machines with an NVIDIA GPU, node-llama-cpp’s prebuilt CUDA backend does this intermittently mid-generation, printing ggml-cuda.cu:106: CUDA error. In-process, that ended the consumer and every result it had computed. In a worker, it ends the worker.

gpu (or NODE_LLAMA_CPP_GPU) Starts on When a worker crashes
unset, or "auto" the best backend this machine has the call is retried on the next one — CUDA, then Vulkan, then the CPU — and the crashed backend is skipped for the rest of the process
"cuda", "vulkan", "metal", or false (CPU) that backend never replaced: the call fails with an InferenceError naming the crash, recorded on run.error

Each switch is announced once on stderr; the messages are in the warnings reference. The CPU is a last resort, because it is about 50× slower. If you would rather fail fast than wait, pin a backend.

Only an abnormal exit counts as a crash. An ordinary error, such as a context overflow, a grammar error or a missing model file, passes through unchanged on the same worker and is never retried on another backend.

The backend is not part of the provider identity or the cache key. Two backends run the same weights under the same grammar, and a cache that missed after every fallback would make you pay for the work twice.

One worker serves every model loaded for a backend. It holds your process open only while a call is in flight, so a script that never calls disposeLlamaModels() still exits when its work is done, and the worker exits with it. If you bundle this package into a single file, the worker cannot be found. The runtime then runs in-process as before and warns that a native crash will end the process.

ADR 01012 records the crash, the options weighed, and why a pre-flight probe cannot replace this.

node-llama-cpp is an optional peer dependency, and npm does not install optional peers. Since detection ends at llama-cpp precisely because it needs no key and no account, a machine with no credentials and no binding would otherwise have nothing left to fall back to.

So the library installs it on demand, into a directory it owns — never your node_modules, your package.json, or your lockfile.

function defaultLlamaRuntimeDirectory(env?: Record<string, string | undefined>): string;
function nodeLlamaCppStatus(options?: RuntimeInstallOptions): Promise<RuntimeStatus>;
function importNodeLlamaCpp(options?: RuntimeInstallOptions): Promise<unknown>;
function resetRuntimeInstall(): void;
type RuntimeStatus =
| { state: "present" }
| { state: "installable"; directory: string }
| { state: "refused"; reason: string };
interface RuntimeInstallOptions {
directory?: string; // defaults to defaultLlamaRuntimeDirectory()
exec?: ExecFn; // test seam
env?: Record<string, string | undefined>;
timeoutMs?: number; // default 15 minutes
importShim?: (url: string) => Promise<unknown>; // test seam
probeImport?: () => Promise<unknown>; // test seam
}

Resolution runs in three steps, stopping at the first that works: your own installed copy, then <runtime dir>/loader.mjs from an earlier run, then an install.

Variable Effect
INFERENCE_RUNTIME_DIR Where the binding is installed. Default ~/.hawkeyexl-inference/runtime.
INFERENCE_NO_AUTO_INSTALL Set to anything non-empty to refuse installing. llama-cpp then reports unavailable instead.

nodeLlamaCppStatus answers whether the runtime is usable without installing it. That is what keeps availableProviders() a query: asking a machine what it can do never changes what it can do. The install happens later, when the provider is actually resolved, and warns once before it starts.

importNodeLlamaCpp is exported so you can pre-fetch deliberately — in a Docker build, say, so the first real run does not pay for it.

The shim is written only after npm exits 0, so an interrupted install leaves a prefix that is retried rather than one that looks ready and is not. Concurrent callers in one process share a single install; concurrent processes serialise on a lock file in the prefix.

src/providers/llama-models.ts, llama-cpp.ts, llama-clean.ts, and llama-install.ts. Catalog invariants, tier selection, every clearing rule, and the install’s ordering and concurrency guarantees are pinned in test/unit/llama-*.test.ts.