Local models reference
Everything exported for the llama-cpp provider.
For using it, see run models locally.
The catalog
Section titled “The catalog”const LLAMA_MODELS: Readonly<Record<string, LlamaModelEntry>>;
interface LlamaModelEntry { uri: string; // an exact blob path, never a :QUANT tag sizeBytes: number; license: string; tier?: LlamaTier; // absent for entries outside the tier ladder notes: string;}| Alias | Size | Tier | License |
|---|---|---|---|
granite-4.1-3b-q2 |
1.41 GB | fast |
Apache-2.0 |
qwen3.5-4b |
2.91 GB | balanced |
Apache-2.0 |
qwen3.5-9b |
5.97 GB | quality |
Apache-2.0 |
gemma-4-e2b |
2.62 GB | — | Apache-2.0 |
gemma-4-e4b |
4.22 GB | — | Apache-2.0 |
gemma-4-12b |
6.72 GB | — | Apache-2.0 |
gemma-4-26b-a4b |
14.25 GB | — | Apache-2.0 |
gemma-4-e2b-q2 |
2.19 GB | — | Apache-2.0 |
All entries are unsloth GGUF builds, ungated. Each is deep-frozen.
The untiered Gemma entries backed the three tiers until ADR 01009 retiered the catalog on measured results. They still resolve by name, so an existing pin keeps working; nothing selects them automatically.
gemma-4-e2b-q2 is kept only so existing pins resolve. Do not choose it. It does not
reliably terminate: on a 12-page benchmark it left 6 pages unfinished at 120 seconds, with one
still running at 400 seconds.
The catalog is a vetted default set, not a whitelist gate — any Hugging Face GGUF reference or
local .gguf path also works.
Selectors and tiers
Section titled “Selectors and tiers”const LLAMA_TIERS: readonly ["fast", "balanced", "quality"];const LLAMA_SELECTORS: readonly ["auto", "fast", "balanced", "quality"];type LlamaTier = (typeof LLAMA_TIERS)[number];type LlamaSelector = (typeof LLAMA_SELECTORS)[number];
function isLlamaSelector(model: string): model is LlamaSelector;function tierForBudget(budgetBytes: number): LlamaTier;function aliasForTier(tier: LlamaTier): string;function uriForTier(tier: LlamaTier): string;function resolveLlamaModelRef(model: string): string;tierForBudget requires sizeBytes * 3.5 <= budget, walks tiers smallest to largest, and floors at
fast.
resolveLlamaModelRef expands an alias to its pinned URI and passes through hf:,
huggingface:, hf.co/, http(s)://, and anything ending .gguf. It throws for a selector,
and rejects a bare user/repo as a mistyped alias rather than guessing.
Model files on disk
Section titled “Model files on disk”function defaultLlamaModelsDirectory(): string;function blobNameFor(model: string): string;function isModelDownloaded(model: string, directory: string): boolean;function clearLlamaModels(options?: ClearLlamaModelsOptions): Promise<ClearLlamaModelsResult>;function disposeLlamaModels(): Promise<void>;
interface ClearLlamaModelsOptions { directory?: string; models?: string[]; // aliases or URIs; omit to clear everything dryRun?: boolean;}
interface ClearLlamaModelsResult { files: { path: string; sizeBytes: number }[]; freedBytes: number; directory: string; dryRun: boolean;}defaultLlamaModelsDirectory() is INFERENCE_MODELS_DIR or ~/.hawkeyexl-inference/models — this
library’s own directory, deliberately not node-llama-cpp’s shared ~/.node-llama-cpp/models.
Selective clearing matches by suffix, because downloads are prefixed hf_<user>_, and removes
every part of a split model. An unknown name is rejected rather than silently matching nothing.
disposeLlamaModels() frees the process-wide weight cache and stops the
local-model workers. Weights load once per process and are
shared across providers naming the same model.
The runtime seam
Section titled “The runtime seam”interface LlamaRuntime { resolveModelFile(uri: string, directory: string): Promise<string>; loadModel(path: string): Promise<LlamaLoadedModel>; getMemoryBudgetBytes(): Promise<number>;}
interface LlamaLoadedModel { createSession(systemPrompt: string, contextSize?: number): Promise<LlamaSession>; dispose(): Promise<void>; readonly trainContextSize?: number; // the model's training context, in tokens countTokens?(text: string): number | Promise<number>; // the model's own tokenizer}
interface LlamaSession { prompt(text: string, options: LlamaPromptOptions): Promise<LlamaPromptResult>; dispose(): Promise<void>; readonly contextSize?: number; // tokens of context actually created}
interface LlamaPromptOptions { schema: Record<string, unknown>; temperature: number; thoughtTokens: number; maxTokens?: number;}
interface LlamaPromptResult { text: string; usage?: TokenUsage; stopReason?: string;}
function defaultLlamaRuntime(options?: { gpu?: LlamaGpu }): LlamaRuntime;
type LlamaGpu = "auto" | "cuda" | "vulkan" | "metal" | false;defaultLlamaRuntime() is lazy: constructing it starts no process and loads no native code. Its
first use finds node-llama-cpp (a missing module raises an InferenceError naming
npm i node-llama-cpp) and starts a worker process to run
it. The built dist/index.js contains only a dynamic import, so consumers without the package
still typecheck.
Inject a fake through ProviderSpec.llamaRuntime to test the whole local path with no weights and
no GPU. See testing your integration.
getMemoryBudgetBytes returns max(free VRAM, totalmem() / 2), falling back to totalmem() / 2
when the GPU probe fails.
The provider passes createSession the context size it computed for the prompt. The real runtime
creates exactly that many tokens of context, rounded up to a multiple of 256 by llama.cpp, and 8192
when no size is passed. trainContextSize and countTokens are optional, so a fake written without
them still satisfies the interface. Without countTokens, the provider uses the 8192-token default
and does not check that the prompt fits. countTokens may return a promise; the real runtime’s does,
because its tokenizer lives in the worker.
LlamaCppProvider
Section titled “LlamaCppProvider”class LlamaCppProvider implements InferenceProvider { constructor(model: string, options?: LlamaCppProviderOptions);}
interface LlamaCppProviderOptions { runtime?: LlamaRuntime; thoughtTokens?: number; // default 0 maxTokens?: number; contextSize?: number; // default: sized to each prompt, 8192 tokens at least modelsDirectory?: string; gpu?: LlamaGpu; // default: NODE_LLAMA_CPP_GPU, else "auto"}A fresh session per call — the contract is single-shot. The schema is compiled to a GBNF grammar
and restated in the system prompt, because the grammar hides descriptions from the model.
Truncation at maxTokens is reported as an error naming the limit rather than a validation failure.
Context size
Section titled “Context size”Each call gets a fresh context, sized once the prompt is known. The context is most of what a call costs in memory beyond the weights, so this is the number that decides whether a run fits a machine.
The provider counts the tokens the call needs with the model’s own tokenizer:
| Part | Tokens |
|---|---|
| System prompt, with the schema restated | counted |
| User prompt | counted |
| Chat-template overhead | 512 |
| Response | maxTokens, or 2048 when unset, plus thoughtTokens |
contextSize |
Context created | When the call does not fit |
|---|---|---|
| unset (default) | 8192 tokens, or the count when it is larger, up to the model’s training context | error naming the training context |
| a number | exactly that many tokens | error naming llamaCpp.contextSize |
A call that does not fit is never sent, because llama.cpp would shift the start of the prompt out of
the context. It fails with an InferenceError, which completeValidatedJSON and judge record on
run.error. The messages are in the error reference.
Earlier releases let node-llama-cpp choose the size, and it chose the largest context free
memory allowed. For granite-4.1-3b-q2, trained on 131072 tokens, a one-field prompt of about 4000
tokens took 12.9 GB on CPU. With the 8192-token default the same call takes 2.85 GB.
ADR 01011
records the measurements and the options weighed.
The worker process and the backend
Section titled “The worker process and the backend”The real runtime does not run llama.cpp in your process. It forks a worker, llama-worker.js from
beside the package’s index.js, and sends it each load, session and prompt over IPC. Your process
never initialises a GPU backend.
The reason is that llama.cpp reports a fatal GPU error by aborting the process (GGML_ABORT), and
no try can catch that. On some Windows machines with an NVIDIA GPU, node-llama-cpp’s prebuilt CUDA
backend does this intermittently mid-generation, printing ggml-cuda.cu:106: CUDA error. In-process,
that ended the consumer and every result it had computed. In a worker, it ends the worker.
gpu (or NODE_LLAMA_CPP_GPU) |
Starts on | When a worker crashes |
|---|---|---|
unset, or "auto" |
the best backend this machine has | the call is retried on the next one — CUDA, then Vulkan, then the CPU — and the crashed backend is skipped for the rest of the process |
"cuda", "vulkan", "metal", or false (CPU) |
that backend | never replaced: the call fails with an InferenceError naming the crash, recorded on run.error |
Each switch is announced once on stderr; the messages are in the warnings reference. The CPU is a last resort, because it is about 50× slower. If you would rather fail fast than wait, pin a backend.
Only an abnormal exit counts as a crash. An ordinary error, such as a context overflow, a grammar error or a missing model file, passes through unchanged on the same worker and is never retried on another backend.
The backend is not part of the provider identity or the cache key. Two backends run the same weights under the same grammar, and a cache that missed after every fallback would make you pay for the work twice.
One worker serves every model loaded for a backend. It holds your process open only while a call is
in flight, so a script that never calls disposeLlamaModels() still exits when its work is done, and
the worker exits with it. If you bundle this package into a single file, the worker cannot be found.
The runtime then runs in-process as before and warns that a native crash will end the process.
ADR 01012 records the crash, the options weighed, and why a pre-flight probe cannot replace this.
Getting the runtime
Section titled “Getting the runtime”node-llama-cpp is an optional peer dependency, and npm does not install optional peers. Since
detection ends at llama-cpp precisely because it needs no key and no account, a machine with no
credentials and no binding would otherwise have nothing left to fall back to.
So the library installs it on demand, into a directory it owns — never your node_modules, your
package.json, or your lockfile.
function defaultLlamaRuntimeDirectory(env?: Record<string, string | undefined>): string;function nodeLlamaCppStatus(options?: RuntimeInstallOptions): Promise<RuntimeStatus>;function importNodeLlamaCpp(options?: RuntimeInstallOptions): Promise<unknown>;function resetRuntimeInstall(): void;
type RuntimeStatus = | { state: "present" } | { state: "installable"; directory: string } | { state: "refused"; reason: string };
interface RuntimeInstallOptions { directory?: string; // defaults to defaultLlamaRuntimeDirectory() exec?: ExecFn; // test seam env?: Record<string, string | undefined>; timeoutMs?: number; // default 15 minutes importShim?: (url: string) => Promise<unknown>; // test seam probeImport?: () => Promise<unknown>; // test seam}Resolution runs in three steps, stopping at the first that works: your own installed copy, then
<runtime dir>/loader.mjs from an earlier run, then an install.
| Variable | Effect |
|---|---|
INFERENCE_RUNTIME_DIR |
Where the binding is installed. Default ~/.hawkeyexl-inference/runtime. |
INFERENCE_NO_AUTO_INSTALL |
Set to anything non-empty to refuse installing. llama-cpp then reports unavailable instead. |
nodeLlamaCppStatus answers whether the runtime is usable without installing it. That is what
keeps availableProviders() a query: asking a machine what it can do never changes what it can do.
The install happens later, when the provider is actually resolved, and warns once before it starts.
importNodeLlamaCpp is exported so you can pre-fetch deliberately — in a Docker build, say, so the
first real run does not pay for it.
The shim is written only after npm exits 0, so an interrupted install leaves a prefix that is retried rather than one that looks ready and is not. Concurrent callers in one process share a single install; concurrent processes serialise on a lock file in the prefix.
Source of truth
Section titled “Source of truth”src/providers/llama-models.ts, llama-cpp.ts, llama-clean.ts, and llama-install.ts. Catalog
invariants, tier selection, every clearing rule, and the install’s ordering and concurrency
guarantees are pinned in test/unit/llama-*.test.ts.