Skip to content

Choosing a model

What auto picks, how to override it, and what stops working once a grammar is doing the constraining.

The full catalog and signatures are in the local models reference.

Kind Example
Selector auto (default), fast, balanced, quality
Curated alias qwen3.5-4b
Hugging Face URI hf:unsloth/Qwen3.5-9B-GGUF/Qwen3.5-9B-UD-Q4_K_XL.gguf
Local file ./models/my-model.gguf

A bare user/repo is rejected rather than guessed at — it is almost always a mistyped alias.

It takes the larger of free GPU VRAM and half of system RAM. All of RAM when there is no GPU, or when the probe fails.

Larger rather than VRAM alone because llama.cpp offloads the layers that fit onto the GPU and keeps the rest in system RAM. A box with a small GPU and plenty of RAM still runs a big model well, and sizing off VRAM alone would leave most of the machine idle.

A model needs 3.5× its file size in that budget. Tiers are walked smallest to largest, and the result floors at fast — you always get something.

The weights are one part. The context is the other: each call gets its own, sized to its prompt. It is 8192 tokens unless the prompt and its response need more, and never more than the model was trained on. On granite-4.1-3b-q2, a call with a 4000-token prompt peaks at about 2.85 GB on CPU.

To pin the size, on a machine you share with other jobs say, set contextSize:

await makeProviderAsync({
provider: "llama-cpp",
llamaCpp: { contextSize: 4096, maxTokens: 512 },
});

A prompt that does not fit is refused with an error that lists its token counts, never truncated. The sizing rule and the messages are in the local models reference.

The catalog is exported, so you can see sizes and licenses before triggering a multi-gigabyte download.

examples/local-catalog.mjs
// Inspecting the curated model catalog before triggering a multi-gigabyte download.
// Runs with no API key and no weights — this reads exported data only.
import {
LLAMA_MODELS,
LLAMA_SELECTORS,
LLAMA_TIERS,
aliasForTier,
defaultLlamaModelsDirectory,
isLlamaSelector,
resolveLlamaModelRef,
tierForBudget,
uriForTier,
} from "@hawkeyexl/inference";
// Decimal GB, matching how the catalog and the library's own download notice report sizes.
const GB = 1_000_000_000;
console.log("selectors:", LLAMA_SELECTORS);
console.log("tiers:", LLAMA_TIERS);
console.log("weights live in:", defaultLlamaModelsDirectory());
console.log("\ncatalog:");
for (const [alias, entry] of Object.entries(LLAMA_MODELS)) {
const size = (entry.sizeBytes / GB).toFixed(2).padStart(5);
console.log(` ${alias.padEnd(18)} ${size} GB ${entry.license} tier=${entry.tier ?? "-"}`);
}
// What `auto` would pick on a given machine. The budget is the larger of free GPU
// VRAM and half of system RAM; a model needs 3.5x its file size in that budget.
console.log("\nwhat auto picks:");
for (const gb of [6, 8, 16, 24, 32]) {
const tier = tierForBudget(gb * GB);
console.log(` ${String(gb).padStart(2)} GB budget -> ${tier.padEnd(8)} -> ${aliasForTier(tier)}`);
}
console.log("\nselector check:", isLlamaSelector("auto"), isLlamaSelector("gemma-4-e2b"));
// Aliases expand to an exact blob path, never a :QUANT tag, so a model cannot
// silently re-point underneath a cache key that already names it.
console.log("\nquality tier resolves to:");
console.log(" ", uriForTier("quality"));
console.log("alias resolves to:");
console.log(" ", resolveLlamaModelRef("gemma-4-e2b"));
catalog:
granite-4.1-3b-q2 1.41 GB Apache-2.0 tier=fast
qwen3.5-4b 2.91 GB Apache-2.0 tier=balanced
qwen3.5-9b 5.97 GB Apache-2.0 tier=quality
gemma-4-e2b 2.62 GB Apache-2.0 tier=-
gemma-4-e4b 4.22 GB Apache-2.0 tier=-
gemma-4-12b 6.72 GB Apache-2.0 tier=-
gemma-4-26b-a4b 14.25 GB Apache-2.0 tier=-
gemma-4-e2b-q2 2.19 GB Apache-2.0 tier=-
what auto picks:
6 GB budget -> fast -> granite-4.1-3b-q2
8 GB budget -> fast -> granite-4.1-3b-q2
16 GB budget -> balanced -> qwen3.5-4b
24 GB budget -> quality -> qwen3.5-9b
32 GB budget -> quality -> qwen3.5-9b

Every catalog entry is Apache-2.0 and ungated, which matters if you operate somewhere that audits model licenses.

tierForBudget is exported, so you can check what a machine you do not have in front of you would get.

Catalog entries pin an exact blob path, never a :QUANT tag:

hf:unsloth/Qwen3.5-9B-GGUF/Qwen3.5-9B-UD-Q4_K_XL.gguf

A tag could be re-pointed upstream, which would silently change the weights behind a cache key that already names the model. Pinning removes that possibility. If you have a populated cache, this is the detail that protects it.

The schema is compiled to a GBNF grammar. Four behaviors differ from a hosted provider, and a schema that works against Anthropic can behave differently here:

  1. required is ignored. Every key in properties is always emitted.
  2. additionalProperties defaults to false.
  3. Numeric bounds are not enforced by the grammar. A minimum/maximum violation comes back as well-formed JSON and is caught by the normal Ajv validation and retry.
  4. descriptions are invisible to the grammar. The provider restates the schema in the system prompt instead, so descriptions still steer the model — they just arrive by a different route.

None of these affect the built-in VERDICT_SCHEMA, which requires all of its fields. They matter if you wrote your own — see custom verdict schema.

A grammar constrains generation from token zero, which cuts a reasoning model off mid-thought. So thinking is disabled unless you budget for it:

await makeProviderAsync({
provider: "llama-cpp",
llamaCpp: { thoughtTokens: 512 },
});

If local output looks worse than you expected, check this before blaming the model.