Choosing a model
What auto picks, how to override it, and what stops working once a grammar is doing the
constraining.
The full catalog and signatures are in the local models reference.
What you can pass as model
Section titled “What you can pass as model”| Kind | Example |
|---|---|
| Selector | auto (default), fast, balanced, quality |
| Curated alias | qwen3.5-4b |
| Hugging Face URI | hf:unsloth/Qwen3.5-9B-GGUF/Qwen3.5-9B-UD-Q4_K_XL.gguf |
| Local file | ./models/my-model.gguf |
A bare user/repo is rejected rather than guessed at — it is almost always a mistyped alias.
How auto measures your machine
Section titled “How auto measures your machine”It takes the larger of free GPU VRAM and half of system RAM. All of RAM when there is no GPU, or when the probe fails.
Larger rather than VRAM alone because llama.cpp offloads the layers that fit onto the GPU and keeps the rest in system RAM. A box with a small GPU and plenty of RAM still runs a big model well, and sizing off VRAM alone would leave most of the machine idle.
A model needs 3.5× its file size in that budget. Tiers are walked smallest to largest, and the
result floors at fast — you always get something.
What a call costs in memory
Section titled “What a call costs in memory”The weights are one part. The context is the other: each call gets its own, sized to its prompt.
It is 8192 tokens unless the prompt and its response need more, and never more than the model was
trained on. On granite-4.1-3b-q2, a call with a 4000-token prompt peaks at about 2.85 GB on CPU.
To pin the size, on a machine you share with other jobs say, set contextSize:
await makeProviderAsync({ provider: "llama-cpp", llamaCpp: { contextSize: 4096, maxTokens: 512 },});A prompt that does not fit is refused with an error that lists its token counts, never truncated. The sizing rule and the messages are in the local models reference.
Inspect before you download
Section titled “Inspect before you download”The catalog is exported, so you can see sizes and licenses before triggering a multi-gigabyte download.
// Inspecting the curated model catalog before triggering a multi-gigabyte download.// Runs with no API key and no weights — this reads exported data only.
import { LLAMA_MODELS, LLAMA_SELECTORS, LLAMA_TIERS, aliasForTier, defaultLlamaModelsDirectory, isLlamaSelector, resolveLlamaModelRef, tierForBudget, uriForTier,} from "@hawkeyexl/inference";
// Decimal GB, matching how the catalog and the library's own download notice report sizes.const GB = 1_000_000_000;
console.log("selectors:", LLAMA_SELECTORS);console.log("tiers:", LLAMA_TIERS);console.log("weights live in:", defaultLlamaModelsDirectory());
console.log("\ncatalog:");for (const [alias, entry] of Object.entries(LLAMA_MODELS)) { const size = (entry.sizeBytes / GB).toFixed(2).padStart(5); console.log(` ${alias.padEnd(18)} ${size} GB ${entry.license} tier=${entry.tier ?? "-"}`);}
// What `auto` would pick on a given machine. The budget is the larger of free GPU// VRAM and half of system RAM; a model needs 3.5x its file size in that budget.console.log("\nwhat auto picks:");for (const gb of [6, 8, 16, 24, 32]) { const tier = tierForBudget(gb * GB); console.log(` ${String(gb).padStart(2)} GB budget -> ${tier.padEnd(8)} -> ${aliasForTier(tier)}`);}
console.log("\nselector check:", isLlamaSelector("auto"), isLlamaSelector("gemma-4-e2b"));
// Aliases expand to an exact blob path, never a :QUANT tag, so a model cannot// silently re-point underneath a cache key that already names it.console.log("\nquality tier resolves to:");console.log(" ", uriForTier("quality"));console.log("alias resolves to:");console.log(" ", resolveLlamaModelRef("gemma-4-e2b"));catalog: granite-4.1-3b-q2 1.41 GB Apache-2.0 tier=fast qwen3.5-4b 2.91 GB Apache-2.0 tier=balanced qwen3.5-9b 5.97 GB Apache-2.0 tier=quality gemma-4-e2b 2.62 GB Apache-2.0 tier=- gemma-4-e4b 4.22 GB Apache-2.0 tier=- gemma-4-12b 6.72 GB Apache-2.0 tier=- gemma-4-26b-a4b 14.25 GB Apache-2.0 tier=- gemma-4-e2b-q2 2.19 GB Apache-2.0 tier=-
what auto picks: 6 GB budget -> fast -> granite-4.1-3b-q2 8 GB budget -> fast -> granite-4.1-3b-q2 16 GB budget -> balanced -> qwen3.5-4b 24 GB budget -> quality -> qwen3.5-9b 32 GB budget -> quality -> qwen3.5-9bEvery catalog entry is Apache-2.0 and ungated, which matters if you operate somewhere that audits model licenses.
tierForBudget is exported, so you can check what a machine you do not have in front of you would
get.
Why blob paths are pinned
Section titled “Why blob paths are pinned”Catalog entries pin an exact blob path, never a :QUANT tag:
hf:unsloth/Qwen3.5-9B-GGUF/Qwen3.5-9B-UD-Q4_K_XL.ggufA tag could be re-pointed upstream, which would silently change the weights behind a cache key that already names the model. Pinning removes that possibility. If you have a populated cache, this is the detail that protects it.
What changes under a grammar
Section titled “What changes under a grammar”The schema is compiled to a GBNF grammar. Four behaviors differ from a hosted provider, and a schema that works against Anthropic can behave differently here:
requiredis ignored. Every key inpropertiesis always emitted.additionalPropertiesdefaults tofalse.- Numeric bounds are not enforced by the grammar. A
minimum/maximumviolation comes back as well-formed JSON and is caught by the normal Ajv validation and retry. descriptions are invisible to the grammar. The provider restates the schema in the system prompt instead, so descriptions still steer the model — they just arrive by a different route.
None of these affect the built-in VERDICT_SCHEMA, which requires all of its fields. They matter if
you wrote your own — see custom verdict schema.
Thinking is off by default
Section titled “Thinking is off by default”A grammar constrains generation from token zero, which cuts a reasoning model off mid-thought. So thinking is disabled unless you budget for it:
await makeProviderAsync({ provider: "llama-cpp", llamaCpp: { thoughtTokens: 512 },});If local output looks worse than you expected, check this before blaming the model.
- Manage model files — where weights live, and reclaiming disk
- Local models reference — the full catalog and every export