Managing model files
Finding multi-gigabyte weights, moving them somewhere with room, and deleting them safely.
Choosing what to download is choosing a model. Signatures are in the local models reference.
Where weights live
Section titled “Where weights live”~/.hawkeyexl-inference/modelsThis library’s own directory — deliberately not node-llama-cpp’s shared
~/.node-llama-cpp/models.
That shared default is used by node-llama-cpp’s own CLI and anything else on the machine using it. Clearing it could destroy weights this library never downloaded. Owning a directory removes the hazard instead of defending against it.
Relocate it
Section titled “Relocate it”Two ways, in precedence order:
// Per provider — wins.await makeProviderAsync({ provider: "llama-cpp", llamaCpp: { modelsDirectory: "/mnt/models" } });# Process-wide.export INFERENCE_MODELS_DIR=/mnt/modelsdefaultLlamaModelsDirectory() reports the effective value, which is the fastest way to answer
“where did it actually put them.”
// Finding, inspecting, and safely clearing local model weights.// Runs with no API key and no real weights: it operates on a throwaway directory.
import { mkdtempSync, readdirSync, rmSync, writeFileSync } from "node:fs";import { tmpdir } from "node:os";import { join } from "node:path";
import { blobNameFor, clearLlamaModels, defaultLlamaModelsDirectory, isModelDownloaded,} from "@hawkeyexl/inference";
// Weights live in this library's OWN directory, not node-llama-cpp's shared one,// so clearing can never destroy models something else downloaded.// Override with INFERENCE_MODELS_DIR, or llamaCpp.modelsDirectory per provider.console.log("default directory:", defaultLlamaModelsDirectory());
// The on-disk filename an alias resolves to. Downloads are prefixed hf_<user>_.console.log("blob for gemma-4-e2b:", blobNameFor("gemma-4-e2b"));
// A throwaway directory standing in for a populated models directory.const directory = mkdtempSync(join(tmpdir(), "inference-models-"));const fake = (name, bytes) => writeFileSync(join(directory, name), Buffer.alloc(bytes));
fake(`hf_unsloth_${blobNameFor("gemma-4-e2b")}`, 2048);fake(`hf_unsloth_${blobNameFor("gemma-4-12b")}`, 4096);fake(`hf_unsloth_${blobNameFor("gemma-4-e4b")}.ipull`, 512); // an interrupted downloadfake("notes.txt", 64); // not a model — must survive
console.log("is gemma-4-e2b downloaded:", isModelDownloaded("gemma-4-e2b", directory));console.log("is gemma-4-e4b downloaded:", isModelDownloaded("gemma-4-e4b", directory));
// Always preview first. dryRun reports and deletes nothing.const preview = await clearLlamaModels({ directory, dryRun: true });console.log("would remove:", preview.files.length, "files,", preview.freedBytes, "bytes");console.log("dryRun deleted nothing:", readdirSync(directory).length === 4);
// Clear one model by alias. Only that model's blobs go.const one = await clearLlamaModels({ directory, models: ["gemma-4-12b"] });console.log("removed by alias:", one.files.length, "remaining:", readdirSync(directory).length);
// Clear the rest. Interrupted .ipull partials are removed too.const rest = await clearLlamaModels({ directory });console.log("removed the rest:", rest.files.length);
// Only .gguf and .gguf.ipull are ever touched, top level only, never recursing.console.log("survivors:", readdirSync(directory));
// An unknown name is rejected rather than silently matching nothing.try { await clearLlamaModels({ directory, models: ["not-a-real-model"] });} catch (error) { console.log("unknown model rejected:", error.message.includes("not-a-real-model"));}
rmSync(directory, { recursive: true, force: true });default directory: C:\Users\you\.hawkeyexl-inference\modelsblob for gemma-4-e2b: gemma-4-E2B-it-qat-UD-Q4_K_XL.ggufis gemma-4-e2b downloaded: trueis gemma-4-e4b downloaded: falsewould remove: 3 files, 6656 bytesdryRun deleted nothing: trueremoved by alias: 1 remaining: 3removed the rest: 2survivors: [ 'notes.txt' ]unknown model rejected: truePreview first
Section titled “Preview first”-
dryRunreportsfilesandfreedBytesand deletes nothing.const { files, freedBytes } = await clearLlamaModels({ dryRun: true }); -
Then clear, once the report matches what you expected.
await clearLlamaModels(); // everythingawait clearLlamaModels({ models: ["gemma-4-12b"] }); // one model, all its partsawait clearLlamaModels({ directory: "/mnt/models" }); // a non-default directory
Lead with dryRun every time you point this at a directory you have not cleared before. It costs
one line.
The guarantees
Section titled “The guarantees”You are being asked to run a delete over a directory holding gigabytes. These hold regardless of which directory you point at:
- Only
.ggufand.gguf.ipullfiles are ever touched.notes.txtin the sample survives. - Top level only. Subdirectories are never walked.
- Loaded weights are disposed first — a memory-mapped model cannot be deleted on Windows.
- A file held open is skipped, not forced.
- An unknown model name is rejected, rather than silently matching nothing and reporting success.
Interrupted .ipull partial downloads and every part of a split model go together, so you never end
up with half a model on disk.
Free memory without deleting
Section titled “Free memory without deleting”await disposeLlamaModels();Weights load once per process and are shared across every provider naming the same model. In a long-running process — a server, a watch mode — this releases them without touching disk.
It is a free function rather than a method because the cache is process-wide, not per-provider: disposing one provider’s model would break another provider that named the same weights.
- Choosing a model — what to download in the first place
- Local models reference — full signatures
- ADR 01003 — why the directory is owned rather than shared