Skip to content

Managing model files

Finding multi-gigabyte weights, moving them somewhere with room, and deleting them safely.

Choosing what to download is choosing a model. Signatures are in the local models reference.

~/.hawkeyexl-inference/models

This library’s own directory — deliberately not node-llama-cpp’s shared ~/.node-llama-cpp/models.

That shared default is used by node-llama-cpp’s own CLI and anything else on the machine using it. Clearing it could destroy weights this library never downloaded. Owning a directory removes the hazard instead of defending against it.

Two ways, in precedence order:

// Per provider — wins.
await makeProviderAsync({ provider: "llama-cpp", llamaCpp: { modelsDirectory: "/mnt/models" } });
Terminal window
# Process-wide.
export INFERENCE_MODELS_DIR=/mnt/models

defaultLlamaModelsDirectory() reports the effective value, which is the fastest way to answer “where did it actually put them.”

examples/manage-model-files.mjs
// Finding, inspecting, and safely clearing local model weights.
// Runs with no API key and no real weights: it operates on a throwaway directory.
import { mkdtempSync, readdirSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import {
blobNameFor,
clearLlamaModels,
defaultLlamaModelsDirectory,
isModelDownloaded,
} from "@hawkeyexl/inference";
// Weights live in this library's OWN directory, not node-llama-cpp's shared one,
// so clearing can never destroy models something else downloaded.
// Override with INFERENCE_MODELS_DIR, or llamaCpp.modelsDirectory per provider.
console.log("default directory:", defaultLlamaModelsDirectory());
// The on-disk filename an alias resolves to. Downloads are prefixed hf_<user>_.
console.log("blob for gemma-4-e2b:", blobNameFor("gemma-4-e2b"));
// A throwaway directory standing in for a populated models directory.
const directory = mkdtempSync(join(tmpdir(), "inference-models-"));
const fake = (name, bytes) => writeFileSync(join(directory, name), Buffer.alloc(bytes));
fake(`hf_unsloth_${blobNameFor("gemma-4-e2b")}`, 2048);
fake(`hf_unsloth_${blobNameFor("gemma-4-12b")}`, 4096);
fake(`hf_unsloth_${blobNameFor("gemma-4-e4b")}.ipull`, 512); // an interrupted download
fake("notes.txt", 64); // not a model — must survive
console.log("is gemma-4-e2b downloaded:", isModelDownloaded("gemma-4-e2b", directory));
console.log("is gemma-4-e4b downloaded:", isModelDownloaded("gemma-4-e4b", directory));
// Always preview first. dryRun reports and deletes nothing.
const preview = await clearLlamaModels({ directory, dryRun: true });
console.log("would remove:", preview.files.length, "files,", preview.freedBytes, "bytes");
console.log("dryRun deleted nothing:", readdirSync(directory).length === 4);
// Clear one model by alias. Only that model's blobs go.
const one = await clearLlamaModels({ directory, models: ["gemma-4-12b"] });
console.log("removed by alias:", one.files.length, "remaining:", readdirSync(directory).length);
// Clear the rest. Interrupted .ipull partials are removed too.
const rest = await clearLlamaModels({ directory });
console.log("removed the rest:", rest.files.length);
// Only .gguf and .gguf.ipull are ever touched, top level only, never recursing.
console.log("survivors:", readdirSync(directory));
// An unknown name is rejected rather than silently matching nothing.
try {
await clearLlamaModels({ directory, models: ["not-a-real-model"] });
} catch (error) {
console.log("unknown model rejected:", error.message.includes("not-a-real-model"));
}
rmSync(directory, { recursive: true, force: true });
default directory: C:\Users\you\.hawkeyexl-inference\models
blob for gemma-4-e2b: gemma-4-E2B-it-qat-UD-Q4_K_XL.gguf
is gemma-4-e2b downloaded: true
is gemma-4-e4b downloaded: false
would remove: 3 files, 6656 bytes
dryRun deleted nothing: true
removed by alias: 1 remaining: 3
removed the rest: 2
survivors: [ 'notes.txt' ]
unknown model rejected: true
  1. dryRun reports files and freedBytes and deletes nothing.

    const { files, freedBytes } = await clearLlamaModels({ dryRun: true });
  2. Then clear, once the report matches what you expected.

    await clearLlamaModels(); // everything
    await clearLlamaModels({ models: ["gemma-4-12b"] }); // one model, all its parts
    await clearLlamaModels({ directory: "/mnt/models" }); // a non-default directory

Lead with dryRun every time you point this at a directory you have not cleared before. It costs one line.

You are being asked to run a delete over a directory holding gigabytes. These hold regardless of which directory you point at:

  • Only .gguf and .gguf.ipull files are ever touched. notes.txt in the sample survives.
  • Top level only. Subdirectories are never walked.
  • Loaded weights are disposed first — a memory-mapped model cannot be deleted on Windows.
  • A file held open is skipped, not forced.
  • An unknown model name is rejected, rather than silently matching nothing and reporting success.

Interrupted .ipull partial downloads and every part of a split model go together, so you never end up with half a model on disk.

await disposeLlamaModels();

Weights load once per process and are shared across every provider naming the same model. In a long-running process — a server, a watch mode — this releases them without touching disk.

It is a free function rather than a method because the cache is process-wide, not per-provider: disposing one provider’s model would break another provider that named the same weights.