Skip to content

Caching

Turn a second run over unchanged input into a free replay.

Signatures and the on-disk format are in the cache reference.

This is the whole lesson, and everything else follows from it:

This library ships no domain prompt text and no PROMPT_VERSION. buildCacheKey hashes the parts you name. It does not decide what should invalidate an entry.

That is deliberate. Every consumer has a different notion of what changed — a page body, a prompt revision, an ensemble size, a requested field set — and any library-chosen answer would be wrong for all of them.

A key that is too narrow serves stale verdicts. Too wide and it never hits. Only you know which facts matter.

examples/cache-replay.mjs
// A cached ensemble: the second run replays from disk, calls nothing, and costs nothing.
// Runs with no API key: MockProvider stands in for a real provider.
import { mkdtempSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import {
JsonCache,
MockProvider,
buildCacheKey,
costOfRuns,
mockVerdict,
pricingFor,
runEnsemble,
sha256,
} from "@hawkeyexl/inference";
const cacheDir = mkdtempSync(join(tmpdir(), "inference-example-"));
const provider = new MockProvider(
[mockVerdict("pass", 0.95), mockVerdict("pass", 0.93), mockVerdict("pass", 0.97)],
"claude-sonnet-4-5",
);
const system = "You evaluate whether a page satisfies an assertion.";
const pageBody = "# Authentication\nUse a bearer token. Refresh it every 24 hours.";
const user = `# Assertion\nThe page documents authentication.\n\n# Page\n${pageBody}`;
const runs = 3;
// You compose the key, because only you know what should invalidate an entry.
// The library hashes the parts you name; it ships no PROMPT_VERSION of its own.
const MY_PROMPT_VERSION = 4;
const cacheKey = buildCacheKey([
provider.provider(),
provider.modelName(),
`v${MY_PROMPT_VERSION}`,
`r${runs}`,
sha256(pageBody), // pre-hash long parts; key parts should stay short
]);
const cache = new JsonCache(cacheDir, true, "my-tool");
const pricing = pricingFor(provider.modelName());
const first = await runEnsemble({ provider, system, user, runs, cache, cacheKey, label: "my-tool" });
const second = await runEnsemble({ provider, system, user, runs, cache, cacheKey, label: "my-tool" });
console.log("first cached flags:", first.map((r) => r.cached));
console.log("second cached flags:", second.map((r) => r.cached));
// The provider saw three requests in total, not six.
console.log("provider calls:", provider.requests.length);
// Verdicts are identical, and the replay is free.
console.log("verdicts match:", JSON.stringify(first.map((r) => r.verdict?.match)) === JSON.stringify(second.map((r) => r.verdict?.match)));
console.log("first cost usd:", costOfRuns(first, pricing).toFixed(6));
console.log("replay cost usd:", costOfRuns(second, pricing).toFixed(6));
rmSync(cacheDir, { recursive: true, force: true });
first cached flags: [ false, false, false ]
second cached flags: [ true, true, true ]
provider calls: 3
verdicts match: true
first cost usd: 0.009000
replay cost usd: 0.000000

Three calls, not six. Identical verdicts. The replay is free.

buildCacheKey takes an array of strings and returns a hash. Two properties are worth knowing:

Parts are length-prefixed before joining. So ["a|b", "c"] and ["a", "b|c"] produce different keys. You can compose from user-controlled strings without worrying about separator collisions.

Pre-hash anything long. Key parts should stay short, so run a page body or a document through sha256 first rather than passing it whole.

const cacheKey = buildCacheKey([
provider.provider(),
provider.modelName(),
`v${MY_PROMPT_VERSION}`, // your own — bump when your prompt changes
`r${runs}`, // ensemble size changes the result
sha256(pageBody), // pre-hash long parts
]);

JsonCache deliberately does not know what shape you store. It handles the failures it can see:

  • An unparseable file is a miss, not a crash.
  • A cache hit that is not an array is a miss for runEnsemble, not a crash.

What it cannot catch is an entry that is well-formed and obsolete — written by an older version of your own schema. That one replays happily.

Every existing consumer of this package solved it the same way: wrap the cache and re-validate on read.

class VerdictCache extends JsonCache {
get(key) {
const entry = super.get(key);
if (!Array.isArray(entry)) return undefined;
// Reject entries written before your current verdict shape.
if (!entry.every((run) => run.verdict === undefined || isValidVerdict(run.verdict))) {
return undefined;
}
return entry;
}
}

Three independent teams inventing the same wrapper is a documentation gap, not a coincidence. Use the pattern rather than rediscovering it.

The cache is an optimization, never a dependency.

A failed write warns once per instance — using the label you passed to the constructor — and the run continues. A read-only workspace, a full disk, or a permissions problem must not abort work whose inference already succeeded and was already paid for.

const cache = new JsonCache(".mytool/cache", enabled, "mytool");

The second argument disables the cache entirely; get then returns undefined and set does nothing.

JsonCache takes a directory and makes no assumptions about it. The layout every existing consumer converged on:

.mytool/
├── cache/ # the judge cache
└── cache/fill/ # a separate cache for the extraction command

A dot-directory named after your tool, and one subdirectory per command.

Gitignore it, and mind the anchoring:

# Depth-agnostic on purpose. A leading-path pattern like `.mytool/cache/` is
# anchored to this file's directory, so a run inside test/fixtures/ would leak
# its cache into a tracked fixture. `**/` matches wherever the command ran.
**/.mytool/

That comment is a scar from a real repository. A command run from a subdirectory creates its cache there, and an anchored pattern does not catch it.

Replayed runs come back flagged cached: true, and costOfRuns skips them. That is why the sample above prints 0.000000 for the second run. See cost and budgets.