Skip to content

Testing your integration

Every boundary that would otherwise leave your process has an injection seam. Use them and your unit tests need no network, no API key, and no model weights.

Boundary Seam Injected via
The provider contract MockProvider constructed directly, or ProviderSpec.mockResponses
A subprocess ExecFn ProviderSpec.exec
Native local inference LlamaRuntime ProviderSpec.llamaRuntime or llamaCpp.runtime

Each is the narrowest thing that could be faked at that point. This library uses exactly these seams on itself — every unit test runs through one of them, and the only live tests are gated on environment variables and skipped by default.

examples/testing-seams.mjs
// The three injection seams: test an entire integration with no network,
// no credential, and no GGUF weights.
import {
MockProvider,
completeValidatedJSON,
judge,
makeProvider,
makeProviderAsync,
mockVerdict,
} from "@hawkeyexl/inference";
const system = "You evaluate whether a page satisfies an assertion.";
const user = "# Assertion\nThe page documents authentication.\n\n# Page\nUse a bearer token.";
// ---------------------------------------------------------------------------
// Seam 1 — MockProvider, for anything above the provider contract.
// ---------------------------------------------------------------------------
const provider = new MockProvider([mockVerdict("pass", 0.95)]); // cycles when exhausted
const consensus = await judge({ provider, system, user, runs: 3 });
console.log("1. verdict:", consensus.verdict, "zone:", consensus.zone);
// requests records every CompleteJSONRequest, in order. This is how you assert on
// the part you actually own: that your prompt and schema were composed correctly.
console.log(" requests seen:", provider.requests.length);
console.log(" system prompt matched:", provider.requests[0].system === system);
console.log(" temperature:", provider.requests[0].temperature);
// An { error } entry rejects, which is how you prove your failure paths.
const flaky = new MockProvider([{ error: "429 rate limited" }]);
const failed = await completeValidatedJSON({
provider: flaky,
system,
user,
schema: { type: "object", required: ["ok"], properties: { ok: { type: "boolean" } } },
});
console.log(" scripted failure recorded:", failed.error);
// ---------------------------------------------------------------------------
// Seam 2 — ExecFn, for the claude-cli provider's subprocess.
// ---------------------------------------------------------------------------
const calls = [];
const fakeExec = async (cmd, opts) => {
calls.push({ cmd, input: opts?.input });
// The CLI wraps its answer in a { result: string } envelope.
return {
code: 0,
stdout: JSON.stringify({
result: JSON.stringify({
claim: "The page documents authentication.",
observed: "The page describes bearer tokens.",
match: "pass",
confidence: 0.88,
reasoning: "The token type is stated explicitly.",
}),
}),
stderr: "",
timedOut: false,
};
};
const cli = makeProvider({ provider: "claude-cli", exec: fakeExec });
const cliResult = await judge({ provider: cli, system, user, runs: 1 });
console.log("2. verdict:", cliResult.verdict, "zone:", cliResult.zone);
console.log(" argv:", calls[0].cmd.slice(0, 2).join(" "), "...");
// The prompt goes over stdin, never argv — user content routinely exceeds the
// ~32K Windows command-line limit.
console.log(" prompt sent over stdin:", calls[0].input.includes("bearer token"));
console.log(" prompt absent from argv:", !calls[0].cmd.join(" ").includes("bearer token"));
// ---------------------------------------------------------------------------
// Seam 3 — LlamaRuntime, for in-process local inference.
// ---------------------------------------------------------------------------
const fakeRuntime = {
async getMemoryBudgetBytes() {
return 16_000_000_000;
},
async resolveModelFile(uri, directory) {
return `${directory}/${uri.split("/").pop()}`;
},
async loadModel() {
return {
async createSession() {
return {
async prompt() {
return {
text: JSON.stringify({
claim: "The page documents authentication.",
observed: "The page describes bearer tokens.",
match: "pass",
confidence: 0.9,
reasoning: "Stated directly.",
}),
usage: { inputTokens: 400, outputTokens: 60 },
stopReason: "stop",
};
},
async dispose() {},
};
},
async dispose() {},
};
},
};
const local = await makeProviderAsync({ provider: "llama-cpp", llamaRuntime: fakeRuntime });
const localResult = await judge({ provider: local, system, user, runs: 1 });
console.log("3. model:", local.modelName());
console.log(" verdict:", localResult.verdict, "zone:", localResult.zone);
console.log(" no weights downloaded, no GPU touched");
1. verdict: pass zone: auto-pass
requests seen: 3
system prompt matched: true
temperature: 0
scripted failure recorded: 429 rate limited
2. verdict: pass zone: auto-pass
argv: claude -p ...
prompt sent over stdin: true
prompt absent from argv: true
3. model: qwen3.5-4b
verdict: pass zone: auto-pass
no weights downloaded, no GPU touched
const provider = new MockProvider([mockVerdict("pass", 0.95)]);

Responses cycle when exhausted, so a one-element script answers every call. mockVerdict(match, confidence, overrides?) builds a VERDICT_SCHEMA-shaped payload; for other schemas pass { json: ... } directly.

provider.requests records every CompleteJSONRequest, in order. This is how you test the part you actually own — that your prompt and schema were composed correctly.

expect(provider.requests).toHaveLength(3);
expect(provider.requests[0].system).toBe(MY_SYSTEM_PROMPT);
expect(provider.requests[0].schema).toBe(mySchema);
expect(provider.requests[0].temperature).toBe(0);
const flaky = new MockProvider([{ error: "429 rate limited" }]);

An { error } entry rejects. This is how you prove the guarantees you are relying on:

  • an errored run forces human-review, however confident the others were;
  • completeValidatedJSON returns a run with error set rather than throwing or coercing;
  • your budget gate stops before the next call rather than after the overspend.

None of that can be triggered against a live provider on demand.

For the claude-cli provider’s subprocess:

const fakeExec = async (cmd, opts) => ({
code: 0,
stdout: JSON.stringify({ result: JSON.stringify(myVerdict) }),
stderr: "",
timedOut: false,
});
const provider = makeProvider({ provider: "claude-cli", exec: fakeExec });

The CLI wraps its answer in a { result: string } envelope, so your fake must too.

Assert on cmd for the argv, and on opts.input for the piped stdin. The sample above checks the property that matters most: the prompt goes over stdin and never appears in argv, because user content routinely exceeds the ~32K Windows command-line limit.

For local inference, with no weights and no GPU:

const fakeRuntime = {
async getMemoryBudgetBytes() { return 16_000_000_000; },
async resolveModelFile(uri, directory) { return `${directory}/${uri.split("/").pop()}`; },
async loadModel() {
return {
async createSession() {
return {
async prompt() { return { text: JSON.stringify(myVerdict), usage: { inputTokens: 400, outputTokens: 60 } }; },
async dispose() {},
};
},
async dispose() {},
};
},
};
const provider = await makeProviderAsync({ provider: "llama-cpp", llamaRuntime: fakeRuntime });

getMemoryBudgetBytes is what auto probes, so returning a fixed number makes tier selection deterministic in tests.

resetTemperatureWarning() clears the once-per-process warning that fires when temperature > 0. Call it in a beforeEach if you assert on warnings, or a test that runs second will not see one.

The same pattern, applied to documentation. Every runnable sample here lives in examples/, uses MockProvider or a fake runtime, and is:

  • rendered onto the page from that file with a ?raw import, so a page cannot show code that differs from the file, and
  • executed on every pull request by Doc Detective, which runs the file and asserts on its output.

A broken example fails CI like any other test. See ADR 01005.