Testing your integration
Every boundary that would otherwise leave your process has an injection seam. Use them and your unit tests need no network, no API key, and no model weights.
One seam per boundary
Section titled “One seam per boundary”| Boundary | Seam | Injected via |
|---|---|---|
| The provider contract | MockProvider |
constructed directly, or ProviderSpec.mockResponses |
| A subprocess | ExecFn |
ProviderSpec.exec |
| Native local inference | LlamaRuntime |
ProviderSpec.llamaRuntime or llamaCpp.runtime |
Each is the narrowest thing that could be faked at that point. This library uses exactly these seams on itself — every unit test runs through one of them, and the only live tests are gated on environment variables and skipped by default.
All three
Section titled “All three”// The three injection seams: test an entire integration with no network,// no credential, and no GGUF weights.
import { MockProvider, completeValidatedJSON, judge, makeProvider, makeProviderAsync, mockVerdict,} from "@hawkeyexl/inference";
const system = "You evaluate whether a page satisfies an assertion.";const user = "# Assertion\nThe page documents authentication.\n\n# Page\nUse a bearer token.";
// ---------------------------------------------------------------------------// Seam 1 — MockProvider, for anything above the provider contract.// ---------------------------------------------------------------------------
const provider = new MockProvider([mockVerdict("pass", 0.95)]); // cycles when exhaustedconst consensus = await judge({ provider, system, user, runs: 3 });console.log("1. verdict:", consensus.verdict, "zone:", consensus.zone);
// requests records every CompleteJSONRequest, in order. This is how you assert on// the part you actually own: that your prompt and schema were composed correctly.console.log(" requests seen:", provider.requests.length);console.log(" system prompt matched:", provider.requests[0].system === system);console.log(" temperature:", provider.requests[0].temperature);
// An { error } entry rejects, which is how you prove your failure paths.const flaky = new MockProvider([{ error: "429 rate limited" }]);const failed = await completeValidatedJSON({ provider: flaky, system, user, schema: { type: "object", required: ["ok"], properties: { ok: { type: "boolean" } } },});console.log(" scripted failure recorded:", failed.error);
// ---------------------------------------------------------------------------// Seam 2 — ExecFn, for the claude-cli provider's subprocess.// ---------------------------------------------------------------------------
const calls = [];const fakeExec = async (cmd, opts) => { calls.push({ cmd, input: opts?.input }); // The CLI wraps its answer in a { result: string } envelope. return { code: 0, stdout: JSON.stringify({ result: JSON.stringify({ claim: "The page documents authentication.", observed: "The page describes bearer tokens.", match: "pass", confidence: 0.88, reasoning: "The token type is stated explicitly.", }), }), stderr: "", timedOut: false, };};
const cli = makeProvider({ provider: "claude-cli", exec: fakeExec });const cliResult = await judge({ provider: cli, system, user, runs: 1 });
console.log("2. verdict:", cliResult.verdict, "zone:", cliResult.zone);console.log(" argv:", calls[0].cmd.slice(0, 2).join(" "), "...");// The prompt goes over stdin, never argv — user content routinely exceeds the// ~32K Windows command-line limit.console.log(" prompt sent over stdin:", calls[0].input.includes("bearer token"));console.log(" prompt absent from argv:", !calls[0].cmd.join(" ").includes("bearer token"));
// ---------------------------------------------------------------------------// Seam 3 — LlamaRuntime, for in-process local inference.// ---------------------------------------------------------------------------
const fakeRuntime = { async getMemoryBudgetBytes() { return 16_000_000_000; }, async resolveModelFile(uri, directory) { return `${directory}/${uri.split("/").pop()}`; }, async loadModel() { return { async createSession() { return { async prompt() { return { text: JSON.stringify({ claim: "The page documents authentication.", observed: "The page describes bearer tokens.", match: "pass", confidence: 0.9, reasoning: "Stated directly.", }), usage: { inputTokens: 400, outputTokens: 60 }, stopReason: "stop", }; }, async dispose() {}, }; }, async dispose() {}, }; },};
const local = await makeProviderAsync({ provider: "llama-cpp", llamaRuntime: fakeRuntime });const localResult = await judge({ provider: local, system, user, runs: 1 });
console.log("3. model:", local.modelName());console.log(" verdict:", localResult.verdict, "zone:", localResult.zone);console.log(" no weights downloaded, no GPU touched");1. verdict: pass zone: auto-pass requests seen: 3 system prompt matched: true temperature: 0 scripted failure recorded: 429 rate limited2. verdict: pass zone: auto-pass argv: claude -p ... prompt sent over stdin: true prompt absent from argv: true3. model: qwen3.5-4b verdict: pass zone: auto-pass no weights downloaded, no GPU touchedMockProvider
Section titled “MockProvider”const provider = new MockProvider([mockVerdict("pass", 0.95)]);Responses cycle when exhausted, so a one-element script answers every call. mockVerdict(match, confidence, overrides?) builds a VERDICT_SCHEMA-shaped payload; for other schemas pass
{ json: ... } directly.
Assert on what you sent
Section titled “Assert on what you sent”provider.requests records every CompleteJSONRequest, in order. This is how you test the part you
actually own — that your prompt and schema were composed correctly.
expect(provider.requests).toHaveLength(3);expect(provider.requests[0].system).toBe(MY_SYSTEM_PROMPT);expect(provider.requests[0].schema).toBe(mySchema);expect(provider.requests[0].temperature).toBe(0);Script a failure
Section titled “Script a failure”const flaky = new MockProvider([{ error: "429 rate limited" }]);An { error } entry rejects. This is how you prove the guarantees you are relying on:
- an errored run forces
human-review, however confident the others were; completeValidatedJSONreturns a run witherrorset rather than throwing or coercing;- your budget gate stops before the next call rather than after the overspend.
None of that can be triggered against a live provider on demand.
ExecFn
Section titled “ExecFn”For the claude-cli provider’s subprocess:
const fakeExec = async (cmd, opts) => ({ code: 0, stdout: JSON.stringify({ result: JSON.stringify(myVerdict) }), stderr: "", timedOut: false,});
const provider = makeProvider({ provider: "claude-cli", exec: fakeExec });The CLI wraps its answer in a { result: string } envelope, so your fake must too.
Assert on cmd for the argv, and on opts.input for the piped stdin. The sample above checks the
property that matters most: the prompt goes over stdin and never appears in argv, because user
content routinely exceeds the ~32K Windows command-line limit.
LlamaRuntime
Section titled “LlamaRuntime”For local inference, with no weights and no GPU:
const fakeRuntime = { async getMemoryBudgetBytes() { return 16_000_000_000; }, async resolveModelFile(uri, directory) { return `${directory}/${uri.split("/").pop()}`; }, async loadModel() { return { async createSession() { return { async prompt() { return { text: JSON.stringify(myVerdict), usage: { inputTokens: 400, outputTokens: 60 } }; }, async dispose() {}, }; }, async dispose() {}, }; },};
const provider = await makeProviderAsync({ provider: "llama-cpp", llamaRuntime: fakeRuntime });getMemoryBudgetBytes is what auto probes, so returning a fixed number makes tier selection
deterministic in tests.
Test isolation
Section titled “Test isolation”resetTemperatureWarning() clears the once-per-process warning that fires when temperature > 0.
Call it in a beforeEach if you assert on warnings, or a test that runs second will not see one.
How these docs test themselves
Section titled “How these docs test themselves”The same pattern, applied to documentation. Every runnable sample here lives in
examples/, uses MockProvider or a
fake runtime, and is:
- rendered onto the page from that file with a
?rawimport, so a page cannot show code that differs from the file, and - executed on every pull request by Doc Detective, which runs the file and asserts on its output.
A broken example fails CI like any other test. See ADR 01005.