Skip to content

Choose a provider

Pick among five providers and configure the one you chose — or let the library pick for you.

Full option tables are in the providers reference; this page is about making the decision.

Omit provider and the library detects the highest-priority one this machine can actually use, ending at the free local model. It works with no API keys at all.

const provider = await makeProviderAsync({}); // detects
examples/detect-provider.mjs
// Letting the library pick a provider this machine can actually use.
// Runs with no API key — detection ends at the free local model.
import {
DETECTION_ORDER,
InferenceError,
detectProvider,
makeProvider,
} from "@hawkeyexl/inference";
// Priority order, highest first. `mock` is deliberately absent: it answers {} unless
// scripted, which would sail through as a real result.
console.log("detection order:", DETECTION_ORDER);
console.log("mock is never auto-selected:", !DETECTION_ORDER.includes("mock"));
// Detection probes the environment, then the Claude CLI, then the local runtime —
// so it is async. The synchronous factory refuses rather than emitting a
// non-concrete identity that would poison a cache key.
try {
makeProvider({});
} catch (error) {
console.log("sync factory refused:", error instanceof InferenceError);
}
// The highest-priority provider this machine can use right now.
const detected = await detectProvider();
console.log("detected:", detected);
console.log("detected is concrete:", DETECTION_ORDER.includes(detected));
// Omitting `provider` in a spec is identical to passing "auto".
// const provider = await makeProviderAsync({});
detection order: [ 'anthropic', 'openai', 'claude-cli', 'llama-cpp' ]
mock is never auto-selected: true
sync factory refused: true
detected: claude-cli
detected is concrete: true
Order Available when
anthropic ANTHROPIC_API_KEY is non-empty
openai OPENAI_API_KEY is non-empty, or baseUrl is set (a keyless local server)
claude-cli claude --version runs and exits 0
llama-cpp node-llama-cpp is installed

mock is never auto-selected. It answers { json: {} } unless scripted, which would sail through as a real result — the opposite of the guarantee everything else here is built on. Ask for it by name.

Two one-time warnings are emitted: which provider was auto-selected — an env var moving between runs would otherwise silently change what an eval measured — and, if llama-cpp wins and its weights are absent, the model and download size before the download starts.

Choosing explicitly is not a commitment either

Section titled “Choosing explicitly is not a commitment either”

makeProvider takes a flat, library-owned ProviderSpec. Swapping providers is a one-line change, and nothing above the provider contract knows or cares which one you picked. Do not over-deliberate this — you can change it later.

One part of the choice is consequential, and it is the easiest column to skim past.

provider Structured output via Credential Reports usage
anthropic a forced tool call ANTHROPIC_API_KEY yes
openai strict json_schema, falling back to json_object OPENAI_API_KEY yes
claude-cli schema in the prompt, --output-format json your local claude auth no
llama-cpp a GBNF grammar compiled from the schema none — runs locally yes
mock responses you script none synthetic

Map your own config into ProviderSpec rather than passing your config object through. The library owns this shape so it does not have to know anything about yours — see ADR 01000.

interface ProviderSpec {
// Omit, or pass "auto", to detect. Detection needs makeProviderAsync.
provider?: "anthropic" | "openai" | "claude-cli" | "llama-cpp" | "mock" | "auto";
model?: string | null; // null or undefined -> the per-provider default
apiKeyEnv?: string | null; // default ANTHROPIC_API_KEY / OPENAI_API_KEY
baseUrl?: string; // openai only, default https://api.openai.com/v1
command?: string; // claude-cli only, default "claude"
timeoutMs?: number; // claude-cli only, default 180000
pricing?: Pricing; // override the built-in price table
anthropic?: AnthropicProviderOptions; // toolName, toolDescription, maxTokens
openai?: OpenAICompatProviderOptions; // schemaName
llamaCpp?: LlamaCppProviderOptions; // thoughtTokens, maxTokens, contextSize, modelsDirectory
exec?: ExecFn; // test seam for claude-cli
llamaRuntime?: LlamaRuntime; // test seam for llama-cpp
mockResponses?: MockResponse[];
}

null and omission mean the same thing: use the default. If your config carries explicit nulls, strip them rather than forwarding them.

Cache keys and price lookups need to know which provider and model — but they do not need a client, and a run served entirely from cache should not demand an API key.

resolveProviderIdentity(spec) returns { provider, model } while constructing nothing and reading no credential. Pair it with a lazy thunk so a fully-cached run never needs a key at all.

examples/choose-provider.mjs
// Reading a provider's identity without constructing it — and without a credential.
// Runs with no API key.
import {
DEFAULT_MODELS,
DEFAULT_OPENAI_BASE_URL,
InferenceError,
makeProvider,
resolveProviderIdentity,
} from "@hawkeyexl/inference";
// Cache keys and price lookups need the provider's identity. They do not need a client,
// and a fully-cached run should not demand an API key. resolveProviderIdentity constructs
// nothing and reads no credential.
console.log("anthropic default:", resolveProviderIdentity({ provider: "anthropic" }));
console.log("explicit model: ", resolveProviderIdentity({ provider: "anthropic", model: "claude-haiku-4-5" }));
console.log("openai default: ", resolveProviderIdentity({ provider: "openai" }));
console.log("per-provider defaults:", DEFAULT_MODELS);
console.log("openai base url:", DEFAULT_OPENAI_BASE_URL);
// Construction is where a credential is required. A missing key is an *operational*
// failure, thrown as InferenceError — not a model failure recorded on a run.
try {
makeProvider({ provider: "anthropic", apiKeyEnv: "EXAMPLE_KEY_THAT_IS_NOT_SET" });
} catch (error) {
console.log("construction threw:", error instanceof InferenceError);
console.log("message:", error.message);
}
// So defer construction until you actually need to call out. A run served
// entirely from cache never touches this thunk.
let cached;
const getProvider = () => (cached ??= makeProvider({ provider: "mock" }));
const identity = resolveProviderIdentity({ provider: "mock" });
console.log("identity without constructing:", identity);
console.log("provider constructed yet:", cached !== undefined);
console.log("now constructing:", getProvider().modelName());
console.log("provider constructed yet:", cached !== undefined);
anthropic default: { provider: 'anthropic', model: 'claude-sonnet-4-5' }
explicit model: { provider: 'anthropic', model: 'claude-haiku-4-5' }
openai default: { provider: 'openai', model: 'gpt-4o-mini' }
construction threw: true
message: Anthropic provider needs EXAMPLE_KEY_THAT_IS_NOT_SET set (or choose another provider)
identity without constructing: { provider: 'mock', model: 'mock-model' }
provider constructed yet: false
Class How it arrives Examples
Operational thrown as InferenceError missing API key, unknown provider name, unresolved local selector
Model returned on run.error schema validation failed after both attempts, provider API error, timeout

Catch InferenceError at your boundary and translate it into your own error type, or your CLI’s exit-code mapping will not recognise it and a fixable config mistake becomes an unhandled stack trace.

Targets any /chat/completions server — OpenAI, Azure, Ollama, Groq, Together. It prefers strict json_schema mode, which requires rewriting your schema into a restricted subset: every property listed in required, optionality expressed as a null type union, and unsupported keywords dropped. Nulls are stripped back out of the response, so you see the schema you wrote.

If the server rejects response_format, it permanently falls back to json_object with the schema restated in the prompt. Keyless servers are allowed — only api.openai.com requires a key.

Uses your local Claude CLI authentication, so there is no API key to manage. The prompt goes over stdin, never argv — user content routinely exceeds the ~32K Windows command-line limit.

Reports no usage. See the caution above.

Needs makeProviderAsync whenever the model is a selector like auto, because resolving one reads GPU memory. Start at run models locally.

Not just for tests — it is the fastest way to develop against this package before choosing anything. See testing your integration.