Model providers
Configure an OpenAI-compatible endpoint or a local model.
Bi talks to models through OpenAI-compatible APIs, so most services plug straight in.
Add a provider
Fill in:
- Base URL — the endpoint, e.g.
https://api.openai.com/v1 - API Key — your key
- Model — e.g.
gpt-4o - Kind —
openaifor the OpenAI-compatible protocol,ollamafor Ollama’s native API
Each model can also declare its capabilities (vision / tool / reason / think) and context window, used for display and filtering.
Where the key lives
Keys are stored locally in .bi/llm.json, which is in .gitignore. Bi has no server of its own — keys only ever travel to the endpoint you configured. When upgrading from an older .bi/config.json, its api_key is migrated across once.
Local models
Point at a local OpenAI-compatible endpoint to keep content entirely on your machine. Common choices:
- Ollama
- LM Studio
- vLLM
Anything that exposes /v1/chat/completions will work. Ollama additionally has native protocol support (kind: ollama), which lists local models directly and accepts provider-specific parameters such as keep_alive and num_ctx through extra.
Model selection and failover
default takes one of two forms:
"auto"— let the routing policy choose"<providerID>/<modelID>"— pin one specific model
What auto means is worth being precise about, because it is not rotation:
- No rotation — every call takes the first healthy model in the candidate pool, and it stays that way for as long as it keeps working
- Switch only on failure — a connection error, timeout, 401, 429 or 5xx marks the model unhealthy, starts a cooldown, and retries that one call against the next healthy candidate
- Automatic return — once the cooldown expires it re-enters the pool, and being first it comes back naturally
Two limitations to keep in mind:
- Failover only applies in
automode. A pinned model never switches. - No switching once a token has been emitted. Retrying a call that already began streaming would produce a second answer, so the error is raised to the caller instead. The agent re-issues a call each loop iteration, so the next turn still has a chance to switch.
The candidate pool comes from each model’s in_auto flag, or from an explicit ordered auto.pool list. The default policy is priority with 2 retries and a 120-second cooldown.
Context cost
See context management for how long sessions behave.