Bi

Model providers

Configure an OpenAI-compatible endpoint or a local model.

Raw markdown

Bi talks to models through OpenAI-compatible APIs, so most services plug straight in.

Add a provider

Fill in:

  • Base URL — the endpoint, e.g. https://api.openai.com/v1
  • API Key — your key
  • Model — e.g. gpt-4o
  • Kind — openai for the OpenAI-compatible protocol, ollama for Ollama’s native API

Each model can also declare its capabilities (vision / tool / reason / think) and context window, used for display and filtering.

Where the key lives

Keys are stored locally in .bi/llm.json, which is in .gitignore. Bi has no server of its own — keys only ever travel to the endpoint you configured. When upgrading from an older .bi/config.json, its api_key is migrated across once.

Local models

Point at a local OpenAI-compatible endpoint to keep content entirely on your machine. Common choices:

  • Ollama
  • LM Studio
  • vLLM

Anything that exposes /v1/chat/completions will work. Ollama additionally has native protocol support (kind: ollama), which lists local models directly and accepts provider-specific parameters such as keep_alive and num_ctx through extra.

Model selection and failover

default takes one of two forms:

  • "auto" — let the routing policy choose
  • "<providerID>/<modelID>" — pin one specific model

What auto means is worth being precise about, because it is not rotation:

  • No rotation — every call takes the first healthy model in the candidate pool, and it stays that way for as long as it keeps working
  • Switch only on failure — a connection error, timeout, 401, 429 or 5xx marks the model unhealthy, starts a cooldown, and retries that one call against the next healthy candidate
  • Automatic return — once the cooldown expires it re-enters the pool, and being first it comes back naturally

Two limitations to keep in mind:

  1. Failover only applies in auto mode. A pinned model never switches.
  2. No switching once a token has been emitted. Retrying a call that already began streaming would produce a second answer, so the error is raised to the caller instead. The agent re-issues a call each loop iteration, so the next turn still has a chance to switch.

The candidate pool comes from each model’s in_auto flag, or from an explicit ordered auto.pool list. The default policy is priority with 2 retries and a 120-second cooldown.

Context cost

See context management for how long sessions behave.