---
title: "Model providers"
description: "Configure an OpenAI-compatible endpoint or a local model."
---

Bi talks to models through OpenAI-compatible APIs, so most services plug straight in.

## Add a provider

Fill in:

- **Base URL** — the endpoint, e.g. `https://api.openai.com/v1`
- **API Key** — your key
- **Model** — e.g. `gpt-4o`
- **Kind** — `openai` for the OpenAI-compatible protocol, `ollama` for Ollama's native API

Each model can also declare its capabilities (`vision` / `tool` / `reason` / `think`) and context window, used for display and filtering.

## Where the key lives

Keys are stored locally in `.bi/llm.json`, which is in `.gitignore`. **Bi has no server of its own** — keys only ever travel to the endpoint you configured. When upgrading from an older `.bi/config.json`, its `api_key` is migrated across once.

## Local models

Point at a local OpenAI-compatible endpoint to keep content entirely on your machine. Common choices:

- Ollama
- LM Studio
- vLLM

Anything that exposes `/v1/chat/completions` will work. Ollama additionally has native protocol support (`kind: ollama`), which lists local models directly and accepts provider-specific parameters such as `keep_alive` and `num_ctx` through `extra`.

## Model selection and failover

`default` takes one of two forms:

- `"auto"` — let the routing policy choose
- `"<providerID>/<modelID>"` — pin one specific model

What `auto` means is worth being precise about, because it is **not** rotation:

- **No rotation** — every call takes the first healthy model in the candidate pool, and it stays that way for as long as it keeps working
- **Switch only on failure** — a connection error, timeout, 401, 429 or 5xx marks the model unhealthy, starts a cooldown, and retries that one call against the next healthy candidate
- **Automatic return** — once the cooldown expires it re-enters the pool, and being first it comes back naturally

Two limitations to keep in mind:

1. **Failover only applies in `auto` mode.** A pinned model never switches.
2. **No switching once a token has been emitted.** Retrying a call that already began streaming would produce a second answer, so the error is raised to the caller instead. The agent re-issues a call each loop iteration, so the next turn still has a chance to switch.

The candidate pool comes from each model's `in_auto` flag, or from an explicit ordered `auto.pool` list. The default policy is `priority` with 2 retries and a 120-second cooldown.

## Context cost

See [context management](/en/docs/advanced/context) for how long sessions behave.
