---
title: "Context management"
description: "How long sessions carry turns forward, and how the request prefix stays stable."
---

Every request builds a fresh agent, so the model has no memory of what was said earlier in the session. Bi repacks the stored conversation into a model message sequence on the server, which is what makes multi-turn conversation work at all.

## Cross-turn packing

There is one rule: **the most recent turns are rebuilt in full, older turns are summarized.**

- The **last `recent_turns_full` turns** (default 6) are rebuilt completely — `user`, `assistant` and `tool` messages alike. Tool blocks stored on an assistant row are restored into `assistant.toolCalls` plus the following `role: "tool"` messages, so `tool_call` / `tool` pairing stays intact. That pairing is a hard requirement of OpenAI-compatible APIs, and a break will get the turn rejected by the gateway.
- **Older turns** collapse into a summary inside `<earlier_conversation_summary>`, holding the user's key points, the assistant's conclusions and the list of changed files. The summary is **assembled purely by rules, with no model call**: no extra cost, no added latency, and reproducible output.
- History beyond **40 turns is dropped** rather than reaching further back.

**Compaction is counted in turns, not tokens.** Many turns does not necessarily mean you will hit the ceiling, and a single very long turn can overflow early. To control it precisely, adjust `recent_turns_full`: raise it to keep more verbatim text, lower it to start summarizing sooner.

An interrupted or timed-out tool can produce an empty result string. In that case Bi substitutes a placeholder, `(no output: tool interrupted or timed out)`. It tells the model the step produced nothing, and avoids an empty `content` tripping strict gateways.

## Prefix stability

Summaries are fingerprinted on "count of older turns + id of the last message". As long as the older history is unchanged — which is the normal case when appending to a conversation — the summary is frozen and reused, keeping the request prefix byte-for-byte identical.

There is a slack layer on top: the freeze point only advances once the fully-injected portion exceeds `recent_turns_full + 2` turns. The purpose is to space out prefix invalidation. Without it, every added turn would shift the summary position and the cache would essentially never hit.

**This exists to serve the model provider's prompt cache.** The cache lives on the provider's side; Bi stores nothing. What Bi does is keep the head bytes of every request identical so the provider can keep hitting it.

## Tuning

`recent_turns_full` lives in `.bi/config.json`. A value ≤ 0 or a missing key falls back to the default of 6.

- Short tasks, API calls, running scripts: lowering it, or leaving it alone, costs nothing — the summary barely participates.
- Long-document analysis, iterative work: raise it to 10–15 to keep more verbatim text and reduce the drift that aggressive summarization introduces.
- Expensive contexts: lower it to save tokens, accepting that older detail is summarized away sooner.
