Context & auto compaction
This layer answers two questions: how much context is in use, and when history must be compacted. Its division of labour with the compression pipeline is explicit: the pipeline compresses a single text, while this compresses the message sequence of a session and persists the summary on the session.
How usage is computed
| Part | Method |
| History | Counts only messages not yet covered by a summary: everything up to ContextSummaryThroughMessageId is excluded |
| Summary | The stored summary text converted to tokens |
| System prompt | A derived value: the last reported PromptTokens minus history and summary; when unavailable it falls back to an empirical value (1500 chat / 3500 code sessions) |
| Window | The request's contextWindow first, then the model's ContextWindow, otherwise 128000 |
An empty session reports 0 usage but still returns all three parts (system / history / summary), so the UI never jumps between "two parts" and "three parts".
When auto compaction triggers
| Constant | Default | Meaning |
AutoCompactRatio | 0.8 | Compaction only triggers at ≥ 80% of the window; otherwise null is returned and the caller ignores it |
DefaultKeepRecent | 6 | The newest 6 messages are never summarized |
MaxTranscriptChars | 60000 | At most the last 60k characters of history are sent to the summarizer, so the summary prompt cannot blow up |
MaxSummaryAttempts | 3 | Up to three attempts: an empty or too-short summary retries immediately (no fixed delay) |
What compaction actually changes
- History starts after the last summarized message, keeps the newest
KeepRecent entries, and hands the rest to the summary model.
- The summary is written back to
ContextSummary, and ContextSummaryThroughMessageId advances to the last included message id.
- Original messages are untouched: compaction only changes what is sent to the model next; history stays complete for review and export.
- A
notice(kind=compacted) event is emitted and the UI shows a notice above the input.
Division of labour with the compression pipeline
| Aspect | Compression pipeline | Auto compaction (this page) |
| Applies to | One text (a history slice, a tool output) | The whole message sequence of a session |
| Triggered by | Explicit calls (before an iteration, before one-shot tasks) | Automatically at ≥ 80% window usage |
| Persistence | Replaces only the text sent this turn | Summary persisted in ContextSummary + ThroughMessageId |
| On failure | Falls back to the original text, no side effects | Simply does not inject a summary; the flow continues |
Failure and degradation
- Missing topic / session: returns
notFound.
- Nothing compactable (
toSummarize.Count == 0): returns context.noCompressibleHistory.
- Summary model unavailable: returns null and writes no summary.
- The summary call throws: a warning is logged and null is returned — auto compaction never blocks the conversation.
- When it fails the UI shows no notice; usage keeps growing and compaction can be triggered manually from the context panel.
Observability
- Log prefix
[Context]: "attempt x/y returned nothing", "summary too short", "compaction done: N messages → M characters (about before → after tokens)".
- UI: the context panel in chat, the usage ring and compaction notice in code sessions.
- Usage: first-iteration input characters and post-compression characters are recorded so the savings are measurable.
Practical advice
- If a long session "forgets" earlier content, check usage first: at 80% compaction is by design, not a bug.
- To keep more original wording, raise
KeepRecent; the cost is faster context growth.
- Poor summaries: point the setting at a better summary model, or lower
MaxTranscriptChars so the summary stays focused.
- Need the full record (audit, post-mortem)? Messages are never modified — read or export the session history.