Context & auto compaction

This layer answers two questions: how much context is in use, and when history must be compacted. Its division of labour with the compression pipeline is explicit: the pipeline compresses a single text, while this compresses the message sequence of a session and persists the summary on the session.

How usage is computed

PartMethod
HistoryCounts only messages not yet covered by a summary: everything up to ContextSummaryThroughMessageId is excluded
SummaryThe stored summary text converted to tokens
System promptA derived value: the last reported PromptTokens minus history and summary; when unavailable it falls back to an empirical value (1500 chat / 3500 code sessions)
WindowThe request's contextWindow first, then the model's ContextWindow, otherwise 128000
An empty session reports 0 usage but still returns all three parts (system / history / summary), so the UI never jumps between "two parts" and "three parts".

When auto compaction triggers

ConstantDefaultMeaning
AutoCompactRatio0.8Compaction only triggers at ≥ 80% of the window; otherwise null is returned and the caller ignores it
DefaultKeepRecent6The newest 6 messages are never summarized
MaxTranscriptChars60000At most the last 60k characters of history are sent to the summarizer, so the summary prompt cannot blow up
MaxSummaryAttempts3Up to three attempts: an empty or too-short summary retries immediately (no fixed delay)

What compaction actually changes

  1. History starts after the last summarized message, keeps the newest KeepRecent entries, and hands the rest to the summary model.
  2. The summary is written back to ContextSummary, and ContextSummaryThroughMessageId advances to the last included message id.
  3. Original messages are untouched: compaction only changes what is sent to the model next; history stays complete for review and export.
  4. A notice(kind=compacted) event is emitted and the UI shows a notice above the input.

Division of labour with the compression pipeline

AspectCompression pipelineAuto compaction (this page)
Applies toOne text (a history slice, a tool output)The whole message sequence of a session
Triggered byExplicit calls (before an iteration, before one-shot tasks)Automatically at ≥ 80% window usage
PersistenceReplaces only the text sent this turnSummary persisted in ContextSummary + ThroughMessageId
On failureFalls back to the original text, no side effectsSimply does not inject a summary; the flow continues

Failure and degradation

  • Missing topic / session: returns notFound.
  • Nothing compactable (toSummarize.Count == 0): returns context.noCompressibleHistory.
  • Summary model unavailable: returns null and writes no summary.
  • The summary call throws: a warning is logged and null is returned — auto compaction never blocks the conversation.
  • When it fails the UI shows no notice; usage keeps growing and compaction can be triggered manually from the context panel.

Observability

  • Log prefix [Context]: "attempt x/y returned nothing", "summary too short", "compaction done: N messages → M characters (about before → after tokens)".
  • UI: the context panel in chat, the usage ring and compaction notice in code sessions.
  • Usage: first-iteration input characters and post-compression characters are recorded so the savings are measurable.

Practical advice

  • If a long session "forgets" earlier content, check usage first: at 80% compaction is by design, not a bug.
  • To keep more original wording, raise KeepRecent; the cost is faster context growth.
  • Poor summaries: point the setting at a better summary model, or lower MaxTranscriptChars so the summary stays focused.
  • Need the full record (audit, post-mortem)? Messages are never modified — read or export the session history.