The compression pipeline shrinks long text before it reaches a model: ten nodes run in a fixed order, algorithms first with an LLM fallback, and one hard guard — if the result is not shorter, the original text is kept. It is the first layer of context cost control, paired with the second layer, session-level auto compaction.
Node order and default enablement come from code; users only choose what stays enabled. New nodes appear in existing configs after an upgrade.
Two entry points
Entry
Behaviour
Used by
CompressAsync
Runs the whole pipeline, including the LLM summary node
Chat / code session history compression
CompressAlgorithmicAsync
Algorithm nodes only (allowLlm=false): no model call even when the mode is llm or hybrid
Single-shot inputs: Wiki generation, graph extraction
Why the distinction? A one-shot long input is compressed once, so paying for an extra model call does not pay off. Chat history is re-sent on every turn, so LLM summarization amortizes.
Three modes
Mode
Meaning
algorithmic (default)
Algorithm nodes only; llm_summary is skipped even when enabled (with a log line explaining why)
llm
Algorithm nodes plus the LLM summary (subject to the length threshold)
hybrid
Same as llm: algorithms strip obvious redundancy first, then the model handles semantic compression
The ten nodes (execution order)
#
Key
Name
What it does
Default
1
structured
Structured folding
Folds long payloads — base64, HTML, JSON, tables — keeping a structural skeleton and samples
on
2
dedup
Duplicate merge
Removes duplicate paragraphs, lines and inline sentences, keeping the first occurrence
on
3
whitespace
Whitespace
Trims redundant spaces, newlines and indentation
on
4
log_dedup
Log dedup
Detects and collapses repeated log patterns
on
5
near_dup
Near duplicates
SimHash similarity so repeated content differing only by timestamp / id is merged
on
6
number_normalize
Number normalization
Replaces numbers with placeholders to reduce variation
on
7
stopwords
Stop word filter
Removes common low-signal words (Chinese and English)
off
8
keywords
Key line extraction
Scores lines by information content and keeps errors, conclusions and paths, folding the rest
on
9
headtail
Head / tail
For very long output, keeps the head and tail and folds the middle (build logs, large files)
on
10
llm_summary
LLM summary
Model-based semantic summary, gated as described below
off
Node order is fixed in code (Order 1–10); configuration only chooses which nodes stay enabled. The order matters: structural noise first, then duplicates, then numeric/key-line reduction, with the LLM summary always last.
Four gates for the LLM summary
Gate
Value
When unmet
Entry allows it
allowLlm=true (CompressAsync)
Single-shot tasks skip it
Mode
llm or hybrid
Skipped under algorithmic
Length threshold
LlmThreshold ≥ 500 characters (configurable)
Short text is skipped to save a call
Model available
The configured summary model, falling back to the default chat model
Neither available: log a warning and keep the original
The default summary prompt asks the model to keep conclusions, numbers, paths and file names, commands, todos and constraints; drop repetition and pleasantries; never invent facts; and answer in the source language.
Guard and degradation
Not shorter → keep the original: if the pipeline output is not shorter than the input, the input is returned (nodes such as number normalization and duplicate annotation can make certain text longer).
Empty input / pipeline off / no enabled nodes: passthrough, no processing.
LLM failure: a missing provider only logs a warning and returns the original; it never blocks the main flow.
Every node logs input/output character counts, and the run logs the final compression ratio.
Configuration storage and upgrade semantics
Item
Details
Storage key
App setting CompressionConfig (JSON)
Master switch
Enabled, off by default — the whole pipeline is skipped when off
Merge strategy
Node set, order and labels come from code; only the user's enable/disable choice is preserved, so nodes added in a release show up in existing configs instead of staying invisible
Historic nodes
Nodes present in a stored config but removed from code are kept at the end (Order + 100) for easier rollback and debugging
Self-inspection
DescribeAsync() reports the nodes that will actually run, the mode, and whether llm_summary is skipped and why
How it differs from session-level auto compaction
Aspect
Compression pipeline (this page)
Session auto compaction
Applies to
A single text (a history slice, a tool output)
The whole message sequence of a session
Triggered by
Explicit calls (before each iteration, before one-shot tasks)
Automatically when usage reaches 80% of the window
Result
Replaces only the text sent to the model
Writes a summary into ContextSummary / ContextSummaryThroughMessageId; original messages are untouched
Knobs
Node toggles, mode, threshold, summary model and prompt
Recent-message count, summary model, auto ratio
Tuning advice grounded in the implementation
Turn the master switch on first — the pipeline ships disabled; enabling it with default nodes already helps.
Long logs / build output: keep structured, log_dedup and headtail on; those three handle CLI noise best.
De-noise without touching semantics: stay in algorithmic mode and leave llm_summary off.
Long-session cost: switch to llm / hybrid, keep the threshold at 500+ characters, and pick a cheap summary model.
Enable with care:number_normalize replaces digits and stopwords deletes words — turn both off for tasks that depend on exact values or wording.
Troubleshooting: grep the backend log for [Compression]; it lists per-node before/after sizes, skip reasons and the final ratio.