Compression pipeline

The compression pipeline shrinks long text before it reaches a model: ten nodes run in a fixed order, algorithms first with an LLM fallback, and one hard guard — if the result is not shorter, the original text is kept. It is the first layer of context cost control, paired with the second layer, session-level auto compaction.

Input text history / tool output structured folding base64 / HTML / JSON / tables dedup paragraph / line / sentence level whitespace spaces, newlines, indentation log_dedup collapse repeated log patterns near_dup SimHash similarity number_normalize numbers to placeholders keywords keep high-signal lines headtail keep head and tail stopwords off by default LLM summary only in llm / hybrid mode and length >= threshold (500) Guard not shorter → keep original
Node order and default enablement come from code; users only choose what stays enabled. New nodes appear in existing configs after an upgrade.

Two entry points

EntryBehaviourUsed by
CompressAsyncRuns the whole pipeline, including the LLM summary nodeChat / code session history compression
CompressAlgorithmicAsyncAlgorithm nodes only (allowLlm=false): no model call even when the mode is llm or hybridSingle-shot inputs: Wiki generation, graph extraction
Why the distinction? A one-shot long input is compressed once, so paying for an extra model call does not pay off. Chat history is re-sent on every turn, so LLM summarization amortizes.

Three modes

ModeMeaning
algorithmic (default)Algorithm nodes only; llm_summary is skipped even when enabled (with a log line explaining why)
llmAlgorithm nodes plus the LLM summary (subject to the length threshold)
hybridSame as llm: algorithms strip obvious redundancy first, then the model handles semantic compression

The ten nodes (execution order)

#KeyNameWhat it doesDefault
1structuredStructured foldingFolds long payloads — base64, HTML, JSON, tables — keeping a structural skeleton and sampleson
2dedupDuplicate mergeRemoves duplicate paragraphs, lines and inline sentences, keeping the first occurrenceon
3whitespaceWhitespaceTrims redundant spaces, newlines and indentationon
4log_dedupLog dedupDetects and collapses repeated log patternson
5near_dupNear duplicatesSimHash similarity so repeated content differing only by timestamp / id is mergedon
6number_normalizeNumber normalizationReplaces numbers with placeholders to reduce variationon
7stopwordsStop word filterRemoves common low-signal words (Chinese and English)off
8keywordsKey line extractionScores lines by information content and keeps errors, conclusions and paths, folding the reston
9headtailHead / tailFor very long output, keeps the head and tail and folds the middle (build logs, large files)on
10llm_summaryLLM summaryModel-based semantic summary, gated as described belowoff
Node order is fixed in code (Order 1–10); configuration only chooses which nodes stay enabled. The order matters: structural noise first, then duplicates, then numeric/key-line reduction, with the LLM summary always last.

Four gates for the LLM summary

GateValueWhen unmet
Entry allows itallowLlm=true (CompressAsync)Single-shot tasks skip it
Modellm or hybridSkipped under algorithmic
Length thresholdLlmThreshold ≥ 500 characters (configurable)Short text is skipped to save a call
Model availableThe configured summary model, falling back to the default chat modelNeither available: log a warning and keep the original

The default summary prompt asks the model to keep conclusions, numbers, paths and file names, commands, todos and constraints; drop repetition and pleasantries; never invent facts; and answer in the source language.

Guard and degradation

  • Not shorter → keep the original: if the pipeline output is not shorter than the input, the input is returned (nodes such as number normalization and duplicate annotation can make certain text longer).
  • Empty input / pipeline off / no enabled nodes: passthrough, no processing.
  • LLM failure: a missing provider only logs a warning and returns the original; it never blocks the main flow.
  • Every node logs input/output character counts, and the run logs the final compression ratio.

Configuration storage and upgrade semantics

ItemDetails
Storage keyApp setting CompressionConfig (JSON)
Master switchEnabled, off by default — the whole pipeline is skipped when off
Merge strategyNode set, order and labels come from code; only the user's enable/disable choice is preserved, so nodes added in a release show up in existing configs instead of staying invisible
Historic nodesNodes present in a stored config but removed from code are kept at the end (Order + 100) for easier rollback and debugging
Self-inspectionDescribeAsync() reports the nodes that will actually run, the mode, and whether llm_summary is skipped and why

How it differs from session-level auto compaction

AspectCompression pipeline (this page)Session auto compaction
Applies toA single text (a history slice, a tool output)The whole message sequence of a session
Triggered byExplicit calls (before each iteration, before one-shot tasks)Automatically when usage reaches 80% of the window
ResultReplaces only the text sent to the modelWrites a summary into ContextSummary / ContextSummaryThroughMessageId; original messages are untouched
KnobsNode toggles, mode, threshold, summary model and promptRecent-message count, summary model, auto ratio

Tuning advice grounded in the implementation

  • Turn the master switch on first — the pipeline ships disabled; enabling it with default nodes already helps.
  • Long logs / build output: keep structured, log_dedup and headtail on; those three handle CLI noise best.
  • De-noise without touching semantics: stay in algorithmic mode and leave llm_summary off.
  • Long-session cost: switch to llm / hybrid, keep the threshold at 500+ characters, and pick a cheap summary model.
  • Enable with care: number_normalize replaces digits and stopwords deletes words — turn both off for tasks that depend on exact values or wording.
  • Troubleshooting: grep the backend log for [Compression]; it lists per-node before/after sizes, skip reasons and the final ratio.