Wiki generation

Wiki generation is a long job: collect material → compress → plan the outline → generate pages concurrently → write the overview → persist the whole set. The difficulty is not the prose of one page but the engineering chain around it: 4–10 pages per run, each possibly continued to completion, and partial failures that must not sink the set.

Collect material project dir · code context 10 semantic hits / page Compress BaseContext rawCodeContext Plan 2–4 chapters · 1–3 pages each 4–10 pages · maxTokens 3072 Generate topic pages concurrency 4 · 8192 tokens / page continued up to 4 rounds Overview page must include "Overall architecture" + Mermaid ends with a documentation index Persist as one set shared SetId · overview SortOrder = 0 · topics = plan order + 1 job status: 0 queued / 1 running / 2 done / 3 failed
The stage order is fixed: planning precedes generation and the overview is written last; a failed page lands in the warning set instead of aborting the run.

Triggers and job model

EndpointPurpose
POST /api/wiki/projects/{projectId}/generateEnqueue a full set with { modelId? }
POST /api/wiki/{id}/regenerateRegenerate one page, reusing its planned Brief
GET /api/wiki · GET /api/wiki/sets · GET /api/wiki/{id}Page list / set list / single page, all filterable by projectId
GET /api/wiki/jobs · GET /api/wiki/jobs/{id}Job list and detail (stage, progress, errors, warnings)
GET /api/wiki/sets/{setId}/exportZip export: one Markdown per page (00-…md, 01-…md) plus a README.md index
DELETE /api/wiki/{id} · DELETE /api/wiki/sets/{setId}Delete a page or a whole set

The background task type is WikiGenerate and the model id travels in the task metadata (Guid.TryParse(item.Metadata)). If a project already has a queued or running job, EnqueueGenerateAsync returns that job instead of queueing another. Entering the frontend at /wiki?project=…&generate=1 triggers one generation and clears the parameter; the model selector lists non-embedding models only, and an empty selection means the default completion model.

Material collection and compression

  • The collector turns a managed project into material: a local directory or a remote work project, base context and code context.
  • When the code semantic index is available it supplies fragments per page — roughly 10 hits (SemanticHitsPerPage = 10); otherwise it falls back to sampling source files.
  • Before the model call, both BaseContext and rawCodeContext pass through the compression pipeline (so Enabled, Mode and node configuration all matter).
  • Collection failures abort immediately and never reach the background queue: missing local directory, failed remote connection, unknown work project and empty directory each have their own message.

Planning

ItemValue
Return shape{ chapters: [{ title, pages: [{ title, brief }] }] }
Size constraints2–4 chapters, 1–3 pages each, 4–10 pages total (cap constant MaxModulePages = 10)
Plan token budget3072
Planning failureThe exception is caught and the run falls back to four default pages (project overview & stack / directory layout / core modules / quick start & configuration) without blocking generation

Brief is persisted and becomes the writing brief for that page; single-page regeneration reuses it, so a regeneration stays aligned with the original outline.

Generation

MechanismValue / behaviour
ConcurrencyMaxPageConcurrency = 4
Per-page output capContentMaxTokens = 8192
Continue triggerFinish reason length / max_tokens, or an unbalanced code fence (odd fence count)
Continue limitMaxContinueRounds = 4; the continue prompt carries everything generated so far
Continue guardContinuations containing meta phrases such as "existing content", "not provided" or "cannot determine" are dropped
Page retryPageRetryCount = 1 with an 800ms backoff
Body requirementsMarkdown, Chinese, strictly grounded in the material, Mermaid diagrams for structure/flow/module relations, text code blocks for trees, inline code for paths and types
Cross-page contextThe prompt lists sibling pages in the same set so pages can reference each other

The overview page is generated separately and must contain a project intro, key features, tech stack, an ## Overall architecture section (with a Mermaid graph / flowchart) and a final documentation index listing chapters verbatim. If it fails, a plain-text overview (intro + documentation index) is written instead.

Persistence, sets and staleness

Field / ruleNotes
WikiDocumentProjectId · SetId · SortOrder · Title · Chapter · Brief · Content
WikiGenerationJobStatus (0 queued / 1 running / 2 done / 3 failed) · Stage · Progress · TotalPages · DonePages · ErrorMessage · WarningMessage · ModelId · SetId · CompletedAt
Write orderTopic pages are generated and collected first, then the overview, and only then does a single write happen: overview SortOrder = 0, topic pages = plan order + 1
Set aggregationGetAllAsync groups by SetId and sorts by SortOrder, so the overview is always first
StalenessGetSetsAsync asks the context collector to compute staleness, filling IsStale and StaleFileCount (project files changed → suggest regenerating)

Failure and degradation

SituationBehaviour
No model availableJob fails with wiki.modelNotConfigured
Empty page contentWarning + retry per PageRetryCount; still empty → the page fails and lands in failedPages
Some pages failThe set is still persisted; failures are reported in WarningMessage (orange banner in the UI)
All topic pages failThrows wiki.allPagesFailed
Overview failsFalls back to a plain-text overview (intro + documentation index)
Single-page regeneration failsReturns wiki.regenerateFailed
Background job interrupted / cancelledSame as graph jobs: shutdown requeues, explicit cancel marks failure

Observability

  • Frontend job bar: stage + donePages/totalPages + percentage; failures show a red bar with errorMessage, partial failures an orange notice with warningMessage.
  • Stage labels come from localisation keys: wiki.stageQueued, stageStarting, stageCollecting, stagePlanning, stageGeneratingPages, stageOverview, stageCompleted.
  • Usage source Wiki: input tokens, compressed tokens, latency.
  • There is no per-page timing field: only job-level stage and progress, so page-level triage means reading logs.

Practice notes

  • Plan quality sets the ceiling: a vague plan makes every page vague. To change structure, change the planning prompt and its size constraints first.
  • Continuation completes rather than rewrites: it resends everything generated so far, so a smaller ContentMaxTokens means more rounds and a bigger context bill.
  • Partial failure is not an incident: the set stays usable and failing pages can be regenerated individually.
  • Exported filenames start with an index (00-, 01-) so file-manager order matches the in-app order.
  • Changing compression settings changes the material sent to the model, which changes Wiki output; factor that in when content quality regresses.