Documentation
Hetu is a local-first AI knowledge and agent workspace. Each page below documents one feature area: where it lives, how it works, which settings it reads and how to debug it.
New here? Read Getting started → Notes → Agent workspace → Code sessions. For internals, go straight to Architecture.
Pages
| Page | Covers |
|---|---|
| Getting started | Three ways to run, first provider and index, data directory locations |
| Notes | Dual-view editor, selection AI actions, notebooks and tags, version history, share links, trash, export |
| Agent workspace | Sessions and topics, SSE events, input toolbelt, permission and run modes, queue and steering, distill to note |
| Code sessions | Projects, worktrees and branches, Git sync, the seven panels, agent tools, PR flow |
| Knowledge base | Sources and chunking, vector pipeline, coverage, index management, search test, dimensions |
| Graph & Wiki | Entity and relation extraction, types and visualization, Wiki inputs and stages |
| Memories & Dream | Memory fields and scopes, recall triggers, memory graph, consolidation parameters |
| Kanban & automation | Board columns, task fields, agent execution, workflow nodes, background job types, cron, inbox |
| Extensions & MCP | Agent presets and tool policies, skill sources, built-in tools, MCP configuration, embedded apps |
| Models & usage | Provider protocols, model purposes, per-scenario defaults, proxy endpoints, usage metrics, compression |
| Desktop app | Channels, updater and mirrors, tray, sidecar lifecycle, worktree root and cleanup, packaging |
| Data & security | Data directory, backup and migration, PostgreSQL switch, vector storage, security boundaries |
| Architecture | Layers, request path, data model, repository API, SQLite limits, performance practice |
| FAQ | Install failures, model errors, stuck indexing, worktrees, update errors, performance |
| Development | Repository layout, build and verify commands, gotchas, release flow |
Internals
For readers who change code or behaviour: the concrete mechanism, configuration and defaults, guards and degradation, and the debugging path of each module.
| Module | Covers |
|---|---|
| Compression pipeline | Two entry points, three modes, the ten nodes with order and defaults, the four gates on the LLM summary, the not-shorter guard and config merge semantics |
| Agent loop | The six steps of a run, iteration and tool limits, where steering is injected, history compression boundaries, MCP loading, event outlets |
| Permission modes & approvals | What each mode does, the three-layer decision order, rule fields and match scoring, the effect of autopilot, entry-point differences |
| Context & auto compaction | Usage parts and window resolution, the 0.8 trigger, keeping recent messages, summary retries and the 60k cap |
| SSE event protocol | Three frame shapes, the full event list with fields, workflow events, stream differences, adding an event safely |
| Chunking & vector index | LLM vs structural chunking, write order and DELETE+INSERT, vector reuse, dimension resolution and rebuild, degradation paths |
| Semantic retrieval | Note and chunk recall with over-fetch factors, dedup merge, projection, the four consumers, degradation and quality practice |
| Knowledge graph extraction | The single-pass path, JSON contract and type whitelists, normalisation and dedupe, per-note cleanup and restore, entity merge, failure paths |
| Wiki generation | Collection and compression, planning constraints and fallback, concurrency 4 with continuation, overview requirements, set persistence and staleness |
| Memory write & recall | Memory model and scopes, the three write paths, the five scoring components, topK and side effects, storage backends |
| Dream consolidation | Triggers and preconditions, defaults and constraints, greedy clustering and keeper selection, merge bookkeeping, decay and forgetting |
| Task board & automation | Transition whitelist and timestamps, card model and automation predicate, trigger points and gate, comment loops, enrichment, conditional polling |
| Workflow engine | Definition model and JSON storage, nine endpoints, graph validation, ten node executors, run limits and status machine, human review, SSE events |
| Background & scheduled tasks | Unique index and enqueue dedupe, single consumer with a 2-minute sweep, five task branches; scheduled fields, interval/cron scheduling, four executors |
| Inbox & notifications | Notification model with level and category constants, aggregation by CategoryKey, endpoints, state and deletion rules, frontend behaviour |
What problem it solves
| Situation | How Hetu handles it |
|---|---|
| Notes, chat logs and code changes live in separate tools | All three write one database; retrieval, graph, memories and Wiki are built on top of that data |
| AI gives advice, you still copy-paste the work | Code sessions operate in the repository: isolated worktree, line-level diff, checkpoint rollback, commits and PRs |
| Long sessions explode context and cost | Compression pipeline plus usage accounting (input / compressed / output / cached tokens, latency, per-model and per-source split) |
| Worried about data and keys leaving the machine | No account, no telemetry; keys encrypted with DataProtection; requests only go to providers you configure |