Documentation

Hetu is a local-first AI knowledge and agent workspace. Each page below documents one feature area: where it lives, how it works, which settings it reads and how to debug it.

New here? Read Getting started → Notes → Agent workspace → Code sessions. For internals, go straight to Architecture.
Desktop shell (Tauri 2) window · tray · updater · sidecar Browser (dev mode) http://localhost:5174 React 19 frontend pages / components / Zustand / TanStack Query SSE rendering · send queue · panels i18n (zh / en) · light and dark themes ASP.NET Core API REST (ApiResponse / PagedResult) + SSE /scalar/v1 reference · /api/health probe Domain services (Hetu.Core) notes / chat / workspace / knowledge / graph / Wiki / memories repository interfaces · task coordinator · agent loop Model providers OpenAI-compatible · Anthropic API keys encrypted at rest MCP servers stdio · JSON-RPC 2.0 · tools/list + tools/call Git and CLIs worktrees · gh / glab · real PTY External search web search (toggle) SQLite + sqlite-vec (default) PostgreSQL + pgvector
A request travels through the shell or browser → React frontend → ASP.NET Core API → domain services → repositories → the local database; models, MCP, Git and search are external dependencies.

Pages

PageCovers
Getting startedThree ways to run, first provider and index, data directory locations
NotesDual-view editor, selection AI actions, notebooks and tags, version history, share links, trash, export
Agent workspaceSessions and topics, SSE events, input toolbelt, permission and run modes, queue and steering, distill to note
Code sessionsProjects, worktrees and branches, Git sync, the seven panels, agent tools, PR flow
Knowledge baseSources and chunking, vector pipeline, coverage, index management, search test, dimensions
Graph & WikiEntity and relation extraction, types and visualization, Wiki inputs and stages
Memories & DreamMemory fields and scopes, recall triggers, memory graph, consolidation parameters
Kanban & automationBoard columns, task fields, agent execution, workflow nodes, background job types, cron, inbox
Extensions & MCPAgent presets and tool policies, skill sources, built-in tools, MCP configuration, embedded apps
Models & usageProvider protocols, model purposes, per-scenario defaults, proxy endpoints, usage metrics, compression
Desktop appChannels, updater and mirrors, tray, sidecar lifecycle, worktree root and cleanup, packaging
Data & securityData directory, backup and migration, PostgreSQL switch, vector storage, security boundaries
ArchitectureLayers, request path, data model, repository API, SQLite limits, performance practice
FAQInstall failures, model errors, stuck indexing, worktrees, update errors, performance
DevelopmentRepository layout, build and verify commands, gotchas, release flow

Internals

For readers who change code or behaviour: the concrete mechanism, configuration and defaults, guards and degradation, and the debugging path of each module.

ModuleCovers
Compression pipelineTwo entry points, three modes, the ten nodes with order and defaults, the four gates on the LLM summary, the not-shorter guard and config merge semantics
Agent loopThe six steps of a run, iteration and tool limits, where steering is injected, history compression boundaries, MCP loading, event outlets
Permission modes & approvalsWhat each mode does, the three-layer decision order, rule fields and match scoring, the effect of autopilot, entry-point differences
Context & auto compactionUsage parts and window resolution, the 0.8 trigger, keeping recent messages, summary retries and the 60k cap
SSE event protocolThree frame shapes, the full event list with fields, workflow events, stream differences, adding an event safely
Chunking & vector indexLLM vs structural chunking, write order and DELETE+INSERT, vector reuse, dimension resolution and rebuild, degradation paths
Semantic retrievalNote and chunk recall with over-fetch factors, dedup merge, projection, the four consumers, degradation and quality practice
Knowledge graph extractionThe single-pass path, JSON contract and type whitelists, normalisation and dedupe, per-note cleanup and restore, entity merge, failure paths
Wiki generationCollection and compression, planning constraints and fallback, concurrency 4 with continuation, overview requirements, set persistence and staleness
Memory write & recallMemory model and scopes, the three write paths, the five scoring components, topK and side effects, storage backends
Dream consolidationTriggers and preconditions, defaults and constraints, greedy clustering and keeper selection, merge bookkeeping, decay and forgetting
Task board & automationTransition whitelist and timestamps, card model and automation predicate, trigger points and gate, comment loops, enrichment, conditional polling
Workflow engineDefinition model and JSON storage, nine endpoints, graph validation, ten node executors, run limits and status machine, human review, SSE events
Background & scheduled tasksUnique index and enqueue dedupe, single consumer with a 2-minute sweep, five task branches; scheduled fields, interval/cron scheduling, four executors
Inbox & notificationsNotification model with level and category constants, aggregation by CategoryKey, endpoints, state and deletion rules, frontend behaviour

What problem it solves

SituationHow Hetu handles it
Notes, chat logs and code changes live in separate toolsAll three write one database; retrieval, graph, memories and Wiki are built on top of that data
AI gives advice, you still copy-paste the workCode sessions operate in the repository: isolated worktree, line-level diff, checkpoint rollback, commits and PRs
Long sessions explode context and costCompression pipeline plus usage accounting (input / compressed / output / cached tokens, latency, per-model and per-source split)
Worried about data and keys leaving the machineNo account, no telemetry; keys encrypted with DataProtection; requests only go to providers you configure