Knowledge base

Notes, files and URLs are chunked and embedded so chat and semantic search can retrieve them. Entry: Knowledge base (route /knowledge-base).

Notes Files (upload) URLs (fetch text) Chunking split by structure, record index and method Embedding model assigned to the embedding purpose Vectors and metadata SQLite: vec0 virtual table PostgreSQL: pgvector column dimensions = Embedding:Dimensions Retrieval consumers chat toolbelt · search test · Wiki generation · model-invoked search tools
Indexing pipeline and consumers: chunk → embed → write vectors and metadata; chat, the search test, Wiki generation and model tools all recall from the same vectors.

Sources

SourceHow to addNotes
Notes“Generate index” in the editor toolbar, or the batch action on the knowledge base pageChunked along Markdown structure
FilesUpload on the knowledge base pageFile name, size and MIME are stored; text is extracted then chunked
URLsPaste a URLPage text (and title) is fetched, then chunked

Indexing pipeline

  1. Chunk: content is split by structure; each chunk stores its index, method and text.
  2. Embed: the model assigned to the embedding purpose produces vectors, written to the vector store (sqlite-vec vec0 virtual table on SQLite, pgvector column on PostgreSQL).
  3. Record metadata: model name and dimensions per chunk power the status view and validation.
  4. Task it: the whole run is a background task (GenerateEmbedding / GenerateKnowledgeItemEmbedding) with duration and failure reason visible on the Tasks page.

The three tabs

TabContent
OverviewTotal items, indexed items, coverage, vector dimensions, split by type (notes / files / URLs), running task count
Index managementPer-item status: chunk count, model, dimensions, last indexed time, whether a task is running; re-index one item or batch-generate indexes for everything unindexed
Search testRun a query against the vector index, adjust Top-K, inspect highlighted chunks — the fastest way to judge index quality
Knowledge base: index coverage, vector dimensions and index management
Click for full size

Index-related tables

TableColumns
KnowledgeItemsType (Note / File / Url), Title, Content, SourceUrl, FileName, FileSize, MimeType, NoteId, IsDeleted
NoteChunksKnowledgeItemId, ChunkIndex, ChunkMethod, Content, Summary
NoteChunkEmbeddingsChunkId, Model, Dimensions, UpdatedAt, Embedding (on SQLite the vector lives in vec_chunk_embeddings)

The status endpoint returns totalItems / indexedItems / unindexedItems / noteCount / fileCount / urlCount / hasEmbeddingProvider / dimensions / runningTaskCount; each list entry carries chunkCount, embeddingModel, embeddingDimensions, embeddingUpdatedAt and hasRunningTask.

How retrieval is used

  • The “Knowledge base” toggle above the input injects recall results as context and shows the hit list.
  • The model can also query it through tools (note search, knowledge retrieval), constrained by permission mode and tool allowlists.
  • Wiki generation can layer semantic search on top of repository material.

Configuration and limits

ItemDetails
Embedding:DimensionsVector dimensions, default 1536; must match the model or SQLite writes fail
Changing modelsSwitching embedding models requires re-indexing (old vectors have the wrong dimensions)
AnthropicNo public embedding API — use an OpenAI-compatible provider
Deletion semanticsMoving a note to trash soft-deletes the knowledge item but keeps chunks and vectors, so restoring it stays indexed

Performance notes

  • Status is aggregated in the database (counts per type with GROUP BY, DISTINCT indexed items, per-item chunk summaries) instead of loading all vector metadata.
  • The page only polls while there are unindexed items or running tasks, and the status endpoints carry a short response cache.