Notes, files and URLs are chunked and embedded so chat and semantic search can retrieve them. Entry: Knowledge base (route /knowledge-base).
Indexing pipeline and consumers: chunk → embed → write vectors and metadata; chat, the search test, Wiki generation and model tools all recall from the same vectors.
Sources
Source
How to add
Notes
Notes
“Generate index” in the editor toolbar, or the batch action on the knowledge base page
Chunked along Markdown structure
Files
Upload on the knowledge base page
File name, size and MIME are stored; text is extracted then chunked
URLs
Paste a URL
Page text (and title) is fetched, then chunked
Indexing pipeline
Chunk: content is split by structure; each chunk stores its index, method and text.
Embed: the model assigned to the embedding purpose produces vectors, written to the vector store (sqlite-vecvec0 virtual table on SQLite, pgvector column on PostgreSQL).
Record metadata: model name and dimensions per chunk power the status view and validation.
Task it: the whole run is a background task (GenerateEmbedding / GenerateKnowledgeItemEmbedding) with duration and failure reason visible on the Tasks page.
The three tabs
Tab
Content
Overview
Total items, indexed items, coverage, vector dimensions, split by type (notes / files / URLs), running task count
Index management
Per-item status: chunk count, model, dimensions, last indexed time, whether a task is running; re-index one item or batch-generate indexes for everything unindexed
Search test
Run a query against the vector index, adjust Top-K, inspect highlighted chunks — the fastest way to judge index quality
ChunkId, Model, Dimensions, UpdatedAt, Embedding (on SQLite the vector lives in vec_chunk_embeddings)
The status endpoint returns totalItems / indexedItems / unindexedItems / noteCount / fileCount / urlCount / hasEmbeddingProvider / dimensions / runningTaskCount; each list entry carries chunkCount, embeddingModel, embeddingDimensions, embeddingUpdatedAt and hasRunningTask.
How retrieval is used
The “Knowledge base” toggle above the input injects recall results as context and shows the hit list.
The model can also query it through tools (note search, knowledge retrieval), constrained by permission mode and tool allowlists.
Wiki generation can layer semantic search on top of repository material.
Configuration and limits
Item
Details
Embedding:Dimensions
Vector dimensions, default 1536; must match the model or SQLite writes fail
Changing models
Switching embedding models requires re-indexing (old vectors have the wrong dimensions)
Anthropic
No public embedding API — use an OpenAI-compatible provider
Deletion semantics
Moving a note to trash soft-deletes the knowledge item but keeps chunks and vectors, so restoring it stays indexed
Performance notes
Status is aggregated in the database (counts per type with GROUP BY, DISTINCT indexed items, per-item chunk summaries) instead of loading all vector metadata.
The page only polls while there are unindexed items or running tasks, and the status endpoints carry a short response cache.