Semantic retrieval
Retrieval runs on two paths: approximate search in sqlite-vec / pgvector (fast) and in-memory cosine similarity (fallback). Both optimise for "never break the flow" before "be exact".
Two recall paths
| Path | Implementation |
| Note retrieval | Whole-note vectors first (vec_note_embeddings), then chunk vectors (vec_chunk_embeddings); each path takes topK × 3 candidates |
| Code retrieval | Dedicated vec_work_code_chunks; on SQLite the candidate count is clamp(topK × 5, 20, 200) to absorb per-project filtering, on PostgreSQL ORDER BY Vector <=> @query LIMIT topK |
| Merge | The same knowledge item can be hit by both paths: deduplicated by id, keeping the first occurrence, then truncated to topK |
| Fallback | Missing extension, missing virtual table, dimension mismatch or no pgvector → in-memory cosine similarity |
Result projection
| Source | Shown |
| Chunk hit | The chunk Summary when present, otherwise a trimmed content snippet |
| Note hit | The first 120 characters of the note |
| Shape | Id / Title / ContentSnippet / UpdatedAt — rendered directly as the hit list, clickable to the source |
Who consumes it
- The "Knowledge base" toggle in chat: hits are injected as context and listed in the UI.
- The search test on the knowledge base page: run a query, adjust
TopK (default 10), inspect highlighted chunks — the fastest index-quality check.
- Wiki generation: code retrieval supplies related snippets per module page.
- Model-initiated search: retrieval tools let the model decide when to search, constrained by permission mode and tool allowlists.
Failure and degradation
| Case | Behaviour |
Empty query / topK ≤ 0 | Returns an empty result immediately |
| Vector store unavailable | Falls back to in-memory computation (slower, still correct) |
| Count endpoint fails | Returns -1 for "unknown" instead of blocking the page |
| Cancellation | Propagates; never swallowed |
Quality practice
- Bad hits? Check index management first: is the item indexed, and does its chunk count make sense (a single chunk usually means very short or merged content)?
- Over-fetching then truncating is deliberate (×3 for notes, ×5 for code): deduplication and business filters consume candidates.
- The merge does not re-rank by score: order reflects each path's own similarity, avoiding jitter from incomparable scores.
- Chunk summary quality directly shapes the readability of hit snippets — that is the main payoff of LLM chunking.