Semantic retrieval

Retrieval runs on two paths: approximate search in sqlite-vec / pgvector (fast) and in-memory cosine similarity (fallback). Both optimise for "never break the flow" before "be exact".

Two recall paths

PathImplementation
Note retrievalWhole-note vectors first (vec_note_embeddings), then chunk vectors (vec_chunk_embeddings); each path takes topK × 3 candidates
Code retrievalDedicated vec_work_code_chunks; on SQLite the candidate count is clamp(topK × 5, 20, 200) to absorb per-project filtering, on PostgreSQL ORDER BY Vector <=> @query LIMIT topK
MergeThe same knowledge item can be hit by both paths: deduplicated by id, keeping the first occurrence, then truncated to topK
FallbackMissing extension, missing virtual table, dimension mismatch or no pgvector → in-memory cosine similarity

Result projection

SourceShown
Chunk hitThe chunk Summary when present, otherwise a trimmed content snippet
Note hitThe first 120 characters of the note
ShapeId / Title / ContentSnippet / UpdatedAt — rendered directly as the hit list, clickable to the source

Who consumes it

  • The "Knowledge base" toggle in chat: hits are injected as context and listed in the UI.
  • The search test on the knowledge base page: run a query, adjust TopK (default 10), inspect highlighted chunks — the fastest index-quality check.
  • Wiki generation: code retrieval supplies related snippets per module page.
  • Model-initiated search: retrieval tools let the model decide when to search, constrained by permission mode and tool allowlists.

Failure and degradation

CaseBehaviour
Empty query / topK ≤ 0Returns an empty result immediately
Vector store unavailableFalls back to in-memory computation (slower, still correct)
Count endpoint failsReturns -1 for "unknown" instead of blocking the page
CancellationPropagates; never swallowed

Quality practice

  • Bad hits? Check index management first: is the item indexed, and does its chunk count make sense (a single chunk usually means very short or merged content)?
  • Over-fetching then truncating is deliberate (×3 for notes, ×5 for code): deduplication and business filters consume candidates.
  • The merge does not re-rank by score: order reflects each path's own similarity, avoiding jitter from incomparable scores.
  • Chunk summary quality directly shapes the readability of hit snippets — that is the main payoff of LLM chunking.