Knowledge graph extraction

Graph extraction turns one note into nodes and edges in a single-pass job: run algorithmic compression to shorten the input, ask the model for JSON only, then normalise names, dedupe by name, dedupe relations by (source, target, type) and persist. There is no multi-turn conversation and no tool calling, so the design stays repeatable, cleanable and locally repairable.

Note text one note or a batch of ids Algorithmic compression CompressAlgorithmic Graph / chat model JSON output only Normalise & dedupe trim + lowercase · name · triple Persist GraphEntity / GraphRelation Confidence = 0.8 (extracted) Outlets Canvas: entities with ≥1 relation and edges with both ends present SSE: meta / entities / relations / done
The path stays single-pass: compress, structured output, normalise and dedupe, persist; both the canvas and the SSE stream read the same result.

Entry points

EntryBehaviour
POST /api/graph/extract/{noteId}Extract one note synchronously, no request body
POST /api/graph/extract/batchSynchronous batch with { noteIds: [] }; fits small selections
POST /api/graph/extract/batch-queueEnqueues background jobs, after de-duplicating NoteIds; use this for large batches
Background task typeGraphExtract, dispatched by the background processor; status and failures appear on the background tasks page
Frontend entryKnowledge graph page → "Extract from notes", with note search, notebook grouping and select-all
Enqueue dedupe works on two levels: (Type, EntityId) pairs are deduped inside a batch, and targets that already have a queued or running job are skipped. Clicking extract repeatedly therefore does not pile up duplicate jobs.

Prompt contract

ItemValue
Return shape{ entities: [{ name, type, description }], relations: [{ source, target, relation, description }] }
Entity typesconcept / person / organization / technology / project, plus custom for manual entries
Relation typesbelong_to / related_to / depends_on / contains / compared_with, plus custom
Aliasesorg → organization, tech → technology
DefaultsUnknown entity type becomes concept; unknown relation type becomes related_to
Parsing toleranceStrip ```json fences, then take the outermost {...} / [...]; on failure return null and surface graph.parseFailed

Model resolution is "graph purpose provider" → fall back to "chat purpose provider"; with neither available the call returns graph.modelUnavailable and writes nothing. Usage is recorded under source Graph with input tokens, compressed tokens, latency and a content preview (the note title).

Dedupe, merge and cleanup

ObjectRule
Entity dedupeCompare Trim()-ed lowercased names; reuse an existing entity instead of creating a second one, then dedupe again within the batch via createdEntities
Relation dedupeKey on (SourceEntityId, TargetEntityId, normalised type); skip when present
ConfidenceExtracted relations are always 0.8; manually created ones are 1.0
Cleanup by noteCleanUpByNoteIdAsync soft-deletes the note's relations, removes entities left isolated and also sweeps legacy orphans
Restore by noteRestoreByNoteIdAsync clears the soft-delete flag on that note's relations (useful after re-indexing)
Entity mergePOST /api/graph/merge repoints related edges to the survivor, drops self-loops (source = target) and deletes the merged entity

Rendering adds one more filter: only entities with at least one relation and edges whose both ends still exist are returned. Isolated nodes count as dirty data and are hidden — a common reason for "extraction succeeded but the canvas did not change" when the model produced no edges.

Manual maintenance endpoints

  • POST /api/graph/entities · PUT /api/graph/entities/{id} · DELETE /api/graph/entities/{id}
  • POST /api/graph/relations · DELETE /api/graph/relations/{id}
  • POST /api/graph/merge · merge duplicate entities

How the compression pipeline is involved

Extraction calls CompressAlgorithmicAsync, i.e. algorithmic nodes only, no LLM summary. This is a single-pass task that already spends one model call; stacking an LLM compression on top adds cost and summaries drop entity names. Pipeline defaults are Enabled = false, Mode = algorithmic, LlmThreshold = 500; the node list and order live in the compression pipeline page.

When extraction quality drops, check two things first: the before/after lengths in the log (over-compression cuts long-tail entities) and whether the call fell back to a weaker chat model.

Failure and degradation

SituationBehaviour
No model availableReturns graph.modelUnavailable, nothing is written
Output not parsable as JSONReturns graph.parseFailed, nothing is written
Entities returned without relationsWrite succeeds but the canvas stays empty (isolated entities are filtered)
Background job interrupted by shutdownReset to queued and enqueued again; not counted as a failure
Background job cancelled on purposeMarked failed with the message "task cancelled"

Observability

  • Usage source Graph: input tokens, compressed tokens, latency, content preview (note title).
  • The SSE stream emits meta / entities / relations / done; unknown event types are preserved for forward compatibility.
  • Page state: selectedNoteIds, isExtracting, extractResults, extractQueued, reported per note.
  • The graph service has no unified log prefix; compression logs come from [Compression].

Practice notes

  • Extraction is per-note and effectively overwrites previous edges: cleaning up before re-extracting a changed note keeps entity names convergent.
  • Normalisation is only trim + lowercase, with no synonym merging; aliases are a human task via the merge endpoint, and it is ongoing maintenance.
  • Extracted relations always carry 0.8 confidence regardless of evidence, so do not treat it as a quality ranking.
  • Always use batch-queue for large batches: the synchronous batch blocks the HTTP request and has no requeue protection.