Models & usage

Hetu ships no models of its own: it adapts protocols and assigns capability to scenarios, then records what every call cost.

Providers and models

ItemDetails
ProtocolsOpenAI-compatible (including Ollama / LM Studio / vLLM) and Anthropic
ConnectionBase URL plus API key; keys are encrypted with DataProtection
Purposeschat (conversation and agent loop) / embedding (indexing and retrieval) / completion (light rewrites)
Model IDSent as-is in requests; must match what the provider expects

Per-scenario default models

ScenarioTypical choice
Chat / code sessionsYour strongest reasoning model
Wiki generation / graph extractionA mid-tier model with good long-context behaviour
Topic organizing / compression summariesA cheap, fast model
Note AI / query rewritingLow-latency small model
EmbeddingA vector model whose dimensions match Embedding:Dimensions

Proxy service

Entry: /proxy. Exposes your configured models through local compatible endpoints so editors, CLIs and scripts reuse the same keys and the same usage accounting.

# OpenAI-compatible endpoint
Base URL: http://localhost:5000/api/proxy/openai/v1
API key : generated on the Proxy page

# Anthropic-compatible endpoint
Base URL: http://localhost:5000/api/proxy/anthropic
FeatureDetails
RoutingControls which models are visible through the proxy
Shadow modelSends the same request to a second model for output comparison
Usage attributionProxied calls are counted too, tagged with the proxy source

Calling the proxy

# list the models exposed by the proxy
curl http://localhost:5000/api/proxy/openai/v1/models \
  -H "Authorization: Bearer <local proxy key>"

# run a completion through it
curl http://localhost:5000/api/proxy/openai/v1/chat/completions \
  -H "Authorization: Bearer <local proxy key>" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hi"}]}'

Proxied calls are recorded in usage with the proxy source; model visibility and the shadow model are configured on the Proxy page.

Provider and model fields

LevelFields
ProviderName, protocol (OpenAI-compatible / Anthropic), Base URL, API key (encrypted), enabled flag
ModelModelId (sent verbatim), display name, purpose (chat / embedding / completion), default flag per purpose, context metadata
CompressionTrigger threshold, target size, summary model; both pre- and post-compression tokens land in usage

Usage and cost

MetricDetails
OverviewTotal messages, input / compressed / output / cached tokens, average latency, active days, today's usage
Trends7 / 30 / 90 day curves, plus week×hour and year×day heatmaps
DistributionBy model and by source (chat / code / kanban / Wiki / graph / scheduled / proxy …)
Detail logPaged entries with time, model, token segments, latency, source and content preview; filterable by source and business reference
Usage page: token trend, distribution by model and source, heatmaps
Click for full size

Compression pipeline

Entry: Settings → Cost control. Long conversations cannot grow forever, so the pipeline summarizes at a threshold.

ParameterDetails
Trigger thresholdContext usage ratio at which compression starts
Target sizeHow much content to keep afterwards
Summary modelWhich model writes the summary (falls back to the default chat model)
AccountingBoth pre- and post-compression tokens are recorded so savings are measurable

Keeping cost down

PracticeWhy
Use cheap models for indexing, organizing and taggingHigh volume, low reasoning demand
Only enable knowledge base / memory toggles when neededAvoids injecting retrieval results into every request
Use isolated worktrees and define task boundaries in code sessionsFewer trial-and-error turns
Autopilot for long tasks, interactive for quick questionsThe former saves confirmation rounds, the latter avoids pointless tool calls