Latency & SLA
Honest, per-endpoint latency guidance and the production pattern for keeping memory off your user's response path. Retrieval is fast; fact extraction is LLM-bound, so run it asynchronously.
/v1/memories/process in synchronous mode.01. What is fast, what is LLM-bound
- Fast (retrieval path). Search, raw writes, and reads combine retrieval, fusion, filters, and optional ranking stages. The fast-search release objective is p95 at or below 750ms; it is an engineering target, not a per-request or contractual guarantee.
- LLM-bound (write/reasoning path). Fact extraction, conflict resolution, and dialectic reasoning call a reasoning model. These are seconds, not milliseconds, and they are exactly the operations you should run in the background.
02. Per-endpoint guidance
Engineering objectives and endpoint behavior, not a universal observed latency distribution or contractual guarantee. Cold starts, dependency delays, geography and large batches can run higher.
| Endpoint | Objective / behavior | Class | Mode |
|---|---|---|---|
POST /v1/search (fast) | ≤750ms p95 objective | Retrieval | Inline (hot path) |
POST /v1/learning/decisions | ≤1000ms p95 engineering objective | Scoped evidence + selection | No LLM call; not a contractual SLA |
POST /v1/memories/raw (async index) | ≤750ms p50 objective | Durable write | Returns searchable + status_url |
POST /v1/memories/raw (wait_for_index) | Up to 5s default index wait, plus processing | Write + index | 201 if ready, otherwise 202 |
GET /v1/memories/{id} | sub-second | Read | Inline |
POST /v1/chat/completions (non-stream) | 30s generation / 45s total default deadlines | LLM + retrieval | Timeout fails with 504; not fabricated success |
POST /v1/chat/completions (stream) | <1500ms first-content p95 objective | LLM + retrieval | Stream tokens |
POST /v1/memories/process (sync) | ~7–22s | LLM extraction | Use async → |
POST /v1/memories/process (async default) | sub-second 202 target | Queued | Background |
POST /v1/profile/dialectic | ~10–14s | LLM reasoning | Background / await |
PATCH /v1/memories/{id} | ≤1.5s response objective | Durable update + re-index | 200 if ready, otherwise 202 + status_url |
No sub-50ms end-to-end search guarantee is made. An isolated lookup timing does not include hybrid retrieval, reranking, networking or generation. Measure against your own clients using Server-Timing, X-Process-Time, X-Hebbrix-TTFT-Target-Ms, and X-Hebbrix-Degraded-Stages response headers.
Measure large histories separately
For procurement, qualify your actual collection sizes and relevant decision density: record owner/collection/context scope, history size (including 10k/100k memory workloads), configured comparable window, top-k, filters, fast versus quality search, optional feature model, cache state, concurrency, client region and deployed build. Report client round-trip p50/p95/p99 separately from server processing, retrieval and generation. Include errors and throttled requests rather than dropping them.
The decision evidence window is bounded at 2,000 comparable records, not a guarantee of constant SQL latency over arbitrarily large histories. More feature categories add computation; dense concurrent writes and dependency delays can exceed objectives. No large-history latency distribution or historical uptime percentage is claimed by this page. Arrange a workload-specific capacity and latency review via contact@hebbrix.com; see current status and support.
03. Production pattern: async extraction
async_dispatch defaults to true on /v1/memories/process. You get a 202 with a job_id; poll /v1/memories/jobs/{job_id} for completion. Your user already has their answer.
Raw ingestion has a separate durability/readiness contract. With wait_for_index: false (the default), the response acknowledges the atomic PostgreSQL and indexing-outbox commit; inspect searchable and poll status_url. With wait_for_index: true, the server-default wait for searchability is up to 5 seconds, plus request processing. It returns 201 if indexing completes or 202 with Retry-After, Location, and outbox_event_id when the durable write is still converging.
import os, time, requests
BASE = "https://api.hebbrix.com/v1"
H = {"Authorization": f"Bearer {os.environ['HEBBRIX_API_KEY']}"}
# 1) Inline: fetch context for THIS turn (fast, hot path)
ctx = requests.post(f"{BASE}/search",
headers=H, json={"query": user_message, "collection_id": "coll_123", "limit": 5}).json()
# 2) Respond to the user with YOUR LLM (using ctx) ... already done here ...
# 3) Off the hot path: learn from the turn ASYNCHRONOUSLY
job = requests.post(f"{BASE}/memories/process", headers=H, json={
"messages": [
{"role": "user", "content": user_message},
{"role": "assistant", "content": assistant_reply},
],
"collection_id": "coll_123",
"async_dispatch": True, # explicit; true is currently the default
}).json()
# 4) (Optional) poll for completion in a background worker
job_id = job["job_id"]
while True:
s = requests.get(f"{BASE}/memories/jobs/{job_id}", headers=H).json()
if s["status"] in ("completed", "failed"):
break
time.sleep(1)04. SLA & status
- What we publish.
GET /v1/slois the canonical public contract. It currently describes best-effort service with p95/p99 and availability engineering objectives, not a generic contractual SLA. Signed plan or enterprise terms control. - Observability. Every response carries
X-Process-Time(server ms) andX-Request-ID. Chat additionally reports stage timings and any degraded stages. Measure real latency from your own clients and include the request id and deployed build header when reporting a slow call.