Docs K  Search
Docs/Benchmarks
Evaluation

LoCoMo benchmark

91.8%Company-reported LoCoMo score.

01. Current reported score

Hebbrix reports a LoCoMo score of 91.8%, stated on September 2, 2026. The website and documentation use this same score. It is a company-reported figure, not an independent audit.

02. Evidence and scope

The current score statement does not include a corresponding run artifact, question count or judge configuration. Those details are needed to reproduce this figure and compare it fairly with another system.

The February 2026 historical run artifact is a separate evaluation with its own measured result, methodology, category breakdown and failures. It is preserved unchanged and is not presented as evidence for the current 91.8% figure.

03. What the score does not establish

Results depend on the dataset, question categories, answer model, retrieval configuration and judge. A conversational-memory score is not a guarantee of production recall, latency, safety or Outcome Memory task-success lift. Evaluate representative workloads with a controlled baseline before relying on the score for deployment decisions.

Ask the docs
reading · this page

Hi! I'm the Hebbrix docs assistant. Ask me anything about this page: setup, code examples, endpoints, pricing, or integrations.