Membase Benchmarks

Membase· Research

Benchmarking
long-term memory
for agents.

Measured on LoCoMo, LongMemEval and DMR. Powered by episodic extraction and multi-round retrieval that sends the reader a few thousand tokens instead of the whole history.

LoCoMo
93.1
LongMemEval
92.6
DMR
92.2

Mean tokens per retrieval call

  • Unibase Memory~6,500
  • Full-context26,000+

4× fewer tokens on LoCoMo. On LongMemEval the gap is 13×: ~9,000 tokens against a 115,000-token history.

Benchmark deep-dives

Accuracy and mean context tokens per question on each benchmark. Single pass over the full question set, graded by the benchmark's own judge.

LoCoMo

1,540 questions, 4 categories.
Single-hop, multi-hop, open-domain and temporal recall across multi-session conversations spanning months.

93.1%

Accuracy

6,562

Mean context tokens are 4× below the full history.

Gold session reached the reader for 96–99% of questions. Swapping the reader model moves the score by less than 0.1 points.

Accuracy by category

OverallCategory
100%
90%
80%
70%
60%
  • 94.6%
  • 93.6%
  • 91.6%
  • 83.3%
  • 93.1%
  • Single-hopn=841
  • Multi-hopn=282
  • Temporaln=321
  • Open-domainn=96
  • Overalln=1,540
  • Single-hopn=84194.6%
  • Multi-hopn=28293.6%
  • Temporaln=32191.6%
  • Open-domainn=9683.3%
  • Overalln=1,54093.1%

Methodology

Why the numbers look this way

Each score traces back to a specific part of the architecture, not just asserted.

Recall correctness

Every session is narrated into timestamped episodes that keep who, what and when together, so a fact is retrieved with its context. The retrieval decider reads the first hits and asks follow-up questions before settling, which is why the gold session is in context for 99.95% of LongMemEval questions.

Context footprint

Keyword and vector search are fused by reciprocal rank and only the top twenty episodes enter the prompt. That keeps a LongMemEval call at about 9,000 tokens against a 115k-token history, and a LoCoMo call at about 6,500 against 26k.

Response time

Search runs in 1.1–2.5 s at the median, including one to three decider rounds. End-to-end median is 3–15 s depending on the reader model; the memory layer is not the bottleneck.

Performance

Latency and tokens

Search and end-to-end timings, plus how much context each call actually sends to the reader.

LoCoMoLongMemEvalDMR
search latency, p50 / p951.67 s / 7.02 s2.53 s / 6.11 s1.13 s / 1.71 s
end-to-end, p50 / p958.30 s / 18.0 s14.7 s / 30.2 s3.21 s / 6.34 s
context tokens per question6,5628,9701,602
full history per question~26k~115k—
token reduction4×13×—
reader modelgpt-4.1-minigpt-5.5gpt-4o-mini

ARCHITECTURE

What's inside Membase

Four pieces working together, from how a session is cut up to where the memories live.

  1. 01 · SEGMENT

    Boundary detection

    An LLM pass splits each session into topical cells before extraction, so one episode never straddles two subjects.

  2. 02 · EXTRACT

    Episodic extraction

    Each cell becomes a titled, timestamped narrative from the user's point of view. Optional profile, fact and foresight layers sit beside it.

  3. 03 · QUERY

    Multi-round retrieval

    Hybrid search per sub-query, fused by reciprocal rank; a decider marks core evidence and issues new queries for up to three rounds.

  4. 04 · PERSIST

    Local-first store

    SQLite and FAISS on disk, scoped per user, with any OpenAI-compatible model for extraction, retrieval and answering.

Research Blog

The posts behind the numbers and architecture above.

Research posts from the Unibase Research Team are coming soon.

View all research posts→
Membase · benchmark reportSeptember 2026