Proposal: A Real Benchmark for Long-Term AI Memory Systems
The Problem Nearly every AI memory system is publishing scores on benchmarks that don't adequately measure what they claim to measure. We audited LoCoMo and found 6.4% of the answer key is factually wrong (99 errors in 1,540 questions), the LLM judge accepts 63% of intentionally wrong answers, and 56% of per-category system comparisons are statistically indistinguishable from noise. LongMemEval-S uses ~115K tokens per question — every frontier model can hold that in context. It's a better co...
Original Source
Read the full article at Dev →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.