MemoryAgentBench
★ 459ICLR 2026 research benchmark for evaluating memory in LLM agents across four competencies: accurate retrieval, test-time learning, long-range understanding, and conflict resolution. Ships harnesses for long-context, RAG, and agentic memory methods (mem0, Zep, Letta, Cognee) plus LLM-as-judge scripts.
AI Tools | Python · agent-memory · benchmark
View Project →