MemoryAgentBench
View on GitHubOpen source code for ICLR 2026 Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
ICLR 2026 research benchmark for evaluating memory in LLM agents across four competencies: accurate retrieval, test-time learning, long-range understanding, and conflict resolution. Ships harnesses for long-context, RAG, and agentic memory methods (mem0, Zep, Letta, Cognee) plus LLM-as-judge scripts.
Use Cases
Benchmark LLM agent memory across multi-turn interactionsCompare long-context agents vs RAG agentsEvaluate accurate retrieval (AR) performanceMeasure test-time learning (TTL) from injected contextTest long-range understanding (LRU) over conversationsEvaluate conflict resolution (CR) on contradictory factsRun LLM-as-judge scoring for memory QA and summarizationReproduce baselines for mem0, Zep, Letta, and Cognee memory methodsAblate chunk size effects in RAG memory pipelinesBuild chunked multi-turn datasets for memory evaluationEvaluate graph RAG, self-RAG, and RAPTOR retrieversBenchmark recommendation and summarization memory tasks
Built With
- Language
- Python
- Frameworks
- LangChain · PyTorch · Transformers · sentence-transformers · FAISS · Qdrant · LanceDB · pgvector · mem0 · Zep · Letta · Cognee · DeepEval · FastAPI · HippoRAG
Tags
agent-memory · benchmark · llm-agents · evaluation · multi-turn · long-context · rag · retrieval · test-time-learning · conflict-resolution · llm-as-judge · memory-systems · research · iclr-2026