Vibe Coding Discover

AI Tools

MemoryAgentBench

View on GitHub

Open source code for ICLR 2026 Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

★ 45963 forksPythonMITHUST-AI-HYZ

ICLR 2026 research benchmark for evaluating memory in LLM agents across four competencies: accurate retrieval, test-time learning, long-range understanding, and conflict resolution. Ships harnesses for long-context, RAG, and agentic memory methods (mem0, Zep, Letta, Cognee) plus LLM-as-judge scripts.

Use Cases

Benchmark LLM agent memory across multi-turn interactionsCompare long-context agents vs RAG agentsEvaluate accurate retrieval (AR) performanceMeasure test-time learning (TTL) from injected contextTest long-range understanding (LRU) over conversationsEvaluate conflict resolution (CR) on contradictory factsRun LLM-as-judge scoring for memory QA and summarizationReproduce baselines for mem0, Zep, Letta, and Cognee memory methodsAblate chunk size effects in RAG memory pipelinesBuild chunked multi-turn datasets for memory evaluationEvaluate graph RAG, self-RAG, and RAPTOR retrieversBenchmark recommendation and summarization memory tasks

Built With

Language
Python
Frameworks
LangChain · PyTorch · Transformers · sentence-transformers · FAISS · Qdrant · LanceDB · pgvector · mem0 · Zep · Letta · Cognee · DeepEval · FastAPI · HippoRAG

Tags

agent-memory · benchmark · llm-agents · evaluation · multi-turn · long-context · rag · retrieval · test-time-learning · conflict-resolution · llm-as-judge · memory-systems · research · iclr-2026