AMA-Bench
View on GitHub[ICML 26] An evaluation framework assessing long-context retention and long-horizon memory performance for agentic applications (AMA-bench).
AMA-Bench is an ICML 2026 evaluation framework for agentic memory: methods build memory from long agent trajectories, retrieve evidence, and answer QA scored by LLM-as-judge. Includes vLLM/API pipelines, cross-judge validation, and a HF leaderboard.
Use Cases
Benchmarking long-horizon memory in agentic applicationsEvaluating memory construction and retrieval from agent trajectoriesComparing memory methods (BM25, embedding, AMA-Agent) on open-ended QALLM-as-judge scoring with multi-model cross-validation and agreement statsTesting long-context retention of LLMsRunning local vLLM vs remote API evaluation pipelinesEvaluating tool-using coding agents (Codex, Claude Code) on trajectory QASubmitting results to a Hugging Face leaderboard
Built With
- Language
- Python
- Frameworks
- vLLM · PyTorch · Hugging Face Transformers · FAISS · Ray · Starlette · Gymnasium · MiniGrid · TextWorld · ALFWorld · Hugging Face Datasets · tiktoken
Tags
agent-memory · long-horizon · long-context · benchmark · evaluation · llm-as-judge · memory-retrieval · embeddings · bm25 · vllm · leaderboard · agent-trajectories · icml · python