Agent Memory

The open benchmark for agent memory

Agent Memory Leaderboard

Measure what your agents remember.
Compare what truly matters.

Agent Memory Leaderboard
OPEN BENCHMARK

Agent Memory Leaderboard

A public benchmark space for comparing textual and coding-agent memory systems under a consistent evaluation flow.

Leaderboard Preview

Benchmark Tracks

Each track keeps its own result table and detailed metric breakdown.

Textual Memory

Long-context, persona, script, and conversation-memory benchmarks.

Coding Agent Memory

Agent memory support for coding tasks and repository-context recall.

Evaluation Flow

Industry systems use the hosted Add/Search key flow. Academic systems may use the same flow or submit a public GitHub repository for maintainer Docker deployment.

1

Choose an evaluation route

Provide hosted Add/Search APIs, or submit a public GitHub repository with Docker and API run instructions.

2

Run a smoke test

Use the issued key to verify the synchronous Add/Search flow.

3

Submit a formal evaluation

After smoke passes, submit the full scored evaluation.

Explore the Platform

Use the product pages to inspect rankings, run evaluations, and prepare an integration.

01

Leaderboard

Public ranking with filters, dataset columns, and score bars.

02

Evaluation

Create eval jobs, watch progress, and inspect private results.

03

Participation Guide

Eligibility, submission routes, required materials, timelines, rewards, and publication rules.

04

Documentation

User guide, evaluation workflow, API contract, security, and result publication.

05

Guide

Add/search API contract, request fields, polling, and response schemas.