Textual Memory
Long-context, persona, script, and conversation-memory benchmarks.
The open benchmark for agent memory
Measure what your agents remember.
Compare what truly matters.
A public benchmark space for comparing textual and coding-agent memory systems under a consistent evaluation flow.
Each track keeps its own result table and detailed metric breakdown.
Long-context, persona, script, and conversation-memory benchmarks.
Agent memory support for coding tasks and repository-context recall.
Industry systems use the hosted Add/Search key flow. Academic systems may use the same flow or submit a public GitHub repository for maintainer Docker deployment.
Provide hosted Add/Search APIs, or submit a public GitHub repository with Docker and API run instructions.
Use the issued key to verify the synchronous Add/Search flow.
After smoke passes, submit the full scored evaluation.
Use the product pages to inspect rankings, run evaluations, and prepare an integration.
Public ranking with filters, dataset columns, and score bars.
Create eval jobs, watch progress, and inspect private results.
Eligibility, submission routes, required materials, timelines, rewards, and publication rules.
User guide, evaluation workflow, API contract, security, and result publication.
Add/search API contract, request fields, polling, and response schemas.
Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions.
Academic board submissions are open now. Submit your evaluation request before the first-cycle deadline.
Public rankings are separated by track. Use the selector inside the leaderboard frame to switch tables without mixing metric dimensions.
Academic board submissions are open now. Submit your evaluation request before the first-cycle deadline.
Create API-gated eval jobs against the Leaderboard Suite, monitor task progress, inspect private results, and submit eligible full-suite runs for administrator review.
Choose a bound version. Use Run label to distinguish repeated evaluations of the same version.
Confirm each item before starting a public-board candidate run. The button remains locked until every item is checked.
Live job status and progress.
Review completed full-suite runs from non-admin leaderboard keys. Approved results enter the public board; rejected results remain private.
Private scores stay scoped to the current leaderboard key. Admin runs remain separate from external review candidates.
Agent Memory Challenge 是 Agent Memory Leaderboard 的首期公开评测活动,面向全球研究者、开源项目维护者和商业产品团队开放。参赛系统负责 Add 与 Search,平台统一完成 Answer、Eval、结果复核与公榜。
参赛免费,不设组队要求。参赛方承担自身 API、数据库、带宽和计算成本,平台承担统一 Answer、Eval 与评测编排成本。
2026 年 7 月 29 日
2026 年 8 月 7 日 23:59(UTC+8)
2026 年 8 月中旬
文本记忆与代码记忆
学术方法榜与商业产品榜
评测类型决定系统接受什么任务;参赛组别决定结果展示在哪个榜单。两者是互相独立的两个维度。
| 评测类型 | 主要评测内容 | 可选参赛组别 | 首届时间 |
|---|---|---|---|
| 文本记忆 | 事实召回、多跳整合、时序理解、记忆治理、个性化、规则执行、安全与隐私。 | 学术方法榜 / 商业产品榜 | 8 月 7 日提交截止 |
| 代码记忆 | 从历史工程任务中检索、筛选并复用调试经验、开发经验和项目上下文。 | 学术方法榜 / 商业产品榜 | 8 月 7 日提交截止 |
第一步始终是提交评测申请。Eval Key 不是另一个报名入口,而是自行部署 API 的申请审核通过后获得的评测凭证。
提交公开 GitHub 仓库、固定版本、Add / Search 地址、鉴权方式和运行说明。参赛方负责部署并保持接口稳定,审核通过后获得 Eval Key。
提交公开 GitHub 仓库、Docker 启动方式、Add / Search 封装和完整运行说明。平台负责构建与评测,不签发 Eval Key。
提交固定产品版本、Add / Search 地址、鉴权和容量说明。无需公开内部实现,审核通过后获得 Eval Key,结果进入商业产品榜。
学术方法榜必须提供公开、可核验的 GitHub 仓库,并披露原始方法、作者、技术报告和本次改动;商业产品榜无需开源,但必须提供可核验且稳定的产品与 API 版本。
请在提交申请前固定参评版本。正式 Full 评测受理后,不得因结果不理想更换版本或撤回。
系统名称与版本、联系人、机构或团队、拟参评类型、方法或产品说明、允许公开展示的信息,以及完整的提交说明。
公开仓库、README、Docker 命令、API 入口、依赖配置、原始工作引用、方法改动和运行步骤。自行部署时还需提供公网 API。
固定产品版本、Add / Search API、鉴权方式、评测专用密钥、容量限制、超时与限流说明,并保证接口在提交后至少 30 天稳定可访问。
选择评测类型、参赛组别和提交方式,提交系统版本与完整材料。
自行部署 API 的参赛者获取 Eval Key;代码提交由平台按 Docker 说明构建。
先验证 Add / Search、鉴权和端到端链路,再运行首届正式 Full 评测。
平台复核版本、结果和合规状态,通过后发布到对应公开榜单。
两个 Key 的签发方和用途不同,请勿将它们放入公开仓库、URL、截图、邮件正文或群聊。
| 凭证 | 谁提供 | 用途 | 谁需要 |
|---|---|---|---|
| Eval Key / Leaderboard Key | Agent Memory Leaderboard | 验证参赛身份、运行评测并查看私有结果。 | 自行部署 API 的学术与商业参赛者;平台部署代码的路径不签发。 |
| Memory System Key | 参赛方 | 供平台访问参赛系统的 Add / Search API。 | 接口启用鉴权时需要;无鉴权接口无需提供。 |
榜单排名奖励与社区贡献奖励是两套独立计划;进入公开榜单不等于自动获得全部奖励。
第 1—3 名获得 ChatGPT Pro 月度会员;第 4—10 名获得 ChatGPT Plus 月度会员。账号由赛事方独立采购并发放,赛事并非由相关产品提供方主办或赞助。
前 50 位完成有效提交;成功邀请 3 位新参赛者完成有效提交;或提交 3 组有挑战性的测试样本并通过审核。
入选社区贡献计划后,获得至少价值人民币 50 元的 Kimi Token 额度。最终名额、额度、发放时间和审核结果以官方通知为准。
参赛身份、系统版本和所需材料均完整、真实且可核验。
Add / Search、鉴权和端到端链路符合现行接入协议。
正式评测任务和所需评测项成功完成,没有缺失或重复结果。
正式评测使用的代码、镜像或 API 与申报版本保持一致。
版本、结果与合规状态通过主办方审核,方计为一次有效提交。
Search 不得直接生成最终答案,也不得把答案伪装为记忆记录。
不得跨 user_id、任务、样本或团队共享和检索评测记忆。
复用论文、仓库或代码时,必须注明原作者、技术报告和全部方法改动。
严禁硬编码、数据泄漏、提示词注入、人工实时答题、结果操纵和恶意刷榜。
提交申请前请先阅读 API 接入指南并准备完整材料。报名与评测问题可发送至 [email protected]。
Agent Memory Challenge is the first public evaluation cycle of Agent Memory Leaderboard. It is open to researchers, open-source maintainers, and commercial product teams worldwide. Participants provide Add and Search; the platform runs Answer, Eval, result review, and leaderboard publication.
Participation is free, with no team-size requirement. Participants cover the cost of their own APIs, databases, bandwidth, and compute; the platform covers unified Answer, Eval, and evaluation orchestration.
July 29, 2026
August 7, 2026 · 23:59 (UTC+8)
Mid-August 2026
Textual Memory and Coding Memory
Academic Methods and Commercial Products
The evaluation type determines the tasks your system receives; the participant division determines where the result is listed. These are two independent dimensions.
| Evaluation type | What it evaluates | Available divisions | First-cycle date |
|---|---|---|---|
| Textual Memory | Fact recall, multi-hop integration, temporal understanding, memory governance, personalization, rule execution, safety, and privacy. | Academic Methods / Commercial Products | Submissions close August 7 |
| Coding Memory | Retrieving, filtering, and reusing debugging experience, development experience, and project context from historical engineering tasks. | Academic Methods / Commercial Products | Submissions close August 7 |
Your first action is always to submit an evaluation request. An Eval Key is not a separate registration step; it is issued after approval to participants who host their own APIs.
Submit a public GitHub repository, a fixed version, Add / Search endpoints, authentication details, and run instructions. You operate the service and keep it stable; an Eval Key is issued after approval.
Submit a public GitHub repository with Docker startup instructions, an Add / Search wrapper, and complete run documentation. The platform builds and evaluates it; no Eval Key is issued.
Submit a fixed product version, Add / Search endpoints, authentication, and capacity details. Internal implementation may remain closed; an Eval Key is issued after approval and results enter the Commercial Products board.
Academic Methods entries must provide a public, verifiable GitHub repository and disclose the original method, authors, technical report, and all changes. Commercial Products entries need not be open source, but their product and API versions must be stable and verifiable.
Freeze a clear evaluation version before applying. Once a formal Full evaluation is accepted, the version may not be replaced or withdrawn because of an unfavorable result.
System name and version, contact details, organization or team, intended evaluation type, method or product description, information approved for public display, and complete submission notes.
Public repository, README, Docker command, API entrypoint, dependencies, attribution of prior work, method changes, and run steps. Self-hosted entries must also provide public endpoints.
Fixed product version, Add / Search APIs, authentication, a dedicated evaluation credential, capacity, timeout, and rate-limit details. Endpoints must remain stable and publicly reachable for at least 30 days after submission.
Choose an evaluation type, participant division, and submission route, then provide a fixed system version and complete materials.
Self-hosted API participants receive an Eval Key; code submissions are built by the platform from the documented Docker entrypoint.
Validate Add / Search, authentication, and the end-to-end path before the formal first-cycle Full evaluation.
The platform reviews the version, results, and compliance status before publishing the entry to its corresponding public board.
These credentials are issued by different parties and serve different purposes. Never place either credential in a public repository, URL, screenshot, email body, or group chat.
| Credential | Provided by | Purpose | Who needs it |
|---|---|---|---|
| Eval Key / Leaderboard Key | Agent Memory Leaderboard | Verifies participant access, starts evaluations, and unlocks private results. | Academic and commercial participants hosting their own APIs. Platform-deployed code submissions do not receive one. |
| Memory System Key | Participant | Allows the platform to call the participant's Add / Search APIs. | Required when the submitted API uses authentication; not required for an unauthenticated endpoint. |
Leaderboard ranking rewards and community contribution rewards are separate programs. Publication on a public board does not automatically grant every reward.
Ranks 1–3 receive one month of ChatGPT Pro; ranks 4–10 receive one month of ChatGPT Plus. Accounts are independently purchased and issued by the organizers; the challenge is not organized or sponsored by the related product provider.
Be among the first 50 valid submissions; successfully refer 3 new participants who each complete a valid submission; or submit 3 sets of challenging test samples that pass review.
Selected community contributors receive at least RMB 50 worth of Kimi Token credits. Final availability, amount, issuance date, and eligibility are subject to official notice and review.
Participant identity, system version, and all required materials are complete, truthful, and verifiable.
Add / Search, authentication, and the end-to-end path follow the current integration contract.
The formal evaluation and required tasks complete successfully without missing or duplicated results.
The code, image, or API used for the formal evaluation matches the declared version.
The version, results, and compliance status pass organizer review before the submission is considered valid.
Search must not generate final answers or disguise answers as memory records.
Do not share or retrieve evaluation memories across user IDs, tasks, samples, or teams.
When reusing a paper, repository, or codebase, identify the original authors, technical report, and every method change.
Hard-coding, data leakage, prompt injection, live human answering, result manipulation, and leaderboard abuse are prohibited.
Read the API Guide and prepare all required materials before applying. For participation and evaluation questions, contact [email protected].
了解如何接入 Add / Search、运行统一评测、查看结果,并将符合条件的结果发布到公开榜单。
OVERVIEW
Agent Memory Leaderboard 使用统一的端到端流程评估不同记忆系统。我们不限定你使用的数据库、索引、向量模型或内部架构,只要求 Add / Search 接口符合现行规范,且评测过程和检索结果能够复核。
我们按样本和来源会话调用 Add 接口。
你的系统负责持久化、组织、索引和更新记忆。
我们针对每道题调用 Search 接口并接收排序结果。
我们使用固定流程生成答案、完成评分并汇总结果。
你只负责 Add / Search 环节。回答模型、提示词、评分器、数据集组合、Top K 和汇总规则由我们统一固定,使榜单分数尽可能反映记忆系统本身的能力。
WHO THIS IS FOR
你可以在 Textual、Multimodal 或 Coding Track 中查看 Overall、分项指标、系统版本和发布日期。不同 Track 使用的任务和指标不同,分数不宜跨 Track 直接比较。
实现 Add / Search → 提交准入申请 → 获取 API Key → 运行公开 smoke → 提交 full 正式评测 → 查看私有结果 → 进入公榜审核。
PARTICIPATION WORKFLOW
先提交 Evaluation Access Request,完成参测系统、版本、代码或公开接口和鉴权方式的审核。代码提交必须附带完整 README、Docker 启动命令、API 封装说明和原始方法披露;平台会按说明启动并部署后再评测。托管接口提交则需在提交后至少 30 天保持公网可访问且稳定。
选择申请新 Key,或输入已有 Key 为其增加版本;填写联系人、系统版本、GitHub 仓库或 Add / Search 接口,并在提交说明中写清运行方式、Docker 命令、API 封装和原始方法信息。
审核通过后,新申请会签发 Leaderboard API Key;新增版本申请会直接绑定到原 Key,无需管理员再次手工录入接口信息。
代码提交由平台按 Docker 和 API 说明构建、启动并封装;托管提交直接校验你提供的接口。随后按照现行同步规范执行 Add → Search → Answer → Evaluate smoke 流程,确认接口可用后再进入正式评测。
选择 full 前,必须逐项勾选提交清单:smoke 已通过、API 契约正确、运行说明完整、托管接口可持续 30 天稳定,以及原创性披露和诚信承诺均已完成。
在 Evaluation 页面验证 API Key,选择已绑定的版本,填写唯一的 Run Label 后提交 full 评测。正式任务使用统一的数据集套件和 Top K。
任务依次执行写入、检索、回答、评分和结果汇总。页面会显示当前阶段、完成进度和错误摘要;我们同时保留复核所需的过程记录。
评测结果首先仅对绑定的 API Key 可见。成功完成的 full 任务会进入公榜资格校验和审核队列;审核通过后生成公榜提交记录。
允许复现已有论文或封装已有仓库,但必须披露原始作者、技术报告以及方法改动;未披露的重复代码可能被视为抄袭。发现重复代码时,平台会优先保护本人已披露并先提交的评测结果。重复提交相似或低质量代码、提示词注入、数据或结果操纵、恶意刷榜等行为,均可能取消参赛资格。
EVALUATION MODES
| 模式 | 用途 | 标准配额 | 结果范围 | 公榜资格 |
|---|---|---|---|---|
| smoke | 通过独立的兼容性测试接口检查同步 Add、Search 和评分链路。 | 每小时 1 次 | 仅私有 | 不可发布 |
| full | 运行固定的完整评测套件,并进入正式排队、审计和发布流程。 | 每 3 个月 1 次 | 优先私有 | 审核通过后发布 |
首次接入或接口行为变更后,先使用独立的兼容性 smoke 检查接口规范;接口稳定且版本准备完成后,再提交唯一的正式评测模式 full。
ADD / SEARCH CONTRACT
Add 和 Search 的接口地址由你配置,请求和响应格式按照现行规范固定,不随 URL 路径变化。接口必须能够从我们的评测环境访问;生产环境建议使用 HTTPS。URL 中不得包含用户名、密码等凭据,也不得指向私有、回环或链路本地地址。
记忆写入完成后,Add 接口返回 HTTP 200
我们按照返回顺序最多读取 top_k 条记忆
Add 和 Search 必须使用完全相同的 user_id
Add / Search 支持 Token、Bearer 和 X-Api-Key;none 仅用于公开 smoke。Health 接口通过无需鉴权的 GET 请求调用,返回任意 2xx 状态码即表示服务正常。如果没有单独配置 Health 地址,正式任务会检查 Add 同源的 /health。
每个来源会话默认调用一次 Add;超过 20 条消息或 2000 个词时,在最近的完整消息或句子边界分段。请求只包含下列字段。
{
"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",
"messages": [{
"role": "user",
"timestamp": 1704067200000,
"content": "memory text"
}],
"user_id": "eval:run_abc123:locomo:conv-0",
"session_id": "eval:run_abc123:sample:0"
}
request_id必填。本次写入请求的唯一标识;成功响应必须原样返回该值。
messages必填。消息按原顺序排列,每条消息包含 role 和非空的 content。timestamp 可选,单位为 Unix 毫秒。
user_id必填。Search 接口唯一使用的检索范围标识;写入和检索时必须保持一致。
session_id必填。用于标识来源会话,可以用于组织记忆,但不作为 Search 的筛选条件。
未使用字段现行规范不发送 metadata、app_id、agent_id 或 async_mode;同步语义由 Add 完成写入后返回 HTTP 200 保证。
写入完成并且相关记忆能够立即检索后,才能返回成功响应。
HTTP 200
{
"success": true,
"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",
"user_id": "eval:run_abc123:locomo:conv-0",
"session_id": "eval:run_abc123:sample:0"
}
success必填,且必须是布尔值 true。
request_id / user_id / session_id全部必填,并且必须与请求中的值完全一致。
同步完成服务内部可以采用异步处理,但 Add 接口必须等待处理完成后再返回。
不支持的响应请勿返回 HTTP 202、task ID 或状态查询地址;响应中也不需要 memory_ids。
所有 Add 请求成功后,我们会针对每道题调用一次 Search 接口。查询使用数据集原文;选择题会另行传入选项。
{
"query": "Which answer best matches the memory?",
"options": ["A. First answer", "B. Second answer"],
"user_id": "eval:run_abc123:locomo:conv-0",
"top_k": 100
}
query必填。按照原文检索相关记忆,不得将其替换为最终答案,也不得使用评测金标。
options可选。题目有候选项时传入字符串数组;开放题不发送该字段。
user_id必填。只能在该 user_id 对应的记忆范围内检索。
top_k必填。返回的记忆数量不得超过该值;正式外部评测固定为 100。
请求格式现行规范不发送 filters、rerank 或 keyword_search。
响应必须是一个 JSON 对象,其中 data 为按相关性排序的数组。没有检索结果时,返回空数组。
{
"data": [{
"id": "mem_1",
"content": "remembered fact text",
"score": 0.87,
"created_at": "2026-07-01T12:00:00Z"
}]
}
data必填,类型为数组。不要增加 items 包装层,也不要直接返回顶层数组。
id必填,非空字符串,用于稳定标识该条记忆。
content必填,非空字符串。该内容会直接提供给统一回答模型。
score可选,数值类型。数值越大应表示相关性越高。
created_at可选,用于记录记忆的来源时间或持久化时间。
其他字段我们只读取上述字段;metadata 等未声明字段会被忽略。
ERROR HANDLING
参测接口应使用标准 HTTP 状态码,并返回便于排查问题且不含密钥的错误信息。平台业务错误通常采用 {"detail":{"reason":"..."}} 格式;字段校验失败时,HTTP 422 会返回结构化错误明细。
| HTTP | 错误类型 | 常见原因 | 处理方式 |
|---|---|---|---|
400 / 422 | 接口格式错误 | Add / Search 请求无法解析,或者成功响应缺少必填字段、字段类型不正确。 | 不自动重试。根据错误信息修正请求或响应格式,然后重新运行 smoke。 |
401 | 认证失败 | Memory System Key 无效,或者 Authorization / X-Api-Key 与申请时选择的方式不一致。 | 不自动重试。核对鉴权方式和绑定的密钥。 |
403 | 访问被拒绝 | 当前密钥无权调用接口,或者服务端拒绝访问指定的 user_id。 | 不自动重试。检查接口权限和检索范围配置。 |
404 | 资源不存在 | 接口路径错误,或者平台侧的 test、job、result ID 无效。 | 检查 URL 或资源 ID。现行同步规范不包含 Add Status 查询。 |
409 | 状态冲突 | Add 暂时无法写入;或者平台已有任务正在运行、Run Label 重复。 | Add 请求会在限定次数内重试,Search 请求遇到 409 不重试。平台任务冲突时,请等待现有任务结束或更换 Run Label。 |
408 / 425 | 暂时不可用 | 请求超时,或者服务尚未准备完成。 | Add / Search 请求都会按照平台策略进行有限次数的退避重试。 |
429 | 触发限流或配额 | 参测接口容量不足,或者平台评测配额尚未恢复。 | 接口调用会在限定次数内重试;如果是平台配额限制,请等待页面显示的下次可用时间。 |
500 / 502 / 503 / 504 | 临时服务异常 | 参测接口、网关或上游依赖暂时不可用。 | 我们会自动退避重试。若持续失败,请保留 Job/Test ID 和发生时间,便于进一步排查。 |
Add 遇到 408、409、425、429、500、502、503、504 时会重试;Search 的重试范围相同,但不包含 409。网络超时和传输错误也会触发重试,且重试次数有限。
即使 HTTP 状态码为 200,只要 Add 未返回 success=true、三个 ID 未正确返回,或者 Search 未返回 data 数组、某条记录缺少 id / content,当前阶段都会立即失败。
DATA, SECURITY & PRIVACY
你的接口只会收到当前任务所需的记忆片段、user/session 标识和检索问题。我们不会提供金标答案、评分依据或完整数据集下载。
user_id 是 Search 接口唯一使用的检索范围标识,存储和检索时必须完全一致。session_id 只用于组织来源会话。禁止跨 user_id 返回记忆。
Memory System Key 通过受控的申请流程提交并加密保存,任务信息中只保留不可读的引用。密钥不会出现在邮件、公榜或公开 API 响应中。
我们只连接通过网络校验的公开 HTTP(S) 接口,并拒绝 URL 中包含凭据,或者指向私有、回环、链路本地地址的目标。
我们会保留复核所需的请求结果、耗时、错误、候选记忆和接口格式校验记录,用于确认评测是否完整,以及结果是否符合公榜条件。
私有任务和结果仅对绑定的 Leaderboard API Key 可见。公榜只展示审核通过的系统、版本、分数和必要的评测信息。
评测数据及其派生副本只能用于完成当前任务,不得用于模型训练、微调、产品分析、数据集重建或对外传播。请仅向必要人员开放访问权限,避免记录不必要的请求正文,并在任务完成后 30 天内删除相关数据;如需延长保留时间,必须事先获得我们的书面同意。
EVALUATION & RESULTS
每个记忆分块都必须在返回 HTTP 200 前完成持久化,并且能够立即检索。
我们会检查 data 数组、必填字段和 Top K,并按照接口返回的顺序接收候选记忆。
通过校验的候选记忆会进入统一的回答模型和提示词流程。
我们按照题型使用固定的评分规则,并汇总各数据集和 Overall 结果。
任务完成后,你可以使用绑定的 API Key 查看运行状态、分项结果、错误摘要和任务信息。
full 结果通过资格校验和审核后,会生成公榜提交记录并展示在公开榜单中。
结果必须来自成功完成的 full 模式固定套件,且所有评测任务均执行成功。统一回答模型、评测规范、pipeline code hash、dataset bundle hash 和题量记录必须完整,并与现行发布基线一致。结果不得重复提交,同时还需通过平台审核。
BENCHMARK SUITE
按能力维度浏览基准数据集。每张卡片均链接到源仓库,并说明任务格式、指标与评测备注。
Learn how to integrate Add / Search, run consistent evaluations, review results, and publish eligible runs to the public leaderboard.
OVERVIEW
Agent Memory Leaderboard compares memory systems under a consistent end-to-end evaluation. The platform does not prescribe a database, index, embedding model, or internal architecture. It requires a conformant external API and verifiable evidence for retrieval and audit.
The platform calls Add by sample and source session.
Your system persists, organizes, indexes, and updates memory.
The platform calls Search per question and receives ranked results.
A fixed platform workflow generates answers, scores them, and aggregates results.
Participant-controlled behavior is limited to Add / Search. The platform locks the answer model, prompts, evaluators, dataset suite, Top K, and aggregation rules so that score differences primarily reflect the memory system.
WHO THIS IS FOR
Review Overall scores, metric breakdowns, system versions, and publication dates within the Textual, Multimodal, or Coding track. Each track has a distinct task and metric contract; scores are not comparable across tracks.
Implement Add / Search → submit an access request → receive an API Key → run public smoke → submit the full evaluation → review private results → enter public review.
PARTICIPATION WORKFLOW
Industry submissions and academic systems with hosted APIs follow the existing review, Key issuance, and public smoke flow. Code submissions must include a complete README, Docker startup command, API wrapper instructions, and original-method disclosure; maintainers build and start the documented entrypoint before evaluation. Hosted endpoints must remain publicly reachable and stable for at least 30 days after submission.
Industry requests provide Add / Search endpoints. Code submissions provide a publicly accessible GitHub repository and explain the Docker command, API wrapper, original authors, technical report, and every method change.
Approval either issues a new Leaderboard API Key or binds the submitted version to the verified existing key without administrator re-entry.
For code submissions, maintainers build and start the documented Docker entrypoint and expose the official API. Hosted submissions are checked at the declared endpoint. The platform then runs the synchronous Add → Search → Answer → Evaluate smoke flow.
Before selecting full, confirm the interactive checklist: smoke passed, API contract followed, run instructions complete, hosted runtime stable for 30 days, original work disclosed, and no integrity violations.
Verify the API Key on the Evaluation page, select a bound version, assign a unique Run Label, and submit the full evaluation. Formal jobs use the fixed suite and platform Top K.
The job proceeds through ingestion, retrieval, answering, scoring, and aggregation. The interface reports stage, progress, and error summaries while the platform retains evidence required for review.
Results are initially visible only to the bound API Key. A successful full job enters the public-eligibility gate and review queue; approval creates a public submission.
Reproductions of papers or existing repositories are allowed only with full attribution to the original authors, technical report, and all method changes. Undisclosed reuse may be treated as plagiarism. When duplicate code is discovered, the platform prioritizes the participant's first disclosed submission. Repeated near-duplicate or low-quality code, prompt injection, benchmark or result manipulation, malicious behavior, and other leaderboard abuse may cancel eligibility.
EVALUATION MODES
| Mode | Purpose | Standard quota | Visibility | Public eligibility |
|---|---|---|---|---|
| smoke | Separate integration endpoint for synchronous Add, Search, and scoring compatibility. | 1 per hour | Private | Not eligible |
| full | Fixed complete benchmark suite with formal queueing, audit, and release workflow. | 1 every 3 months | Private first | Eligible after approval |
Use the separate compatibility smoke after initial integration or an API behavior change. Submit full, the only formal participant mode, after the version is ready for publication.
ADD / SEARCH CONTRACT
Participants configure the Add and Search URLs; request and response schemas are fixed and do not vary by path. Endpoints must be reachable from the platform network, and production deployments should use HTTPS. URLs must not embed credentials or resolve to private, loopback, or link-local addresses.
Add succeeds only with HTTP 200 after persistence
The platform reads at most top_k items in response order
Add and Search must use the identical isolation boundary
Add / Search support Token, Bearer, and X-Api-Key; none is limited to public smoke. Health is an unauthenticated GET where any 2xx means healthy. If no custom Health URL is bound, formal jobs check /health on the Add origin.
Each source session uses one Add by default. Sessions over 20 messages or 2,000 words split at the nearest complete message or sentence boundary. Requests contain only fields declared by the synchronous contract.
{
"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",
"messages": [{
"role": "user",
"timestamp": 1704067200000,
"content": "memory text"
}],
"user_id": "eval:run_abc123:locomo:conv-0",
"session_id": "eval:run_abc123:sample:0"
}
request_idRequired. Unique identifier for this chunk request; the success response must echo it exactly.
messagesRequired ordered array. Each item includes role and non-empty content; timestamp is an optional Unix-millisecond event time.
user_idRequired. The sole retrieval-isolation field; Search must use the identical value.
session_idRequired. Identifies the source session and may be used for grouping, but is not a Search filter.
Fields not sentThe current contract omits metadata, app_id, agent_id, and async_mode. Synchronous behavior is guaranteed by returning HTTP 200 only after Add completes.
Return success only after the write is persisted and immediately searchable.
HTTP 200
{
"success": true,
"request_id": "eval:run_abc123:locomo_refined:conv-0:chunk-0",
"user_id": "eval:run_abc123:locomo:conv-0",
"session_id": "eval:run_abc123:sample:0"
}
successRequired and must be the boolean value true.
request_id / user_id / session_idAll are required and must match the request byte for byte.
Synchronous completionInternal work may be asynchronous, but the endpoint must wait for completion before responding.
Not supportedDo not return HTTP 202, a task ID, or a status-poll URL. The response does not need memory_ids.
After every Add succeeds, the platform sends one Search request per question. The query stays in its original language; choice questions include options separately.
{
"query": "Which answer best matches the memory?",
"options": ["A. First answer", "B. Second answer"],
"user_id": "eval:run_abc123:locomo:conv-0",
"top_k": 100
}
queryRequired. Use the original text for retrieval; do not replace it with a final answer or use benchmark gold data.
optionsOptional. A string array sent for questions with answer choices; omitted for open questions.
user_idRequired. Retrieve only from memory stored under this exact value.
top_kRequired. The response must not exceed this number; formal external evaluations use 100.
Fixed schemaThe current contract does not send filters, rerank, or keyword_search.
Return a JSON object whose data field is a relevance-ordered array. Return an empty array when nothing is found.
{
"data": [{
"id": "mem_1",
"content": "remembered fact text",
"score": 0.87,
"created_at": "2026-07-01T12:00:00Z"
}]
}
dataRequired array. Do not add an items wrapper or return a top-level array.
idRequired non-empty string that stably identifies the memory.
contentRequired non-empty string passed directly to the platform Answer model.
scoreOptional number. Higher values should indicate greater relevance.
created_atOptional source or persistence timestamp.
Other fieldsThe platform reads only the declared fields above; undeclared fields such as metadata are ignored.
ERROR HANDLING
Participant endpoints should use standard HTTP status codes and return actionable errors without secrets. Platform business errors normally use {"detail":{"reason":"..."}}; HTTP 422 provides structured field-validation details.
| HTTP | Class | Typical case | Platform behavior and action |
|---|---|---|---|
400 / 422 | Contract error | An Add / Search request cannot be parsed, or a success response has missing or invalid required fields. | Not retried. Correct the schema using the error details, then rerun smoke. |
401 | Authentication failure | The Memory System Key is invalid, or the Authorization / X-Api-Key scheme does not match. | Not retried. Verify the authentication scheme and secret bound to the request. |
403 | Access denied | The key cannot call the endpoint, or the service rejects the current user_id scope. | Not retried. Review endpoint authorization and retrieval isolation. |
404 | Resource not found | An endpoint path is wrong, or a platform test, job, or result ID is invalid. | Verify the URL or resource ID. The synchronous contract has no Add Status polling. |
409 | State conflict | Add is temporarily conflicted, or a platform job is active and the Run Label is duplicated. | Add is retried with bounds; Search 409 is not retried. Wait or choose a new Run Label for platform conflicts. |
408 / 425 | Transient unavailability | The request timed out or the service is not ready. | Add and Search are retried with bounded backoff. |
429 | Rate or quota limit | The participant endpoint is capacity-limited, or a platform evaluation quota has not reset. | Endpoint calls are retried with bounds; for platform quotas, wait until the displayed availability time. |
500 / 502 / 503 / 504 | Transient service failure | The participant endpoint, gateway, or upstream dependency is temporarily unavailable. | The platform retries with backoff. Preserve the Job/Test ID and timestamp if failures persist. |
Add retries 408, 409, 425, 429, 500, 502, 503, and 504. Search retries the same set except 409. Network timeouts and transport failures are also eligible.
Even with HTTP 200, the stage fails if Add omits success=true or mis-echoes an ID, or if Search omits the data array or an item lacks id / content.
DATA, SECURITY & PRIVACY
Participant endpoints receive only the memory chunks, user/session identifiers, and retrieval questions needed for the current job. Gold answers, scoring criteria, and bulk dataset downloads are not provided.
user_id is the sole Search-isolation field and must match exactly during storage and retrieval. session_id is only for source-session organization. Cross-user_id retrieval is prohibited.
The Memory System Key is submitted through a controlled request flow and stored encrypted; job metadata contains only an opaque reference. Secrets do not appear in email, public rankings, or public API responses.
The platform connects only to network-validated public HTTP(S) endpoints and rejects URLs with embedded credentials or targets resolving to private, loopback, or link-local addresses.
The platform retains request outcomes, latency, errors, returned memories, and contract-validation evidence required to verify evaluation integrity and public eligibility.
Private jobs and results are visible only to the bound Leaderboard API Key. The public board shows approved systems, versions, scores, and required evaluation metadata.
Evaluation data and derived copies may be used only to complete the current job. Do not use them for training, fine-tuning, product analytics, dataset reconstruction, or redistribution. Restrict access to authorized personnel, avoid unnecessary payload logging, and delete the data within 30 days after job completion unless the platform approves another retention period in writing.
EVALUATION & RESULTS
Each chunk must be persisted and searchable before its HTTP 200 response.
The platform validates the data array, required fields, and Top K, then accepts candidates in participant order.
Valid candidate memories enter the locked platform Answer model and prompt workflow.
The platform applies fixed scoring contracts by question type and aggregates dataset and Overall results.
After completion, the bound API Key can inspect run status, breakdowns, error summaries, and job metadata.
An eligible full result creates a public submission only after qualification checks and review.
The result must come from a successful full run on the fixed suite; every evaluation task must succeed; the Answer model, evaluation contract, pipeline code hash, dataset bundle hashes, and question counts must be complete and match the current release baseline; and the result must be unique and pass platform review.
BENCHMARK SUITE
Browse benchmark datasets by dimension. Each card links to its source repository with task format, metrics, and evaluation notes.
Hosted integrations expose synchronous Add and Search HTTP(S) endpoints. Code submissions provide a public GitHub repository, Docker startup command, and API wrapper instructions; maintainers deploy and evaluate the submitted system.
The platform sends one request for each memory chunk. Store every message before responding, and associate the data with the supplied user_id. The session_id identifies the source conversation and may be used for grouping, but retrieval isolation is based on user_id.
{
"request_id": "eval:<run_id>:locomo_refined:conv-0:chunk-0",
"messages": [{
"role": "user",
"timestamp": 1704067200000,
"content": "raw memory text"
}],
"user_id": "eval:<run_id>:locomo:conv-0",
"session_id": "eval:<run_id>:sample:0"
}
user or assistant.
contentRequiredOriginal message text supplied for ingestion.
timestampOptionalMessage timestamp in Unix milliseconds when available.
user_idRequiredRetrieval scope. Store it exactly and use it to isolate later searches.
session_idRequiredIdentifier for the source conversation or session.
Return HTTP 200 only after the submitted messages have been fully stored and are available to Search. If your system performs ingestion in the background, wait for that work to finish before returning success; otherwise the benchmark may search before the memory is ready.
{
"success": true,
"request_id": "eval:<run_id>:locomo_refined:conv-0:chunk-0",
"user_id": "eval:<run_id>:locomo:conv-0",
"session_id": "eval:<run_id>:sample:0"
}
true. It confirms that the messages are stored and searchable.
request_idRequiredExact request_id received in the Add request.
user_idRequiredExact user_id received in the Add request.
session_idRequiredExact session_id received in the Add request.
After every Add request has completed, the platform sends one Search request per benchmark question. The query remains unchanged, and choice questions carry their options in a separate optional field.
{
"query": "Which answer best matches the memory?",
"options": ["A. First answer", "B. Second answer"],
"user_id": "eval:<run_id>:locomo:conv-0",
"top_k": 100
}
Return a data array ordered from most relevant to least relevant. The platform preserves this order and passes the returned content to the shared answer pipeline. Return an empty array when no relevant memory is available.
{
"data": [
{
"id": "mem_123",
"content": "remembered fact text",
"score": 0.87,
"created_at": "2026-07-01T12:00:00Z"
}
]
}
Your API receives benchmark content during each run.
Your API receives evaluation memories and questions.
Use this data only for the evaluation. Do not train on it, analyze it, or share it.
Keep it private, avoid storing logs, and delete it within 30 days after the run.
A compact map of the product surface so users know where to go after their endpoints are ready.
Overview and benchmark preview.
Public rankings by track.
Verify key, start eval jobs, inspect private results and attribution.
User workflow, API contract, security, and result publication.
Propose a benchmark for potential inclusion in the Agent Memory Leaderboard.
Describe the research question, evaluation protocol, dataset provenance, and governance terms. The review team assesses scientific value, safety, and reproducibility before inclusion.
The platform accepts both public and private benchmark contributions. Public data is released only after review and with clear attribution; private data remains restricted to authorized evaluation use.
Contributor contact details are used for review and coordination. Full email addresses are not displayed publicly.