The harnesses for the benchmarks reported on past.dev/benchmarks, published so anyone can read the method and run it. They use only past.dev's public API, the same calls you would make to add memory to your own product.
| Folder | Benchmark |
|---|---|
beam/ |
BEAM, long-conversation memory at 100K, 500K, 1M and 10M tokens |
To run one, you need a past.dev account (sign up), its organization key, and an OpenRouter key. Each conversation gets its own project, and a new account allows fewer projects than a full split needs; the folder's README says how to run within that.
The saved baseline covers all 100 conversations and 2,000 questions. The combined score, weighted by question count, is 90.02%.
| Split | Conversations | Questions | Score | Files |
|---|---|---|---|---|
| 100K | 20 | 400 | 92.08% | Summary · Results |
| 500K | 35 | 700 | 89.63% | Summary · Results |
| 1M | 35 | 700 | 90.65% | Summary · Results |
| 10M | 10 | 200 | 85.03% | Summary · Results |
See the BEAM README for the category breakdown and instructions for verifying the saved results or running the benchmark.
The code is MIT licensed (LICENSE). Benchmarks, datasets and third-party prompts keep
their own terms, credited in each folder's NOTICE.md.