Ling Yang and colleagues introduced PAST-Bench, a benchmark probing whether personal agents improve from accumulated experience. The announcement, shared via X, targets a gap in evaluating long-term adaptation rather than one-shot task success.
Key facts
- PAST-Bench announced by Ling Yang via X thread
- Benchmark targets experience-driven agent evolution
- No baseline scores or task list published yet
- Amplified by @rohanpaul_ai on social media
- Framed against static single-turn evaluation methods
What PAST-Bench Measures
PAST-Bench, announced by Ling Yang via a thread on X and amplified by @rohanpaul_ai, asks a deceptively simple question: do personal agents actually improve as they accumulate interactions with a user? According to the announcement, the benchmark is designed to test "experience-driven evolution" — whether an agent's behavior, responses, and planning get better over repeated engagements.
This is a notable departure from standard evaluation. Most benchmarks — SWE-Bench, GAIA, or the various agentic tool-use suites — score a model on isolated tasks. PAST-Bench instead wants to capture the compounding effect of memory and adaptation, which is closer to how a personal assistant actually operates in production.
The announcement is thin on specifics. No task list, no baseline scores, no scoring rubric has been published yet. The thread reads as a call for community engagement rather than a finished artifact. That limits what can be verified today, but the framing itself is the signal.
Why This Matters Now
The timing aligns with a broader industry push toward persistent memory and agentic personalization. In the past 90 days, several major labs have shipped or previewed memory features — long-context retrieval, user-profile embeddings, and session-persistence layers. PAST-Bench is an attempt to put a measurable yardstick on that trend.
The structural observation here is that evaluation is lagging capability. Models can now hold 1M+ token contexts and recall user preferences across sessions, but there is no agreed-upon way to score whether that recall actually produces better outcomes. PAST-Bench, if it matures into a real suite, would be the first widely-cited attempt to close that gap.
There is also a contrarian angle worth noting. The name itself — "PAST" — implies that the field has been ignoring history. That is a fair charge against much of the agentic evaluation literature, which has favored static, reproducible task sets over dynamic, user-specific ones. Dynamic benchmarks are harder to build and harder to compare across labs, which is likely why they have been avoided.
The source does not disclose whether the benchmark will be open-sourced, what compute budget was used to design it, or whether any baseline models have been evaluated. Those are material omissions for a benchmark proposal, and the community should push for them before treating PAST-Bench as a standard.
What to Watch
Watch for the release of the full PAST-Bench paper or repository. If it lands with baseline numbers across frontier models — GPT-5, Claude Opus, Gemini 2.5 — it will immediately become a reference point for memory-adaptive agent claims. If it stays a thread, it will join a graveyard of benchmark teasers.






