StreamArena, announced via @HuggingPapers, introduces 243 full-length videos averaging 88.8 minutes for hour-scale interactive streaming video understanding. The benchmark's 3,646 open-ended tasks test real-time perception, historical recall, tool use, and proactive interaction.
Key facts
- 243 full-length videos in StreamArena
- Average video length: 88.8 minutes
- 3,646 open-ended tasks total
- Four task types: perception, recall, tool use, interaction
- Announced via @HuggingPapers on X
StreamArena, a new benchmark announced via @HuggingPapers, targets hour-scale, interactive streaming video understanding with 243 full-length videos averaging 88.8 minutes each. The dataset includes 3,646 open-ended tasks spanning four skill areas: real-time perception, historical recall, tool use, and proactive interaction.
Why hour-scale matters
Most video benchmarks rely on short clips—typically 10-30 seconds—that fail to stress long-context memory and real-time reasoning. StreamArena's 88.8-minute average video length marks a departure from that convention, pushing models to maintain state and answer queries that require recalling events from earlier in the video. This mirrors real-world applications like live surveillance, sports broadcasting, and long-form content moderation, where models must track evolving scenes and respond to user interruptions.
The open-ended task design is also notable. Rather than multiple-choice or clip-level classification, tasks require generating free-form responses, which better reflect interactive use cases such as asking a model "what happened just before the goal?" or "summarize the last 20 minutes." Tool use tasks likely involve querying external systems—like timestamps or metadata—while proactive interaction tasks may require the model to initiate updates without being prompted.
The announcement does not disclose model performance numbers or baseline results. It is not yet clear which models have been evaluated or how they fare on the 3,646 tasks. The source provides no leaderboard, no scoring methodology, and no comparison to prior benchmarks like Video-MME or LongVideoBench, leaving open questions about difficulty and saturation.
What makes StreamArena different
StreamArena's focus on interactivity—not just passive understanding—sets it apart. Most long-video benchmarks, such as Video-MME or EgoSchema, evaluate a model's ability to answer questions after watching a video. StreamArena, by contrast, emphasizes interactive streaming, where a model must handle real-time queries, recall past events, and potentially use tools mid-stream. This aligns with the growing deployment of AI assistants in live environments, where latency and memory are as critical as accuracy.
The 243-video scale is modest compared to some datasets, but the hour-length videos mean the total duration exceeds 360 hours of footage. That is a significant compute and annotation effort, and it suggests the benchmark is designed for depth over breadth.
Open questions
No evaluation results are provided, so it is impossible to gauge current model performance. The source does not specify whether videos are from a single domain or diverse sources, nor does it detail the annotation process for the 3,646 tasks. The absence of baselines means the benchmark's difficulty is unverified. As with any new benchmark, watch for overfitting or task leakage in future submissions.
StreamArena is a timely contribution, addressing the gap between short-clip evaluation and real-world streaming demands. Whether it becomes a standard reference will depend on its accessibility, reproducibility, and the transparency of its scoring pipeline—details that the announcement does not yet provide.
What to watch
Watch for the release of the dataset and scoring code, plus any baseline evaluations from the authors. If StreamArena gains traction, expect third-party model runs and comparisons against Video-MME or LongVideoBench. A public leaderboard would be the next concrete signal of adoption.








