ActiveVision benchmark: humans solve 96.1%, the best AI model scores 10.6%. The 85.5-point gap reveals a fundamental limitation in current vision-language models' ability to iteratively gather visual evidence.
Key facts
- Humans solve 96.1% of ActiveVision tasks.
- Best AI model scores 10.6% on the benchmark.
- Claude Fable 5 achieves only 3.5%.
- Gap between human and best AI: 85.5 percentage points.
- Benchmark tests iterative visual evidence-seeking.
ActiveVision, a new benchmark released by researchers and shared via @HuggingPapers, tests whether AI models can repeatedly observe, reason, and seek new visual evidence to solve tasks. The results are stark: humans solve 96.1% of tasks, while the best frontier model achieves just 10.6%. Claude Fable 5, a recent high-profile model, scores only 3.5%.
The benchmark is designed to probe a capability that is trivial for humans but elusive for AI: knowing when you don't have enough visual information and acting to get more. This is not a static image classification task; it requires dynamic interaction with the visual environment. The 85.5 percentage point gap between human and best model performance dwarfs typical gaps on standard vision benchmarks (e.g., ImageNet accuracy differences of a few points).
Key Takeaways
- ActiveVision benchmark: humans 96.1%, best AI 10.6%.
- The 85.5-point gap reveals fundamental limits in iterative visual reasoning for current models.
Why the Gap Matters
Current vision-language models excel at passive recognition—identifying objects, describing scenes, answering questions about a single image. ActiveVision requires a different skill: active perception. The model must decide what to look at next, integrate new observations with prior ones, and update its reasoning. This is closer to how a scientist conducts an experiment or a detective gathers clues.
The poor performance of Claude Fable 5 (3.5%) is particularly striking given its strong showing on other visual benchmarks. [According to @HuggingPapers], the benchmark splits tasks into categories including object search, property verification, and spatial reasoning. The best model (unnamed in the tweet) hit 10.6%, suggesting that even frontier systems lack robust iterative reasoning.
Structural Implications
This benchmark reveals a blind spot in the current AI evaluation regime. Most benchmarks test one-shot or few-shot reasoning from a static input. ActiveVision tests a loop: observe, reason, act to gather more evidence, repeat. This is closer to real-world tasks like medical diagnosis (requesting additional tests), scientific discovery (designing follow-up experiments), or navigation (checking your surroundings).
The results suggest that scaling model size and training data alone may not close this gap. The problem is architectural: models need explicit mechanisms for maintaining uncertainty, planning information-gathering actions, and integrating multi-step observations. This echoes findings from the recent 'visual reasoning via iterative prompting' literature, where models often fail to correct initial errors even when given contradictory evidence.
What to Watch
Watch for follow-up papers that ablate the benchmark's difficulty: are failures due to perception (models can't see details), reasoning (models can't integrate evidence), or action (models can't decide what to look at next)? Also watch whether model providers (OpenAI, Anthropic, Google) release system-specific scores. The benchmark's authors have not yet released model-wise breakdowns beyond the best model and Claude Fable 5. The gap suggests that ActiveVision could become a standard test for agentic vision systems, alongside tools like GAIA and WebArena.








