A new benchmark called 'Show, Don't Tell' tests spatial cognition by requiring image-generation models to answer in pixels. GPT Image 2 solves 37% of spatial cases missed by GPT-5.4, revealing a gap in text-based spatial reasoning.
Key facts
- Benchmark called 'Show, Don't Tell'.
- GPT Image 2 solves 37% of spatial cases missed by GPT-5.4.
- Models answer in pixels, not text.
- Tests spatial cognition like object placement and orientation.
- Source is a single tweet from @HuggingPapers.
A new benchmark called 'Show, Don't Tell' tests spatial cognition by requiring image-generation models to answer in pixels, not text. According to @HuggingPapers, the benchmark reveals that GPT Image 2 solves 37% of spatial cases missed by GPT-5.4. This suggests that text-based spatial reasoning may miss key visual understanding that pixel-level outputs capture.
The benchmark evaluates models on spatial tasks such as object placement, orientation, and relative positioning. By requiring pixel-level outputs, it forces models to demonstrate understanding through generation rather than description. This approach may reveal gaps in models that perform well on text-based spatial reasoning tasks but fail to produce accurate spatial imagery.
Key Takeaways
- New 'Show, Don't Tell' benchmark tests spatial cognition via pixel-level outputs.
- GPT Image 2 solves 37% of cases missed by GPT-5.4, highlighting a gap in text-based spatial reasoning.
Why pixel-level evaluation matters
Text-based spatial reasoning benchmarks often rely on multiple-choice or free-text answers, which can mask incomplete understanding. The 'Show, Don't Tell' benchmark addresses this by requiring models to generate images that reflect spatial relationships. GPT Image 2's performance suggests that image-generation models can capture spatial details that text-only models miss, potentially because they are trained to produce coherent pixel outputs.
The 37% improvement on previously missed cases indicates a non-trivial gap. However, the benchmark does not disclose the number of test cases or the difficulty distribution, making it hard to compare across models or tasks. The source tweet from @HuggingPapers does not provide details on the dataset size or the specific spatial categories tested.
Limitations and open questions
The benchmark's small sample size and lack of published methodology limit its reliability. No code or dataset has been released, and the results are based on a single tweet. The claim that GPT Image 2 solves 37% of cases missed by GPT-5.4 is notable, but without a full paper or reproducible experiments, it remains an interesting data point rather than a definitive finding.
Additionally, the benchmark does not control for model training data overlap or prompt engineering. GPT Image 2 may have been trained on similar spatial tasks, while GPT-5.4 may not have been optimized for image generation. The comparison is therefore not apples-to-apples, and the results should be interpreted cautiously.
What to watch
Watch for a full paper or dataset release from the benchmark authors, which would allow reproducibility tests and broader model comparisons. If the benchmark gains traction, expect new spatial reasoning leaderboards for image-generation models.







