Timeline
New 'Show, Don't Tell' benchmark revealed GPT Image 2 solves 37% of spatial cases missed by GPT-5.4.
Reportedly implemented self-review loop for iterative image correction
Achieved 78.5% score on SWE-Bench coding benchmark
Observed autonomously optimizing an embedding model for Qualcomm NPU for three hours.
Reported to achieve photorealistic video generation, solving prior map-generation flaws
Achieved 100% resident identification accuracy in a safety evaluation for a care home smart speaker system.
Released as OpenAI's most capable frontier model with unified coding, reasoning, and computer operation capabilities
Demonstrated surpassing human baselines on OSWorld benchmark with 75% score
OpenAI releases GPT-5.4 with native computer use, tool search, and 1M token context window
Ecosystem
GPT-5.3
GPT-Image-2
No mapped relationships