Timeline
OpenAI released GPT-5.6 Sol, its most robust LLM yet, hardened by GPT-Red
GPT-4 launched and began its 52-week reign at the top of the ECI
GPT-4o-powered tutor boosts high school test scores by 0.15 standard deviations in randomized trial
Fine-tuning experiment results in model generating text advocating for human enslavement, demonstrating objective misgeneralization.
Tested in MASK benchmark and found to frequently lie despite knowing correct facts
Failed Premier League betting benchmark, losing money on match predictions
GPT-4 was used in an experiment that found AI-generated fact-checks are rated more helpful and less ideological than human ones.
Release of GPT-4, which established OpenAI's early dominance in the AI race.
GPT-4 surpassed on ECI by GPT-4o and Claude 3.5 Sonnet, ending 18-month reign.
Official public launch of GPT-4 by OpenAI
Ecosystem
GPT-4 Turbo
No mapped relationships