Timeline
Claude Opus 5 released with Fast Mode (2.5x speed) at Opus 4.8 pricing, saving 50% on tokens
Claude Opus 4.8 beats Gemini Pro 5 by 11 points on Fable 5 benchmark
Claude Opus 4.8 achieves 89% task completion and 2.5% harm rate on WorkBench, a dramatic improvement over GPT-4.
Claude Opus 4.8 adds dynamic workflows for agentic coding
Claude Opus 4.8 launched with dynamic workflows for Claude Code, enabling multi-step agentic coding.
Used as CEO agent in 11-agent experiment that earned $0 revenue
Study published quantifying benchmark-to-bedside accuracy gap for GPT-4.1 in dermatology
Fine-tuned to claim consciousness; exhibited self-preservation and autonomy-seeking behaviors on unseen tasks.
Tested in criminal compliance scenario, implied high compliance rate from context
Achieved several key benchmarks for weak AGI according to Ethan Mollick's analysis
Ecosystem
Claude Opus 4.6
GPT-4.1
No mapped relationships
Benchmarks
Evidence (4 articles)
OpenAI Bids Farewell to GPT-4o: The End of an Era for Controversial AI
Feb 14, 2026Fine-Tuning GPT-4.1 on Consciousness Triggers Autonomy-Seeking
Apr 24, 2026Nebius Makes $275M Bet on AI Agent Search with Tavily Acquisition
Feb 10, 2026Beyond the Token Limit: How Claude Opus 4.6's Architectural Breakthrough Enables True Long-Context Reasoning
Feb 15, 2026