opus 4 8
30 articles about opus 4 8 in AI news
Opus 4.8 Still Owns 28% of Anthropic Spend
Ramp data shows Opus 4.8 at 28% spend vs Opus 5's 3.5%. Pin your Claude Code model to opus-4-8 for consistency, or use sonnet-4-6 for cheaper tasks.
Grok 4.6 Matches GPT-5.6, Opus 5; 4.7 (2.1T) Weeks Away
Grok 4.6 matches GPT-5.6 and Opus 5 on many benchmarks per @kimmonismus, with 2.1T-parameter Grok 4.7 arriving in weeks. The 1.5T model ships with Cursor integration.
How to Use Claude Opus 4.6 as Your Claude Code Orchestrator with Opus 5
Claude Code users can set Opus 4.6/4.8 as the main model and Opus 5 as subagents via `/model` and agent configs, balancing stability with intelligence for coding tasks.
DeepSeek-V4-Flash-Vision-Exp Matches Opus 4.8 on Visual-Agent Benchmarks
DeepSeek released V4-Flash-Vision-Exp, a small multimodal agent model that approaches or beats Opus 4.8 on visual benchmarks, per @kimmonismus.
27B Agent Beats Claude Opus 4.8, GPT-5.5 on Research Replication
A 27B agent named Replica reportedly beat Claude Opus 4.8 and GPT-5.5 on held-out research replication, per @omarsar0. No methodology or scores disclosed, so the claim is unverified but suggests efficiency can rival scale.
GLM-5.2 Nears Claude Opus 4.8 on Cyber-Defense Tests — Here's Why Claude
GLM-5.2 rivals Mythos 5 on cyber-defense. For Claude Code users: expect price cuts, better Opus 4.8 security, and new MCP options. Test your workflows with /model.
Claude Opus 5 Is Too Verbose: The Two-Model Split That Fixes It
Anthropic documented Opus 5's verbosity on launch day. Use the two-model split: Fable 5 plans, Opus 5 executes. This halves costs and restores readability.
Octopus Deploy's MCP Server: The Missing Onboarding Tool for Kubernetes Teams
Install Octopus Deploy's MCP server via `claude mcp add` to let Claude Code query environments, inspect releases, and trigger Kubernetes deployments—cutting onboarding friction and context-switching for ops teams.
Claude Code: When Should You Use Opus Max vs. Sonnet Low? A Cost-Per-Token
The key takeaway: match model strength and effort to task complexity. Sonnet Max beats Opus Low for structured work; Opus Max wins on novel problems. Use /model and --max-effort to optimize.
Prime Intellect's Prime Agent Hits 95.5% on ARC-AGI-3 With Opus 5
Prime Intellect's open-source Prime Agent scored 95.5% on ARC-AGI-3 with Opus 5, exceeding the human baseline via a self-improving RLM harness.
Claude Tool Use: Fable 5 Beats Opus 4.8 at 1.00 Calls
SemiAnalysis analyzed 2.27M Claude responses, finding Fable 5 averages 1.00 tool calls per response versus 0.76 for Opus 4.8. The Opus line shows a downward trend.
Opus 5 Backlash Signals Anthropic's Souring Developer Mood
Opus 5 launch draws harsh developer criticism per X post, signaling Anthropic's eroding technical brand amid competitive pressure.
Opus 5 Generates Playable Game for $423 in One Prompt
Opus 5 generated a playable game from one prompt, using 690M tokens at $423, per @kimmonismus. One person replaced a dev team.
Anthropic Ships Opus 5: Near-Fable Coding at Half Price
Anthropic launched Claude Opus 5 at $5/$25 per million tokens, matching Fable 5 on CursorBench at half cost. ARC-AGI 3 score 3x next-best; Claude Code v2.1.219 adds ops features.
Open-Weight Models Just Matched Claude Opus 4.6 — Here's How to Run Them
Route Claude Code to open-weight models like Kimi K3 and DeepSeek V4 Flash via ANTHROPIC_BASE_URL or LiteLLM. Test cheap models before spending Opus 4.6 credits. The open-weight revolution is now a Claude Code workflow decision.
Cut Your CLAUDE.md Rules 57% with This Opus 5 Audit Procedure
Audit CLAUDE.md rules against Claude Opus 5's system prompt and tool descriptions. One user cut 5,789 words to 2,463 by removing conflicts and redundancies.
Anthropic Ships Claude Opus 5: Fable-Level Intelligence at Half the Price
Anthropic released Claude Opus 5 on July 24 with a 1M token context, 128k output, and Fable-5-approaching intelligence at half the price, unchanged from Opus 4.8.
Opus 5 Hits 0% Prompt Injection Rate in Browser Agents
Anthropic's Opus 5 with Auto Mode achieved 0% prompt injection success across 129 tests, challenging OpenAI's view that the problem is unsolvable.
GPT-5.6 Sol Leads DeepSWE at 72.7%, Beating Opus 5's 68.8%
GPT-5.6 Sol scores 72.7% on DeepSWE, beating Opus 5's 68.8%. The undocumented benchmark tests autonomous SWE agents.
Claude Opus 5 Is Now in Claude Code: How to Use Fast Mode and Save 50% on Tokens
Claude Opus 5 is now in Claude Code with Fast Mode (2.5x speed) at Opus 4.8 pricing. Run `claude code --model opus-5` to start saving 50% on tokens immediately.
Traders Bet Claude Opus 4.8 Launch Imminent as Options Spike
Traders bet Anthropic will launch Claude Opus 4.8 within days, based on options market activity. The model would succeed Opus 4.7 (69.2% SWE-bench Pro) and compete with GPT-5.
Claude Opus 4.8 Now Beats Gemini Pro 5 in Coding Benchmarks — What It
Claude Opus 4.8 beats Gemini Pro 5 by 11 points on Fable 5. Claude Code users should run `claude code --model opus-4.8` for complex coding tasks.
Databricks Defaults to Chinese Model GLM 5.2, Matches Opus at $1.28/Task
Databricks defaulted to GLM 5.2 after it matched Opus 4.8 at $1.28/task vs $1.94. The move signals enterprises building custom benchmarks and multi-vendor AI stacks.
Anthropic Claims Claude Opus 4.7 Hits 92% Honesty, Cuts Sycophancy
Anthropic's Claude Opus 4.7 scores 92% on internal honesty benchmark, reducing sycophancy. The model also improves SWE-Bench to 79.8, up from 71.2.
Claude Sonnet 5 Beats Opus 4.8 on Knowledge Work at Lower Cost
Anthropic released Claude Sonnet 5, which beats Sonnet 4.6 across all benchmarks and edges past Opus 4.8 on GDPval-AA v2 with a score of 1,618.
GLM-5.2 matches Opus 4.7 at 1/5 the price in Snowflake coding test
Zhipu AI's GLM-5.2 matched Claude Opus 4.7 on a Snowflake coding benchmark at one-fifth the cost, threatening Western AI lab pricing and IPO valuations.
WorkBench Revisited: Claude Opus 4.8 Hits 89% Task Completion
Claude Opus 4.8 completes 89% of WorkBench tasks with 2.5% harm rate, up from GPT-4's 43% and 26% in 2024, showing capability and safety align.
Fable 5: Claude's Biggest Leap Since Opus 4.5, Says Beta Tester
Beta tester says Fable 5 is Claude's biggest leap since Opus 4.5, with emergent debugging and design capabilities.
Anthropic Releases Claude Mythos Publicly as 'Fable' at 2x Opus Price
Anthropic released Claude Mythos publicly as 'Fable' at 2x Opus pricing, targeting agent workflows with strong safety limits.
Claude Opus 4.7 Matches Dedicated NMR Software on Chemistry Tasks
Claude Opus 4.7 matches NMR software on chemistry tasks per Anthropic blog, but methodology and benchmarks undisclosed.