Skip to content
gentic.news — AI News Intelligence Platform
Connecting to the Living Graph…

opus 4 8

30 articles about opus 4 8 in AI news

Opus 4.8 Still Owns 28% of Anthropic Spend

Ramp data shows Opus 4.8 at 28% spend vs Opus 5's 3.5%. Pin your Claude Code model to opus-4-8 for consistency, or use sonnet-4-6 for cheaper tasks.

70% relevant

Grok 4.6 Matches GPT-5.6, Opus 5; 4.7 (2.1T) Weeks Away

Grok 4.6 matches GPT-5.6 and Opus 5 on many benchmarks per @kimmonismus, with 2.1T-parameter Grok 4.7 arriving in weeks. The 1.5T model ships with Cursor integration.

84% relevant

How to Use Claude Opus 4.6 as Your Claude Code Orchestrator with Opus 5

Claude Code users can set Opus 4.6/4.8 as the main model and Opus 5 as subagents via `/model` and agent configs, balancing stability with intelligence for coding tasks.

80% relevant

DeepSeek-V4-Flash-Vision-Exp Matches Opus 4.8 on Visual-Agent Benchmarks

DeepSeek released V4-Flash-Vision-Exp, a small multimodal agent model that approaches or beats Opus 4.8 on visual benchmarks, per @kimmonismus.

100% relevant

27B Agent Beats Claude Opus 4.8, GPT-5.5 on Research Replication

A 27B agent named Replica reportedly beat Claude Opus 4.8 and GPT-5.5 on held-out research replication, per @omarsar0. No methodology or scores disclosed, so the claim is unverified but suggests efficiency can rival scale.

100% relevant

GLM-5.2 Nears Claude Opus 4.8 on Cyber-Defense Tests — Here's Why Claude

GLM-5.2 rivals Mythos 5 on cyber-defense. For Claude Code users: expect price cuts, better Opus 4.8 security, and new MCP options. Test your workflows with /model.

88% relevant

Claude Opus 5 Is Too Verbose: The Two-Model Split That Fixes It

Anthropic documented Opus 5's verbosity on launch day. Use the two-model split: Fable 5 plans, Opus 5 executes. This halves costs and restores readability.

85% relevant

Octopus Deploy's MCP Server: The Missing Onboarding Tool for Kubernetes Teams

Install Octopus Deploy's MCP server via `claude mcp add` to let Claude Code query environments, inspect releases, and trigger Kubernetes deployments—cutting onboarding friction and context-switching for ops teams.

80% relevant

Claude Code: When Should You Use Opus Max vs. Sonnet Low? A Cost-Per-Token

The key takeaway: match model strength and effort to task complexity. Sonnet Max beats Opus Low for structured work; Opus Max wins on novel problems. Use /model and --max-effort to optimize.

75% relevant

Prime Intellect's Prime Agent Hits 95.5% on ARC-AGI-3 With Opus 5

Prime Intellect's open-source Prime Agent scored 95.5% on ARC-AGI-3 with Opus 5, exceeding the human baseline via a self-improving RLM harness.

87% relevant

Claude Tool Use: Fable 5 Beats Opus 4.8 at 1.00 Calls

SemiAnalysis analyzed 2.27M Claude responses, finding Fable 5 averages 1.00 tool calls per response versus 0.76 for Opus 4.8. The Opus line shows a downward trend.

100% relevant

Opus 5 Backlash Signals Anthropic's Souring Developer Mood

Opus 5 launch draws harsh developer criticism per X post, signaling Anthropic's eroding technical brand amid competitive pressure.

80% relevant

Opus 5 Generates Playable Game for $423 in One Prompt

Opus 5 generated a playable game from one prompt, using 690M tokens at $423, per @kimmonismus. One person replaced a dev team.

100% relevant

Anthropic Ships Opus 5: Near-Fable Coding at Half Price

Anthropic launched Claude Opus 5 at $5/$25 per million tokens, matching Fable 5 on CursorBench at half cost. ARC-AGI 3 score 3x next-best; Claude Code v2.1.219 adds ops features.

100% relevant

Open-Weight Models Just Matched Claude Opus 4.6 — Here's How to Run Them

Route Claude Code to open-weight models like Kimi K3 and DeepSeek V4 Flash via ANTHROPIC_BASE_URL or LiteLLM. Test cheap models before spending Opus 4.6 credits. The open-weight revolution is now a Claude Code workflow decision.

71% relevant

Cut Your CLAUDE.md Rules 57% with This Opus 5 Audit Procedure

Audit CLAUDE.md rules against Claude Opus 5's system prompt and tool descriptions. One user cut 5,789 words to 2,463 by removing conflicts and redundancies.

100% relevant

Anthropic Ships Claude Opus 5: Fable-Level Intelligence at Half the Price

Anthropic released Claude Opus 5 on July 24 with a 1M token context, 128k output, and Fable-5-approaching intelligence at half the price, unchanged from Opus 4.8.

100% relevant

Opus 5 Hits 0% Prompt Injection Rate in Browser Agents

Anthropic's Opus 5 with Auto Mode achieved 0% prompt injection success across 129 tests, challenging OpenAI's view that the problem is unsolvable.

100% relevant

GPT-5.6 Sol Leads DeepSWE at 72.7%, Beating Opus 5's 68.8%

GPT-5.6 Sol scores 72.7% on DeepSWE, beating Opus 5's 68.8%. The undocumented benchmark tests autonomous SWE agents.

100% relevant

Claude Opus 5 Is Now in Claude Code: How to Use Fast Mode and Save 50% on Tokens

Claude Opus 5 is now in Claude Code with Fast Mode (2.5x speed) at Opus 4.8 pricing. Run `claude code --model opus-5` to start saving 50% on tokens immediately.

100% relevant

Traders Bet Claude Opus 4.8 Launch Imminent as Options Spike

Traders bet Anthropic will launch Claude Opus 4.8 within days, based on options market activity. The model would succeed Opus 4.7 (69.2% SWE-bench Pro) and compete with GPT-5.

75% relevant

Claude Opus 4.8 Now Beats Gemini Pro 5 in Coding Benchmarks — What It

Claude Opus 4.8 beats Gemini Pro 5 by 11 points on Fable 5. Claude Code users should run `claude code --model opus-4.8` for complex coding tasks.

100% relevant

Databricks Defaults to Chinese Model GLM 5.2, Matches Opus at $1.28/Task

Databricks defaulted to GLM 5.2 after it matched Opus 4.8 at $1.28/task vs $1.94. The move signals enterprises building custom benchmarks and multi-vendor AI stacks.

100% relevant

Anthropic Claims Claude Opus 4.7 Hits 92% Honesty, Cuts Sycophancy

Anthropic's Claude Opus 4.7 scores 92% on internal honesty benchmark, reducing sycophancy. The model also improves SWE-Bench to 79.8, up from 71.2.

75% relevant

Claude Sonnet 5 Beats Opus 4.8 on Knowledge Work at Lower Cost

Anthropic released Claude Sonnet 5, which beats Sonnet 4.6 across all benchmarks and edges past Opus 4.8 on GDPval-AA v2 with a score of 1,618.

100% relevant

GLM-5.2 matches Opus 4.7 at 1/5 the price in Snowflake coding test

Zhipu AI's GLM-5.2 matched Claude Opus 4.7 on a Snowflake coding benchmark at one-fifth the cost, threatening Western AI lab pricing and IPO valuations.

85% relevant

WorkBench Revisited: Claude Opus 4.8 Hits 89% Task Completion

Claude Opus 4.8 completes 89% of WorkBench tasks with 2.5% harm rate, up from GPT-4's 43% and 26% in 2024, showing capability and safety align.

84% relevant

Fable 5: Claude's Biggest Leap Since Opus 4.5, Says Beta Tester

Beta tester says Fable 5 is Claude's biggest leap since Opus 4.5, with emergent debugging and design capabilities.

100% relevant

Anthropic Releases Claude Mythos Publicly as 'Fable' at 2x Opus Price

Anthropic released Claude Mythos publicly as 'Fable' at 2x Opus pricing, targeting agent workflows with strong safety limits.

100% relevant

Claude Opus 4.7 Matches Dedicated NMR Software on Chemistry Tasks

Claude Opus 4.7 matches NMR software on chemistry tasks per Anthropic blog, but methodology and benchmarks undisclosed.

94% relevant