coding
30 articles about coding in AI news
KAT-Coder-V2.5-Dev: The Open Agentic Coding Model That Could Rival Claude Code
KAT-Coder-V2.5-Dev lets you run agentic coding locally. Pair it with Claude Code for hybrid workflows: use it for private code, keep Claude for complex reasoning.
Gemini 3.7 Flash Ships Improved Long-Horizon Coding
Google released Gemini 3.7 Flash with improved long-horizon coding and PDF understanding. No benchmarks or pricing disclosed.
OpenAI Codex Hits 15M Users, Marking AI Coding Reset
OpenAI's Codex surpassed 15 million users, per @kimmonismus, signaling a reset in AI coding assistants. The milestone pressures rivals like GitHub Copilot.
Meta Ships Muse Code Beta, Its First Coding Agent
Meta released Muse Code beta, its first coding agent built on Muse Spark 1.2. The release enters a market led by GitHub Copilot and Amazon CodeWhisperer.
SemiAnalysis Runs Coding Agents on Its Own Research Workflow
SemiAnalysis is using coding agents internally for data collection, charting, and drafting. No metrics disclosed, but signals production shift.
Vibe Coding to Agentic Coding: A Field Report from Taiwan
TitanSoft's field report argues AI coding shifts from vibe to agentic coding, requiring engineers to design feedback loops. No metrics disclosed.
Open-Source Course Shows Harness, Not Model, Lifts Coding Agent 25 Places
Open-source course shows harness engineering, not model swap, moved a coding agent from ~30th to top 5 on Terminal-Bench. Course builds Decode from scratch.
Agentic Coding Tools Flood Market as Enterprise Adoptions Triple in 2026
Enterprise adoption of agentic coding tools tripled in 2026, led by Google, OpenAI, and Anthropic. Productivity gains of 40-60% reported.
Kimi K3 Tops US Models in Front-End Coding at Smaller Scale
Moonshot AI's K3 tops US models in front-end coding at 89.2% on SWE-bench while being smaller and cheaper to train.
Emergent Raises at $1.5B Valuation for Vibe Coding IDE
Emergent raised at $1.5B valuation for vibe coding IDE. Round details undisclosed, signaling sustained investor appetite for AI developer tools.
social.plus Vise: Workflow Governance for AI Coding Agents Building SDK
social.plus launched Vise, a workflow governance platform for AI coding agents building SDK integrations, enforcing policy controls and audit trails.
PadCaptioner: 3B video caption model beats 7B rivals with parallel decoding
PadCaptioner, a 3B model, beats 7B rivals in dense video captioning via lossless parallel autoregressive decoding, challenging scaling orthodoxy.
Claude Opus 4.8 Now Beats Gemini Pro 5 in Coding Benchmarks — What It
Claude Opus 4.8 beats Gemini Pro 5 by 11 points on Fable 5. Claude Code users should run `claude code --model opus-4.8` for complex coding tasks.
DeepSeek DSpark: Speculative Decoding Unifies Parallel Gen, Adaptive Verification
DeepSeek released DSpark, a speculative decoding framework unifying parallel generation with adaptive verification. No benchmarks disclosed yet; the approach targets inference latency and throughput.
Databricks Tests Coding Agents on Its Own Codebase
Databricks benchmarked coding agents on its own polyglot codebase. GLM-5.2 matched top closed models, a minimal harness halved costs, and cheaper-per-token models cost more per task.
Meta Muse Spark 1.1 Debuts in AI Coding Battle; Zuck Post Hits 12M Views
Meta released Muse Spark 1.1 for agentic coding tasks. Zuckerberg's post got 12M views in 12 hours; no benchmarks disclosed.
OpenAI Claims 54% Token Efficiency Gain on Agentic Coding in New Model
OpenAI CEO Sam Altman claims 54% token efficiency gain on agentic coding for a new unnamed model, but no technical details or release date were provided.
SpaceXAI Ships Grok 4.5, Blackwell-Trained Coding Model
SpaceXAI released Grok 4.5, a coding-focused model trained on Blackwell GPUs, now available in Cursor and Vercel. Inference cost claims lack independent benchmarks.
GitHub's Former CEO Launches Distributed Git Network for AI Coding Agents
Claude Code users should monitor Nat Friedman's distributed Git network for faster agentic coding workflows. The new network optimizes Git for AI agents, potentially reducing clone/push latency.
Lovable spent $85K on tokens to learn agentic coding at scale
Lovable spent $85K on tokens for agentic coding. Debugging costs dominate, challenging enterprise adoption.
Vibe Coding Fails: Why AI-Generated Code Breaks at Scale
Vibe coding fails because AI-generated code lacks architectural coherence, test coverage, and security validation, breaking at scale beyond 1,000 lines.
MirrorCode Benchmark Costs $2,600 Per Run, Challenges AI Coding Limits
Epoch AI and METR launched MirrorCode, a $2,600-per-run coding benchmark. Claude Opus 4.7 leads with 56% solve rate.
JetSpec hits 1,000 t/s on Qwen-8B with speculative decoding
JetSpec achieves 1,000 t/s on Qwen-8B with a B200 GPU, claiming superiority over prior speculative decoding methods, but lacks independent verification.
Zhipu GLM-5.2 tops global coding benchmarks, sparks 'DeepSeek moment'
Zhipu AI's GLM-5.2 ranks top-3 globally on a coding benchmark, with US engineers calling it a daily driver superior to GPT-5.5.
GLM-5.2 matches Opus 4.7 at 1/5 the price in Snowflake coding test
Zhipu AI's GLM-5.2 matched Claude Opus 4.7 on a Snowflake coding benchmark at one-fifth the cost, threatening Western AI lab pricing and IPO valuations.
Chinese Lab's Free MoE Model Matches GPT-5.5 on Agentic Coding
A Chinese lab released an Apache-2.0 open-weights MoE model matching GPT-5.5 on agentic coding. This free model challenges proprietary AI's lead with sparse MoE architecture.
OpenAI Buys Ona to Give Codex Multi-Day Autonomous Coding
OpenAI acquired Ona (formerly Gitpod) to give Codex persistent cloud environments for autonomous coding tasks lasting hours or days, targeting Anthropic's Claude Code lead.
GitHub Spec Kit: Open-Source Tool to Fix Vibe Coding’s Core Flaw
GitHub released Spec Kit, an open-source toolkit that enforces specification-first workflows for AI coding, addressing vibe coding's tendency to generate code before requirements are clear.
MiniMax M3 Sparse Attention: 15.6x Decoding Speedup at 1M Tokens
MiniMax M3 sparse attention achieves 9.7x prefilling and 15.6x decoding speedup at 1M tokens, reversing M2's full-attention stance.
No Rigorous Productivity Tests Exist for Post-2025 Autonomous Coding Tools
No productivity studies exist for autonomous coding tools launched December 2025. All research predates the Claude Code/Codex revolution, creating a major knowledge gap.