models
30 articles about models in AI news
SemiAnalysis: Open Models Still Trail Closed Frontier by 1-2 Years
SemiAnalysis argues open models still trail closed frontier by 1-2 years, with post-training and inference-time compute as key differentiators. The gap persists across all training eras.
Alibaba's ABSeeker Lets 4B Agent Match 30B Search Models
Alibaba's ABSeeker adds step-level credit assignment, letting a 4B search agent match ~30B models. Backtracking from answers densifies reward signals.
Chinese Models Claim Top-3 Front-End Design Spots
Chinese models Kimi and Qwen now hold two of the top three front-end design spots, sharing the podium with Anthropic's Claude. Parity signals a narrowing capability gap.
Open-Weight Models Just Matched Claude Opus 4.6 — Here's How to Run Them
Route Claude Code to open-weight models like Kimi K3 and DeepSeek V4 Flash via ANTHROPIC_BASE_URL or LiteLLM. Test cheap models before spending Opus 4.6 credits. The open-weight revolution is now a Claude Code workflow decision.
NVIDIA's Molt: 9.2K-Line RL Framework Scales to 1T-Parameter MoE Models
NVIDIA released Molt, a 9.2K-line PyTorch RL framework scaling to 1T-parameter MoE models via vLLM, targeting agentic tasks with fully-async rollout.
Cohere Open-Sources Three AI Models Under Apache 2.0
Cohere released three open-source AI models under Apache 2.0 in 2025, expanding its enterprise portfolio with speech, language, and code capabilities.
Jensen Huang: DeepSeek, Kimi open models boost Nvidia sales
Jensen Huang says Chinese open models DeepSeek and Kimi boost Nvidia GPU demand, not threaten it. Market misunderstood their impact twice.
Alibaba Releases RynnBrain 1.1 Embodied AI Models at 2B-122B Scales
Alibaba released RynnBrain 1.1 on Hugging Face with 2B, 9B, and 122B-A10B MoE models for robot manipulation, but disclosed no benchmarks.
Decoy Font Tricks AI Vision Models With Dual-Layer Glyphs
Mixfont's Decoy Font hides text from AI vision models by layering two characters into one glyph, exploiting a tokenization blind spot in ChatGPT and Gemini.
Google Ships 3 Flash Models as 3.5 Pro Remains Missing
Google shipped three Gemini Flash models but 3.5 Pro remains delayed. Efficiency gains don't close the frontier gap with OpenAI and Anthropic.
Vercel Data: Open Models Spend Collapses to All-Time Low
Closed AI models hit 97.09% spend share via Vercel; open models at all-time low over past 5 days.
Trump Weighs Restrictions on US Firms Using Chinese AI Models
Trump admin weighs restrictions on US firms using Chinese AI models, per Axios. Businesses already adopting cheaper Chinese alternatives, creating policy tension.
Kimi K3 Tops US Models in Front-End Coding at Smaller Scale
Moonshot AI's K3 tops US models in front-end coding at 89.2% on SWE-bench while being smaller and cheaper to train.
Murati's Thinking Machines Ships 975B Inkling — Leads US Open Models
Murati's Thinking Machines releases Inkling, a 975B-parameter MoE model that leads US open models but trails Chinese rivals on benchmarks and cost.
Function-Aware Fill-in-the-Middle Boosts SWE-Bench by +5.4 on 14B Models
Function-aware FIM mid-training boosts SWE-Bench by +2.8 to +5.4 on 7B-14B models, preserving general abilities. Six checkpoints and 400K dataset open-sourced.
Google alone ships full any-to-any multimodal models
Mollick notes Google alone ships full any-to-any multimodal models; OpenAI and Anthropic lag. This gives Google a structural advantage in agentic workflows.
Open-weight models now run 29% of gateway tokens, up from 11% in April
Open-weight models now handle 29% of gateway tokens, up from 11% in April. The 18-point jump signals accelerating enterprise adoption of open architectures like Llama 3 and Mistral.
Nvidia, Hugging Face Open-Source Robot Models to Democratize Physical AI
Nvidia and Hugging Face open-sourced robot models to democratize physical AI, providing pre-trained models and simulation tools on the Hugging Face hub.
GPT-4 Held Top Spot 52 Weeks; Today's Models Last 7
GPT-4 dominated the ECI for a year. Today's top models last 7 weeks median, with 17 leadership changes since Feb 2024.
Mira Murati's Thinking Machines beats frontier models by 29.8% with Bridgewater's expert judgments
Thinking Machines beat frontier models by 29.8% fewer errors using Bridgewater's expert judgments, at 13.8x lower inference cost.
OpenAI Cuts Inference Costs by Half on Some Models
OpenAI cut inference costs by 50%+ on some models for logged-out ChatGPT users, per The Information. The move reduces operational expenses.
Epoch AI's EBR-Bench: Top Models Score 30-50% on Experience-Based Reasoning
Epoch AI's EBR-Bench tests experience-based reasoning. Top models score 30-50%, with Google Gemini 3 Pro leading at 48.2%, revealing a gap between pattern matching and true learning.
Trump Lifts Export Ban on Anthropic’s Mythos, Fable Models
U.S. lifted export ban on Anthropic's Mythos and Fable models June 30. Anthropic restores access July 1 under deal requiring proactive security risk detection.
CLI-Universe: Qwen3-32B fine-tuned on 6K trajectories beats models 10x larger on Terminal-Bench 2.0
CLI-Universe synthesizes terminal-agent tasks; Qwen3-32B fine-tuned on 6K trajectories hits 33.4% on Terminal-Bench 2.0, beating models 10x larger.
World Action Models Survey Unifies 100+ Methods Under One Taxonomy
A survey reviews 100+ world action models, unifying world models, video generation, and VLA policies under one taxonomy.
BeliefDiffusion Uses Diffusion Models for Robot Navigation in Partially
BeliefDiffusion combines diffusion models with MPC for robot navigation in partially observable environments, outperforming model-free RL and generative baselines in synthetic maps.
MA-ProofBench: GPT-5.5 Hits 16% on Math Analysis, Most Models Near 0%
MA-ProofBench, a new theorem-proving benchmark for mathematical analysis, shows GPT-5.5 achieving 16% on undergraduate problems and 5% on PhD-level, with most models near 0% on the harder set.
US Gov’t Orders Anthropic to Shut Down Strongest Claude Models
US ordered Anthropic to shut down strongest Claude models via export controls. No official confirmation yet.
Trillion Labs Builds Industrial World Models on NVIDIA Omnibus
Trillion Labs announced Industrial World Models for AI Factories using NVIDIA Omniverse and Nemotron to optimize data centers and power plants.
Larger models learn rare skills by forgetting them less, new paper shows
New paper from Stanford, MIT, Harvard, and Anthropic shows larger models learn rare skills because they forget them less during training, tested on OLMo models from 4M to 4B parameters.