model
30 articles about model in AI news
Robots Learn Self-Supervised Progress Tracking via Reward Modeling Survey
Survey unifies progress reward modeling for robots to self-assess advancement, stagnation, or regression during tasks, replacing binary success signals.
NVIDIA's Molt: 9.2K-Line RL Framework Scales to 1T-Parameter MoE Models
NVIDIA released Molt, a 9.2K-line PyTorch RL framework scaling to 1T-parameter MoE models via vLLM, targeting agentic tasks with fully-async rollout.
Cohere Open-Sources Three AI Models Under Apache 2.0
Cohere released three open-source AI models under Apache 2.0 in 2025, expanding its enterprise portfolio with speech, language, and code capabilities.
Jensen Huang: DeepSeek, Kimi open models boost Nvidia sales
Jensen Huang says Chinese open models DeepSeek and Kimi boost Nvidia GPU demand, not threaten it. Market misunderstood their impact twice.
Alibaba Releases RynnBrain 1.1 Embodied AI Models at 2B-122B Scales
Alibaba released RynnBrain 1.1 on Hugging Face with 2B, 9B, and 122B-A10B MoE models for robot manipulation, but disclosed no benchmarks.
Offloop's D1 dispatcher model fixes multi-agent chaos
Offloop's D1 dispatcher model prevents multi-agent channel noise by assigning turns and escalating stuck tasks to humans, as shown in an overnight benchmark run.
Decoy Font Tricks AI Vision Models With Dual-Layer Glyphs
Mixfont's Decoy Font hides text from AI vision models by layering two characters into one glyph, exploiting a tokenization blind spot in ChatGPT and Gemini.
Open-Source Course Shows Harness, Not Model, Lifts Coding Agent 25 Places
Open-source course shows harness engineering, not model swap, moved a coding agent from ~30th to top 5 on Terminal-Bench. Course builds Decode from scratch.
Google Ships 3 Flash Models as 3.5 Pro Remains Missing
Google shipped three Gemini Flash models but 3.5 Pro remains delayed. Efficiency gains don't close the frontier gap with OpenAI and Anthropic.
Vercel Data: Open Models Spend Collapses to All-Time Low
Closed AI models hit 97.09% spend share via Vercel; open models at all-time low over past 5 days.
Trump Weighs Restrictions on US Firms Using Chinese AI Models
Trump admin weighs restrictions on US firms using Chinese AI models, per Axios. Businesses already adopting cheaper Chinese alternatives, creating policy tension.
Alibaba Qwen3.8: 2.4T Parameter Open-Weight Model Incoming
Alibaba's Qwen3.8, a 2.4T parameter open-weight model, was announced. It would be the largest open-weight model ever, but lacks benchmark details.
90 Hours of Black Myth: Wukong Fuel New World Model Benchmark
A new survey and benchmark rethinks interactive world models as game engines, with a data engine collecting over 90 hours of Black Myth: Wukong gameplay.
Kimi K3 Tops US Models in Front-End Coding at Smaller Scale
Moonshot AI's K3 tops US models in front-end coding at 89.2% on SWE-bench while being smaller and cheaper to train.
Murati's Thinking Machines Ships 975B Inkling — Leads US Open Models
Murati's Thinking Machines releases Inkling, a 975B-parameter MoE model that leads US open models but trails Chinese rivals on benchmarks and cost.
Cursor Doubles Model Usage on All Plans, Adds Grok 4.5
Cursor doubled included model usage on all plans, adding Grok 4.5 and Composer 2.5 without price changes, pressuring competitors like GitHub Copilot.
Function-Aware Fill-in-the-Middle Boosts SWE-Bench by +5.4 on 14B Models
Function-aware FIM mid-training boosts SWE-Bench by +2.8 to +5.4 on 7B-14B models, preserving general abilities. Six checkpoints and 400K dataset open-sourced.
Google alone ships full any-to-any multimodal models
Mollick notes Google alone ships full any-to-any multimodal models; OpenAI and Anthropic lag. This gives Google a structural advantage in agentic workflows.
Open-weight models now run 29% of gateway tokens, up from 11% in April
Open-weight models now handle 29% of gateway tokens, up from 11% in April. The 18-point jump signals accelerating enterprise adoption of open architectures like Llama 3 and Mistral.
Colibri Runs 744B-Parameter Model on 25GB RAM, No GPU
Colibri claims to run a 744B-parameter model on 25GB RAM without GPU, but lacks evidence. If true, it could democratize large-model inference.
Soofi S 30B-A3B: German open model tops English, German benchmarks
German consortium releases Soofi S 30B-A3B, an open MoE model beating OLMo 3 and Apertus 70B on English and German benchmarks while activating only 3.2B of 31.6B parameters.
PadCaptioner: 3B video caption model beats 7B rivals with parallel decoding
PadCaptioner, a 3B model, beats 7B rivals in dense video captioning via lossless parallel autoregressive decoding, challenging scaling orthodoxy.
NVIDIA Drops 30B Nemotron Audex Audio Model with MoE
NVIDIA released Nemotron Audex 30B-A3B, a 30B-parameter MoE audio model unifying ASR, understanding, and TTS with 3B active parameters.
BAAI Orca World Model Matches π0.5 With No Action Labels
BAAI's Orca world model matches specialized π0.5 on five robotics tasks, trained on 125,000 hours of video without action labels, predicting abstract world states.
OpenAI Claims 54% Token Efficiency Gain on Agentic Coding in New Model
OpenAI CEO Sam Altman claims 54% token efficiency gain on agentic coding for a new unnamed model, but no technical details or release date were provided.
SpaceXAI Ships Grok 4.5, Blackwell-Trained Coding Model
SpaceXAI released Grok 4.5, a coding-focused model trained on Blackwell GPUs, now available in Cursor and Vercel. Inference cost claims lack independent benchmarks.
Nvidia, Hugging Face Open-Source Robot Models to Democratize Physical AI
Nvidia and Hugging Face open-sourced robot models to democratize physical AI, providing pre-trained models and simulation tools on the Hugging Face hub.
Cohere Releases Arabic Speech Recognition Model Under Apache 2.0
Cohere released Arabic speech recognition model under Apache 2.0, claiming world's most accurate open-source version. No benchmark data or training details disclosed.
Lilian Weng Argues Harness Design, Not Model Rewrites, Is Path to RSI
Lilian Weng argues RSI starts with harness design, not model rewrites, citing Sakana AI's The AI Scientist in Nature 2026 and two other projects.
Anthropic Shows Models Detect Brain Surgery Interventions Mid-Reasoning
Anthropic's J-space paper proves causal reasoning control and model detection of interventions, raising alignment and eval awareness questions.