video
30 articles about video in AI news
RA-Bench: Crisis Video Detectors Fail at 1.4% FakeR
RA-Bench shows AI crisis video detectors fail, with MLLMs dropping to 1.4% FakeR after social dissemination. No detector family generalizes across 9 generators.
StreamArena: 243-Video Benchmark for Hour-Long Streaming AI
StreamArena launches with 243 hour-long videos and 3,646 open-ended tasks for interactive streaming video understanding, pushing beyond short-clip benchmarks. No baselines yet.
Dyna-2 World-Action Model Trained on 1M Hours Video
Dyna-2, trained on 1M+ hours of egocentric video, jointly predicts future video and actions. Claims new scaling laws but no benchmarks or technical details released.
MiniMax H3 Video Model Beats Seedance 2.0, Opens Weights
MiniMax launched H3 video model, ranking #1 in editing benchmarks while opening weights to challenge ByteDance's Seedance 2.0 and Google's Gemini Omni Flash.
ByteDance Seedance 2.5 Generates 30-Second Video in One Shot
ByteDance's Seedance 2.5 generates 30-second video in one pass with 30 image, 10 video, 10 audio references. Follows Seedance 2.0 from February 2026.
Gemini Robotics ER 2 Hits 60% Video Completeness, Beats 1.6
Google's Gemini Robotics 2.0 ships ER 2 VLM with 60% video accuracy, but dexterity and safety models stay unreleased.
Flux 3 Beats Seedance 2.0 in BFL Tests, Adds Native Audio to Video
Black Forest Labs released Flux 3, a multimodal model generating 20-second video with native audio. Internal tests claim 52% preference over Seedance 2.0, but independent results are pending.
KeyFrame-Compass Benchmark Targets Keyframe Video Generation Gaps
KeyFrame-Compass is the first benchmark for keyframe-conditioned video generation, with 386 samples and six metrics.
Drive After Effects from Claude Code: Generate Production Videos via MCP
aftr is an open-source MCP server that lets Claude Code drive After Effects programmatically via JSON commands over WebSocket, enabling automated video rendering without manual work.
GenCeption: Video Diffusion Backbone Beats Specialists on 5 Vision Tasks
GenCeption uses video diffusion as a vision backbone, matching specialists with 7-500x less data and generalizing from synthetic to real footage.
PadCaptioner: 3B video caption model beats 7B rivals with parallel decoding
PadCaptioner, a 3B model, beats 7B rivals in dense video captioning via lossless parallel autoregressive decoding, challenging scaling orthodoxy.
Microsoft ResearchStudio-Reel Turns PDF Into Poster, Video, Blog
Microsoft ResearchStudio-Reel converts PDFs into poster, video, blog, and reel using Claude Code and Codex, with editable Office outputs.
Browser-use open-sources Claude-powered video editor
Browser-use open-sourced a video editor inside Claude Computer Use, replacing manual editing with natural language commands and challenging Adobe Premiere.
Google Launches $0.034 Image Model, Video API for Gemini
Google launched Nano Banana 2 Lite ($0.034/image, 4-second generation) and Gemini Omni Flash ($0.10/second video API), targeting high-throughput developer pipelines.
Bluezoo Launches AI Agent for In-Store Video Advertising
Bluezoo launched an AI agent for in-store video advertising that uses computer vision to analyze shopper engagement and optimize ad content in real time, promising improved ad effectiveness for retailers.
AI editor matches pro on 84% of video cuts in blind test
AI editor matched pro on 84% of video cuts in blind test of 4-hour project. Suggests editorial judgment is partially automatable.
Mirage: Microsoft's 10.57x faster video gen skips RGB render loop
Microsoft's Mirage stores 3D scenes as latent tokens, achieving 10.57x faster video generation and 55x less memory, with SOTA WorldScore consistency.
LTX Studio Turns AI Video Clips Into Editable Scenes
LTX Studio + LTX-2.3 lets users edit AI video scenes, not just generate clips. This shifts AI video from demo to production tool.
Kling AI Video Enters Hollywood Production with 'House of David'
Kling AI video used in 'House of David', first Hollywood production at industrial scale. Show reached 44M+ viewers, #1 on Prime Video U.S.
HAVEN Benchmark Exposes MLLM Gap Between Fluency and Video Understanding
HAVEN benchmark tests MLLMs on hierarchical video understanding across frame, shot, and video levels. Results show top models lack grounded multimodal reasoning despite fluent text generation.
POV Shopping Videos Threaten Luxury Brand Control, BoF Warns
BoF warns POV shopping videos risk luxury brand exclusivity by prioritizing authenticity over controlled imagery, with no disclosed revenue impact.
Tavus Debuts AI Avatars Without Source Video Footage
Tavus announced AI avatars no longer need source video, enabling generation from images or text. The shift lowers barriers for enterprise video production.
Pollo AI Underprices Seedance 2.0 at $0.11/Video
Pollo AI offers Seedance 2.0 at $0.11/video, 5-10x below Seedance's API rates, signaling a pricing war in AI video generation.
Luma Labs Opens Uni-1.1 API for Production — Image, Not Video, and #1 ELO Comes With a Caveat
Luma Labs has shipped the Uni-1.1 API for production — an image-generation model (not video) with two REST endpoints, Python and JavaScript SDKs, and support for up to nine reference images per call. The widely-cited '#1 Human Preference ELO' is from Luma's own internal pairwise evaluation; on pure text-to-image Luma reports #2 behind Google Nano Banana. Pricing: ~$0.09 per 2K image, 10–30% below Nano Banana 2 / Pro.
UniVidX Generates Video From 1,000 Samples, SIGGRAPH 2026
UniVidX generates omni-directional video from <1,000 training samples, using diffusion priors with stochastic masking, accepted at SIGGRAPH 2026.
Google DeepMind Launches Real-Time Video AI Co-Clinician
Google DeepMind launched AI Co-Clinician, a real-time video analysis system for triadic care, claiming 30% fewer diagnostic errors in early tests.
NVIDIA Nemotron 3 Nano Omni: Open Multimodal Model Unifies Video, Audio, Image, Text
NVIDIA announced Nemotron 3 Nano Omni, an open multimodal model that processes video, audio, images, and text in a unified architecture, expanding accessibility for multimodal AI research.
Microsoft World-R1: RL Aligns Text-to-Video with 3D Physics
Microsoft's World-R1 framework applies reinforcement learning with feedback from pre-trained 3D foundation models to align text-to-video outputs with physical 3D constraints, improving structural coherence without modifying the underlying video diffusion architecture.
Mirage's Cappy Edits Video via Text Message with No App
Mirage launched Cappy, a text-based video editing service that delivers fully edited videos via SMS. This first-of-its-kind approach eliminates traditional editing interfaces entirely.
OpenAI Teases 'Not a Screenshot' AI Video Model
OpenAI posted a cryptic tweet stating 'This is not a screenshot' with a video link, strongly hinting at a new AI video generation model. This marks a direct move into a space currently led by rivals like Runway and Pika.