-
ByteDance Unveils Seedance 2.5: Native 30-Second 4K AI Video with 50 Multimodal Inputs
ByteDance
video
-
Google DeepMind Invests $75M in A24, Forms First AI Research Partnership with a Film Studio
Google DeepMind
industry
-
Vidu S1: Real-Time Interactive Video Generation at 42 FPS on Consumer GPUs
Shengshu Technology / Tsinghua University
research
-
MiniMax Launches H3, an Omni-Modal Model Generating 2K Video With Native Stereo Audio
MiniMax
video
-
UniVidX: One Diffusion Backbone for RGB, Intrinsic Maps, and RGBA Video Generation
research
-
MiniMax Hailuo 2.3 Launches with Media Agent and 50% Cheaper Batch Video Generation
MiniMax
video
-
Lance: 3B Unified Multimodal Model for Understanding, Generation, and Editing (314 HF upvotes)
ByteDance Research
research
-
xAI Grok Imagine Video 1.5: Image-to-Video with Native Audio Tops Arena Leaderboard, API Now Live
xAI
video
-
Google Veo 3.1 Brings Audio to All Flow Editing Modes and New Insert/Remove Tools
Google DeepMind
video
-
Lionsgate Takes Equity Stake in Runway, Plans AI Short-Form Episodic Series
Runway
industry
-
xAI Releases Grok Imagine Video 1.5: #1 on Video Arena Leaderboard at $4.20/min
xAI
video
-
Kling AI Launches 3.0 Turbo and 3.0 Omni: Fast Previews and 4K Editing with Character Consistency
Kuaishou
video
-
ByteDance Seedance 2.5 Opens Enterprise Beta: 30-Second Native Video Generation
ByteDance
video
-
Vidu S1: A Real-Time Interactive Video Generation Model
ShengShu Technology
research
-
Video Generation Models are General-Purpose Vision Learners
Google DeepMind
research
-
LongLive-2.0: NVFP4 Parallel Infrastructure for Long Video Generation (NVIDIA, 1,220 HF upvotes)
NVIDIA
research
-
Multiplayer Interactive World Models with Representation Autoencoders
research
-
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
Alibaba AMAP CV Lab
research
-
Self Gradient Forcing: Native Long Video Extrapolation
research
-
Causal Forcing++: 2-Step Distillation Enables Real-Time Interactive Video Generation
Tsinghua University
research
-
SANA-WM: Minute-Scale 720p World Modeling on a Single GPU
NVIDIA
research
-
Echo-Infinity: Real-Time Infinite Video Generation via Learnable Memory Query
research
-
Flow-DPPO: Principled RL Alignment for Flow Matching Image and Video Models
Tencent Hunyuan
research
-
ElevenLabs Launches Avatars in ElevenCreative: TTS-Native AI Talking-Head Video
ElevenLabs
video
-
DreamX-World 1.0: General-Purpose Interactive World Model with 6DoF Camera Control
AMAP-ML (Alibaba Maps AI Lab)
research
-
Google Releases Gemini Omni Flash for Video Generation via API
Google DeepMind
video
-
From Pixels to States: Rethinking Interactive World Models as Game Engines
research
-
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
MCG-NJU (Nanjing University)
research
-
PhiZero: A World Model Built Around Physical Language
NLPR, Institute of Automation, Chinese Academy of Sciences
research
-
MiniMax H3 Open-Weight Model Released on Hugging Face
MiniMax
video
-
AnyFlow: Any-Step Video Diffusion with On-Policy Flow Map Distillation
MIT / NVIDIA
research
-
WorldDirector: Controllable World Simulator with Persistent Dynamic Object Memory
research
-
Elon Musk Declares Grok Imagine Development Complete
xAI
tools
-
Google's Gemini 'Omni' Video Model Surfaces in Early Demos Ahead of I/O 2026
Google DeepMind
video
-
Gemini Omni Video Model Surfaces Ahead of Google I/O 2026
Google DeepMind
video
-
ShengShu Technology Launches Vidu Claw: AI-Powered End-to-End Ad Production Platform
ShengShu Technology
video
-
VideoKR: 315K-Example Training Corpus for Knowledge- and Reasoning-Intensive Video Understanding
Yale University
research
-
Echo-Memory: Controlled Study of Memory Mechanisms in Action-Conditioned Video World Models
Microsoft Research
research
-
SCAIL-2: End-to-End Character Animation via In-Context Conditioning
Tsinghua University
research
-
PhysisForcing: Physics-Reinforced World Models Improve Robot Manipulation Success by 50%
Peking University / NVIDIA
research