-
ByteDance Unveils Seedance 2.5: Native 30-Second 4K AI Video with 50 Multimodal Inputs
ByteDance
video
-
Google DeepMind Invests $75M in A24, Forms First AI Research Partnership with a Film Studio
Google DeepMind
industry
-
Vidu S1: Real-Time Interactive Video Generation at 42 FPS on Consumer GPUs
Shengshu Technology / Tsinghua University
research
-
MiniMax Launches H3, an Omni-Modal Model Generating 2K Video With Native Stereo Audio
MiniMax
video
-
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
research
-
AI-Generated Video Attacks on Real-World Crisis Events: RA-Bench shows detectors fail on crisis-media forgery
research
-
Alibaba launches Wan 3.0 video model with 30-second single-pass generation and native audio
Alibaba (Tongyi Lab)
video
-
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
research
-
UniVidX: One Diffusion Backbone for RGB, Intrinsic Maps, and RGBA Video Generation
research
-
MiniMax Hailuo 2.3 Launches with Media Agent and 50% Cheaper Batch Video Generation
MiniMax
video
-
Lance: 3B Unified Multimodal Model for Understanding, Generation, and Editing (314 HF upvotes)
ByteDance Research
research
-
xAI Grok Imagine Video 1.5: Image-to-Video with Native Audio Tops Arena Leaderboard, API Now Live
xAI
video
-
Google Veo 3.1 Brings Audio to All Flow Editing Modes and New Insert/Remove Tools
Google DeepMind
video
-
Lionsgate Takes Equity Stake in Runway, Plans AI Short-Form Episodic Series
Runway
industry
-
xAI Releases Grok Imagine Video 1.5: #1 on Video Arena Leaderboard at $4.20/min
xAI
video
-
Kling AI Launches 3.0 Turbo and 3.0 Omni: Fast Previews and 4K Editing with Character Consistency
Kuaishou
video
-
ByteDance Seedance 2.5 Opens Enterprise Beta: 30-Second Native Video Generation
ByteDance
video
-
Vidu S1: A Real-Time Interactive Video Generation Model
ShengShu Technology
research
-
Video Generation Models are General-Purpose Vision Learners
Google DeepMind
research
-
DreamX-Phi 1.0 wins WorldArena 2.0 Track 1 with action-conditioned video world model for robotic manipulation
Alibaba
research
-
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
Zhejiang University
research
-
EchoWM: Open and Enterable Omnimodal World Models
JD.com
research
-
VGI-Bench: Probing Visual Intelligence in Video Generation Models
research
-
LongLive-2.0: NVFP4 Parallel Infrastructure for Long Video Generation (NVIDIA, 1,220 HF upvotes)
NVIDIA
research
-
Multiplayer Interactive World Models with Representation Autoencoders
research
-
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
Alibaba AMAP CV Lab
research
-
Self Gradient Forcing: Native Long Video Extrapolation
research
-
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
NUS
research
-
Vidu S2: real-time interactive, editable, and spatial video generation
ShengShu Intelligence
research
-
Causal Forcing++: 2-Step Distillation Enables Real-Time Interactive Video Generation
Tsinghua University
research
-
SANA-WM: Minute-Scale 720p World Modeling on a Single GPU
NVIDIA
research
-
Echo-Infinity: Real-Time Infinite Video Generation via Learnable Memory Query
research
-
Flow-DPPO: Principled RL Alignment for Flow Matching Image and Video Models
Tencent Hunyuan
research
-
ElevenLabs Launches Avatars in ElevenCreative: TTS-Native AI Talking-Head Video
ElevenLabs
video
-
DreamX-World 1.0: General-Purpose Interactive World Model with 6DoF Camera Control
AMAP-ML (Alibaba Maps AI Lab)
research
-
Google Releases Gemini Omni Flash for Video Generation via API
Google DeepMind
video
-
From Pixels to States: Rethinking Interactive World Models as Game Engines
research
-
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
MCG-NJU (Nanjing University)
research
-
PhiZero: A World Model Built Around Physical Language
NLPR, Institute of Automation, Chinese Academy of Sciences
research
-
MiniMax H3 Open-Weight Model Released on Hugging Face
MiniMax
video
-
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
JD (Jingdong)
research
-
Runway introduces GWM Worlds 2, a real-time interactive world model
Runway
video
-
Programmable World Model: agent-programmed persistent state for interactive video
Alaya Lab
research
-
AnyFlow: Any-Step Video Diffusion with On-Policy Flow Map Distillation
MIT / NVIDIA
research
-
WorldDirector: Controllable World Simulator with Persistent Dynamic Object Memory
research
-
Elon Musk Declares Grok Imagine Development Complete
xAI
tools
-
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
research
-
ReWorld: An Interactive World Model with Long-Horizon Memory
TongyiLab (Alibaba)
research
-
InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter
Alibaba (Tongyi)
research
-
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
research
-
Sber's Kandinsky team open-sources Time Adapter for temporally consistent video generation
Sber
video
-
Google's Gemini 'Omni' Video Model Surfaces in Early Demos Ahead of I/O 2026
Google DeepMind
video
-
Gemini Omni Video Model Surfaces Ahead of Google I/O 2026
Google DeepMind
video
-
ShengShu Technology Launches Vidu Claw: AI-Powered End-to-End Ad Production Platform
ShengShu Technology
video
-
VideoKR: 315K-Example Training Corpus for Knowledge- and Reasoning-Intensive Video Understanding
Yale University
research
-
Echo-Memory: Controlled Study of Memory Mechanisms in Action-Conditioned Video World Models
Microsoft Research
research
-
SCAIL-2: End-to-End Character Animation via In-Context Conditioning
Tsinghua University
research
-
PhysisForcing: Physics-Reinforced World Models Improve Robot Manipulation Success by 50%
Peking University / NVIDIA
research
-
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
research
-
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Alibaba Group
research
-
Krea adds Wan 3.0, MiniMax H3 Max and Gemini Omni Flash 1.1 in two-day model wave
Krea
video
-
DreamX-Creator: native joint audio-video generation at 2K resolution
AMAP-ML (Alibaba)
research
-
Krea connects to ChatGPT for image and video generation inside conversations
Krea
tools
-
ByteDance opens Seedance 2.5 developer API to the public
ByteDance
video
-
Google courts Hollywood studios as OpenAI's licensing push collapses
Google
industry
-
ByteDance founder personally leads real-time spatial-video world model, launch possible as soon as October
ByteDance
video
-
Katzenberg and former Sora head Bill Peebles launch AI video startup
video