Accelerating Text-to-Video Generation with Calibrated Sparse Attention

Authors Shai Yehezkel†**, Shahar Yadin, Noam Elata, Yaron Ostrovsky-Berman, Bahjat Kawar

View publication:https://arxiv.org/abs/2603.05503

VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization

July 2, 2026 research area Computer Vision:/research/?domain=Computer%20Vision conference ICML:/research/?event=ICML

Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what information is preserved and how it is organized. A de facto standard approach to video tokenization is to represent a video as a spatiotemporal 3D grid of tokens, each capturing the corresponding local information in the original signal. This requires the downstream model that consumes the…

STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows

April 30, 2026 research area Computer Vision:/research/?domain=Computer%20Vision, research area Methods and Algorithms:/research/?domain=Methods%20and%20Algorithms conference CVPR:/research/?event=CVPR

Normalizing flows (NFs) are end-to-end likelihood-based generative models for continuous data, and have recently regained attention with encouraging progress on image generation. Yet in the video generation domain, where spatiotemporal complexity and computational cost are substantially higher, state-of-the-art systems almost exclusively rely on diffusion-based models. In this work, we revisit this design space by presenting STARFlow-V, a…

Bottom banner

Our research in machine learning breaks new ground every day.