Accelerating Text-to-Video Generation with Calibrated Sparse Attention
Authors Shai Yehezkel†**, Shahar Yadin, Noam Elata, Yaron Ostrovsky-Berman, Bahjat Kawar
View publication:https://arxiv.org/abs/2603.05503
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
July 2, 2026 research area Computer Vision:/research/?domain=Computer%20Vision conference ICML:/research/?event=ICML
Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what information is preserved and how it is organized. A de facto standard approach to video tokenization is to represent a video as a spatiotemporal 3D grid of tokens, each capturing the corresponding local information in the original signal. This requires the downstream model that consumes the…
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
April 30, 2026 research area Computer Vision:/research/?domain=Computer%20Vision, research area Methods and Algorithms:/research/?domain=Methods%20and%20Algorithms conference CVPR:/research/?event=CVPR
Normalizing flows (NFs) are end-to-end likelihood-based generative models for continuous data, and have recently regained attention with encouraging progress on image generation. Yet in the video generation domain, where spatiotemporal complexity and computational cost are substantially higher, state-of-the-art systems almost exclusively rely on diffusion-based models. In this work, we revisit this design space by presenting STARFlow-V, a…

Our research in machine learning breaks new ground every day.
