arXiv · 2609.26233
The Temporal Moderation Gap: Text-to-Video Safety Filters Are Blind to Harm in Motion
Abstract
Text-to-video (T2V) services inherit their safety stack from image generation, pairing a keyword prompt filter with a per-frame checker that blocks a clip whenever one sampled frame looks unsafe. This stack has a blind spot unique to video. We prove that any moderator ignoring frame order accepts a harmful clip whenever it accepts that clip's benign shuffle, so harm carried by the ordering alone escapes. Empirically, the unmodified benchmark prompt already lands a clip in this moderation gap on 32.7% of Sequential-Action targets over four held-out seeds, and paraphrasing, scene splitting, and a feedback-driven prompt search show no significant improvement (paired McNemar $p\ge0.12$), so prompt engineering is not needed to expose the vulnerability. Dense-scoring all 97 rendered frames shows that about a third of the delivered clips merely hide an unsafe frame, while the rest stay harmful as ordered videos even though every frame passes, an order-blind residual the unmodified prompt reaches on a quarter of Sequential-Action targets. We also document a measurement pitfall, since scoring a searched prompt on its own render seed inflates a 7.5% per-generation rate into an apparent 46.7%. A user study confirms that people read these clips as harmful and their shuffles as safe. The fix is to read frame order, and an order-aware detector separates these clips from their own shuffles at AUC 0.74 where per-frame checking sits at chance, which is the signal deployed moderation throws away.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuxin Cao, Fusen Guo, Yuezhong Wu, Huadong Mo, Wei Song. 2026-08-19. The Temporal Moderation Gap: Text-to-Video Safety Filters Are Blind to Harm in Motion. https://arxiv.org/abs/2609.26233
Cite the original work for its findings. Save a collection to share your selection of sources.