Run VBench End to End
Key Insight
Evaluating video generation is even harder than evaluating images, and a single number like FVD (Fréchet Video Distance) correlates poorly with what people actually like. VBench breaks the vague question "is this video good?" into many separate axes — subject consistency, motion smoothness, aesthetic quality, text alignment, and more — and scores each one, so you learn which aspect a model is weak at instead of getting one blurry verdict. This project runs an open text-to-video (T2V) model through the full VBench suite and reproduces a published leaderboard number — which is itself a lesson in how fragile "just reproduce the benchmark" turns out to be once prompt wording, sampling settings, and clip counts all start to matter.