Most public benchmarks collapse model performance into one broad preference signal.
That makes it hard to understand which capabilities differentiate between models. It's also almost impossible to inspect the evidence behind it. So @RapidataAI is releasing Benchmark.AI.
We started with an SVG generation benchmark including 42 models, 500 prompts, 1.9M+ human judgements, 300K+ match-ups.
We evaluate models separately on Preference, Alignment and Coherence, while making the prompts, outputs, match-ups and methodology public.
Two months ago, we benchmarked @google’s Veo2 model. It fell short, struggling with style consistency and temporal coherence, trailing behind Runway, Pika, @tencent, and even @alibaba-pai.
That’s changed.
We just wrapped up benchmarking Veo3, and the improvements are substantial. It outperformed every other model by a wide margin across all key metrics. Not just better, dominating across style, coherence, and prompt adherence. It's rare to see such a clear lead in today’s hyper-competitive T2V landscape.