A Human-Judged Reward for Video-Audio Generation
Models that generate video with synchronized sound are improving quickly, but the scoring systems used to refine them may steer them away from what people actually prefer. Researchers from Fudan University, Zhejiang University, and industry partners report that the metrics commonly used as training rewards often disagree with human viewers. They propose VA-Judger, a reward model trained on human preference judgments, which in their experiments matched people more closely than existing metrics did and improved a video-audio generator when used for post-training.
The authors argue that current reinforcement learning pipelines combine separate expert metrics: VideoAlign for visual quality, Audiobox Aesthetics for audio, CLAP for audio-text match, and DeSync for synchronization. A model tuned against such a combination can learn to score well on individual metrics while producing content that looks or sounds incoherent to people. The authors describe this as reward hacking.
VA-Judger is a chain-of-thought model, meaning it writes out its reasoning before giving an answer. Built on Qwen3-Omni, it compares two clips generated from the same prompt, scores each from 1 to 10 on five dimensions (prompt alignment, video-audio consistency, audio quality, video quality, and completeness), and names a winner. Training has three stages. First, the model learns the output format from easy pairs with clear quality gaps, using reasoning generated by Gemini 3.1 Pro. Second, for close pairs, the team replaced Gemini’s choices with human labels, because a pilot study found Gemini agreed with annotators on only 60% of hard pairs. The team kept 4,436 reasoning traces whose conclusions and dimension scores matched the human annotations. Third, reinforcement learning rewards both a correct final answer and score orderings that agree with the dimensions annotators cited.
The paper offers one clear example. For a prompt describing a sunlit living room, the expert metrics favored Video 1. VA-Judger instead scored Video 1 at 20 and Video 2 at 35 out of a possible total, noting that Video 1’s audio was distorted while Video 2’s was clear. The authors present this as a case where separate metrics missed a judgment that a holistic reading captured.
On VA-Judger-Bench, which contains 1,150 human-labeled pairs, VA-Judger reached 68.43% overall accuracy against human preferences. That is 10.60 percentage points above the untuned Qwen3-Omni Instruct model with chain-of-thought (57.83%) and nearly 10 points above the best metric ensemble (58.70%). On the out-of-domain split, which uses outputs from closed-source models such as Veo 3.1 and Sora 2 excluded from training, VA-Judger scored 63.40%, compared with 55.40% for the best single metric.
When the reward was used to post-train LTX-2 on 200 prompts, the resulting model ranked first on 17 of 20 automatic metrics. Video quality rose from 2.248 to 3.942, and JavisScore, a joint audio-video measure, rose from 0.074 to 0.230. In a human study, 20 raters made 4,000 selections across three-way comparisons, preferring the VA-Judger-tuned outputs 62.30% of the time, compared with 27.63% for the OmniNFT-tuned model and 10.08% for base LTX-2.
The results carry caveats. VA-Judger did not lead on every measure: OmniNFT recorded the lowest DeSync value (0.226, where lower is better), against 0.592 for VA-Judger. The authors attribute this to OmniNFT directly optimizing DeSync during training, which they describe as possible metric-specific reward hacking. Human labeling is also expensive, and the in-domain test uses generators seen during training. The post-training experiment was small, and it updated only LoRA parameters on a frozen LTX-2 backbone. It remains unclear how well the reward transfers to other generators or larger training runs.
Still, the study suggests that holistic human judgments can supply usable training signals for joint audio-video generation where narrow metrics fall short. It is a single research group’s evaluation, and its broader significance will depend on replication across more models and settings.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.