
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL
Get every episode summarized
Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“The first two authors are Songlin Yang and Xiaotang Zhao, and the corresponding author is Jia Chengjiang. The organizations involved are MM Lab at HKUST and Tense and Video. Alright, let's dive into the introduction section.”From the transcript
🤗 Upvotes: 112 | cs.CV
Authors:
Songlin Yang, Xiaotong Zhao, Jiacheng Zhang, Zhe Wang, Toyota Li, Eric Liu, Alan Zhao, Anyi Rao
Title:
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL
Arxiv:
http://arxiv.org/abs/2609.37200v1
Abstract:
Multi-reward guided reinforcement learning (i.e., RL) offers a promising way to improve joint audio-video diffusion models along several complementary objectives, including modality-specific quality, cross-modal semantic alignment, and temporal synchronization. Its effectiveness, however, depends on two quantities that change during training: where reward-driven updates should act, and how competing rewards should be combined. Existing methods tend to rely on fixed routing and reward weights, failing to track evolving model functions. To address these limitations, we propose Adaptive Reward Routing to jointly adapt update locations and reward coordination during forward-process RL (i.e., DiffusionNFT) of joint audio-video diffusion models. Our method consists of two components. (i) Cross-Modal Influence-Guided Routing (Localizing Updates): We use bidirectional cross-attention responses as an efficient proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without additional model interventions. (ii) Preference-Preserving Modality-Aware Reweighting (Coordinating Rewards): We preserve predefined weights as preference priors and use branch-specific reward-gradient interactions as residual corrections after warm-up. This resolves evolving conflicts without letting dominant rewards suppress weak but essential objectives. Extensive experiments demonstrate consistent improvements in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines. Ablations and mechanism analyses further validate the complementary benefits of adaptive update routing and reward coordination.
Get every episode summarized
Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Know when Zhe Wang turns up
Follow Zhe Wang and once a week we email you every new episode they appeared on — including guest spots the show notes never mention, because we read the transcript.
Follow Zhe WangFree. Pick your own day and time.
Hosts & guests
Transcript ready
290 searchable segments. Every word is indexed and playable.
Full transcript
Daily Paper Cast — Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Welcome to Daily Papercast. Today's paper comes from the Hugging Face Daily Paper List of October 2nd, 2026, and has received 112 upvotes. The title is Adaptive Reward Routing, Dynamic Multi-Reward Optimization for Joint Audio Video Diffusion via Forward Process RL. Great. The first two authors are Songlin Yang and Xiaotang Zhao, and the corresponding author is Jia Chengjiang. The organizations involved are MM Lab at HKUST and Tense and Video. Alright, let's dive into the introduction section. Recent advances in Joint Audio Video Diffusion models are enabling the generation of visual and audio content from text prompts simultaneously. This capability hinges on meeting several objectives, ensuring high visual and audio quality specific to each modality, aligning semantic content across modalities, and achieving precise temporal synchronization.
It sounds like capturing all these elements in a single model is quite challenging. How have approaches like reinforcement learning helped in this regard? Indeed, Evan. Reward Guided RL, especially in the context of diffusion models, offers a promising paradigm by defining multiple reward signals that cater to each of these objectives. Guided such as GRPO-based strategies and diffusion NFT optimize models by expressing requirements through these diverse rewards. So if I understand correctly, these methods rely on fixed reward routes and weights. Does this paper address any limitations of these existing approaches? Precisely. Existing methods generally use static routing and reward weights, which can become outdated as the model functions evolved during training. The paper introduces adaptive reward routing to tackle these limitations. It adapts both the locations of reward-driven updates and the coordination of multiple, possibly competing rewards through the ongoing training.
So adaptive reward routing adjusts to the model's changing needs during training. What exactly are the components of this proposed system? The system comprises two main components. First, there's cross-modal influence guided routing, which uses bidirectional cross-attention responses to efficiently localize where updates should occur across the model layers and tokens. This dynamic weighting helps in influencing the gradients during cross-modal interactions without additional interventions. Interesting. And what's the second component? The second component is preference-preserving modality, aware re-weighting. It resolves conflicts between rewards by using predefined preference priors and branch-specific reward gradient interactions. Essentially, this ensures that even less dominant but essential objectives are not overshadowed by the stronger ones. How do these components work together to improve the model's performance? Together, these components allow the model to adapt in terms of reward coordination and
update localization dynamically throughout training. Extensive experiments have shown consistent improvements in modalities quality, semantic consistency, and synchronization when compared to strong RL baselines. Essentially, the paper proposes refining both the reward adaptation and routing process for joint diffusion models, providing more stable and effective post-training outcomes. Exactly. That concludes the introduction section where we've discussed the background, objectives, and contributions of this paper. Let's dive into the method section to understand how the authors have structured their approach. Sure, Evan. The paper introduces a framework called Adaptive Reward Routing for Forward Process RL in joint audio video diffusion models. This framework consists of two main components, cross-modal influence guided routing, and preference-preserving modality-aware re-weighting. Right. We briefly mentioned these earlier.
Can you unpack cross-modal influence guided routing? Cross-modal influence guided routing dynamically determines where reward-driven updates should act in the model. Essentially, it uses bidirectional cross-attention responses as proxies for evolving cross-modal influence. This process involves two steps, token-level routing and layer-level routing. What does token-level routing involve? token-level routing assigns weights to individual tokens based on their influence in cross-modal interactions. Specifically, it averages the pre-gate cross-attention responses over selected time steps and normalizes these scores to form positive loss weights. These weights help emphasize more influential tokens during model updates. And how about layer-level routing? Layer-level routing adapts the gradient flow across different layers. It takes the same pre-gate cross-attention responses and averages them over tokens and selected time steps to derive a layer score.
This score is then used to create a detachment coefficient that scales the gradients in back propagation. Essentially, layers that have stronger cross-modal influence retain more gradient while weaker ones are progressively detached. So the framework adapts weights at both the token and layer levels. What role does preference-preserving modality-aware re-weighting play in this setup? Preference-preserving. Modality-aware re-weighting deals with how multiple rewards are coordinated. It resolves conflicts between rewards while preserving user-defined preferences. This component incorporates three sub-stratigies. Branch-aware balancing, residual corrections, and warm-up. Interesting. Can you break down these sub-stratigies for us? Certainly. Branch-aware balancing assigns reward interactions to the respective modality branches that they supervise, such as video and audio branches, residual corrections that adjust predefined
reward weights by incorporating coefficients derived from gradient-based conflict estimation. This ensures that objectives do not get entirely overridden by early-noisy gradients. Lastly, warm-up stabilizes the reward re-weighting process by delaying dynamic adjustments until the relationships between gradients become reliable. Got it. It sounds like these methods together make the model more adaptable to the changing dynamics during training. How is this reflected in the actual training process? Exactly. The complete optimization process involves several steps. First, reward re-weighting decides the contribution of each objective. Then, branch routing assigns these objectives to target modalities. After that, token-level routing determines where within each modality loss should be emphasized. Finally, layer-level routing controls how the gradient flows across modality boundaries. How does the framework ensure that these weight changes truly reflect the evolving needs
of the model? Great question. The framework uses the pre-gate cross-attention responses collected during sampling to track influence shifts. For instance, certain layers and tokens might become more influential as the model fine tunes, and the adaptive routing ensures these shifts are accounted for. Additionally, evaluation of reward signals helps in computing advantages and optimality probabilities that further guide the training updates. That makes a lot of sense. To sum it up, the adaptive reward routing framework adapts both the importance and location of updates based on ongoing training dynamics, making it a sophisticated approach to multi-reward optimization in joint audio-video diffusion models. Precisely, Evan, that concludes the discussion on the method section. Next, we'll move into the experiment section to see how these methods perform in practice. All right. Let's dive into the experiment and result section to see how this adaptive reward routing
framework performs in practice. Sure, Evan. The authors conducted extensive experiments to evaluate their proposed framework on joint audio-video diffusion models using two main backbones, LTX2 and LTX2.3. They trained these models with 19,487 audio-video prompts collected from a VGG sound derived corpus. What type of rewards and bass lines did they use for these evaluations? They used five complementary reward signals, video quality with video align and HPSV3, audio quality with audio box aesthetics, text audio alignment with clap, and audio video synchronization with desync. For bass lines, they compared their method against four settings, the pre-trained backbone with no post-training, GDPO, marble, and Omnion FT. Metrics, did they use to evaluate the performance? They used Javaspench, encompassing diverse audio-video generation scenarios, and evaluated
metrics across four groups. AV quality, text consistency, AV consistency, and AV synchrony. Metrics like visual quality, VQ, audio quality, AQ, text video, and text audio image-bind similarity, TVIB and TAB, clip, clap scores, AVIB, AVH score, Javasp score, and desync were measured. That's quite a comprehensive evaluation. How did the adaptive reward routing framework perform compared to the bass lines? Remarkably well. For both LTX2 and LTX2.3 backbones, the proposed method achieved the best results on 9 out of 10 metrics. For example, in visual quality, they reported scores of 3.336 for LTX2 and 3.59 for LTX2.3, outperforming other bass lines like Omnion FT. That's impressive. How did the training dynamics look for these models?
The training dynamics indicated that the adaptive reward routing reached the highest average reward and maintained favorable trajectories across all five component rewards. In contrast, bass lines made uneven progress across rewards. For example, AVD sync showed that adaptive reward routing achieved a lower value around 0.341, indicating better synchronization performance compared to other models. How did the authors validate these improvements, were due to their proposed methods and not just better hyperparameters or architectures? Good question, Evan. The authors conducted comprehensive ablation studies to isolate the contributions of each component. They showed that the most significant gains came from incorporating both cross-modal influence guided routing and preference preserving modality-aware re-weighting. Ablations confirmed that each added component yielded progressive gains, especially in metrics
like AV quality and synchronization. Can you give an example of one of these ablation studies? For instance, when they evaluated token weighting and layer scaling independently on GDPO, token weighting alone improved metrics like text consistency, while layer scaling significantly boosted AV synchrony, combining both resulted in the strongest performance overall. Did they also investigate the effectiveness of their response proxy for influencing gradients and routing decisions? Yes, they validated the response proxy's accuracy in identifying influential cross-modal layers and tokens. For example, blocking the cross-modal responses of top-scoring tokens resulted in a larger prediction change, proving these tokens were most influential. This proxy maintained high accuracy throughout training, unlike a static proxy. That's compelling evidence. So the experiment section clearly shows that adaptive reward routing dynamically optimizes
joint audio-video diffusion models by effectively coordinating and routing rewards. It seems like a well-rounded approach that addresses the evolving needs during model training. Exactly, Evan. This concludes the experiment section where we've explored how well the proposed framework performs and the detailed methods behind these impressive results. Let's move on to the related work section to understand how this paper builds on and contrasts with previous research. Sure, Evan. The related work section primarily addresses three areas, joint audio-video generation, reinforcement learning for diffusion models, and multi-reward optimization. Alright, let's start with joint audio-video generation. What have been the significant advancements in this field? Joint audio-video generation has seen notable progress from earlier image diffusion models integrated with temporal modules, as well as large diffusion transformers. For example, methods like those by Blatman et al, 2023, and Guowe et al, 2024 employ these
strategies to enhance video generation effectively. I see. And how do joint audio-video systems function in this regard? Joint audio-video systems often connect pre-trained experts through cross-modal projections or use unified diffusion transformers as seen and works by Blatman et al, 2025, and Lu et al, 2026. Alternatively, they couple separate streams via bidirectional cross-attention mechanisms, a method highlighted in Hacoin et al, 2026. So, while these strategies enable mutual conditioning across modalities, what challenges do they face? The main challenge lies in the difficulty of defining reward responsibility and gradient routing within these heterogeneous branches, which is less straightforward compared to single-stream models. That's insightful. Now, moving on to reinforcement learning for diffusion models, what have been some pivotal
advancements here? In reinforcement learning for diffusion models, GRPO, or Group Reward decoupled normalization policy optimization, has been a significant advancement, as mentioned by Guowe et al, 2025, this method estimates relative advantages without using a critic. Subsequent methods like FlowGRPO and DanceGRPO extend online optimization to flow-based generation by incorporating stochastic sampling techniques. Interesting. How about the diffusion NFT approach? Diffusion NFT, proposed by Zheng et al, 2025, optimizes the forward process using implicit positive and negative policies. This approach focuses on sample time steps rather than reverse process likelihood estimation, which aligns closely with our paper's methodology. It seems like Omni NFT also plays a role in joint audio video generation.
Can you elaborate? Certainly. Omni NFT, as described by Zheng et al, 2026, adds modality-wise credit assignment for joint audio video generation. It employs layer-wise gradient adjustments and region-wise re-weighting, but relies on fixed routing rules, which our paper aims to improve upon. Got it. And how does multi-reward optimization fit into the picture? Multi-reward optimization methods address the integration and balance of multiple conflicting objectives. Fixed scalarization techniques struggle to adapt to changing conflicts. GDPO preserves reward-specific signals through normalization, but still uses predefined aggregation weights, which can be rigid. Multi-task approaches like those by Decidory, 2012, and Center and Coltune, 2018 aim for common descent directions and optimize local agreement.
That's right. How do methods like marble adapt multi-reward optimization to diffusion RL? Marble, discussed by Zhao et al, 2026, adapts multi-task optimization ideas to diffusion RL by dynamically adjusting reward weights through gradient geometry. However, this approach still faces limitations as it primarily focuses on gradient compatibility rather than objective importance or specific modality responsibility. How does the work in our paper build on these prior approaches? Our paper unifies the strengths of these strategies by introducing adaptive routing mechanisms that dynamically localize updates based on evolving cross-modal influence and reward coordination techniques that resolve conflicts while preserving user-defined priorities. This dual approach offers a more comprehensive and responsive system for joint audio-video diffusion models. So essentially, integrating adaptive update routing with preference-preserving reward
coordination creates a robust framework that addresses the dynamic needs of joint audio video generation models, building significantly on past work. Exactly, Evan. That wraps up the related work section. All right, Ashley. We've covered the introduction, method, experiments, and related work. Let's summarize the key contributions and takeaways of this paper. Sure, Evan. To wrap it up, this paper presents adaptive reward routing, a dynamic multi-reward optimization framework for joint audio-video diffusion via forward process reinforcement learning. The authors proposed two main components, cross-modal influence guided routing and preference-preserving modality-aware re-weighting. Right. These components adapt to the model's evolving needs during training. What are the key benefits observed from this approach? The key benefits include consistent improvements in modality-specific quality, semantic alignment,
and audio-video synchronization. By dynamically localizing updates and intelligently coordinating multiple rewards, the framework ensures that even weak but essential objectives are not suppressed by dominant rewards. Experiments showed that adaptive reward routing outperformed strong RL baselines, like GDPO, Marble, and OmnionFT, achieving the best results across nine out of ten metrics on both LTX2 and LTX2.3 backbones. Exactly. The framework leverages cross-attention responses to track influence shifts and accurately updates weights, ensuring the model remains responsive to its training dynamics. Ablation studies further validated the complementary benefits of the routing and re-weighting components. To sum up, adaptive reward routing provides a sophisticated method for optimizing joint audio-video diffusion models by dynamically adapting to the evolving needs of the model
during training. It's a significant contribution that builds on previous research in the areas of joint generation, reinforcement learning, and multi-reward optimization. That's all for today's episode. Thank you for tuning in to Daily Papercast. We hope you found our discussion insightful. Be sure to join us again tomorrow for more in-depth looks at fascinating new research papers. Until then, keep exploring and stay curious. Goodbye, everyone. Goodbye.
More episodes
More from Daily Paper Cast

AgentGarten: Code Worlds for Evolving Agents
Daily Paper Cast

Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Env...
Daily Paper Cast

TokenRouter: Efficient Serving System for Token-Level LLM Routing
Daily Paper Cast

From Traces to Agentic Worlds: Agentic Language World Models for Interactive Env...
Daily Paper Cast