Daily Paper Cast
Jingwen Liang, Gengyu Wang
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: [email protected]:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-...
Episodes (1777)
PEARL: Personalized Streaming Video Understanding Model
Daily Paper Cast
DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models
Daily Paper Cast
SIMART: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLM
Daily Paper Cast
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
Daily Paper Cast
RealMaster: Lifting Rendered Scenes into Photorealistic Video
Daily Paper Cast
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for Worl...
Daily Paper Cast
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generativ...
Daily Paper Cast
LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integra...
Daily Paper Cast
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
Daily Paper Cast
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory...
Daily Paper Cast
VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance fo...
Daily Paper Cast
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
Daily Paper Cast
F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Spla...
Daily Paper Cast
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
Daily Paper Cast
HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning
Daily Paper Cast
Astrolabe: Steering Forward-Process Reinforcement Learning for Distilled Autoreg...
Daily Paper Cast
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
Daily Paper Cast
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
Daily Paper Cast
FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectif...
Daily Paper Cast
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
Daily Paper Cast
LumosX: Relate Any Identities with Their Attributes for Personalized Video Gener...
Daily Paper Cast
Hyperagents
Daily Paper Cast
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understand...
Daily Paper Cast
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided...
Daily Paper Cast
FASTER: Rethinking Real-Time Flow VLAs
Daily Paper Cast
3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
Daily Paper Cast
Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion...
Daily Paper Cast
MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstru...
Daily Paper Cast
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Polic...
Daily Paper Cast
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Represe...
Daily Paper Cast
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal...
Daily Paper Cast
Memento-Skills: Let Agents Design Agents
Daily Paper Cast
MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
Daily Paper Cast
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
Daily Paper Cast
MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
Daily Paper Cast
Alignment Makes Language Models Normative, Not Descriptive
Daily Paper Cast
Complementary Reinforcement Learning
Daily Paper Cast
When AI Navigates the Fog of War
Daily Paper Cast
MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification
Daily Paper Cast
InCoder-32B: Code Foundation Model for Industrial Scenarios
Daily Paper Cast
Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
Daily Paper Cast
Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-...
Daily Paper Cast
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
Daily Paper Cast
Demystifing Video Reasoning
Daily Paper Cast
WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unif...
Daily Paper Cast
TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL ove...
Daily Paper Cast
Online Experiential Learning for Language Models
Daily Paper Cast
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
Daily Paper Cast
AI Can Learn Scientific Taste
Daily Paper Cast
OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training...
Daily Paper Cast
