Daily Paper Cast
Jingwen Liang, Gengyu Wang
We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: [email protected]:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-...
Episodes (1777)
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
Daily Paper Cast
Part-X-MLLM: Part-aware 3D Multimodal Large Language Model
Daily Paper Cast
MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Ed...
Daily Paper Cast
GroupRank: A Groupwise Reranking Paradigm Driven by Reinforcement Learning
Daily Paper Cast
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
Daily Paper Cast
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
Daily Paper Cast
GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Model...
Daily Paper Cast
DoPE: Denoising Rotary Position Embedding
Daily Paper Cast
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and...
Daily Paper Cast
UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code...
Daily Paper Cast
AIonopedia: an LLM agent orchestrating multimodal learning for ionic liquid disc...
Daily Paper Cast
LiteAttention: A Temporal Sparse Attention for Diffusion Transformers
Daily Paper Cast
Virtual Width Networks
Daily Paper Cast
One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter...
Daily Paper Cast
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
Daily Paper Cast
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalis...
Daily Paper Cast
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
Daily Paper Cast
DeepEyesV2: Toward Agentic Multimodal Model
Daily Paper Cast
Visual Spatial Tuning
Daily Paper Cast
VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Chec...
Daily Paper Cast
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradi...
Daily Paper Cast
V-Thinker: Interactive Thinking with Images
Daily Paper Cast
Scaling Agent Learning via Experience Synthesis
Daily Paper Cast
Diffusion Language Models are Super Data Learners
Daily Paper Cast
LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environme...
Daily Paper Cast
UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interac...
Daily Paper Cast
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
Daily Paper Cast
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
Daily Paper Cast
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Ch...
Daily Paper Cast
Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language F...
Daily Paper Cast
Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
Daily Paper Cast
The Underappreciated Power of Vision Models for Graph Structural Understanding
Daily Paper Cast
UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Fee...
Daily Paper Cast
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
Daily Paper Cast
PHUMA: Physically-Grounded Humanoid Locomotion Dataset
Daily Paper Cast
UniREditBench: A Unified Reasoning-based Image Editing Benchmark
Daily Paper Cast
World Simulation with Video Foundation Models for Physical AI
Daily Paper Cast
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reaso...
Daily Paper Cast
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats
Daily Paper Cast
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement...
Daily Paper Cast
The End of Manual Decoding: Towards Truly End-to-End Language Models
Daily Paper Cast
Kimi Linear: An Expressive, Efficient Attention Architecture
Daily Paper Cast
Surfer 2: The Next Generation of Cross-Platform Computer Use Agents
Daily Paper Cast
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-C...
Daily Paper Cast
The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
Daily Paper Cast
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
Daily Paper Cast
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reaso...
Daily Paper Cast
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy...
Daily Paper Cast
LoongRL:Reinforcement Learning for Advanced Reasoning over Long Contexts
Daily Paper Cast
Language Models are Injective and Hence Invertible
Daily Paper Cast
