Skip to content
TrackPodcasts

Daily Paper Cast

Jingwen Liang, Gengyu Wang

EN1777 episodes
sciencetechnology

We update every weekday to discuss highest-voted papers from Huggingface Daily Paper (https://huggingface.co/papers). Both the podcast scripts and audio are generated by AI. Feedback and suggestions are welcome! Email us: [email protected]:Jingwen Liang, 3D ML, https://www.linkedin.com/in/jingwen-liang/Gengyu Wang, LLM ML, http://wanggengyu.comListen on: Spotify: https://open.spotify.com/show/21nrhmdaA8qoBiH8q03NXLApple Podcast: https://podcasts.apple.com/us/podcast/daily-paper-...

Episodes (1777)

Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance

Daily Paper Cast

Nov 19, 202523:57pending

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

Daily Paper Cast

Nov 19, 202525:57pending

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Ed...

Daily Paper Cast

Nov 19, 202520:43pending

GroupRank: A Groupwise Reranking Paradigm Driven by Reinforcement Learning

Daily Paper Cast

Nov 19, 202523:49pending

TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models

Daily Paper Cast

Nov 19, 202523:11pending

PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image

Daily Paper Cast

Nov 19, 202525:25pending

GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Model...

Daily Paper Cast

Nov 18, 202521:58pending

DoPE: Denoising Rotary Position Embedding

Daily Paper Cast

Nov 18, 202519:24pending

WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and...

Daily Paper Cast

Nov 18, 202525:25pending

UI2Code^N: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code...

Daily Paper Cast

Nov 18, 202525:46pending

AIonopedia: an LLM agent orchestrating multimodal learning for ionic liquid disc...

Daily Paper Cast

Nov 18, 202528:52pending

LiteAttention: A Temporal Sparse Attention for Diffusion Transformers

Daily Paper Cast

Nov 18, 202521:16pending

Virtual Width Networks

Daily Paper Cast

Nov 18, 202522:36pending

One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter...

Daily Paper Cast

Nov 15, 202522:00pending

PAN: A World Model for General, Interactable, and Long-Horizon World Simulation

Daily Paper Cast

Nov 15, 202525:56pending

UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalis...

Daily Paper Cast

Nov 15, 202527:35pending

Too Good to be Bad: On the Failure of LLMs to Role-Play Villains

Daily Paper Cast

Nov 11, 202525:22pending

DeepEyesV2: Toward Agentic Multimodal Model

Daily Paper Cast

Nov 11, 202526:21pending

Visual Spatial Tuning

Daily Paper Cast

Nov 11, 202524:14pending

VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Chec...

Daily Paper Cast

Nov 11, 202522:55pending

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradi...

Daily Paper Cast

Nov 8, 202524:22pending

V-Thinker: Interactive Thinking with Images

Daily Paper Cast

Nov 8, 202521:01pending

Scaling Agent Learning via Experience Synthesis

Daily Paper Cast

Nov 8, 202523:00pending

Diffusion Language Models are Super Data Learners

Daily Paper Cast

Nov 7, 202522:27pending

LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environme...

Daily Paper Cast

Nov 7, 202525:51pending

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interac...

Daily Paper Cast

Nov 7, 202523:19pending

Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

Daily Paper Cast

Nov 6, 202528:20pending

VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation

Daily Paper Cast

Nov 6, 202521:54pending

When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Ch...

Daily Paper Cast

Nov 6, 202524:13pending

Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language F...

Daily Paper Cast

Nov 5, 202524:08pending

Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph

Daily Paper Cast

Nov 5, 202522:50pending

The Underappreciated Power of Vision Models for Graph Structural Understanding

Daily Paper Cast

Nov 5, 202525:56pending

UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Fee...

Daily Paper Cast

Nov 5, 202523:59pending

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation

Daily Paper Cast

Nov 5, 202525:34pending

PHUMA: Physically-Grounded Humanoid Locomotion Dataset

Daily Paper Cast

Nov 5, 202521:26pending

UniREditBench: A Unified Reasoning-based Image Editing Benchmark

Daily Paper Cast

Nov 5, 202523:01pending

World Simulation with Video Foundation Models for Physical AI

Daily Paper Cast

Nov 5, 202529:01pending

ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reaso...

Daily Paper Cast

Nov 4, 202522:48pending

INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats

Daily Paper Cast

Nov 4, 202520:54pending

Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement...

Daily Paper Cast

Nov 4, 202525:27pending

The End of Manual Decoding: Towards Truly End-to-End Language Models

Daily Paper Cast

Nov 1, 202522:34pending

Kimi Linear: An Expressive, Efficient Attention Architecture

Daily Paper Cast

Nov 1, 202523:11pending

Surfer 2: The Next Generation of Cross-Platform Computer Use Agents

Daily Paper Cast

Nov 1, 202523:45pending

Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-C...

Daily Paper Cast

Nov 1, 202525:05pending

The Quest for Generalizable Motion Generation: Data, Model, and Evaluation

Daily Paper Cast

Nov 1, 202522:48pending

Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

Daily Paper Cast

Oct 29, 202523:09pending

Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reaso...

Daily Paper Cast

Oct 24, 202523:06pending

BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy...

Daily Paper Cast

Oct 24, 202521:40pending

LoongRL:Reinforcement Learning for Advanced Reasoning over Long Contexts

Daily Paper Cast

Oct 24, 202521:23pending

Language Models are Injective and Hence Invertible

Daily Paper Cast

Oct 24, 202524:03pending
......