Skip to content
TrackPodcasts
scienceDec 10, 202524:05pending

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning

About this episode

🤗 Upvotes: 55 | cs.CL

Authors:
Tong Wu, Yang Liu, Jun Bai, Zixia Jia, Shuyi Zhang, Ziyong Lin, Yanting Wang, Song-Chun Zhu, Zilong Zheng

Title:
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning

Arxiv:
http://arxiv.org/abs/2512.07461v1

Abstract:
We introduce Native Parallel Reasoner (NPR), a teacher-free framework that enables Large Language Models (LLMs) to self-evolve genuine parallel reasoning capabilities. NPR transforms the model from sequential emulation to native parallel cognition through three key innovations: 1) a self-distilled progressive training paradigm that transitions from ``cold-start'' format discovery to strict topological constraints without external supervision; 2) a novel Parallel-Aware Policy Optimization (PAPO) algorithm that optimizes branching policies directly within the execution graph, allowing the model to learn adaptive decomposition via trial and error; and 3) a robust NPR Engine that refactors memory management and flow control of SGLang to enable stable, large-scale parallel RL training. Across eight reasoning benchmarks, NPR trained on Qwen3-4B achieves performance gains of up to 24.5% and inference speedups up to 4.6x. Unlike prior baselines that often fall back to autoregressive decoding, NPR demonstrates 100% genuine parallel execution, establishing a new standard for self-evolving, efficient, and scalable agentic reasoning.

Get every episode summarized

Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning

Daily Paper Cast

0:00
24:05

More episodes

More from Daily Paper Cast

View all episodes →