Research Papers research paper arxiv machine-learning deep-learning

On-Policy Self-Distillation for Reasoning Compression

arXivMarch 31, 202610 min read0 views

arXiv:2603.05433v4 Announce Type: replace Abstract: Reasoning models think out loud, but much of what they say is noise. We introduce OPSDC (On-Policy Self-Distillation for Reasoning Compression), a method that teaches models to reason more concisely by distilling their own concise behavior back into themselves. The entire approach reduces to one idea: condition the same model on a "be concise" instruction to obtain teacher logits, and minimize per-token reverse KL on the student's own rollouts. No ground-truth answers, no token budgets, no difficulty estimators. Just self-distillation. Yet th — Hejian Sang, Yuanda Xu, Zhengze Zhou, Ran He, Zhipeng Wang, Jiachen Sun

View PDF HTML (experimental)

Abstract:Reasoning models think out loud, but much of what they say is noise. We introduce OPSDC (On-Policy Self-Distillation for Reasoning Compression), a method that teaches models to reason more concisely by distilling their own concise behavior back into themselves. The entire approach reduces to one idea: condition the same model on a "be concise" instruction to obtain teacher logits, and minimize per-token reverse KL on the student's own rollouts. No ground-truth answers, no token budgets, no difficulty estimators. Just self-distillation. Yet this simplicity belies surprising sophistication: OPSDC automatically compresses easy problems aggressively while preserving the deliberation needed for hard ones. On Qwen3-8B and Qwen3-14B, we achieve 57-59% token reduction on MATH-500 while improving accuracy by 9-16 points absolute. On AIME 2024, the 14B model gains 10 points with 41% compression. The secret? Much of what reasoning models produce is not just redundant-it is actively harmful, compounding errors with every unnecessary token. Code is available at this https URL.

Subjects:

Machine Learning (cs.LG)

Cite as: arXiv:2603.05433 [cs.LG]

(or arXiv:2603.05433v4 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2603.05433

arXiv-issued DOI via DataCite

Submission history

From: Hejian Sang [view email] [v1] Thu, 5 Mar 2026 17:54:40 UTC (571 KB) [v2] Sun, 8 Mar 2026 06:29:26 UTC (570 KB) [v3] Tue, 17 Mar 2026 05:05:03 UTC (570 KB) [v4] Sat, 28 Mar 2026 03:56:28 UTC (591 KB)

Original source

arXiv

https://arxiv.org/abs/2603.05433

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

researchpaperarxiv

CountriesRecent

Artificial Intelligence at JPMorgan Chase - Emerj Artificial Intelligence Research

<a href="https://news.google.com/rss/articles/CBMibEFVX3lxTE12bUwyd1dkamZPZExOVHJvb1MxTkZDaml4ak1PbDBvdXlrODBFdmtFVnBMVkhiS1RHTy0yWVRqVmYzQng2NG9VUkcwYVo0R0txZHhMUjFmTDh6NG00N3E2R0RIZWxUb2d4X0dJcw?oc=5" target="_blank">Artificial Intelligence at JPMorgan Chase</a> Emerj Artificial Intelligence Research

Google News: AI

1m1 day ago

ModelsFresh

Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models - WSJ

<a href="https://news.google.com/rss/articles/CBMiuANBVV95cUxNYk90NlRFVDRuRDQxRlFGY3o1SHhHSWdXR3Z3eGJkZjE4blJGSzdKZUNlMlNXR1lUUU5ydGhZQ2ZCS1ItUi12MjBMMEdDc3VfNTE1bUpPYjgxTUI1YU8wZjNZQ3F5RmFyVThObXlZMG9VM1FqQ0xUaThidHNYU3k5dzRBQ2FKcnNLY3FZMjBKcjFUZlFJcVd6dFoyRUd5QlVsVDdCWGVBZk9KXzg4WWotZVdqMUpGS0xUbDBYRmwtWWwxLXRsYU4zSDBLVVhFby12SXFqSVVxWU5YUkMtaVh5b1NPS2tBYkdiR0JuLXR0TEp5MHg0Y1dRR1EyOXV5STdkSzF0U0t2Z0V4UlBJUXkzbDNDNTZvZWotN0Z1UFZ4d2lNY0RMVWo3TEI1MHFrTG11aUZ1bmEtRExzZlhncFg0elYwOTd1RTBvS0t4dGQxcmpvV2JmRU9zWWxMSjVnbW15YklFeG83cWJZNHhEN3JNZXp3WFNGaDdtdDVvNFdTNlJnODFsWlZBTDE1VmRFWGI4SzdFMWxGUFZKUDR5RFNsUGJiaHZnYWlJQmJvTGRRRXdTS3FBVWpIaA?oc=5" target="_blank">Exclusive | Caltech Researchers Claim Radical Compression of High-Fidelity AI Models</a> WSJ

Google News: LLM

1mabout 4 hours ago

ProductsRecent

US may reassess Nato ties after Iran war ends, Rubio says

Secretary of State Marco Rubio said the US may need to reassess its relationship with Nato after the Iran war is finished, calling the military alliance’s alleged lack of support during the Middle East conflict “very disappointing”. Rubio assailed Nato members for denying access to military bases, following prior criticism from US President Donald Trump that partners in the security bloc are “cowards” and that the alliance is a “paper tiger”. “The president and our country will have to...

SCMP Tech (Asia AI)

1mabout 17 hours ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 147 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

More in Research Papers

Research PapersLive

A Retrospective on the ICLR 2026 Review Process

The selection of papers for ICLR 2026 has fully concluded. We extend our congratulations to the authors whose work will appear at the conference. Creating ICLR’s technical program requires immense effort from the authors, reviewers, and area chairs, and we thank you for your contributions and service. For researchers whose work was rejected, we hope […]

blog.iclr.cc

1m31 minutes ago

Research Papers

Vector Researchers present papers at ACL 2024

Vector researchers will be well represented at the 62nd Annual Meeting of the Association for Computational Linguistics in Bangkok, Thailand this year. 14 papers co-authored by Vector-affiliated researchers are being […] The post Vector Researchers present papers at ACL 2024 appeared first on Vector Institute for Artificial Intelligence .

Vector Institute

1mover 1 year ago

Research Papers

Yann LeCun's Team's New Paper: AI Development Mimicking Human Intelligence Hits a Dead End - eu.36kr.com

<a href="https://news.google.com/rss/articles/CBMiU0FVX3lxTFBkbTRhNlhtRnY0cVBERld2OTdWNkRGMXBEaG9Vc21janRUcjJaUlJ4YzZRajVmMGQxNGJYTFB6M3lleUFNakUtWElHdGwzTXBQZjNZ?oc=5" target="_blank">Yann LeCun's Team's New Paper: AI Development Mimicking Human Intelligence Hits a Dead End</a> eu.36kr.com

GNews AI AGI

1m23 days ago

Research Papers

Plans must be made for the welfare of sentient AI, animal consciousness researchers argue - The Hill

<a href="https://news.google.com/rss/articles/CBMiiAFBVV95cUxNNzVaUTkzYkFUaVRsNGtnQVRXS2xsQVZfd1dFQ01RUlNZWUdDbjBNLUNycll2enl2NHp4Z0Ficm9HUnNWUnlvSGFrR3lDVUVxT1QyeE03QWhWcHFDTVJxV3VUQ0FKT3hiTkY3dWZha3JjcjRIM3l3WUtHZVlBUlhxdVBhLW1tdlJ40gGOAUFVX3lxTFBDQnllcVNNa1NRYVMyYlBtVXVxR0VPeHNjTjNMNWNTMFZXRjRkSU1OeXRFNmxvcENqbXkwSERoU1pGdXJYX2g5c214cFJFdEc1WUlkaEE5TlFDTTNoek5yR18tVi1vWUlGUnl4Tk13VWlFMDhzdUUyOUl3RmhNZ0FobTdiVG51N2h1SmJ5Y3c?oc=5" target="_blank">Plans must be made for the welfare of sentient AI, animal consciousness researchers argue</a> The Hill

GNews AI welfare

1mover 1 year ago