Papers
arxiv:2608.11829

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

Published on Sep 28
Authors:
,
,
,
,
,
,
,
,
,
,

Abstract

On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. Under the reverse KL objective, the idealized optimum of OPD aligns the student distribution with that of the teacher. When the teacher consistently outperforms the student, this naturally suggests that OPD should yield broad improvements over the pre-OPD student. However, do such improvements extend across the entire range of test-time sampling budgets? In this work, we revisit this expectation through the lens of test-time scaling by varying the sampling budget K and evaluating performance with pass@K. Across multiple settings, we observe two distinct patterns: OPD can improve pass@K at both small and large sampling budgets, but it can also improve small-budget performance while reducing large-budget pass@K. We show one condition that guarantees such a reversal and an idealized reverse KL counterexample where it occurs even when the teacher has higher accuracy on every problem. To choose between two candidate teachers at a target sampling budget, we propose the Teacher Advantage Score at K (TAS@K), which can be computed before OPD training to predict which teacher will lead to a larger improvement in pass@K. Across three domains and thirteen benchmarks, the ordering predicted by TAS@K agrees with the observed pass@K improvements of the resulting OPD models in 83.6\% of experiments, providing a useful signal for teacher selection at the target pass@K.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.11829
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 7

Browse 7 models citing this paper

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.11829 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.11829 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.