A paper to improve training efficiency of on-policy distillation has been accepted by ACL 2026 findings.