No Pain, More Gain:
Iterative Merging for Effective Multi-Teacher On-Policy Distillation

Seonghyeon Kim1,*, Chaeyun Jang1,*, Noah Lee2, Boseop Kim2, Juho Lee1
1KAIST    2Kakao
2026

*Indicates Equal Contribution
Overview of MOPD and IM-MOPD: iterative merging of teacher task vectors during on-policy distillation
IM-MOPD starts from a uniform merge of teacher task vectors and, during multi-teacher on-policy distillation, periodically adds fixed task-vector increments for domains whose capability recovery still lags, thereby determining effective teacher contributions progressively rather than through a one-shot search.

News

  • 2026.09.29Paper released.

Overview

Building a unified language model requires integrating capabilities acquired through different data and training procedures. Different domains often benefit from different post-training recipes such as SFT, RL, or their combinations, making a modular workflow of independent specialization and subsequent integration attractive. Multi-teacher on-policy distillation (MOPD) supports this workflow: each domain teacher provides token-level supervision at the prefixes generated by the student.

However, MOPD can struggle when teachers have very different post-training histories, even when they share the same reference model. Because distillation occurs on student-generated prefixes, the student's initialization can strongly affect what MOPD is able to recover.

Key finding. Initial benchmark performance is not a reliable predictor of a good MOPD initialization. Merge initialization can start below SFT warm-up yet finish higher after MOPD.

We propose IM-MOPD (Iterative Merging for MOPD): start from a uniform merge of teacher task vectors, and during distillation, periodically add fixed task-vector increments for domains whose recovery still lags. This replaces a continuous per-domain coefficient search with repeated merge-or-not decisions, letting both the relative domain weights and the total task-vector contribution evolve together with the student.

86.9
IM-MOPD, 4B (norm. score)
59.0
SFT warm-up, 4B
76.0
IM-MOPD, 1.7B
46.5
SFT warm-up, 1.7B

Observations

We study a setting where all teachers originate from a shared reference model θref and are independently specialized through SFT or RLVR before being integrated via MOPD. Let si(θ) denote the score of model θ on domain i. We report the normalized score s̃i(θ) = [si(θ) − si(θref)] / [si(φi) − si(θref)], anchoring the reference model at 0 and the domain teacher at 1.

Obs. 1Merge initialization is an effective starting point for MOPD

MOPD from the shared reference model struggles to recover the SFT-trained Medical and Tool-Use teachers, while approaching teacher-level performance on the RL-trained Finance teacher. This uneven recovery indicates that student initialization strongly affects how effectively different teacher capabilities transfer. SFT warm-up partially helps but adds a training stage and requires designing a mixed-domain data mixture and schedule.

We find that simply initializing the student by merging teacher task vectors, with θ0 = θref + Σi αi τi and τi = φi − θref, substantially improves subsequent MOPD recovery without requiring additional training.

Normalized recovery across five domains for different student initializations before and after MOPD
Open / filled circles denote initial / final normalized scores under MOPD. Merge initialization, and IM-MOPD in particular, recovers more per-domain capability than either the reference-model init or SFT warm-up.
Takeaway. Task-vector merging is a training-free way to prepare the student for MOPD, and it can outperform an SFT warm-up that adds an entire extra stage.

Obs. 2A stronger student is not necessarily a better MOPD initialization

In the figure above, SFT warm-up yields higher initial performance than a scaled merge initialization (λ = 2.1) on Tool-Use, yet the ordering reverses after MOPD. IM-MOPD reaches the highest final recovery despite not starting from the strongest benchmark score.

The same pattern appears for merge interventions during training: from a shared MOPD checkpoint, adding a task-vector merge can temporarily reduce normalized performance, but after 50 more MOPD updates the intervened branch is higher than the branch that skipped the intervention. Immediate benchmark performance therefore does not fully characterize a good state for subsequent MOPD.

A prefix-level view corroborates this. Evaluating how well each frozen teacher can continue from prefixes generated by different students, we find that teachers score higher when continuing from merge-generated prefixes than from SFT-warm-up prefixes, even when the merged student itself scores similarly or worse on the benchmark.

Teacher continuation accuracy on MedQA as a function of student prefix fraction
Medical (MedQA).
Teacher continuation task success on Tool-Use as a function of student prefix fraction
Tool-Use (τ2-Telecom).
Takeaway. "Good MOPD init" ≠ "high initial benchmark score." What matters is the learning it enables under teacher supervision on student-generated prefixes.

Obs. 3Coefficient search gets harder as domains grow: both ratio and scale matter

Even with two domains (Medical, IF), a coefficient sweep produces very different post-MOPD trade-offs, and the best initialization by average normalized score is not the best after MOPD. Evaluating a candidate therefore requires simulating the learning it enables, not just inspecting the initial score.

With 5 domains, both the domain-wise ratios and the global merge scale λ = Σi αi affect recovery. Non-uniform weighting improves recovery at larger scale, while normalizing the coefficients back to sum to 1 weakens the result. In other words, some strong configurations lie outside the simplex of convex parameter averaging (λ > 1). Increasing the uniform scale helps but does not match a well-chosen non-uniform configuration.

5-domain normalized score under different merge ratios and global scales lambda
5-domain recovery over MOPD steps under different merge initializations. Both merge ratios and the global scale λ matter, and the best non-uniform configuration lies outside the λ = 1 simplex.
Takeaway. Choosing good merge coefficients requires jointly picking ratios and scale and evaluating each candidate through MOPD. This motivates replacing the fixed one-shot search with a progressive, merge-or-not decision made during training, which is the idea behind IM-MOPD.

Results

In a 5-domain setting (Medical, Tool-Use, Finance, Law, IF), IM-MOPD achieves higher average normalized recovery than MOPD with either uniform merge initialization or SFT warm-up. The gain is largest on the domains that MOPD from the reference model struggles with, without sacrificing the domains it already handles.

Method MedQA CaseHOLD FinQA IFBench τ2 Norm.
Qwen3-4B
Reference (Qwen3-4B-OT3)69.561.859.124.95.30.0
Domain Teacher80.773.475.059.466.7100.0
MOPD69.164.774.356.97.942.8
SFT Warm-up + MOPD73.968.271.153.431.659.0
Uniform Merge + MOPD72.767.674.753.611.453.8
IM-MOPD (ours)76.670.974.057.869.386.9
Qwen3-1.7B
Reference (Qwen3-1.7B-OT3)44.144.642.418.47.90.0
Domain Teacher58.569.662.446.428.1100.0
MOPD49.052.259.136.80.034.9
SFT Warm-up + MOPD52.756.654.832.610.546.5
Uniform Merge + MOPD51.259.958.634.64.446.3
IM-MOPD (ours)53.759.455.942.928.176.0

Domain benchmarks: MedQA (Medical), CaseHOLD (Law), FinQA (Finance), IFBench (Instruction Following), τ2 (Tool-Use). Norm. is the average normalized recovery across the five domains (reference: 0, teacher: 100).

On Qwen3-4B, iterative merging lifts normalized recovery from 53.8 (Uniform Merge) to 86.9. On Qwen3-1.7B, IM-MOPD reaches 76.0 vs. 46.5 for SFT warm-up, and recovers Tool-Use back to the teacher level. Across both sizes, IM-MOPD beats SFT warm-up on all five benchmarks. The largest gains are on Medical and Tool-Use, while Finance and IF stay broadly comparable.

BibTeX

@article{kim2026no,
  title={No Pain, More Gain: Iterative Merging for Effective Multi-Teacher On-Policy Distillation},
  author={Kim, Seonghyeon and Jang, Chaeyun and Lee, Noah and Kim, Boseop and Lee, Juho},
  journal={arXiv preprint arXiv:2609.34745},
  year={2026}
}