Canon: Consensus as Privileged Context for Label-Free Self-Distillation

John Gkountouras1 Josip Jukić1 Ivan Titov1,2
1ILLC, University of Amsterdam 2ILCC, University of Edinburgh
Preprint, 2026

Majority voting already tells a model which of its own solutions to trust. Canon puts one of those solutions into the context of a frozen copy of the model and distills what that teacher predicts back into the model, at every token. No labels, no reward model, one epoch.

Four panels: sample rollouts for an unlabeled prompt; take the majority answer and pick the most confident majority solution; condition a frozen teacher on it; distill the teacher into the student with a Jensen-Shannon loss over all rollouts.
Overview. For each unlabeled prompt, the model samples N solutions and takes the majority answer. A frozen snapshot of the model, conditioned on the most confident solution that reaches it, scores every token of every rollout. The model is trained to match those next-token distributions without ever seeing that solution. One generation pass supplies both the vote and the training data.
+10.5
points pass@1 on AMC 2023 (66.0 → 76.5), Qwen3-4B-Instruct-2507
4.8 GPU-h
to train, against 25 to 39 GPU-hours for the label-free RL baselines
34 of 35
model–benchmark cells improve, across four model families
0.4 pts
behind the same recipe with a teacher conditioned on gold-verified solutions (AMC)

Abstract

Sampling multiple solutions and returning the majority answer is among the most reliable ways to improve the reasoning accuracy of large language models without labels, and a growing family of methods converts this consensus signal into training supervision. However, existing approaches use consensus only in restricted forms: as a filter that selects solutions for fine-tuning, as a preference between answers, or as a scalar reward for reinforcement learning, discarding most of the information that the agreeing solutions contain. We present Canon (Consensus-ANchored self-distillatiON), a label-free training method that turns consensus into dense, token-level supervision. For each unlabeled prompt, Canon samples multiple solutions, extracts the majority answer, and conditions a frozen snapshot of the model on a solution that reaches it; this consensus-anchored teacher then supervises the model on its own rollouts at every token.

Experiments on mathematical and scientific reasoning benchmarks show that Canon improves pass@1 by up to 15 points, outperforming label-free reinforcement learning by 6 points at 15% of its compute and approaching a teacher conditioned on gold solutions. Trained on unlabeled data, it transfers to held-out benchmarks, matching training methods that use gold labels. Analysis attributes most of the improvement to the consensus being distilled into the weights and finds no significant narrowing of coverage. On the mathematics benchmarks, Canon's pass@k advantage over the base model persists at every sampling budget measured, reaching 512 samples per prompt, and its majority vote itself becomes more accurate.

Background: learning from your own rollouts

In standard distillation, a student imitates text a teacher wrote. At test time it has to continue from its own text instead, which it never trained on, and small mistakes compound. On-policy distillation (OPD; Agarwal et al., 2024) swaps the roles: the student writes the rollout, and the teacher grades every token by giving its full next-token distribution at that position. The student minimizes a divergence to it. The feedback is as dense as supervised learning and, as in reinforcement learning, it lands on the states the student actually visits.

On-policy self-distillation (OPSD; Zhao et al., 2026; Hübotter et al., 2026) drops the separate teacher. The teacher is the model itself with privileged information in its context, usually the gold solution. Having read the solution, the model predicts better next tokens, and training the model without that context to match them moves what the solution taught into the weights. The privileged information is the catch: it comes from labels.

Canon keeps the OPSD recipe and changes only where the privileged context comes from: a solution the model itself agrees on.

MethodSupervision comes fromSignal per rolloutLabels needed
SFT on filtered samples (RFT, LMSI)the model's own samples that pass a filter, used as targetsone target sequenceRFT: gold answers. LMSI: none
RL with verifiable rewards (GRPO)an answer checkerone scalar rewardgold answers
Test-time RL (TTRL)agreement with the majority voteone scalar rewardnone
On-policy distillationa stronger teacher modela distribution at every tokennone, but a stronger model
On-policy self-distillationthe model itself, shown the gold solutiona distribution at every tokengold solutions
Canona frozen copy of the model, shown a consensus solutiona distribution at every tokennone

Every row learns from the model's own samples. They differ in what grades those samples and in how much signal each one carries.

How Canon works

Everything comes from one round of sampling on unlabeled prompts. The same rollouts that produce the vote are the ones the model is trained on.

  1. 1Sample.

    Draw N = 32 rollouts per prompt and read off each final answer.

  2. 2Vote.

    The majority answer is the consensus. Of the rollouts that reach it, the one with the highest mean token log-probability becomes the consensus solution.

  3. 3Condition.

    A frozen snapshot of the model reads that solution as a reference and scores each rollout in one teacher-forced pass. It never generates.

  4. 4Distill.

    The model, given only the prompt, matches the teacher's next-token distributions on every rollout, for one epoch.

Steps 1–2: why a vote is worth distilling

Take a model that answers "What is the remainder when $2^{100}$ is divided by $7$?" correctly 40% of the time. Most of its single samples are wrong, but they are wrong in different ways, so the majority of 32 samples is right more than 90% of the time. Canon trains on the gap between those two numbers. Draw samples to watch the vote form, then switch the model's answer distribution to see when the consensus fails.

Model's answers
Sample
majority of N samples is correct one sample is correct
Curve: 4,000 simulated votes per N from the answer distribution selected above. Ties go to the answer that appeared first, as in the paper.
$$\mathcal{L}(\theta)=\mathbb{E}_{x\sim\mathcal{D}}\;\frac{1}{N}\sum_{i=1}^{N}\frac{1}{|y^{(i)}|}\sum_{t=1}^{|y^{(i)}|} D_{\mathrm{JS}}\!\Big(\pi_\theta\big(\cdot\mid x,\,y^{(i)}_{<t}\big)\,\Big\Vert\,\bar\pi\big(\cdot\mid x,\,s_x,\,y^{(i)}_{<t}\big)\Big)$$

$s_x$: consensus solution for prompt $x$. $\bar\pi$: frozen teacher, no gradient. The divergence is over the full vocabulary.

Steps 3–4: what the teacher's context adds

Student and teacher are the same weights and score the same rollout. The only difference is what the teacher gets to read first. With nothing added, the two distributions are identical and there is nothing to learn. Add text to the teacher's context and watch where its next-token distributions move away from the student's. The per-token Jensen–Shannon divergence is the training signal; its mean over the rollout is this rollout's term in the loss above.

Teacher reads
Student's rollout
per-token JSD0→0.25 natsdotted: tokens where rollouts A and B differ · click a token to inspect it

student, prompt only teacher, with its context ◂ the token the student sampled

At this token

Loss for the whole rollout (token-mean JSD)

The distributions here are set by hand over a few candidate tokens, with the rest of the vocabulary pooled into one bucket; every number shown is computed from them. They illustrate the mechanics and are not model outputs. A larger divergence is not automatically better: another problem's solution also moves the teacher, but toward the wrong things.

Dense

A majority-vote reward gives one bit per trajectory. The teacher gives a full next-token distribution at every position of every rollout, including rollouts that disagree with the vote.

Frozen

A teacher that shares live weights with the student collapses: loss and generation entropy reach zero within five steps. A fixed snapshot is stable.

Cheap

No reward model, no preference pairs, no generation beyond self-consistency decoding. On AMC that is 2,656 sequences and about 13 optimizer steps.

The teacher's prompt
{problem}
Please reason step by step, and put your final answer within \boxed{}.

A correct solution to this problem is given below for your reference:
<solution>
{consensus solution}
</solution>

Guided by the reference solution, write your own step-by-step solution, and put your final answer within \boxed{}.

Research questions

    Q1

    Can a model's own consensus replace the gold solution as privileged context?

    On mathematics, almost entirely. With everything else fixed, a consensus-conditioned teacher reaches 76.5 on AMC; a teacher conditioned on gold-verified solutions reaches 76.9.

    • Trained per benchmark, Canon recovers 97% of the gold-conditioned oracle's gain on AMC, 78% on AIME 2024 and 58% on GPQA. GPQA is the one benchmark where gold conditioning keeps a clear advantage.
    • Changing only the teacher's context isolates what carries the signal. A content-free placeholder does nothing (+0.6). The consensus final answer alone gives +3.5, and another prompt's consensus solution, solution-shaped text with the wrong content, gives +3.9.
    • The prompt's own consensus solution gives +10.5. Most of the signal is in its content, beyond the answer it reaches.
    Gain over the base model, by teacher context
    Change in avg@32 (points). Qwen3-4B-Instruct-2507, transductive; every arm uses the same recipe and differs only in what the frozen teacher is conditioned on.
    Q2

    How does it compare with label-free reinforcement learning, and at what cost?

    It is the strongest label-free method we tested: 6.2 points above TTRL on AMC, at 15% of its compute.

    • The five label-free RL methods end within 1.4 points of each other (69.6 to 71.0), whether the reward is the majority vote, a gated vote, semantic entropy or token entropy alone. All sit 5 to 7 points below Canon.
    • Gold-reward GRPO overtakes Canon on AMC after 180 steps and 76.8 GPU-hours. That 3.1-point lead is not significant (95% CI [−0.3, +6.9]), and on held-out AIME 2025 and GPQA the same checkpoint transfers worse.
    • After Canon, one greedy sample scores 76.7 on AMC (base: 67.1), approaching the base model's 32-sample majority vote of 83.1. Much of the benefit of voting moves into the weights.
    AMC 2023 accuracy against training compute
    avg@32, Qwen3-4B-Instruct-2507, all methods trained on the same 83 unlabeled AMC prompts. GPU-hours include generation. Hover a point, or a row of the table below, to find it in the other.
    MethodLabelsGPU-hAMC (trained on)AIME24AIME25GPQAAvg
    Base model—066.032.530.143.943.1
    ScPOnone4.066.432.930.644.343.6
    LMSInone5.666.532.529.344.243.1
    TTRLnone32.470.336.731.345.045.8
    EMPOnone25.669.636.931.745.045.8
    SCRLnone39.271.037.331.445.046.2
    TTRL-Guardnone37.670.534.933.245.746.1
    EM-RLnone24.869.737.331.445.946.1
    Canon (ours)none4.876.539.532.648.949.4
    RFTgold3.670.037.031.244.945.8
    GRPO (60 steps)gold25.671.738.431.845.446.8
    GRPO (180 steps, best)gold76.879.639.730.447.649.3
    Oracle distillationgold4.876.940.431.649.949.7

    avg@32, Qwen3-4B-Instruct-2507. Every trained arm uses the same 83 unlabeled AMC prompts; the other columns evaluate the same checkpoints on benchmarks they never trained on. Bold: best label-free result per column. Shaded rows use gold labels.

    Q3

    Does it hold across models, families and benchmarks?

    34 of 35 model–benchmark cells improve, across seven models from four families and 2B to 9B parameters, with one recipe per model.

    Each cell trains only on that benchmark's unlabeled prompts. Large number: change in avg@32; below it, base → Canon. †Mixture-of-experts, 1B active parameters.
    • The largest single gain is outside the Qwen family: Gemma-4-E4B-it gains 15.4 points on AIME 2024.
    • Gains follow headroom. Qwen3-4B-Instruct already answers 89% of MATH500 samples correctly and votes at 95%, and gains 1.8; models with more room gain 3.3 to 6.6 on the same benchmark. The one regression is LFM2.5-8B-A1B on MATH500.
    • RL on Qwen models can improve math even with random rewards (Shao et al.), which recommend checking a new signal across families and against an uninformative control. A content-free teacher gives no gain, and the gains appear in three other families.
    Q4

    Do the improvements transfer to prompts it never trained on?

    Yes. Trained on unlabeled prompts, Canon improves held-out benchmarks and matches gold-reward GRPO trained on the same prompts.

    • Trained on the 83 AMC prompts alone, Canon improves AIME 2024 by +7.0, AIME 2025 by +2.5 and GPQA by +5.0, a transfer from competition math to graduate-level science.
    • Trained on a 339-prompt unlabeled pool, one epoch reaches 41.7 on held-out AIME 2024, above the final checkpoints of TTRL (39.4, 138 GPU-h) and of GRPO with gold rewards (39.0, 142 GPU-h), at about 12% of their compute.
    • TTRL passes through 41.4 at step 220 before regressing, but choosing that checkpoint needs held-out labels. Canon trains for exactly one epoch.
    • Gains shrink with distance from the training data but stay positive: +9.2 on the training slice, +4.9 on harder problems from the same source, +0.6 on another math benchmark, +3.9 on GPQA (Qwen3.5-4B, OmniMath slice).
    • AIME has 30 problems, so one problem moves maj@32 by 3.3 points. Canon-pool's majority vote lands below the base model's here (53.3 against 56.7).
    Held-out AIME 2024 during training on the unlabeled pool
    avg@32, Qwen3-4B-Instruct-2507. One Canon epoch is roughly 55 optimizer steps at this batch size; the axis counts the RL methods' steps.
    Q5

    Is it only sharpening? Does it narrow what the model can solve?

    Coverage rises with accuracy. On AMC, Canon's pass@k stays above the base model's at every k up to 512, and the gap widens.

    • Canon does concentrate the answer distribution (mean vote share 0.68 → 0.81). Concentration alone, as in EM-RL with a pure token-entropy reward, reaches 69.7 and leaves majority vote at the base model's 83.1.
    • Majority-vote accuracy rises on all four hard benchmarks (+2.4 to +3.3), and pass@32 by 4.9 on AMC, 10.0 on AIME 2024 and 16.7 on AIME 2025.
    • With up to 512 samples per prompt, Canon solves 12 of the 29 AMC and AIME problems the base model never solves, and loses none that the base model solves. RL-trained models have been reported to show the opposite, a crossover at large k (Yue et al., 2025).
    • Zero successes in a finite draw bound a probability without ruling it out. These results argue against sharpening as the whole explanation; they do not establish new capability.
    AMC 2023, up to 512 samples per problem
    Two pooled 256-sample draws per problem. At k = 512 the paired difference is +7.2 points (95% CI [+2.4, +13.3]); at no k does the interval favor the base model.
    Benchmarkmax kbase pass@kCanon pass@kΔ [95% CI]crossovermaj@k base / Canon
    AMC 202351289.296.4+7.2 [+2.4, +13.3]none83.1 / 86.8
    AIME 202425673.380.0+6.7 [0.0, +16.7]none63.3 / 66.7
    AIME 202525660.073.3+13.3 [+3.3, +26.7]none53.3 / 56.7
    GPQA-Diamond12884.983.8−1.0 [−6.1, +4.0]k = 12863.1 / 68.2
    MATH500*6498.498.40.0tie (saturated)96.8 / 95.2

    GPQA has four options, so pass@k at large k rewards spreading guesses across letters; the base model spreads more and closes the gap at k = 128, while its majority vote stays 5 points lower. *Paired on the 63 prompts in the base model's large-k run; over all 500 prompts, Canon reaches pass@64 97.4.

    Q6

    When does consensus supervision help?

    When the base model's majority is right but unsure. When the majority is confidently wrong, Canon reinforces the error, but that case is rare.

    Mean change in avg@32 by the state of the base model's consensus, pooled over twelve transductive cells: 1,493 prompts, three models. Confident means a vote share of at least 0.5.
    • The teacher can only be as good as the consensus behind it. Gains peak where the consensus is informative but not yet expressed in single samples, and fade where the vote is already unanimous.
    • A practical rule follows: Canon is worth running when majority-vote accuracy clearly exceeds single-sample accuracy, the usual situation for small instruct models on competition math and multiple-choice science.
    Gain by base vote share
    Q7

    Which parts of the recipe matter?

    A frozen teacher, dense per-token targets, and supervision on every rollout. The sampling budget matters less: N = 4 already recovers two thirds of the gain.

    • Supervised fine-tuning on the same consensus solution gains 1.4 points; distilling toward the teacher conditioned on it gains 10.5. The dense per-token target, not the choice of solution, carries the improvement.
    • Masking out minority-answer rollouts costs 3.5 points, and skipping prompts with vote share below 0.5 costs 4.4. Low-agreement prompts still carry signal.
    • A live teacher that shares weights with the student collapses to 0.0 avg@32 on the development slice, where the frozen snapshot reaches 72.1.
    • More training helps little: a second round with fresh rollouts adds 0.9 points on AMC at twice the cost, and a third gives it back.
    One change at a time, AMC 2023
    Rollouts per prompt
    avg@32 gain over the base model (66.0). Qwen3-4B-Instruct-2507, transductive; training-seed standard deviation of the recipe is 0.3 points.

    Limitations

    BibTeX

    @article{gkountouras2026canon,
      title   = {Consensus as Privileged Context for Label-Free Self-Distillation},
      author  = {Gkountouras, John and Juki{\'c}, Josip and Titov, Ivan},
      journal = {arXiv preprint arXiv:2607.13643},
      year    = {2026}
    }