Majority voting already tells a model which of its own solutions to trust. Canon puts one of those solutions into the context of a frozen copy of the model and distills what that teacher predicts back into the model, at every token. No labels, no reward model, one epoch.
Sampling multiple solutions and returning the majority answer is among the most reliable ways to improve the reasoning accuracy of large language models without labels, and a growing family of methods converts this consensus signal into training supervision. However, existing approaches use consensus only in restricted forms: as a filter that selects solutions for fine-tuning, as a preference between answers, or as a scalar reward for reinforcement learning, discarding most of the information that the agreeing solutions contain. We present Canon (Consensus-ANchored self-distillatiON), a label-free training method that turns consensus into dense, token-level supervision. For each unlabeled prompt, Canon samples multiple solutions, extracts the majority answer, and conditions a frozen snapshot of the model on a solution that reaches it; this consensus-anchored teacher then supervises the model on its own rollouts at every token.
Experiments on mathematical and scientific reasoning benchmarks show that Canon improves pass@1 by up to 15 points, outperforming label-free reinforcement learning by 6 points at 15% of its compute and approaching a teacher conditioned on gold solutions. Trained on unlabeled data, it transfers to held-out benchmarks, matching training methods that use gold labels. Analysis attributes most of the improvement to the consensus being distilled into the weights and finds no significant narrowing of coverage. On the mathematics benchmarks, Canon's pass@k advantage over the base model persists at every sampling budget measured, reaching 512 samples per prompt, and its majority vote itself becomes more accurate.
In standard distillation, a student imitates text a teacher wrote. At test time it has to continue from its own text instead, which it never trained on, and small mistakes compound. On-policy distillation (OPD; Agarwal et al., 2024) swaps the roles: the student writes the rollout, and the teacher grades every token by giving its full next-token distribution at that position. The student minimizes a divergence to it. The feedback is as dense as supervised learning and, as in reinforcement learning, it lands on the states the student actually visits.
On-policy self-distillation (OPSD; Zhao et al., 2026; Hübotter et al., 2026) drops the separate teacher. The teacher is the model itself with privileged information in its context, usually the gold solution. Having read the solution, the model predicts better next tokens, and training the model without that context to match them moves what the solution taught into the weights. The privileged information is the catch: it comes from labels.
Canon keeps the OPSD recipe and changes only where the privileged context comes from: a solution the model itself agrees on.
| Method | Supervision comes from | Signal per rollout | Labels needed |
|---|---|---|---|
| SFT on filtered samples (RFT, LMSI) | the model's own samples that pass a filter, used as targets | one target sequence | RFT: gold answers. LMSI: none |
| RL with verifiable rewards (GRPO) | an answer checker | one scalar reward | gold answers |
| Test-time RL (TTRL) | agreement with the majority vote | one scalar reward | none |
| On-policy distillation | a stronger teacher model | a distribution at every token | none, but a stronger model |
| On-policy self-distillation | the model itself, shown the gold solution | a distribution at every token | gold solutions |
| Canon | a frozen copy of the model, shown a consensus solution | a distribution at every token | none |
Every row learns from the model's own samples. They differ in what grades those samples and in how much signal each one carries.
Everything comes from one round of sampling on unlabeled prompts. The same rollouts that produce the vote are the ones the model is trained on.
Draw N = 32 rollouts per prompt and read off each final answer.
The majority answer is the consensus. Of the rollouts that reach it, the one with the highest mean token log-probability becomes the consensus solution.
A frozen snapshot of the model reads that solution as a reference and scores each rollout in one teacher-forced pass. It never generates.
The model, given only the prompt, matches the teacher's next-token distributions on every rollout, for one epoch.
Take a model that answers "What is the remainder when $2^{100}$ is divided by $7$?" correctly 40% of the time. Most of its single samples are wrong, but they are wrong in different ways, so the majority of 32 samples is right more than 90% of the time. Canon trains on the gap between those two numbers. Draw samples to watch the vote form, then switch the model's answer distribution to see when the consensus fails.
$s_x$: consensus solution for prompt $x$. $\bar\pi$: frozen teacher, no gradient. The divergence is over the full vocabulary.
Student and teacher are the same weights and score the same rollout. The only difference is what the teacher gets to read first. With nothing added, the two distributions are identical and there is nothing to learn. Add text to the teacher's context and watch where its next-token distributions move away from the student's. The per-token Jensen–Shannon divergence is the training signal; its mean over the rollout is this rollout's term in the loss above.
The distributions here are set by hand over a few candidate tokens, with the rest of the vocabulary pooled into one bucket; every number shown is computed from them. They illustrate the mechanics and are not model outputs. A larger divergence is not automatically better: another problem's solution also moves the teacher, but toward the wrong things.
A majority-vote reward gives one bit per trajectory. The teacher gives a full next-token distribution at every position of every rollout, including rollouts that disagree with the vote.
A teacher that shares live weights with the student collapses: loss and generation entropy reach zero within five steps. A fixed snapshot is stable.
No reward model, no preference pairs, no generation beyond self-consistency decoding. On AMC that is 2,656 sequences and about 13 optimizer steps.
{problem}
Please reason step by step, and put your final answer within \boxed{}.
A correct solution to this problem is given below for your reference:
<solution>
{consensus solution}
</solution>
Guided by the reference solution, write your own step-by-step solution, and put your final answer within \boxed{}.
On mathematics, almost entirely. With everything else fixed, a consensus-conditioned teacher reaches 76.5 on AMC; a teacher conditioned on gold-verified solutions reaches 76.9.
It is the strongest label-free method we tested: 6.2 points above TTRL on AMC, at 15% of its compute.
| Method | Labels | GPU-h | AMC (trained on) | AIME24 | AIME25 | GPQA | Avg |
|---|---|---|---|---|---|---|---|
| Base model | — | 0 | 66.0 | 32.5 | 30.1 | 43.9 | 43.1 |
| ScPO | none | 4.0 | 66.4 | 32.9 | 30.6 | 44.3 | 43.6 |
| LMSI | none | 5.6 | 66.5 | 32.5 | 29.3 | 44.2 | 43.1 |
| TTRL | none | 32.4 | 70.3 | 36.7 | 31.3 | 45.0 | 45.8 |
| EMPO | none | 25.6 | 69.6 | 36.9 | 31.7 | 45.0 | 45.8 |
| SCRL | none | 39.2 | 71.0 | 37.3 | 31.4 | 45.0 | 46.2 |
| TTRL-Guard | none | 37.6 | 70.5 | 34.9 | 33.2 | 45.7 | 46.1 |
| EM-RL | none | 24.8 | 69.7 | 37.3 | 31.4 | 45.9 | 46.1 |
| Canon (ours) | none | 4.8 | 76.5 | 39.5 | 32.6 | 48.9 | 49.4 |
| RFT | gold | 3.6 | 70.0 | 37.0 | 31.2 | 44.9 | 45.8 |
| GRPO (60 steps) | gold | 25.6 | 71.7 | 38.4 | 31.8 | 45.4 | 46.8 |
| GRPO (180 steps, best) | gold | 76.8 | 79.6 | 39.7 | 30.4 | 47.6 | 49.3 |
| Oracle distillation | gold | 4.8 | 76.9 | 40.4 | 31.6 | 49.9 | 49.7 |
avg@32, Qwen3-4B-Instruct-2507. Every trained arm uses the same 83 unlabeled AMC prompts; the other columns evaluate the same checkpoints on benchmarks they never trained on. Bold: best label-free result per column. Shaded rows use gold labels.
34 of 35 model–benchmark cells improve, across seven models from four families and 2B to 9B parameters, with one recipe per model.
Yes. Trained on unlabeled prompts, Canon improves held-out benchmarks and matches gold-reward GRPO trained on the same prompts.
Coverage rises with accuracy. On AMC, Canon's pass@k stays above the base model's at every k up to 512, and the gap widens.
| Benchmark | max k | base pass@k | Canon pass@k | Δ [95% CI] | crossover | maj@k base / Canon |
|---|---|---|---|---|---|---|
| AMC 2023 | 512 | 89.2 | 96.4 | +7.2 [+2.4, +13.3] | none | 83.1 / 86.8 |
| AIME 2024 | 256 | 73.3 | 80.0 | +6.7 [0.0, +16.7] | none | 63.3 / 66.7 |
| AIME 2025 | 256 | 60.0 | 73.3 | +13.3 [+3.3, +26.7] | none | 53.3 / 56.7 |
| GPQA-Diamond | 128 | 84.9 | 83.8 | −1.0 [−6.1, +4.0] | k = 128 | 63.1 / 68.2 |
| MATH500* | 64 | 98.4 | 98.4 | 0.0 | tie (saturated) | 96.8 / 95.2 |
GPQA has four options, so pass@k at large k rewards spreading guesses across letters; the base model spreads more and closes the gap at k = 128, while its majority vote stays 5 points lower. *Paired on the 63 prompts in the base model's large-k run; over all 500 prompts, Canon reaches pass@64 97.4.
When the base model's majority is right but unsure. When the majority is confidently wrong, Canon reinforces the error, but that case is rare.
A frozen teacher, dense per-token targets, and supervision on every rollout. The sampling budget matters less: N = 4 already recovers two thirds of the gain.
@article{gkountouras2026canon,
title = {Consensus as Privileged Context for Label-Free Self-Distillation},
author = {Gkountouras, John and Juki{\'c}, Josip and Titov, Ivan},
journal = {arXiv preprint arXiv:2607.13643},
year = {2026}
}