NeurIPS 2026 · Oral
1University of Amsterdam 2University of Edinburgh
Recent text-only models demonstrate remarkable reasoning capabilities. Extending these to visual domains requires vision-language models to translate images into text descriptions. However, current models, trained to produce captions for human readers, often omit the precise details that reasoning systems require. This creates an interface mismatch: reasoners often fail not due to reasoning limitations but because they lack access to critical visual information. We propose Adaptive-Clarification Reinforcement Learning (AC-RL), which, through interaction, teaches vision models what information reasoners need. Our key insight is that clarification requests during training reveal information gaps; by penalizing success that requires clarification, we create pressure for comprehensive initial captions that enable the reasoner to solve the problem in a single pass. AC-RL improves average accuracy by 4.4 points over pretrained baselines across seven visual reasoning benchmarks, and analysis shows it would cut clarification requests by up to 44% if those were allowed. By treating clarification as a form of implicit supervision, AC-RL demonstrates that vision-language interfaces can be effectively learned through interaction alone, without requiring explicit annotations.
Pairing a small vision-language model with a strong text-only reasoner is attractive: the reasoner can be reused as-is, even when it is only reachable through an API. The weak point is the handoff. A caption that would satisfy a human reader can describe a table without its numbers or a diagram without its marked angles, and the reasoner has no way to look again. We call this the interface mismatch: the reasoner fails for lack of a visual fact, not for lack of reasoning.
Qwen2.5-VL-3B with the same prompt, before and after AC-RL; reasoner DeepSeek-R1-Distill-Qwen-32B. The table is the TableVQA-Bench image, re-typeset here; TableVQA-Bench was never used for training. Both captions are verbatim (Markdown rendered). Nobody told the captioner that the numbers matter; it learned that from the reasoner asking for them.
| Year Ended December 31, | |||
|---|---|---|---|
| 2018 | 2017 | 2016 | |
Net income | $1,096 | $1,346 | $ 566 |
Provision (benefit) for income taxes | 380 | (298) | 343 |
Interest expense, net | 481 | 464 | 511 |
Depreciation of rental equipment | 1,363 | 1,124 | 990 |
Non-rental depreciation and amortization | 308 | 259 | 255 |
EBITDA | 3,628 | 2,895 | 2,665 |
Merger related costs (1) | 36 | 50 | — |
Restructuring charge (2) | 31 | 50 | 14 |
Stock compensation expense, net (3) | 102 | 87 | 45 |
Impact of the fair value mark-up of acquired fleet (4) | 66 | 82 | 35 |
Adjusted EBITDA | $3,863 | $3,164 | $2,759 |
The captioner \(\pi_\theta\) sees the image \(I\) and the question \(Q\) and writes a description \(c_0 \sim \pi_\theta(\cdot \mid I, Q)\). A frozen reasoner reads only \(c_0\) and \(Q\), and decides whether it can solve the problem or needs one more piece of visual information. If it asks a question \(q_1\), a frozen reference captioner answers it with \(c_1 \sim \pi_{\mathrm{ref}}(\cdot \mid I, Q, q_1)\), and the reasoner tries again with the extra detail.
The reward separates three outcomes:
The partial reward does two things. It turns many zero-reward episodes into informative ones, since a caption that was nearly sufficient now scores above a useless one. And the gap \(1-\alpha\) is a price on asking, so the captioner is paid to anticipate the request. We use \(\alpha = 0.7\) in all main experiments and maximize
Two choices keep this honest. Gradients reach only the first caption \(c_0\), and the model answering clarifications is frozen, so the captioner cannot raise its reward by getting better at follow-ups or by leaning on the clarifier. At test time there is no clarification at all.
Optimized with BNPO (a GRPO variant with Beta-normalized advantages), 8 captions per problem, LoRA on Qwen2.5-VL-3B or InternVL3-2B, trained on ViRL-39K. The reasoner is DeepSeek-R1-Distill-Qwen-32B, or Gemini 3.1 Flash-Lite through its API. Because the reasoner and \(\pi_{\mathrm{ref}}\) do not depend on \(\theta\), the tiered reward keeps the policy gradient unbiased (App. B).
Accuracy of the reasoner with captions from each trained model (paper, App. E). A weak penalty barely rewards anticipating the question; a strong one gives almost no credit for nearly-sufficient captions and behaves like a binary reward. A schedule that decays \(\alpha\) linearly from 0.9 to 0.5 reached 50.8 / 37.9, below the fixed value.
Real transcripts on MathVision with one clarification allowed. On the left, the pretrained captioner, which is where training starts; on the right, the same model after AC-RL. The reward chips show what each episode would earn during training. Step through, or switch problems.
Six questions, each with its answer up front. Open a question for the evidence.
Each row adds one ingredient to the same Qwen2.5-VL-3B model, all scored single-pass. Letting the small VLM answer on its own caps it at its own reasoning ability: end-to-end RL on it gains only 1 point. Handing the caption to a 32B reasoner helps more, but the captions were never written for that reader. Training the captioner against the reasoner with a binary correctness reward helps a little; the clarification-aware reward helps most, with the largest gains on the benchmarks where extracting values from the figure is the hard part (DynaMath +10.6, LogicVista +5.8, MathVerse +5.2).
Same recipe with a different backbone: +3.3 points on average over the pretrained InternVL3-2B paired with the same reasoner.
A partial reward makes the learning signal denser, and denser signals help RL on their own. We tested whether that is the whole story in two ways.
Binary reward plus image–caption similarity from Qwen3-VL-Embedding-2B (\(\lambda=0.4\)), same data and optimizer. Paper, App. J.
Similarity helps on 4 of 7 benchmarks but trails AC-RL on all 7. On MathVerse it falls below plain binary RL: rewarding whole-image fidelity pulls the captioner away from the few values a sparse diagram question hinges on.
Each run changes exactly one component of AC-RL. Accuracy with the 32B reasoner.
An uninformative clarifier keeps the reward rule and the reasoner's ability to ask, and removes only the content of the answer: LogicVista falls back to the untrained level. Permuting the reward values among the successful rollouts in each group keeps every advantage magnitude identical and removes only which rollout needed help: −5.6 on LogicVista, −2.0 on MathVision, both significant.
The two reward structures also improve different problems. Across five benchmarks (9,035 problems), 49% of AC-RL's improvements over the pretrained captioner are on problems that Binary-RL still gets wrong. On MMMU's non-math subjects, AC-RL gains 1.75 points where Binary-RL loses 2.06.
We re-run evaluation with clarification switched on, only to measure what the captions were missing. Both captioners were trained against the same 32B reasoner; they differ only in the reward (paper, Table 3).
MathVision (500-problem subset). Up to \(R\) rounds of clarification at evaluation time.
The pretrained captioner gains up to 5.1 points from extra rounds (significant at \(R=2\) and \(R=3\)). The AC-RL captioner is flat at every budget, even though it still asks on 39% of problems at \(R=1\): what those questions would fetch is already in its caption.
Single-pass accuracy on all problems, and on only the problems where the reasoner asked for clarification (paper, Table 3).
We split every caption into sentences and labeled each one: description of what is visible, or answer-like content, meaning a stated answer, a solution step, or a directive to the reasoner. Transcribing a value that is printed in the image counts as description. The labels come from the frozen 32B reasoner following a written rubric, validated on sentences with known labels (recall 0.98, false-positive rate 0.008) and against a blind human check of 100 captions (Cohen's \(\kappa = 0.87\)).
MathVerse, vision-only version: the question is printed in the image. Same problem, both captioners, every sentence labeled.
Numeric claims were additionally checked against the image by Qwen2.5-VL-72B.
The untrained model is the one that tends to solve the problem inside its caption. Deleting its answer-like sentences costs it 10.5 points (7.0 for deleting the same number of random sentences); for AC-RL the cost is 4.7 (2.6 random). If AC-RL's gains were carried by leaked answers, its lead would shrink after filtering; it grows. Residual cases do exist, and they are three times more common from the pretrained model (195 captions vs 65 with two or more answer-like sentences).
AC-RL needs only the reasoner's text output, so the reasoner can sit behind an API. We swapped the open 32B model for Gemini 3.1 Flash-Lite during both training and evaluation, keeping the clarifier local. The per-benchmark pattern is the same as with the open reasoner, and WeMath turns from a small loss into a gain.
Paired bootstrap over test problems, 10,000 resamples (paper, Tables 5 and 6). Qwen2.5-VL-3B captioner.
The captioner was trained against the 32B reasoner only, then paired at test time with a smaller open reasoner and with GPT-5 mini (paper, App. K). Pick a reasoner to highlight it; hover a row for the numbers.
AC-RL vs. the pretrained captioner, both with the 32B reasoner (paper, Fig. 2).
Scale the captioner from 3B to 32B without any RL, keeping the reasoner fixed.
WeMath problems are organized around textbook concepts rather than visual extraction. A ten-times larger captioner barely moves it, so better captions of any kind have little to offer there. AC-RL and Binary-RL shift it by the same small amount (−1.5 vs −1.3, not significant), and with the Gemini reasoner AC-RL improves it by 2.8.
Same captioners and 32B reasoner, benchmarks never used in training.
WorldMedQA-V: all 568 English medical-licensing questions with clinical images. The pattern holds outside mathematics when extraction is the bottleneck. On perception-only questions, where a text bottleneck costs more than the reasoner adds, we do not expect gains.
Captions from the pretrained and AC-RL captioners on the same benchmark problem, with the reasoner's verdict for each. Switch to Be the reasoner to answer from the captions alone, the way the reasoner has to.
Accuracy (%) on seven benchmarks, single-pass evaluation. Click a column to sort within each group.
MathVista testmini, MathVision full, MathVerse vision-only, MMMU dev+val, WeMath testmini (strict), DynaMath worst-case, LogicVista. Other models' numbers are from their reports; † testmini/mini subset; * average over available benchmarks only. Our rows pair a 2B or 3B captioner with a frozen reasoner (DeepSeek-R1-Distill-Qwen-32B or Gemini 3.1 Flash-Lite).
The BibTeX entry will be added once the NeurIPS 2026 proceedings are published.