NeurIPS 2026 · Oral

Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces

John Gkountouras1,   Ivan Titov1,2

1University of Amsterdam   2University of Edinburgh

Swipe sideways to see the whole figure.

INPUT AND FIRST DESCRIPTION DIRECT ANSWER CLARIFICATION · TRAINING ONLY A B C D E 130° 80° x What is the value of angle x? Captioner πθ small VLM · trainable A diagram shows triangle ABC. Two line segments, DB and BE, meet at vertex B. Angle DBA is 130°, angle BCA is 80°, and angle BAC is labeled x. first caption c0 never says whether D, B and E lie on one line caption looks sufficient The answer is 60° ✗ correct? no yes R = 0 R = 1 single pass: this answer is final a detail is missing: ask once Are the points D, B, and Ecollinear? q1 Yes. frozen reference captioner answers with c1 The answer is 50° ✓ correct? no yes R = 0 R = α = 0.7 policy gradient reaches only the first caption; the reasoner and the frozen captioner get none trainable frozen gradient not used at test time
AC-RL on one training example. The trainable captioner \(\pi_\theta\) describes the diagram for a frozen, text-only reasoner, but its caption never says whether D, B and E lie on one line. Depending on the caption, the reasoner either answers right away (middle) or asks one question (right). Guessing, it says 60° and the episode earns \(R=0\). Asking, it learns the points are collinear, answers 50°, and the episode earns \(R=\alpha=0.7\) instead of the full \(R=1\). Only the first caption \(c_0\) receives gradients, so the captioner is pushed to state such facts up front. At test time only the direct path exists.

TL;DR  Pipelines that caption an image for a text-only reasoner often fail because the caption leaves out the one number or label the reasoner needed. AC-RL trains the captioner against the reasoner it will serve, and uses the reasoner's own clarification requests to find out what was missing. No caption annotations, no changes to the reasoner, and a single pass at test time.

+4.4
avg. accuracy over the pretrained captioner on 7 benchmarks (3B captioner, frozen 32B reasoner)
+5.2
with Gemini 3.1 Flash-Lite as a black-box API reasoner, about $15 of API calls
−44%
fewer clarification requests on MathVerse, when requests are allowed
13 / 14
benchmark gains significant at 95% across both reasoners (paired bootstrap)

Abstract

Recent text-only models demonstrate remarkable reasoning capabilities. Extending these to visual domains requires vision-language models to translate images into text descriptions. However, current models, trained to produce captions for human readers, often omit the precise details that reasoning systems require. This creates an interface mismatch: reasoners often fail not due to reasoning limitations but because they lack access to critical visual information. We propose Adaptive-Clarification Reinforcement Learning (AC-RL), which, through interaction, teaches vision models what information reasoners need. Our key insight is that clarification requests during training reveal information gaps; by penalizing success that requires clarification, we create pressure for comprehensive initial captions that enable the reasoner to solve the problem in a single pass. AC-RL improves average accuracy by 4.4 points over pretrained baselines across seven visual reasoning benchmarks, and analysis shows it would cut clarification requests by up to 44% if those were allowed. By treating clarification as a form of implicit supervision, AC-RL demonstrates that vision-language interfaces can be effectively learned through interaction alone, without requiring explicit annotations.

Captions are written for people

Pairing a small vision-language model with a strong text-only reasoner is attractive: the reasoner can be reused as-is, even when it is only reachable through an API. The weak point is the handoff. A caption that would satisfy a human reader can describe a table without its numbers or a diagram without its marked angles, and the reasoner has no way to look again. We call this the interface mismatch: the reasoner fails for lack of a visual fact, not for lack of reasoning.

QuestionWhat was the depreciation of rental equipment for the year ended December 31, 2016?  Answer: 990
Pretrained captioner✗ reasoner answers 1,200
The image shows a financial statement with data for three years ending December 31, 2018, 2017, and 2016. The statement includes various financial metrics such as net income, provision (benefit) for income taxes, interest expense, depreciation of rental equipment, non-rental depreciation and amortization, EBITDA, merger related costs, restructuring charge, stock compensation expense, and impact of the fair value mark-up of acquired fleet. The adjusted EBITDA for each year is also provided.
AC-RL captioner✓ reasoner answers 990
The image shows a financial statement titled "Year Ended December 31," detailing various financial metrics for three consecutive years: 2018, 2017, and 2016. The metrics listed include: - Net income: $1,096, $1,346, and $566 respectively. - Provision (benefit) for income taxes: $380, $(298), and $343 respectively. - Interest expense, net: $481, $464, and $511 respectively. - Depreciation of rental equipment: $1,363, $1,124, and $990 respectively. - Non-rental depreciation and amortization: $308, $259, and $255 respectively. - EBITDA: $3,628, $2,895, and $2,665 respectively. [5 more rows]

Qwen2.5-VL-3B with the same prompt, before and after AC-RL; reasoner DeepSeek-R1-Distill-Qwen-32B. The table is the TableVQA-Bench image, re-typeset here; TableVQA-Bench was never used for training. Both captions are verbatim (Markdown rendered). Nobody told the captioner that the numbers matter; it learned that from the reasoner asking for them.

Method

The captioner \(\pi_\theta\) sees the image \(I\) and the question \(Q\) and writes a description \(c_0 \sim \pi_\theta(\cdot \mid I, Q)\). A frozen reasoner reads only \(c_0\) and \(Q\), and decides whether it can solve the problem or needs one more piece of visual information. If it asks a question \(q_1\), a frozen reference captioner answers it with \(c_1 \sim \pi_{\mathrm{ref}}(\cdot \mid I, Q, q_1)\), and the reasoner tries again with the extra detail.

The reward separates three outcomes:

\[ R(\tau) = \begin{cases} 1 & \text{correct without clarification} \\ \alpha & \text{correct after one clarification} \\ 0 & \text{otherwise} \end{cases} \]

The partial reward does two things. It turns many zero-reward episodes into informative ones, since a caption that was nearly sufficient now scores above a useless one. And the gap \(1-\alpha\) is a price on asking, so the captioner is paid to anticipate the request. We use \(\alpha = 0.7\) in all main experiments and maximize

\[ J(\theta) = \mathbb{E}_{(I,Q)\sim\mathcal{D}}\, \mathbb{E}_{\tau\sim\pi_\theta}\big[R(\tau)\big] \;-\; \beta\, D_{\mathrm{KL}}\big(\pi_\theta \,\|\, \pi_{\mathrm{ref}}\big). \]

Two choices keep this honest. Gradients reach only the first caption \(c_0\), and the model answering clarifications is frozen, so the captioner cannot raise its reward by getting better at follow-ups or by leaning on the clarifier. At test time there is no clarification at all.

Optimized with BNPO (a GRPO variant with Beta-normalized advantages), 8 captions per problem, LoRA on Qwen2.5-VL-3B or InternVL3-2B, trained on ViRL-39K. The reasoner is DeepSeek-R1-Distill-Qwen-32B, or Gemini 3.1 Flash-Lite through its API. Because the reasoner and \(\pi_{\mathrm{ref}}\) do not depend on \(\theta\), the tiered reward keeps the policy gradient unbiased (App. B).

How large should the price of asking be?

Correct, no clarification
1.0
Correct after one clarificationpartial credit \(\alpha\)
0.7
Wrong
0
0.3

Accuracy of the reasoner with captions from each trained model (paper, App. E). A weak penalty barely rewards anticipating the question; a strong one gives almost no credit for nearly-sufficient captions and behaves like a binary reward. A schedule that decays \(\alpha\) linearly from 0.9 to 0.5 reached 50.8 / 37.9, below the fixed value.

One problem, two captioners

Real transcripts on MathVision with one clarification allowed. On the left, the pretrained captioner, which is where training starts; on the right, the same model after AC-RL. The reward chips show what each episode would earn during training. Step through, or switch problems.

Findings

Six questions, each with its answer up front. Open a question for the evidence.

RQ1 Does training the caption for the reasoner help? Yes. AC-RL adds 4.4 points on average for a 3B captioner, without touching the reasoner, and 3.1 points more than training the same captioner with a binary reward.

Each row adds one ingredient to the same Qwen2.5-VL-3B model, all scored single-pass. Letting the small VLM answer on its own caps it at its own reasoning ability: end-to-end RL on it gains only 1 point. Handing the caption to a 32B reasoner helps more, but the captions were never written for that reader. Training the captioner against the reasoner with a binary correctness reward helps a little; the clarification-aware reward helps most, with the largest gains on the benchmarks where extracting values from the figure is the hard part (DynaMath +10.6, LogicVista +5.8, MathVerse +5.2).

InternVL3-2B captioner

Same recipe with a different backbone: +3.3 points on average over the pretrained InternVL3-2B paired with the same reasoner.

RQ2 Is it the clarification signal, or just a denser reward? The clarification signal. A denser reward that ignores clarification recovers only part of the gain, and controls that keep the reward density but cut its link to clarification lose most of it.

A partial reward makes the learning signal denser, and denser signals help RL on their own. We tested whether that is the whole story in two ways.

A different way to densify

Binary reward plus image–caption similarity from Qwen3-VL-Embedding-2B (\(\lambda=0.4\)), same data and optimizer. Paper, App. J.

Similarity helps on 4 of 7 benchmarks but trails AC-RL on all 7. On MathVerse it falls below plain binary RL: rewarding whole-image fidelity pulls the captioner away from the few values a sparse diagram question hinges on.

Break one link at a time

Each run changes exactly one component of AC-RL. Accuracy with the 32B reasoner.

An uninformative clarifier keeps the reward rule and the reasoner's ability to ask, and removes only the content of the answer: LogicVista falls back to the untrained level. Permuting the reward values among the successful rollouts in each group keeps every advantage magnitude identical and removes only which rollout needed help: −5.6 on LogicVista, −2.0 on MathVision, both significant.

The two reward structures also improve different problems. Across five benchmarks (9,035 problems), 49% of AC-RL's improvements over the pretrained captioner are on problems that Binary-RL still gets wrong. On MMMU's non-math subjects, AC-RL gains 1.75 points where Binary-RL loses 2.06.

RQ3 Does the captioner actually learn to anticipate the reasoner's questions? Yes. Its captions trigger far fewer questions than a binary-reward captioner's, and once trained, letting the reasoner ask barely changes accuracy: the answers are already in the caption.

What happens if you let the reasoner ask?

We re-run evaluation with clarification switched on, only to measure what the captions were missing. Both captioners were trained against the same 32B reasoner; they differ only in the reward (paper, Table 3).

single pass, as deployedone clarification allowed

More rounds of questions

MathVision (500-problem subset). Up to \(R\) rounds of clarification at evaluation time.

AC-RL captionerPretrained captioner

The pretrained captioner gains up to 5.1 points from extra rounds (significant at \(R=2\) and \(R=3\)). The AC-RL captioner is flat at every budget, even though it still asks on 39% of problems at \(R=1\): what those questions would fetch is already in its caption.

When it does ask, the question matters

Single-pass accuracy on all problems, and on only the problems where the reasoner asked for clarification (paper, Table 3).

all problemsproblems where it asked

RQ4 Is the captioner just solving the problem for the reasoner? No. The reward only checks the final answer, so this was a real risk. AC-RL captions state answers less often than the pretrained model's, contain fewer computed values, and transcribe visible values more accurately.

We split every caption into sentences and labeled each one: description of what is visible, or answer-like content, meaning a stated answer, a solution step, or a directive to the reasoner. Transcribing a value that is printed in the image counts as description. The labels come from the frozen 32B reasoner following a written rubric, validated on sentences with known labels (recall 0.98, false-positive rate 0.008) and against a blind human check of 100 captions (Cohen's \(\kappa = 0.87\)).

Annotated captions

MathVerse, vision-only version: the question is printed in the image. Same problem, both captioners, every sentence labeled.

description states the answer solution step directive to the reasoner

Across all 783 captions per model

Numeric claims were additionally checked against the image by Qwen2.5-VL-72B.

Pretrained captionerAC-RL captioner
27.1 vs 20.9
accuracy after deleting every answer-like sentence from both models' captions (AC-RL vs pretrained)
+9.0
AC-RL lead on the 100 problems where both captioners wrote fully clean captions
−39.7%
caption length relative to the pretrained captioner

The untrained model is the one that tends to solve the problem inside its caption. Deleting its answer-like sentences costs it 10.5 points (7.0 for deleting the same number of random sentences); for AC-RL the cost is 4.7 (2.6 random). If AC-RL's gains were carried by leaked answers, its lead would shrink after filtering; it grows. Residual cases do exist, and they are three times more common from the pretrained model (195 captions vs 65 with two or more answer-like sentences).

RQ5 Do you need an open reasoner? Does the captioner only work with the one it was trained for? Neither. Training entirely through the Gemini API works (+5.2 average, all 7 benchmarks significant), and a captioner trained for one reasoner still helps reasoners it has never seen.

AC-RL needs only the reasoner's text output, so the reasoner can sit behind an API. We swapped the open 32B model for Gemini 3.1 Flash-Lite during both training and evaluation, keeping the clarifier local. The per-benchmark pattern is the same as with the open reasoner, and WeMath turns from a small loss into a gain.

Gain over the pretrained captioner, with 95% confidence intervals

Paired bootstrap over test problems, 10,000 resamples (paper, Tables 5 and 6). Qwen2.5-VL-3B captioner.

DeepSeek-R1-Distill-Qwen-32B (open)Gemini 3.1 Flash-Lite (API only)
~25,000
Gemini calls over 480 training steps; 23% of them on the clarification path
~$15
total API cost for the training run (about 9 hours on 8 GPUs)
\(42\% \to 24\%\)
clarification rate from early to late training

Reasoners never seen in training

The captioner was trained against the 32B reasoner only, then paired at test time with a smaller open reasoner and with GPT-5 mini (paper, App. K). Pick a reasoner to highlight it; hover a row for the numbers.

Pretrained captionerAC-RL captioner
RQ6 Where does it help, and where doesn't it? It helps where the reasoner is missing a visual fact: measurements, table values, spatial structure. AC-RL optimizes what the reasoner sees, not what it knows.

Change in accuracy by subject

AC-RL vs. the pretrained captioner, both with the 32B reasoner (paper, Fig. 2).

The WeMath exception

Scale the captioner from 3B to 32B without any RL, keeping the reasoner fixed.

WeMath problems are organized around textbook concepts rather than visual extraction. A ten-times larger captioner barely moves it, so better captions of any kind have little to offer there. AC-RL and Binary-RL shift it by the same small amount (−1.5 vs −1.3, not significant), and with the Gemini reasoner AC-RL improves it by 2.8.

Beyond math

Same captioners and 32B reasoner, benchmarks never used in training.

WorldMedQA-V: all 568 English medical-licensing questions with clinical images. The pattern holds outside mathematics when extraction is the bottleneck. On perception-only questions, where a text bottleneck costs more than the reasoner adds, we do not expect gains.

Examples

Captions from the pretrained and AC-RL captioners on the same benchmark problem, with the reasoner's verdict for each. Switch to Be the reasoner to answer from the captions alone, the way the reasoner has to.

Main results

Accuracy (%) on seven benchmarks, single-pass evaluation. Click a column to sort within each group.

MathVista testmini, MathVision full, MathVerse vision-only, MMMU dev+val, WeMath testmini (strict), DynaMath worst-case, LogicVista. Other models' numbers are from their reports; † testmini/mini subset; * average over available benchmarks only. Our rows pair a 2B or 3B captioner with a frozen reasoner (DeepSeek-R1-Distill-Qwen-32B or Gemini 3.1 Flash-Lite).

BibTeX

The BibTeX entry will be added once the NeurIPS 2026 proceedings are published.