NeurIPS 2026 · Poster
1ILLC, University of Amsterdam 2Booking.com
Training tool-using agents for planning is challenging when high-quality demonstrations exist only as final outcomes without the intermediate tool traces that produced them. A common response is to impute traces and apply imitation learning, but this approach is indirect and structurally limited to matching the reference rather than improving upon it. We propose anchored comparative feedback: rather than imitating a reference plan, we train agents to surpass it under a comparative LLM judge while satisfying hard constraints. The anchor provides a stable optimization target, and comparing two identically rendered plans neutralizes surface-level judge biases. On TripTailor, a tool-augmented travel planning benchmark, anchored GRPO achieves 80.1% final success rate compared to 43.7% for imitation learning on imputed traces, with gains on both constraint satisfaction and LLM-judged preference that persist across judge families, rubrics, and a held-out reward model. Gains generalize beyond TripTailor, with a fivefold improvement on the held-out TravelPlanner benchmark. These results suggest that anchored comparative feedback offers an effective approach for learning transferable tool-using planners from outcome-only supervision.
TripTailor gives, for each travel request, an itinerary written by a person. It does not record the searches that person made. A model can only imitate that itinerary if someone first invents the tool calls that would have produced it, and imitation then teaches the model to reproduce the reference rather than improve on it. Meanwhile the untrained agent already has a different problem: many of its plans break rules that a program can check.
This is the common failure. Before training, 53.3% of the agent's test plans break at least one constraint; 33.1% give at least one attraction a visit time outside its recommended range and 22.3% repeat a restaurant (PDF §4.4 and Table 12). Some requirements are programmatic like these, others are matters of judgment, such as whether the attractions fit what the traveller asked for. The method treats the two differently.
For a request \(x\), the policy \(\pi_\theta\) interleaves tool calls with their results, \(\tau=(a_{1:T},o_{1:T})\), and ends with a JSON plan \(\hat{y}\). We have a reference plan \(\tilde{y}\) for the same request but not the trace behind it. Instead of asking a judge whether the plan is good, we ask whether it is better than the reference:
The judge scores both plans from 1 to 5 on the same rubric; the plan wins if its score is higher. Constraints that a program can verify (entities exist in the sandbox, required fields present, total cost within budget, no repeated restaurant or attraction, meal prices in the requested tier, visit durations within the recommended range) are not traded against the judge. They gate the reward:
with \(\mathcal{C}(\tau)\in\{0,1\}\) equal to 1 only if all seven constraints pass. GRPO samples \(K\) trajectories per request and normalizes their rewards within the group,
then applies a clipped policy-gradient update to the tokens the model emitted (tool calls and the plan), never to tool responses.
The judge is used differently in training and in evaluation, so that the policy cannot improve its score by exploiting one judge; see How the judge is used below.
Qwen3-4B-Instruct with LoRA, \(K=8\), asymmetric clipping \(\epsilon_{\text{low}}/\epsilon_{\text{high}} = 0.2 / 0.25\), no KL term; 3,145 training requests, human reference plans converted to JSON. No judge and no reference are used at inference. The method needs reference outcomes and programmatically checkable constraints for the target task; where either is missing it does not apply.
Why. LLM judges are known to favour whichever answer sits in a given position (Wang et al., 2024; Shi et al., 2025; Thakur et al., 2025), to favour outputs from their own model family (Zheng et al., 2023; Ye et al., 2025), and to favour longer or more formatted text (Park et al., 2024; Huang et al., 2025; Chen et al., 2024). Each choice above removes one of these as something the policy could learn to exploit. Randomizing the order during training keeps position-dependent tricks from paying off, and averaging both orders at evaluation cancels the position effect in the measurement. Evaluating with judges from other providers than the training judge means a gain has to hold for evaluators the policy was never optimized against. Both plans are emitted as JSON with a fixed schema and rendered to text by one deterministic template, with chain-of-thought disabled, so length and formatting are nearly identical within every comparison.
PDF §3.5 and §4.1. Evaluator robustness beyond these two judges is in RQ3.
Seven questions, most of them raised in review, each with its answer up front. Open a question for the evidence. All numbers are on the 703-request TripTailor test set unless stated.
The SFT baseline gets the most favourable reconstruction available: its traces are built from the reference plans themselves and lead to them by construction. It still reaches only 43.7%. Its judge win rate even drops, from 45.4% to 39.5%. Anchored GRPO improves on both axes at once: rationality from 47.5% to 94.5% and judge wins from 45.4% to 61.6%. It retains 87.1% of the untrained model's successes and converts 76.5% of its failures, a net gain of +323 requests.
Final Surpass = Valid \(\wedge\) (LLM win \(\vee\) RM win), the TripTailor definition. The Workflow row is the best pipeline from the TripTailor paper, run with GPT-4o-mini; Gemini-3.0-Pro runs the same workflow and also produced the synthetic anchors. PDF Table 1.
Each point is one way of using the same supervision. Methods trained with the gate move right, towards valid plans. Height is judged quality, and there the reward matters: the gate with no judge at all (constraint-only GRPO) reaches 55.5% judge wins, a pointwise judge without a reference (Score GRPO) drops to 38.5%, and comparison against the human plan reaches 61.6%. Anchored DPO uses the same comparison against the reference but learns from pairs of final plans, so the tool calls that produced a plan get no signal and the gate cannot shape them: its plans pass the visit-duration check 74.0% of the time, against 97.9% for Anchored GRPO. Sampling more and selecting afterwards also falls short. Best-of-\(N\) with the constraint filter reaches 55.2, 60.3 and 64.6% validity at \(N\) = 2, 4 and 8, at \(N\) times the inference cost, and at \(N=8\) no selector could do better: only 64.6% of requests have any valid candidate in eight draws. An LLM reranker over the same candidates loses validity as \(N\) grows (48.2, 47.5, 46.1%), since it cannot see constraint violations in the text.
Valid = all seven constraints pass. LLM win against the human reference, same judges and protocol for every point. Best-of-\(N\) reranking is shown at its best \(N\).
The training judge is Gemini-2.5-Flash; none of the evaluators below were used in training. Absolute win rates depend on how strict a judge is, and the open-weights judges are stricter with every system, but the ranking of Anchored GRPO does not change. The reward model, a different kind of evaluator, separates the GRPO variants from the rest but not from each other. The strictest automatic view counts a plan only if it is valid and both the LLM judges and the reward model prefer it to the reference: 38.4% of Anchored GRPO's plans qualify, against 26.5% for the best system in the chart and 30.3% for the best of all methods we compared (constraint-only GRPO, RQ2). On the paper's judges, the difference to Gemini-3.0-Pro is not significant (Gemini-3.0-Pro minus Anchored GRPO: −2.0 pp [−6.5, +2.4], paired bootstrap).
The trained agent does search more: 8.73 calls per plan instead of 5.78, most of the increase in train, attraction and restaurant searches, and 97.0% of its trajectories settle into the same four-turn pattern. But at the same budget the gap remains: among plans with 7 to 8 calls, the only bin where both models have many plans, it is more than 40 points. The per-plan cost grows by about a quarter in tokens, and inference needs neither the judge nor the reference.
Share of plans passing all constraints; \(n\) = plans in the bin.
Averages over the test set.
Total calls over the 703 test requests.
Syn-Anchored GRPO is the same method with the human references replaced by plans that Gemini-3.0-Pro generated for the training requests. No human plans are needed, and constraint satisfaction is about as good: synthetic anchors are slightly better on feasibility (1.7 points, \(p<0.05\)) and not significantly different on rationality. What it loses is judged quality: +18.3 points of LLM wins (\(p<0.001\), paired bootstrap), and +5.7 points of Final Surpass (95% CI [2.0, 9.4], \(p<0.01\)).
None of these benchmarks were used in training. On TravelPlanner, final pass rises from 1.1% to 5.6%; the task is hard at this scale (Qwen3-32B reaches 0.6% zero-shot). On TaskBench, Anchored GRPO is the only trained variant that improves how tool calls are chained (edge F1 +1.66 ± 0.77 points), and on UltraTool it improves all three metrics. Imitation fails differently: SFT on imputed traces keeps choosing plausible tools but writes TripTailor-style arguments, and its argument scores collapse (TaskBench argument F1 2.07, UltraTool argument score 16.12). NaturalPlan asks for constraint reasoning over a given context with no tools; no GRPO variant moves it by more than 1.5 points, and SFT drops on calendar scheduling. On \(\tau^2\)-bench, a multi-turn customer-service task, the gain is a modest 4.1 points and the standard-error bars of the two runs overlap.
Same reward, judge protocol and schedule; only the base model changes. Qwen3-8B is an older release than the Qwen3-4B-2507 model used everywhere else and starts lower; its numbers are from a preliminary checkpoint. Gemma 4 starts with a much higher judge win rate and the same validity as Qwen3-4B, so validity is what limits it before training and what training improves most; its run had not plateaued. Under the strict both-evaluators metric, Gemma 4 goes from 26.3% to 56.8%.
Real test requests with the plans of the untrained agent and of Anchored GRPO, verbatim from the evaluation files behind Table 1, with the per-request constraint checks and judge outcomes from the same files. The itinerary view renders the JSON the model emitted; Raw JSON shows it unchanged.
TripTailor test set, 703 requests (PDF Table 1). Click a column to sort within each group.
Route = average distance between locations relative to the human plan (lower is shorter; not optimized in training). Micro = share of sub-constraints satisfied; Macro = share of plans satisfying all of them. LLM and RM = win rate against the human reference under the LLM judges (GPT-4o and Claude-4.5-Haiku, both presentation orders) and under the reward model. Final Surpass = Valid \(\wedge\) (LLM \(\vee\) RM). Bold = best in column (Route is not ranked, as in the paper). † From the TripTailor paper.
The BibTeX entry will be added once the NeurIPS 2026 proceedings are published.