NeurIPS 2026 · Poster

Learning to Surpass: Training Tool-Using Agents with Anchored Feedback

John Gkountouras1,2,   Fengjun Wang2,   Angelantonio Castelli2,   Satendra Kumar2

1ILLC, University of Amsterdam   2Booking.com

Optimization Loop Trainable Frozen LLM Judge Tools and responses User Query Agent Policy ŷ At inference the policy only calls tools and writes the plan: no gate, no judge, no reference plan. Constraint gate Reward calculation Anchored GRPO (ours) Model Plan Ref Plan J(x, ŷ, ỹ) Win (1) | Tie (0.5) | Lose (0) Score GRPO (baseline) Model Plan J(x, ŷ) Score (0-1) Fail R = 0 GRPO Optimizer Optimization Loop User Query Agent Policy ŷ At inference the policy only calls tools and writes the plan: no gate, no judge, no reference plan. Constraint gate Reward calculation Anchored GRPO (ours) Model Plan Ref Plan J(x, ŷ, ỹ) Win (1) | Tie (0.5) | Lose (0) Score GRPO (baseline) Model Plan J(x, ŷ) Score (0-1) Shown for comparison; not used by our method. Fail R = 0 GRPO Optimizer Trainable Frozen LLM Judge Tools and responses
Anchored GRPO training overview (redrawn from the paper's Figure 1). The optimization loop for a single rollout: the agent policy, the only trainable part, calls tools and emits a plan \(\hat{y}\), which passes through a constraint gate before any reward is computed. A plan that fails a constraint takes the fail path and receives \(R=0\), and the judge is never asked. The reward panels show the two formulations; training uses one of them, never both. Anchored GRPO (ours) has a frozen LLM judge compare the model's plan with the reference plan \(\tilde{y}\), \(J(x,\hat{y},\tilde{y})\), giving 1 for a win, 0.5 for a tie and 0 for a loss. Score GRPO, a baseline we compare against, scores the plan alone, \(J(x,\hat{y})\), which can be exploited through superficial artifacts. GRPO turns the rewards of \(K=8\) rollouts of the same request into advantages.

TL;DR  Tool-using agents are commonly trained by imitating tool traces, and when only finished plans exist, those traces have to be reconstructed first. We skip the reconstruction. The agent finds its own tool calls. Its plan earns reward only if it passes every programmatic constraint, and then according to whether an LLM judge prefers it to the human-written plan for the same request: 1 for a win, 0.5 for a tie. On TripTailor, a 4B model trained this way reaches 80.1% Final Surpass, against 43.7% for fine-tuning on reconstructed traces.

80.1%
Final Surpass on TripTailor: valid and preferred to the human plan. Imitation of imputed traces: 43.7%
61.6%
LLM-judge win rate against the human plan, statistically tied with Gemini-3.0-Pro (59.6%) and ahead of every other system
38.4%
of plans valid and preferred by both the LLM judges and a reward model never used in training; next best method: 30.3%
+18.3
points of judge wins from human anchors over synthetic ones from Gemini-3.0-Pro, same method otherwise (\(p<0.001\))

Abstract

Training tool-using agents for planning is challenging when high-quality demonstrations exist only as final outcomes without the intermediate tool traces that produced them. A common response is to impute traces and apply imitation learning, but this approach is indirect and structurally limited to matching the reference rather than improving upon it. We propose anchored comparative feedback: rather than imitating a reference plan, we train agents to surpass it under a comparative LLM judge while satisfying hard constraints. The anchor provides a stable optimization target, and comparing two identically rendered plans neutralizes surface-level judge biases. On TripTailor, a tool-augmented travel planning benchmark, anchored GRPO achieves 80.1% final success rate compared to 43.7% for imitation learning on imputed traces, with gains on both constraint satisfaction and LLM-judged preference that persist across judge families, rubrics, and a held-out reward model. Gains generalize beyond TripTailor, with a fivefold improvement on the held-out TravelPlanner benchmark. These results suggest that anchored comparative feedback offers an effective approach for learning transferable tool-using planners from outcome-only supervision.

An outcome without its trace

TripTailor gives, for each travel request, an itinerary written by a person. It does not record the searches that person made. A model can only imitate that itinerary if someone first invents the tool calls that would have produced it, and imitation then teaches the model to reproduce the reference rather than improve on it. Meanwhile the untrained agent already has a different problem: many of its plans break rules that a program can check.

This is the common failure. Before training, 53.3% of the agent's test plans break at least one constraint; 33.1% give at least one attraction a visit time outside its recommended range and 22.3% repeat a restaurant (PDF §4.4 and Table 12). Some requirements are programmatic like these, others are matters of judgment, such as whether the attractions fit what the traveller asked for. The method treats the two differently.

Method

For a request \(x\), the policy \(\pi_\theta\) interleaves tool calls with their results, \(\tau=(a_{1:T},o_{1:T})\), and ends with a JSON plan \(\hat{y}\). We have a reference plan \(\tilde{y}\) for the same request but not the trace behind it. Instead of asking a judge whether the plan is good, we ask whether it is better than the reference:

\[ s = J(x, \hat{y}, \tilde{y}) \in \{\text{loss},\, \text{tie},\, \text{win}\}. \]

The judge scores both plans from 1 to 5 on the same rubric; the plan wins if its score is higher. Constraints that a program can verify (entities exist in the sandbox, required fields present, total cost within budget, no repeated restaurant or attraction, meal prices in the requested tier, visit durations within the recommended range) are not traded against the judge. They gate the reward:

\[ R(\tau) = \mathcal{C}(\tau)\cdot r_{\text{pref}}(s), \qquad r_{\text{pref}}(s)=\begin{cases}1 & \text{win}\\ 0.5 & \text{tie}\\ 0 & \text{loss}\end{cases} \]

with \(\mathcal{C}(\tau)\in\{0,1\}\) equal to 1 only if all seven constraints pass. GRPO samples \(K\) trajectories per request and normalizes their rewards within the group,

\[ A^{(k)} = \frac{R^{(k)}-\bar{R}}{\sigma_R+\epsilon}, \]

then applies a clipped policy-gradient update to the tokens the model emitted (tool calls and the plan), never to tool responses.

The judge is used differently in training and in evaluation, so that the policy cannot improve its score by exploiting one judge; see How the judge is used below.

Qwen3-4B-Instruct with LoRA, \(K=8\), asymmetric clipping \(\epsilon_{\text{low}}/\epsilon_{\text{high}} = 0.2 / 0.25\), no KL term; 3,145 training requests, human reference plans converted to JSON. No judge and no reference are used at inference. The method needs reference outcomes and programmatically checkable constraints for the target task; where either is missing it does not apply.

Where does the gradient go?

A group of \(K=8\) rollouts for one request. Set each rollout's constraint check and judge outcome; the reward and the advantage follow from the equations in this section.

How the judge is used

In training

  • One judge, Gemini-2.5-Flash. For each rollout that passes the constraint gate, it scores the rollout's plan and the reference plan from 1 to 5 on the TripTailor rubric. A higher score for the rollout is a win (reward 1), an equal score a tie (0.5), a lower one a loss (0).
  • Order swapped at random. Which plan the judge sees first is randomized for every sample, so the policy cannot learn to benefit from a position.

In evaluation

  • Two other judge families, GPT-4o and Claude-4.5-Haiku, neither of them the training judge. Each compares the two plans twice, once in each order.
  • Four judgments, averaged. Each plan's four scores are averaged; the plan counts as a win only if its average is strictly higher than the reference's. Ties are not wins.
  • A second, different evaluator: a reward model trained on TripTailor preferences that never enters training.

Why. LLM judges are known to favour whichever answer sits in a given position (Wang et al., 2024; Shi et al., 2025; Thakur et al., 2025), to favour outputs from their own model family (Zheng et al., 2023; Ye et al., 2025), and to favour longer or more formatted text (Park et al., 2024; Huang et al., 2025; Chen et al., 2024). Each choice above removes one of these as something the policy could learn to exploit. Randomizing the order during training keeps position-dependent tricks from paying off, and averaging both orders at evaluation cancels the position effect in the measurement. Evaluating with judges from other providers than the training judge means a gain has to hold for evaluators the policy was never optimized against. Both plans are emitted as JSON with a fixed schema and rendered to text by one deterministic template, with chain-of-thought disabled, so length and formatting are nearly identical within every comparison.

PDF §3.5 and §4.1. Evaluator robustness beyond these two judges is in RQ3.

  • P. Wang et al. Large Language Models are not Fair Evaluators. ACL 2024. arXiv:2305.17926
  • L. Shi et al. Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge. 2025. arXiv:2406.07791
  • A. S. Thakur et al. Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges. GEM Workshop 2025. arXiv:2406.12624
  • L. Zheng et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023. arXiv:2306.05685
  • J. Ye et al. Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge. ICLR 2025. arXiv:2410.02736
  • J. Park et al. OffsetBias: Leveraging Debiased Data for Tuning Evaluators. Findings of EMNLP 2024. arXiv:2407.06551
  • H. Huang et al. An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4. Findings of ACL 2025. arXiv:2403.02839
  • G. H. Chen et al. Humans or LLMs as the Judge? A Study on Judgement Bias. EMNLP 2024. arXiv:2402.10669

One comparison, four judgments

Findings

Seven questions, most of them raised in review, each with its answer up front. Open a question for the evidence. All numbers are on the 703-request TripTailor test set unless stated.

RQ1 Does training to surpass the reference beat imitating reconstructed traces? Yes. Final Surpass rises to 80.1%, against 43.7% for SFT on imputed traces, and imitation even lowers the judge win rate below the untrained model's.

The SFT baseline gets the most favourable reconstruction available: its traces are built from the reference plans themselves and lead to them by construction. It still reaches only 43.7%. Its judge win rate even drops, from 45.4% to 39.5%. Anchored GRPO improves on both axes at once: rationality from 47.5% to 94.5% and judge wins from 45.4% to 61.6%. It retains 87.1% of the untrained model's successes and converts 76.5% of its failures, a net gain of +323 requests.

Final Surpass = Valid \(\wedge\) (LLM win \(\vee\) RM win), the TripTailor definition. The Workflow row is the best pipeline from the TripTailor paper, run with GPT-4o-mini; Gemini-3.0-Pro runs the same workflow and also produced the synthetic anchors. PDF Table 1.

RQ2 Is it just the constraint gate? No. The gate buys validity; the anchor buys judged quality. Scoring plans without a reference (Score GRPO) reaches 97.9% rationality but wins only 38.5% of judge comparisons.

Each point is one way of using the same supervision. Methods trained with the gate move right, towards valid plans. Height is judged quality, and there the reward matters: the gate with no judge at all (constraint-only GRPO) reaches 55.5% judge wins, a pointwise judge without a reference (Score GRPO) drops to 38.5%, and comparison against the human plan reaches 61.6%. Anchored DPO uses the same comparison against the reference but learns from pairs of final plans, so the tool calls that produced a plan get no signal and the gate cannot shape them: its plans pass the visit-duration check 74.0% of the time, against 97.9% for Anchored GRPO. Sampling more and selecting afterwards also falls short. Best-of-\(N\) with the constraint filter reaches 55.2, 60.3 and 64.6% validity at \(N\) = 2, 4 and 8, at \(N\) times the inference cost, and at \(N=8\) no selector could do better: only 64.6% of requests have any valid candidate in eight draws. An LLM reranker over the same candidates loses validity as \(N\) grows (48.2, 47.5, 46.1%), since it cannot see constraint violations in the text.

Valid = all seven constraints pass. LLM win against the human reference, same judges and protocol for every point. Best-of-\(N\) reranking is shown at its best \(N\).

RQ3 Is the policy gaming the judge? We find no sign of it: it ranks first under all three judge settings, the paper's GPT-4o and Claude-4.5-Haiku pair and two open-weights judges it never saw. The held-out reward model is less decisive and rates it level with the other GRPO variants.

The training judge is Gemini-2.5-Flash; none of the evaluators below were used in training. Absolute win rates depend on how strict a judge is, and the open-weights judges are stricter with every system, but the ranking of Anchored GRPO does not change. The reward model, a different kind of evaluator, separates the GRPO variants from the rest but not from each other. The strictest automatic view counts a plan only if it is valid and both the LLM judges and the reward model prefer it to the reference: 38.4% of Anchored GRPO's plans qualify, against 26.5% for the best system in the chart and 30.3% for the best of all methods we compared (constraint-only GRPO, RQ2). On the paper's judges, the difference to Gemini-3.0-Pro is not significant (Gemini-3.0-Pro minus Anchored GRPO: −2.0 pp [−6.5, +2.4], paired bootstrap).

RQ4 Is it just more tool calls? No. With 7 to 8 tool calls, 89.9% of trained plans are valid against 46.6% of untrained ones, and the untrained model does not get more valid plans by calling more.

The trained agent does search more: 8.73 calls per plan instead of 5.78, most of the increase in train, attraction and restaurant searches, and 97.0% of its trajectories settle into the same four-turn pattern. But at the same budget the gap remains: among plans with 7 to 8 calls, the only bin where both models have many plans, it is more than 40 points. The per-plan cost grows by about a quarter in tokens, and inference needs neither the judge nor the reference.

Valid plans by number of tool calls

Share of plans passing all constraints; \(n\) = plans in the bin.

Cost per plan

Averages over the test set.

Calls by tool

Total calls over the 703 test requests.

RQ5 Does the anchor have to be written by a person? No, but human anchors teach more of what judges reward. Synthetic anchors from Gemini-3.0-Pro give equally valid plans and 43.2% judge wins instead of 61.6%.

Syn-Anchored GRPO is the same method with the human references replaced by plans that Gemini-3.0-Pro generated for the training requests. No human plans are needed, and constraint satisfaction is about as good: synthetic anchors are slightly better on feasibility (1.7 points, \(p<0.05\)) and not significantly different on rationality. What it loses is judged quality: +18.3 points of LLM wins (\(p<0.001\), paired bootstrap), and +5.7 points of Final Surpass (95% CI [2.0, 9.4], \(p<0.01\)).

RQ6 Does it transfer, and where does it not help? It transfers to other tool-planning benchmarks and does nothing for scheduling puzzles without tools. Absolute levels on TravelPlanner stay low.

None of these benchmarks were used in training. On TravelPlanner, final pass rises from 1.1% to 5.6%; the task is hard at this scale (Qwen3-32B reaches 0.6% zero-shot). On TaskBench, Anchored GRPO is the only trained variant that improves how tool calls are chained (edge F1 +1.66 ± 0.77 points), and on UltraTool it improves all three metrics. Imitation fails differently: SFT on imputed traces keeps choosing plausible tools but writes TripTailor-style arguments, and its argument scores collapse (TaskBench argument F1 2.07, UltraTool argument score 16.12). NaturalPlan asks for constraint reasoning over a given context with no tools; no GRPO variant moves it by more than 1.5 points, and SFT drops on calendar scheduling. On \(\tau^2\)-bench, a multi-turn customer-service task, the gain is a modest 4.1 points and the standard-error bars of the two runs overlap.

RQ7 Does the recipe carry over to other base models? On the two we tried, yes: Final Surpass goes from 21.8% to 37.8% on Qwen3-8B and from 42.4% to 81.7% on Gemma 4.

Same reward, judge protocol and schedule; only the base model changes. Qwen3-8B is an older release than the Qwen3-4B-2507 model used everywhere else and starts lower; its numbers are from a preliminary checkpoint. Gemma 4 starts with a much higher judge win rate and the same validity as Qwen3-4B, so validity is what limits it before training and what training improves most; its run had not plateaued. Under the strict both-evaluators metric, Gemma 4 goes from 26.3% to 56.8%.

Examples

Real test requests with the plans of the untrained agent and of Anchored GRPO, verbatim from the evaluation files behind Table 1, with the per-request constraint checks and judge outcomes from the same files. The itinerary view renders the JSON the model emitted; Raw JSON shows it unchanged.

Main results

TripTailor test set, 703 requests (PDF Table 1). Click a column to sort within each group.

Route = average distance between locations relative to the human plan (lower is shorter; not optimized in training). Micro = share of sub-constraints satisfied; Macro = share of plans satisfying all of them. LLM and RM = win rate against the human reference under the LLM judges (GPT-4o and Claude-4.5-Haiku, both presentation orders) and under the reward model. Final Surpass = Valid \(\wedge\) (LLM \(\vee\) RM). Bold = best in column (Route is not ranked, as in the paper). † From the TripTailor paper.

BibTeX

The BibTeX entry will be added once the NeurIPS 2026 proceedings are published.