John Gkountouras
PhD candidate
Institute for Logic, Language and Computation (ILLC)
University of Amsterdam
I am a final-year PhD candidate at the University of Amsterdam (ILLC), advised by Prof. Ivan Titov and Dr. Wilker Aziz. I work on reinforcement learning and self-distillation for language and vision-language models, in particular on turning signal that exists only at training time (clarification requests, reference outcomes, self-consensus) into supervision. I have been a research intern at Booking.com, working on RL for tool-using agents, and at Stochastic (Harvard Innovation Labs), working on LLM compression.
Interests
- RL post-training for LLMs and VLMs
- Self-distillation and label-free learning
- Tool-using agents
- World models for agents
Education
- PhD in Artificial IntelligenceUniversity of Amsterdam, 2023 to present (expected 2028)
- MSc in Artificial Intelligence (cum laude)University of Amsterdam, 2023
- BSc in MathematicsUniversity of Ioannina, 2021
News
- "Clarification as Supervision" accepted as an Oral at NeurIPS 2026. →
- "Learning to Surpass", from my Booking.com internship, accepted at NeurIPS 2026. →
- CANON (label-free self-distillation) released on arXiv. →
- Started a Machine Learning Scientist internship at Booking.com (RL for tool-using agents), through Jan 2026.
- "Language Agents Meet Causality" accepted at ICLR 2025. →
- Started my PhD at the University of Amsterdam with Ivan Titov.
Selected publications
-
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
NeurIPS 2026Oral
A frozen reasoner may ask the vision-language model for clarification during training at a cost, so information gaps become dense reward and the model learns to describe images completely for single-pass inference.
The BibTeX entry will be added once the NeurIPS 2026 proceedings are published.
-
Learning to Surpass: Training Tool-Using Agents with Anchored Feedback
NeurIPS 2026
GRPO for tool-using planning agents with a comparative LLM judge that rewards beating a reference plan, gated by hard constraints.
The BibTeX entry will be added once the NeurIPS 2026 proceedings are published.
-
Consensus as Privileged Context for Label-Free Self-Distillation
arXiv preprint, 2026
Label-free self-distillation in which a frozen snapshot, conditioned on a solution that reaches the majority answer, supervises the model token by token on its own rollouts.
@article{gkountouras2026canon, title = {{Consensus as Privileged Context for Label-Free Self-Distillation}}, author = {Gkountouras, John and Juki{\'c}, Josip and Titov, Ivan}, journal = {arXiv preprint arXiv:2607.13643}, year = {2026}, eprint = {2607.13643}, archivePrefix = {arXiv}, url = {https://arxiv.org/abs/2607.13643} } -
Language Agents Meet Causality: Bridging LLMs and Causal World Models
ICLR 2025
A causal world model with a language interface lets LLM agents simulate state transitions and plan in visual environments.
@inproceedings{gkountouras2025language, title = {{Language Agents Meet Causality: Bridging LLMs and Causal World Models}}, author = {Gkountouras, John and Lindemann, Matthias and Lippe, Phillip and Gavves, Efstratios and Titov, Ivan}, booktitle = {International Conference on Learning Representations (ICLR)}, year = {2025}, eprint = {2410.19923}, archivePrefix = {arXiv}, url = {https://arxiv.org/abs/2410.19923} }
Experience
- Booking.com · Machine Learning Scientist Intern, Amsterdam Sep 2025 to Jan 2026 RL post-training of tool-using LLM agents; led to a NeurIPS 2026 paper.
- Stochastic (Harvard Innovation Labs) · Deep Learning Research Intern, Boston Nov 2022 to Jun 2023 LLM compression; led to INT2.1.
- University Housing B.V. · Lead Data Scientist and Data Scientist, Utrecht 2020 to 2022
Teaching & service
Teaching
- Natural Language Processing 1 and 2 · Teaching Assistant, MSc AI, University of Amsterdam 2024 to 2026
- Foundation Models (6 MSc students) · Research project supervisor, University of Amsterdam 2024 to 2025
Service
- Reviewer · NeurIPS 2026; ICLR 2026; XAI4CV Workshop at CVPR, 2023 to 2026