Portrait of John Gkountouras

John Gkountouras

PhD candidate
Institute for Logic, Language and Computation (ILLC)
University of Amsterdam

I am a final-year PhD candidate at the University of Amsterdam (ILLC), advised by Prof. Ivan Titov and Dr. Wilker Aziz. I work on reinforcement learning and self-distillation for language and vision-language models, in particular on turning signal that exists only at training time (clarification requests, reference outcomes, self-consensus) into supervision. I have been a research intern at Booking.com, working on RL for tool-using agents, and at Stochastic (Harvard Innovation Labs), working on LLM compression.

Interests

  • RL post-training for LLMs and VLMs
  • Self-distillation and label-free learning
  • Tool-using agents
  • World models for agents

Education

  • PhD in Artificial IntelligenceUniversity of Amsterdam, 2023 to present (expected 2028)
  • MSc in Artificial Intelligence (cum laude)University of Amsterdam, 2023
  • BSc in MathematicsUniversity of Ioannina, 2021

News

  • "Clarification as Supervision" accepted as an Oral at NeurIPS 2026. →
  • "Learning to Surpass", from my Booking.com internship, accepted at NeurIPS 2026. →
  • CANON (label-free self-distillation) released on arXiv. →
  • Started a Machine Learning Scientist internship at Booking.com (RL for tool-using agents), through Jan 2026.
  • "Language Agents Meet Causality" accepted at ICLR 2025. →
  • Started my PhD at the University of Amsterdam with Ivan Titov.

Selected publications

  • Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces

    John Gkountouras, Ivan Titov

    NeurIPS 2026Oral

    A frozen reasoner may ask the vision-language model for clarification during training at a cost, so information gaps become dense reward and the model learns to describe images completely for single-pass inference.

    Paper Project
  • Learning to Surpass: Training Tool-Using Agents with Anchored Feedback

    John Gkountouras, Fengjun Wang, Angelantonio Castelli, Satendra Kumar

    NeurIPS 2026

    GRPO for tool-using planning agents with a comparative LLM judge that rewards beating a reference plan, gated by hard constraints.

    Project
  • Consensus as Privileged Context for Label-Free Self-Distillation

    John Gkountouras, Josip Jukić, Ivan Titov

    arXiv preprint, 2026

    Label-free self-distillation in which a frozen snapshot, conditioned on a solution that reaches the majority answer, supervises the model token by token on its own rollouts.

    Paper Project
  • Language Agents Meet Causality: Bridging LLMs and Causal World Models

    John Gkountouras, Matthias Lindemann, Phillip Lippe, Efstratios Gavves, Ivan Titov

    ICLR 2025

    A causal world model with a language interface lets LLM agents simulate state transitions and plan in visual environments.

    Paper Code Project

See all publications →

Experience

  • Booking.com · Machine Learning Scientist Intern, Amsterdam Sep 2025 to Jan 2026 RL post-training of tool-using LLM agents; led to a NeurIPS 2026 paper.
  • Stochastic (Harvard Innovation Labs) · Deep Learning Research Intern, Boston Nov 2022 to Jun 2023 LLM compression; led to INT2.1.
  • University Housing B.V. · Lead Data Scientist and Data Scientist, Utrecht 2020 to 2022

Teaching & service

Teaching

  • Natural Language Processing 1 and 2 · Teaching Assistant, MSc AI, University of Amsterdam 2024 to 2026
  • Foundation Models (6 MSc students) · Research project supervisor, University of Amsterdam 2024 to 2025

Service

  • Reviewer · NeurIPS 2026; ICLR 2026; XAI4CV Workshop at CVPR, 2023 to 2026