Contributed to pre-training the open multilingual model family on 15 trillion tokens.
Apertus / ACL 2026Research Story
From model reasoning to agent training.
My research connects three questions: how do models solve complex problems, how can they learn from feedback, and how do we train and evaluate them at scale?
During my PhD at ETH Zurich (2021-2025), I studied how language models break down problems, learn reasoning from stronger teachers, and judge their own answers. That work led to Socratic Subquestions, reasoning distillation, and SIKeD, which combines teacher demonstrations with a student's own on-policy generations.
At Microsoft Research, SCREWS explored how to generate, revise, and select answers. At Meta FAIR, ART studied when to trust a revision and when to keep the original answer. Later, at Ema AI, Select-then-Route addressed a related systems question: which model should handle each query to balance accuracy, latency, and inference cost?
As part of the Swiss AI Initiative, I contributed to pre-training Apertus, an open 8B and 70B multilingual model family. This added large-scale foundation-model training to my experience in reasoning and post-training.
Today at Datadog, I work on asynchronous reinforcement learning for long-horizon agents. This brings together the environments agents act in, the rewards and verifiers that assess their behavior, and the training systems that use that feedback to improve the model.
Turning research ideas into trained models.
I design learning methods, implement training experiments, and evaluate what changes in model behavior. Working with researchers and engineers, I use those results to improve the next run, whether that means changing the training data, the reward signal, or the evaluation setup.
- 2026-present
Datadog
Research Scientist. Asynchronous RL, post-training, and long-horizon agents.
- 2025
Ema AI
Model-routing research across accuracy, latency, and inference cost.
- 2021-2025
ETH Zurich
PhD research in reasoning, distillation, refinement, and evaluation.
- Research internship
Meta FAIR
Selective refinement and trust with ART.
- Research internship
Microsoft Research
Reasoning with revisions and selection with SCREWS.
- Open-model collaboration
Swiss AI Initiative
Apertus 8B and 70B, released in 2025; technical report published at ACL 2026.