Research Scientist - Datadog - New York

Shridhar

Models that reason, learn, and scale.

I work on LLM post-training, asynchronous reinforcement learning, and long-horizon agents at Datadog.

My experience spans pre-training LLMs to building agentic systems. I contributed to pre-training Apertus, the Swiss AI Initiative’s open foundation models. During my PhD at ETH Zurich, I worked on agentic pipelines spanning planning, reasoning, distillation, and refinement. Previously, I worked at Meta FAIR, Microsoft Research, and Ema AI.

ETH Zurich PhD Meta FAIR Microsoft Research Ema AI Swiss AI Initiative
Shridhar standing on a snowy mountain range at Glacier 3000, Switzerland.
Glacier 3000 - Switzerland
Foundation-model training Apertus 8B & 70B

Contributed to pre-training the open multilingual model family on 15 trillion tokens.

Apertus / ACL 2026
Model systems 4x lower inference cost

Select-then-Route reports 94.3% accuracy versus 91.7% for the best single model across six benchmarks.

Select-then-Route / EMNLP 2025
LLM evaluation Evaluating LLMs

Contributed to collaborative benchmarks spanning expert-level knowledge, diverse language-model capabilities, and execution-based code evaluation.

Research Story

From model reasoning to agent training.

My research connects three questions: how do models solve complex problems, how can they learn from feedback, and how do we train and evaluate them at scale?

During my PhD at ETH Zurich (2021-2025), I studied how language models break down problems, learn reasoning from stronger teachers, and judge their own answers. That work led to Socratic Subquestions, reasoning distillation, and SIKeD, which combines teacher demonstrations with a student's own on-policy generations.

At Microsoft Research, SCREWS explored how to generate, revise, and select answers. At Meta FAIR, ART studied when to trust a revision and when to keep the original answer. Later, at Ema AI, Select-then-Route addressed a related systems question: which model should handle each query to balance accuracy, latency, and inference cost?

As part of the Swiss AI Initiative, I contributed to pre-training Apertus, an open 8B and 70B multilingual model family. This added large-scale foundation-model training to my experience in reasoning and post-training.

Today at Datadog, I work on asynchronous reinforcement learning for long-horizon agents. This brings together the environments agents act in, the rewards and verifiers that assess their behavior, and the training systems that use that feedback to improve the model.

Turning research ideas into trained models.

I design learning methods, implement training experiments, and evaluate what changes in model behavior. Working with researchers and engineers, I use those results to improve the next run, whether that means changing the training data, the reward signal, or the evaluation setup.

  1. 2026-present

    Datadog

    Research Scientist. Asynchronous RL, post-training, and long-horizon agents.

  2. 2025

    Ema AI

    Model-routing research across accuracy, latency, and inference cost.

  3. 2021-2025

    ETH Zurich

    PhD research in reasoning, distillation, refinement, and evaluation.

  4. Research internship

    Meta FAIR

    Selective refinement and trust with ART.

  5. Research internship

    Microsoft Research

    Reasoning with revisions and selection with SCREWS.

  6. Open-model collaboration

    Swiss AI Initiative

    Apertus 8B and 70B, released in 2025; technical report published at ACL 2026.

Selected Publications

The work behind the ideas.

Published as Kumar Shridhar. Selected papers and collaborations; full record on Google Scholar.

Equal contribution

15 works

  1. ACLpeer-reviewed

    Apertus: Democratizing Open and Compliant LLMs for Global Language Environments

    Alejandro Hernández-Cano et al. · includes

    Presents fully open 8B and 70B multilingual models pretrained on 15 trillion tokens, with reproducible artifacts and explicit attention to data compliance.

  2. Naturepeer-reviewed

    A benchmark of expert-level academic questions to assess AI capabilities

    Center for AI Safety, Scale AI, and the HLE Contributors Consortium · Kumar Shridhar is a consortium contributor

    Builds a 2,500-question expert-level multimodal benchmark spanning dozens of subjects to measure frontier capability and calibration.

  3. Findings of ACLpeer-reviewed

    SIKeD: Self-guided Iterative Knowledge Distillation for Mathematical Reasoning

    , , , ,

    Combines teacher data with a smaller model’s own on-policy generations so the student can learn multiple reasoning strategies and select among them task by task.

  4. EMNLP Industrypeer-reviewed

    Select-then-Route: Taxonomy guided Routing for LLMs

    ,

    Narrows candidate models with a learned taxonomy and then routes through a confidence-based cascade, improving accuracy while reducing inference cost across six benchmarks.

  5. AAAIpeer-reviewed

    Calibrating Large Language Models with Sample Consistency

    , , , , , , , ,

    Derives confidence from consistency across sampled generations and evaluates calibration across eleven language models and nine reasoning datasets.

  6. Findings of ACLpeer-reviewed

    First-Step Advantage: Importance of Starting Right in Multi-Step Math Reasoning

    , , ,

    Studies the leverage of the first reasoning step and trains smaller models to ask how a solution should begin before generating the full rationale.

  7. arXivpreprint

    BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution

    Terry Yue Zhuo et al. · includes

    Introduces an execution-backed human evaluation platform for code generation, grounded in more than 14,000 coding conversations and 4,700 multi-turn preferences.

  8. NAACLpeer-reviewed

    The ART of LLM Refinement: Ask, Refine, and Trust

    , , , , , , , ,

    Learns when a language model should revise an answer and whether the revision should be trusted, separating refinement from the decision to accept it.

  9. ICML Workshopsworkshop paper

    Distilling LLMs’ Decomposition Abilities into Compact Language Models

    ,

    Uses offline reinforcement learning and larger-model feedback to teach compact language models how to decompose complex reasoning problems.

  10. arXivpreprint

    SMART: Self-learning Meta-strategy Agent for Reasoning Tasks

    , , , ,

    Models reasoning-strategy selection as a Markov decision process and uses reinforcement learning so a model can internalize which strategy works for each task.

  11. Findings of ACLpeer-reviewed

    Distilling Reasoning Capabilities into Smaller Language Models

    , ,

    Distills Socratic chain-of-thought into cooperating decomposer and solver models, transferring reasoning into substantially smaller language models.

  12. arXivpreprint

    SCREWS: A Modular Framework for Reasoning with Revisions

    , , , , ,

    Separates reasoning with revisions into sampling, conditional resampling, and selection so a system can explore a new path and reject a harmful revision.

  13. TMLRpeer-reviewed

    Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

    Aarohi Srivastava et al. · includes

    Introduces 204 diverse tasks for characterizing scaling behavior, calibration, emerging capabilities, and model limitations.

  14. EMNLPpeer-reviewed

    Automatic Generation of Socratic Subquestions for Teaching Math Word Problems

    , , , , ,

    Uses conditioned generation and reinforcement learning to produce didactically useful Socratic subquestions for mathematical problem solving.

  15. NeurIPSpeer-reviewed

    Learning to Drop Out: An Adversarial Approach to Training Sequence VAEs

    , , , , ,

    An adversarial dropout strategy improves sequence VAE training by addressing posterior collapse and increasing the information captured in latent representations.

Contact

Let's build better models.

For research roles, technical conversations, or collaborations in model training and evaluation, get in touch. A little context about the team and the problem is a great place to start.

Contact Shridhar

Contact Shridhar

Start a conversation.

Open mail app

No data leaves this page until you open your mail app.