Haibo Ding

Principal Applied Scientist & Manager · LLM Post-Training · Agentic AI · Evaluation

haibo-headshot.png

I am a Principal Applied Scientist and Science Manager at AWS AI Labs. I lead research on LLM post-training and agentic AI, from data and benchmark design through model training and production deployment.

Research interests

  • LLM post-training. Adapting LLMs to domain-specific problems through supervised fine-tuning, reinforcement learning with verifiable rewards, reward modeling, and synthetic data design. Applications include training specialized models for agent evaluation, structured output generation, and learning to ideate.

  • Agent evaluation. Designing benchmarks, datasets, and rubrics for offline and online evaluation, and optimizing tool descriptions from agent trajectories (Automating Agent Evaluation, Agent-EvalKit).

  • Model routing. Developing methods to select the most appropriate LLM for each request while balancing quality and cost (IPR paper).

  • Generative and embodied models. Research on diffusion language models, vision-language-action (VLA) models, and world models for generation, embodied decision-making, and planning (Diffusion Language Model Inference with Monte Carlo Tree Search).

Previously, I was a Senior Research Scientist at Bosch Research. I received my Ph.D. in Computer Science from the University of Utah, where I studied semi-supervised learning for natural language understanding.

News

Mar 24, 2026 Two papers accepted at EACL 2026 on inference for diffusion language models and training agents to ideate.
Dec 03, 2025 Open-sourced Agent-EvalKit — an AI assistant toolkit for build-time agent evaluation.
Dec 02, 2025 Launched Amazon Bedrock AgentCore Evaluations (preview) for agent performance monitoring.
Aug 04, 2025 Organized the KDD Workshop on Automatic Prompt Optimization

More news →

Selected Publications

  1. EACL
    Learning to Ideate for Machine Learning Engineering Agents
    Yunxiang Zhang, Kang Zhou, Zhichao Xu, and 5 more authors
    In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers), Mar 2026
  2. EACL Findings
    Diffusion Language Model Inference with Monte Carlo Tree Search
    Zheng Huang, Kiran Ramnath, Yueyan Chen, and 8 more authors
    In Findings of the Association for Computational Linguistics: EACL 2026, Mar 2026
  3. ArXiv
    Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
    Zhichao Xu, Zongyu Wu, Yun Zhou, and 9 more authors
    ArXiv, Mar 2025
  4. EMNLP
    SLOT: Structuring the Output of Large Language Models
    Darren Yow-Bang Wang, Zhengyuan Shen, Soumya Smruti Mishra, and 3 more authors
    ArXiv, Mar 2025
  5. TACL
    How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
    Zhengbao Jiang, Jun Araki, Haibo Ding, and 1 more author
    Transactions of the Association for Computational Linguistics, Mar 2021