What should we measure?
I build task-grounded benchmarks and agentic quality graders for Copilot, including Today. EmailBench evaluates whether agents complete enterprise email and productivity tasks, beyond whether their tool calls succeed.
Human–agent interaction · Evaluation · Post-training
Senior Applied Scientist, Microsoft
I focus on human–agent interaction, agent evaluation, and post-training for Copilot. I build task-grounded benchmarks, agentic quality graders, and synthetic training data for tool use, multimodal reasoning, and adversarial robustness—connecting model improvements to product quality, cost, and latency.
My research combines controlled experiments, behavioral measurement, and production evaluation to align agents with how people actually work. Previously, I developed learning tools at MIT and deployed scholarly discovery systems with AI2. With Conservation X Labs, I built team-formation algorithms for global innovation contests offering over $2M in prize funding.
I earned my Ph.D. in Human–Computer Interaction at Carnegie Mellon.
I build task-grounded benchmarks and agentic quality graders for Copilot, including Today. EmailBench evaluates whether agents complete enterprise email and productivity tasks, beyond whether their tool calls succeed.
I connect evaluation to synthetic data and post-training. VLM-SlideEval studies slide comprehension and perturbation sensitivity; my counterfactual red-teaming work turns everyday tasks into controlled adversarial cases.
I study how behavior can inform better objectives for AI. Users Mispredict examines where stated preferences diverge from choices; my Semantic Scholar work tests interventions with real users.
* Equal contribution: Zana Buçinca and Hyeonsu B. Kang.
Best Paper Honorable Mention
Best Paper Award