What should we measure?
I build task-grounded benchmarks and agentic quality graders for Copilot, including Today. EmailBench evaluates whether agents complete enterprise email and productivity tasks, beyond whether their tool calls succeed.
Human-agent interaction · Evaluation · Post-training
I focus on human-agent interaction, agent evaluation, and post-training for Microsoft Copilot. I build task-grounded benchmarks, agentic quality graders, and synthetic training data for tool use, multimodal reasoning, and adversarial robustness -- connecting model improvements to product quality, cost, and latency.
My research combines controlled experiments, behavioral measurement, and production evaluation to align agents with how people actually work. Previously, I developed learning tools at MIT and deployed scholarly discovery systems with AI2. With Conservation X Labs, I built team-formation algorithms for global innovation contests offering over $2M in prize funding.
I earned my Ph.D. in Human-Computer Interaction at Carnegie Mellon.
I build task-grounded benchmarks and agentic quality graders for Copilot, including Today. EmailBench evaluates whether agents complete enterprise email and productivity tasks, beyond whether their tool calls succeed.
I connect evaluation to synthetic data and post-training. VLM-SlideEval studies slide comprehension and perturbation sensitivity; my synthetic red-teaming work pairs everyday tasks with adversarial variants to test prompt-injection defenses and their effects on legitimate task performance.
I study how behavior can inform better objectives for AI. Users Mispredict examines where stated preferences diverge from choices; my Semantic Scholar work tests interventions with real users.
* Equal contribution: Zana Buçinca and Hyeonsu B. Kang.
Best Paper Honorable Mention
Best Paper Award