Human–agent interaction · Evaluation · Post-training

Hyeonsu B. Kang

Senior Applied Scientist, Microsoft

I focus on human–agent interaction, agent evaluation, and post-training for Copilot. I build task-grounded benchmarks, agentic quality graders, and synthetic training data for tool use, multimodal reasoning, and adversarial robustness—connecting model improvements to product quality, cost, and latency.

My research combines controlled experiments, behavioral measurement, and production evaluation to align agents with how people actually work. Previously, I developed learning tools at MIT and deployed scholarly discovery systems with AI2. With Conservation X Labs, I built team-formation algorithms for global innovation contests offering over $2M in prize funding.

I earned my Ph.D. in Human–Computer Interaction at Carnegie Mellon.

Research

What should we measure?

I build task-grounded benchmarks and agentic quality graders for Copilot, including Today. EmailBench evaluates whether agents complete enterprise email and productivity tasks, beyond whether their tool calls succeed.

How can evaluations improve agents?

I connect evaluation to synthetic data and post-training. VLM-SlideEval studies slide comprehension and perturbation sensitivity; my counterfactual red-teaming work turns everyday tasks into controlled adversarial cases.

What do people actually want?

I study how behavior can inform better objectives for AI. Users Mispredict examines where stated preferences diverge from choices; my Semantic Scholar work tests interventions with real users.

Publications

Google Scholar

2026

  1. Users Mispredict What Drives Their Preferences for AI Writing Assistance

    Vivian Lai, Zana Buçinca*, Hyeonsu B. Kang*, Mo Houtti, Nil-Jana Akpinar, Kevin Chian, Namjoon Suh, Alex C. Williams

    EMNLP 2026 · Main conference · To appearPreprint ↗

    * Equal contribution: Zana Buçinca and Hyeonsu B. Kang.

  2. EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

    Mukul Singh, Mansi Uniyal, Devin Devlin, Wen Xie, Big Thadawasin, Ritam Dutt, Vivian Lai, Hyeonsu B. Kang

    arXiv preprint · 2026Paper ↗

  3. From Everyday Tasks to Red-Team Cases: Counterfactual Generation and Practical Defenses for AI Agents

    Mohit Chandra, Wesley Deng, Ritam Dutt, Alex Williams, Vivian Lai, Hyeonsu Kang

    In submission · 2026

  4. PIKE: Graph-Enhanced Email Retrieval via Critic-Guided Expansion

    Namjoon Suh, Vivian Lai, Aynaz Taheri, Hyeonsu B. Kang, Anastasia Kuzminykh, Alex C. Williams

    In submission · 2026

2025

  1. VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT

    Hyeonsu B. Kang, Emily Bao, Anjan Goswami

    NeurIPS 2025 · Evaluating the Evolving LLM Lifecycle WorkshopPaper ↗

  2. BioSpark: Beyond Analogical Inspiration to LLM-augmented Transfer

    Hyeonsu B. Kang, David Chuan-en Lin, Yan-Ying Chen, Matthew K. Hong, Nikolas Martelaro, Aniket Kittur

    CHI 2025Paper ↗

    Best Paper Honorable Mention

  3. Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching

    David Chuan-En Lin, Hyeonsu B. Kang, Nikolas Martelaro, Aniket Kittur, Yan-Ying Chen, Matthew K. Hong

    CHI 2025Paper ↗

  4. Hypoveil: A Privacy-Utility Trade-off Aware Framework for Collaborative Multi-Agent Reasoning

    Hyeonsu B. Kang, Roshni Kaushik, Koichi Onoue

    CMU Agent Workshop · 2025

  5. Investigating Explainability with Privacy-Preserving Multi-agent Systems

    Roshni Kaushik, Hyeonsu B. Kang, Koichi Onoue

    CMU Agent Workshop · 2025

2024

  1. Mitigating Barriers to Public Social Interaction with Meronymous Communication

    Nouran Soliman, Hyeonsu B. Kang, Matthew Latzke, Jonathan Bragg, Joseph Chee Chang, Amy X. Zhang, David R Karger

    CHI 2024Paper ↗

    Best Paper Award

  2. PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers

    Yoonjoo Lee, Hyeonsu B. Kang, Matthew Latzke, Juho Kim, Jonathan Bragg, Joseph Chee Chang, Pao Siangliulue

    CHI 2024Paper ↗

  3. The Semantic Reader Project: Augmenting Scholarly Documents through AI-Powered Interactive Reading Interfaces

    Kyle Lo, Joseph Chee Chang, Andrew Head, et al. (including Hyeonsu B. Kang)

    Communications of the ACM (2024)Paper ↗

  4. Imitation of Life: A Search Engine for Biologically Inspired Design

    Hen Emuna, Nadav Borenstein, Xin Qian, Hyeonsu Kang, Joel Chan, Aniket Kittur, Dafna Shahaf

    AAAI 2024Paper ↗

2023

  1. Synergi: A Mixed-Initiative System for Scholarly Synthesis and Sensemaking

    Hyeonsu B. Kang, Sherry Tongshuang Wu, Joseph Chee Chang, Aniket Kittur

    UIST 2023Paper ↗

  2. ComLittee: Literature Discovery with Personal Elected Author Committees

    Hyeonsu B. Kang, Nouran Soliman, Matt Latzke, Joseph Chee Chang, Jonathan Bragg

    CHI 2023Paper ↗

  3. BioSpark: An End-to-End Generative System for Biological-Analogical Inspirations and Ideation

    Hyeonsu B. Kang, David Chuan-En Lin, Nikolas Martelaro, Aniket Kittur, Yan-Ying Chen, Matthew K. Hong

    NeurIPS 2023 Creativity WorkshopPaper ↗

2022

  1. Augmenting Scientific Creativity with an Analogical Search Engine

    Hyeonsu B. Kang, Xin Qian, Tom Hope, Dafna Shahaf, Joel Chan, and Aniket Kittur

    TOCHI 2022Paper ↗

  2. Threddy: An Interactive System for Personalized Thread-based Exploration and Organization of Scientific Literature

    Hyeonsu B. Kang, Joseph Chee Chang, Yongsung Kim, Aniket Kittur

    UIST 2022Paper ↗

  3. From Who You Know to What You Read: Augmenting Scientific Recommendations with Implicit Social Networks

    Hyeonsu B. Kang, Rafal Kocielnik, Andrew Head, Jiangjiang Yang, Matt Latzke, Aniket Kittur, Daniel Weld, Doug Downey, and Jonathan Bragg

    CHI 2022Paper ↗

  4. Scaling Creative Inspiration with Fine-Grained Functional Aspects of Ideas

    Tom Hope, Ronen Tamari, Hyeonsu Kang, Daniel Hershcovich, Joel Chan, Aniket Kittur, and Dafna Shahaf

    CHI 2022Paper ↗

  5. Augmenting Scientific Creativity with Retrieval across Knowledge Domains

    Hyeonsu B. Kang*, Sheshera Mysore*, Kevin Huang*, Haw-Shiuan Chang, Thorben Prein, Andrew McCallum, Aniket Kittur, Elsa Olivetti

    NAACL 2022 WorkshopPaper ↗

Earlier publications · 2017–2019

2019

  1. Matching Open Innovation Projects for Analogical Feedback Exchange

    Hyeonsu Kang, Felicia Ng, Aniket Kittur

    Collective Intelligence 2019

2018

  1. Paragon: An Online Gallery for Enhancing Design Feedback with Visual Examples

    Hyeonsu B. Kang, Gabriel Amoako, Neil Sengupta, Steven Dow

    CHI 2018Paper ↗

  2. Custom Blocks in StarLogo Nova: A Template-Based Approach to Abstraction for Improved Ease of Use and Expressive Power

    Hyeonsu Kang, David Wu, David Wendel

    SIGPLAN 2018