Yuyang Jiang

ๆฑŸ ๅฎ‡้˜ณ

Ph.D. student
Department of Computer Science, University of Southern California
Email: kjiang4work@gmail.com

Research Interest: Science of Evaluation

Robust evaluation methodology is essential for guiding the training of reliable systems. It not only measures system performance, but also seeks to understand system behavior and, most importantly, continuously refine alignment rubrics so they reflect rational human intentions and can be distilled into the training process.

  • Static Evaluation: Design granular yet scalable metrics that capture richer task-specific properties and better match real design goals.
  • Human-in-the-Loop Evaluation: (1) Build representative feedback loops under limited budgets; (2) monitor bidirectional risks in human-AI interaction (e.g., humans: over-reliance, manipulation; AI: over-alignment, sycophancy) and develop collaboration paradigms that preserve rationality on both sides.
  • Interactive (Agentic) Evaluation: (1) Study the strengths and limits of foundational structures that emerge in agentic systems; (2) test the robustness of cooperative behaviors under adversarial conditions.

I'm especially interested in applying these ideas to AI safety and healthcare.

Milestones

2023

๐Ÿ‘ฉโ€๐ŸŽ“ B.Econ. Economics & Mathematics Playground โ†’

CUFE CEMA | ๐Ÿ‡จ๐Ÿ‡ณ Beijing | Human behavior modeling

2025

๐Ÿ‘ฉโ€๐ŸŽ“ M.S. Statistics Playground โ†’

UChicago Stats | ๐Ÿ‡บ๐Ÿ‡ธ Chicago | Theoretical machine learning & deep learning

2026

๐Ÿ‘ฉ๐Ÿปโ€๐Ÿ’ป Research Intern

Vector Institute | ๐Ÿ‡จ๐Ÿ‡ฆ Toronto | Agent safety audits

2026

๐Ÿคฉ ICML 2026 Paper Acceptance

ICML 2026 | ๐Ÿ‡ฐ๐Ÿ‡ท Seoul | Multi-agent debate & scalable oversight

๐Ÿค– AI conference . ๐Ÿฉป Healthcare conference . *Equal first co-authors; โ€ Equal second co-authors.

Publications

Resources and Evaluation

CLEAR: A Clinically Grounded Tabular Framework for Radiology Report Evaluation
Yuyang Jiang, Chacha Chen, Shengyuan Wang, Feng Li, Zecong Tang, Benjamin M. Mervak, Lydia Chelala, Christopher M Straus, Reve Chahine, Samuel G. Armato III*, Chenhao Tan*.
EMNLP 2025 ๐Ÿค–. [Paper] [Poster] [Slides] [Dataset submitted to PhysioNet (under review)] [Code]
GPT-4V Cannot Generate Radiology Reports Yet
Yuyang Jiang*, Chacha Chen*, Dang Nguyen, Benjamin M. Mervak, Chenhao Tan.
ML4H 2024 (Poster) ๐Ÿฉป, NAACL 2025 ๐Ÿค–. [Paper] [Poster] [Slides] [Code]

Safety and Alignment

Beyond One-Way Influence: Bidirectional Opinion Dynamics in Multi-Turn Human-LLM Interactions
Yuyang Jiang, Longjie Guoโ€ , Yuchen Wuโ€ , Aylin Caliskan, Tanu Mitra and Hua Shen.
CHI 2026 BiAlign Workshop ๐Ÿค–. [Preprint] [Workshop]
Collaborative Disagreement Resolution for Scalable Oversight
Yuyang Jiang*, Chacha Chen*, Teng Wuโ€ , Liwen Sunโ€ , Han Liu, Shi Feng and Chenhao Tan.
ICML 2026 ๐Ÿค–. [Presentation] [Paper]

Academic Service

Presentation: ICML 2026 ๐Ÿค– (Poster), EMNLP 2025 ๐Ÿค– (Poster), NAACL 2025 ๐Ÿค– (Poster), ML4H 2024 ๐Ÿฉป (Poster), TTIC Multimodal AI Workshop 2024 ๐Ÿค– (Lightning Talk)

Reviewer: CHIL 2026 ๐Ÿฉป, Sage Digital Health ๐Ÿฉป, CHI 2026 ๐Ÿค–

Teaching Experience

BUSN 32200: Artificial Intelligence MBA Course (Course Design Team Member) | Winter 2025 | Instructor: Dacheng Xiu
BUSN 32810: Artificial Intelligence EMBA Course (Course Design Team Member) | Summer 2024 | Instructor: Dacheng Xiu
BUSN 20800: Big Data Undergraduate Course | Winter 2024 | Instructor: Dacheng Xiu