About me

I am a first-year PhD student in the School of Computing and Information Systems, Faculty of Engineering and Information Technology, at The University of Melbourne.

My research focuses on agent safety and trustworthy machine learning β€” building AI systems, and autonomous agents in particular, that behave reliably, safely, and as intended.

I am fortunate to be advised by A/Prof. Xingliang Yuan and Dr. Shaanan Cohney, and I also work closely with Dr. Feng Liu. Previously, during my master’s degree, I conducted research in the TMLR Group, Melbourne, advised by Dr. Feng Liu.

I am always happy to discuss research and potential collaborations. Feel free to reach out at yuhaos1@student.unimelb.edu.au or asymptote1527@gmail.com.

Research interests

  • Agent safety
  • Trustworthy machine learning

News

  • Jun 2026 Released our preprint TRIAD β€” From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents.
  • Feb 2026 BiFTA was accepted to TMLR! πŸŽ‰
  • Jan 2026 Our paper on stealthy fine-tuning data extraction was accepted to ICLR 2026! πŸŽ‰
  • May 2025 SSNI was accepted to ICML 2025! πŸŽ‰

Publications

You can also find my articles on my Google Scholar profile.

In submission From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents

From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents

Yuhao Sun, Jiacheng Zhang, Shaanan Cohney, Zhexin Zhang, Feng Liu, Xingliang Yuan

arXiv preprint, 2026

TMLR 2026 Let's Roll a BiFTA: Bi-refinement for Fine-grained Text-visual Alignment in Vision-Language Models

Let's Roll a BiFTA: Bi-refinement for Fine-grained Text-visual Alignment in Vision-Language Models

Yuhao Sun, Chengyi Cai, Jiacheng Zhang, Zesheng Ye, Xingliang Yuan, Feng Liu

TMLR 2026, 2026

ICLR 2026 Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!

Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!

Zhexin Zhang, Yuhao Sun, Junxiao Yang, Shiyao Cui, Hongning Wang, Minlie Huang

ICLR 2026, 2026

ICML 2025 Sample-Specific Noise Injection For Diffusion-Based Adversarial Purification

Sample-Specific Noise Injection For Diffusion-Based Adversarial Purification

Yuhao Sun*, Jiacheng Zhang*, Zesheng Ye, Chaowei Xiao, Feng Liu (*Equal contribution)

ICML 2025, 2025

Education

  • Ph.D. in Computer Science, The University of Melbourne β€” School of Computing and Information Systems, Faculty of Engineering and Information Technology Β· Aug 2025 – present
  • M.S. in Information Technology (Artificial Intelligence), The University of Melbourne Β· Feb 2023 – Jun 2024 (with Distinction)
  • B.S. in Computer Science, The University of Melbourne Β· Feb 2020 – Nov 2022

Academic service

  • Reviewer, ICML, 2025, 2026
  • Reviewer, ICLR, 2026
  • Reviewer, TMLR
  • Reviewer, Neural Networks