Guanjie Huang
Ph.D. Candidate in Artificial Intelligence at HKUST(GZ)
Guangzhou, China
I am a Ph.D. candidate in Artificial Intelligence at The Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Li Liu and Prof. Danny H. K. Tsang. My research focuses on efficient and reliable multimodal intelligence, especially audio-visual understanding, cued speech recognition, multimodal large language models, and uncertainty estimation.
My recent work studies how language models and specialized agents can reason over visual speech and hand cues, how multimodal systems can learn from limited data, and how their confidence and authenticity can be assessed. I am also interested in audio deepfake detection and physically grounded evaluation of audio-visual generation.
Before my doctoral study, I worked on industrial computer vision, large-scale multimodal dataset construction, UAV tracking, and learning-assisted optimization. I received an M.A.I. from the Australian National University and a B.Eng. in Software Engineering from the University of Electronic Science and Technology of China.
Research interests: audio-visual learning · multimodal large language models and agents · speech and cued speech recognition · trustworthy multimodal AI · model uncertainty · audio deepfake detection
I welcome conversations about research collaboration and multimodal AI opportunities.
news
| May 06, 2026 | Presented STF-ACSR, our semi training-free cued speech recognition work, at ICASSP 2026. Paper · Code |
|---|---|
| Jan 14, 2026 | Our paper Semantic Modulated Prompting for Few-Shot Audio-Visual Classification was published in IEEE/ACM TASLP. Paper |
| Aug 01, 2025 | Cued-Agent was selected for an oral presentation at ACM Multimedia 2025, together with a Student Travel Award. Paper · Code |
selected publications
- ACM MMCued-Agent: A Multi-Agent Framework for Automatic Cued Speech RecognitionIn Proceedings of the ACM International Conference on MultimediaOral presentation, 2025Cued Speech RecognitionMulti-Agent SystemsMultimodal LearningLLM Self-Correction
- CVPRWUnderwater Camouflaged Object Tracking Meets Vision-Language SAM2In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2025Underwater TrackingCamouflaged ObjectsVision-Language ModelsSAM2