Guanjie Huang

Ph.D. Candidate in Artificial Intelligence at HKUST(GZ)

Open to full-time opportunities Multimodal AI research and applied research roles

I am a Ph.D. candidate in Artificial Intelligence at The Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Li Liu and Prof. Danny H. K. Tsang. My research focuses on efficient and reliable multimodal intelligence, especially audio-visual understanding, cued speech recognition, multimodal large language models, and uncertainty estimation.

My recent work studies how language models and specialized agents can reason over visual speech and hand cues, how multimodal systems can learn from limited data, and how their confidence and authenticity can be assessed. I am also interested in audio deepfake detection and physically grounded evaluation of audio-visual generation.

Before my doctoral study, I worked on industrial computer vision, large-scale multimodal dataset construction, UAV tracking, and learning-assisted optimization. I received an M.A.I. from the Australian National University and a B.Eng. in Software Engineering from the University of Electronic Science and Technology of China.

Research interests: audio-visual learning · multimodal large language models and agents · speech and cued speech recognition · trustworthy multimodal AI · model uncertainty · audio deepfake detection

I welcome conversations about research collaboration and multimodal AI opportunities.

news

May 06, 2026 Presented STF-ACSR, our semi training-free cued speech recognition work, at ICASSP 2026. Paper · Code
Jan 14, 2026 Our paper Semantic Modulated Prompting for Few-Shot Audio-Visual Classification was published in IEEE/ACM TASLP. Paper
Aug 01, 2025 Cued-Agent was selected for an oral presentation at ACM Multimedia 2025, together with a Student Travel Award. Paper · Code

selected publications

  1. TASLP
    Semantic Modulated Prompting for Few-Shot Audio-Visual Classification
    Guanjie Huang, Yawen Cui, Danny Hin Kwok Tsang, and 2 more authors
    IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2026
    Audio-Visual LearningFew-Shot LearningPrompt TuningModality Rebalancing
  2. ICASSP
    Lend a Hand: Semi Training-Free Cued Speech Recognition via MLLM-Driven Hand Modeling
    Guanjie Huang, Danny H. K. Tsang, Xiao-Ping Zhang, and 1 more author
    In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026
    Cued Speech RecognitionMultimodal LLMTraining-Free LearningVisual Speech
  3. ACM MM
    Cued-Agent: A Multi-Agent Framework for Automatic Cued Speech Recognition
    Guanjie Huang, Danny H. K. Tsang, Shan Yang, and 2 more authors
    In Proceedings of the ACM International Conference on MultimediaOral presentation, 2025
    Cued Speech RecognitionMulti-Agent SystemsMultimodal LearningLLM Self-Correction
  4. ACM MM
    PhyAVBench: A Physical-Centric Audio-Visual Benchmark for Video Generation
    Tianxin Xie, Wentao Lei, Guanjie Huang, and 1 more author
    In Proceedings of the ACM International Conference on MultimediaOral presentation, 2026
    Audio-Visual GenerationPhysical GroundingEvaluation BenchmarkText-to-Audio-Video
  5. TPAMI
    WebUAV-3M: A Benchmark Unveiling the Power of Million-Scale Deep UAV Tracking
    Chunhui Zhang, Guanjie Huang, Li Liu, and 5 more authors
    IEEE Transactions on Pattern Analysis and Machine IntelligenceCo-first author. , 2023
    UAV TrackingMultimodal TrackingLarge-Scale DatasetVisual Object Tracking
  6. NeurIPS
    WebUOT-1M: Advancing Deep Underwater Object Tracking with a Million-Scale Benchmark
    Chunhui Zhang, Li Liu, Guanjie Huang, and 3 more authors
    In Advances in Neural Information Processing Systems, 2024
    Underwater TrackingLarge-Scale BenchmarkMultimodal TrackingObject Tracking
  7. CVPRW
    Underwater Camouflaged Object Tracking Meets Vision-Language SAM2
    Chunhui Zhang, Li Liu, Guanjie Huang, and 5 more authors
    In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2025
    Underwater TrackingCamouflaged ObjectsVision-Language ModelsSAM2