CV
Research experience, education, selected publications, projects, and awards. A detailed Chinese PDF is available from the download button.
Contact Information
| Name | Guanjie Huang |
| Professional Title | Ph.D. Candidate in Artificial Intelligence |
| ghuang565@connect.hkust-gz.edu.cn |
Professional Summary
Ph.D. candidate at HKUST(GZ) researching efficient and reliable multimodal intelligence, with a focus on audio-visual learning, cued speech recognition, multimodal large language models and agents, uncertainty estimation, and audio deepfake detection.
Experience
-
2023 - present Guangzhou, China
Doctoral Researcher
The Hong Kong University of Science and Technology (Guangzhou)
Research on audio-visual intelligence, multimodal large language models, model reliability, and assistive speech technologies.
- Developed few-shot audio-visual learning methods using semantic prompts, adapters, latent attention, and prototype regularization.
- Built training-free and semi-training-free cued speech recognition systems with multimodal agents and MLLM-driven hand modeling.
- Studied audio deepfake detection, physically grounded audio-visual generation benchmarks, and evidential uncertainty for in-context learning.
- Contributed to multimodal assessment systems for people with hearing impairment using speech, video, text, and fNIRS signals.
-
2022 - 2023 Suzhou, China
Deep Learning Algorithm Engineer
Suzhou INSNEX Intelligent Technology Co., Ltd.
Developed industrial visual inspection algorithms for defect synthesis and detection.
- Designed GAN-based defect generation methods for data augmentation and industrial inspection; the work contributed to a patent.
- Built generation, detection, and multi-task training pipelines for low-data defect scenarios.
-
2021 - 2022 Shenzhen, China
Algorithm Engineer
Shenzhen Research Institute of Big Data
Worked on large-scale multimodal UAV tracking datasets and learning-assisted optimization.
- Co-developed WebUAV-3M, containing 3.3 million frames from 4,500 videos across 223 categories, together with semi-automatic annotation tools and language/audio modalities.
- Explored deep-learning acceleration for mixed-integer linear programming and developed Python evaluation, testing, and visualization tools.
-
2020 - 2020 Canberra, Australia
Master's Researcher
Australian National University
Researched unified image style transfer methods.
-
2019 - 2019 Canberra, Australia
Research Assistant
CSIRO
Developed interactive 3D geological visualization tools with PyVista and reproducible Binder environments.
Education
-
2023 - present Guangzhou, China
Ph.D.
The Hong Kong University of Science and Technology (Guangzhou)
Artificial Intelligence
- Advised by Prof. Li Liu and Prof. Danny H. K. Tsang.
- Research on efficient audio-visual learning, cued speech recognition, audio-visual forgery detection, and language-model uncertainty.
-
2018 - 2020 Canberra, Australia
-
2014 - 2018 Chengdu, China
Bachelor of Engineering
University of Electronic Science and Technology of China
Software Engineering
- Outstanding Graduate.
Selected Projects
Publications
-
2026 Semantic Modulated Prompting for Few-Shot Audio-Visual Classification
IEEE/ACM Transactions on Audio, Speech, and Language Processing
-
2026 Lend a Hand: Semi Training-Free Cued Speech Recognition via MLLM-Driven Hand Modeling
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
-
2025 Cued-Agent: A Multi-Agent Framework for Automatic Cued Speech Recognition
ACM International Conference on Multimedia
-
2026 PhyAVBench: A Physical-Centric Audio-Visual Benchmark for Video Generation
ACM International Conference on Multimedia
-
2023 WebUAV-3M: A Benchmark Unveiling the Power of Million-Scale Deep UAV Tracking
IEEE Transactions on Pattern Analysis and Machine Intelligence
-
2024 WebUOT-1M: Advancing Deep Underwater Object Tracking with a Million-Scale Benchmark
Advances in Neural Information Processing Systems
-
2025 -
2024 Content-Aware Efficient Learner for Audio-Visual Emotion Recognition
International Conference on Social Robotics
Awards
-
2026 4th Place, Efficient Speech Deepfake Detection Challenge 2
IEEE ICME
-
2025 Student Travel Award
ACM Multimedia
-
2025 Outstanding Paper Award, CV4Animals Workshop
CV4Animals
-
2024 Best Student Paper Nomination
ICSR
-
2023 Outstanding Technology Academic Paper
Shenzhen
-
2018 Outstanding Graduate
UESTC
Skills
Research (Advanced): Audio-visual learning, multimodal representation and alignment, speech recognition, multimodal LLMs and agents, model uncertainty, audio forgery detection
Machine Learning (Advanced): Transformers, adapters and parameter-efficient learning, self-supervised learning, attention and prototype methods, CNNs, U-Net, GANs, multi-task learning, evidential learning
Engineering (Advanced): Python, data and semi-automatic annotation pipelines, model training and evaluation, visualization, APIs, automated testing
Languages
Chinese : Native
English : Professional working proficiency
Certificates
- ISTQB Certified Tester - International Software Testing Qualifications Board (2017)