GLAD: Audio Deepfake Detection
Global-local SSL features, sample-adaptive gating, and SaniBoost improve detection under unseen attacks and domain shifts.
GLAD (Global-Local Adaptive Detector) studies speech deepfake detection under unseen attacks and domain shifts, where static feature selection, global-semantic bias, and environment-specific shortcuts often fail to generalize.

Figure 3 from the manuscript: SaniBoost, the Hierarchical Global-Local backbone, and Hierarchical Adaptive Gating form the complete GLAD pipeline.
Core highlights
- Hierarchical global-local encoding: semantic-acoustic cross-attention combines dual-stream self-supervised features, while multi-granularity fusion connects global temporal context with local CNN traces.
- Sample-adaptive gating: dynamically reweights multiple SSL layers and heterogeneous backbones according to the evidence in each attack sample.
- SaniBoost augmentation: noise sanitization and signal normalization reduce shortcut learning from environmental artifacts.
The resulting system placed fourth in the Efficient Speech Deepfake Detection Challenge 2 at IEEE ICME 2026.