ICME 2026 Challenge Session

Grand Challenge 1: ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge

Session Schedule

โฐ 9:30 AM โ€“ 12:30 PM, July 5 ๐Ÿ“ Thai Boromphimarn 4

1. Opening Remarks

9:30 โ€“ 9:40 AM ยท 10 min
Host: Ming Li
  • Welcome and introduction
  • Overview of the challenge/session
  • Introduction of schedule and presentation format

2. Invited Talk & Q&A

9:40 โ€“ 10:40 AM ยท 60 min
Speaker: Xin Wang Host: Ming Li

On Evaluation Metrics of Speech Deepfake Detection โ€“ Lessons Learned from ASVSpoof5

Abstract ASVspoof challenges, which are dedicated to speech anti-spoofing or deepfake detection, have been using discriminative-aware evaluation metrics in the past ten years. While they are useful in measuring how good the detectors are at discriminating real and fake data, they rely on an oracle setup and ignore the goodness of the detectors from the user's point of view. The latter aspect is also known as score calibration. This talk serves as a mini-tutorial on this missing aspect and shares the findings from the latest ASVspoof 5 challenge.
Xin Wang
Bio Xin Wang is a Project Associate Professor and a JST PRESTO researcher at the National Institute of Informatics, Japan. He is one of the organizers of the last three ASVspoof Challenges. He is also a member of the organizing team for the last three VoicePrivacy Challenges, an international joint effort on speech privacy.

3. Oral Presentation Session

3.1 Challenge Summary Paper

10:40 โ€“ 11:00 AM ยท 20 min
Speaker: Xueping Zhang
Paper ID Paper Title
223 Overview of ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge

3.2 Coffee Break

11:00 โ€“ 11:10 AM ยท 10 min

3.3 Accepted Paper Spotlight Presentations

11:10 AM โ€“ 12:00 PM ยท 50 min
Format: 7 min each (5 min presentation + 2 min Q&A)

Certificate presentation and commemorative photos for each team; sponsor acknowledgement during the award presentation for the first-place team.

Order Paper ID Paper Title
1 208 Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection
2 209 MIF: Multi-view Interaction Framework for Speech and Environment Sound Deepfake Detection
3 215 Dual-Stream EAT-XLSR: A Cascade Framework for Component-Aware Audio Deepfake Detection
4 217 Deepfake Audio Detection Using Self-supervised Fusion Representations
5 221 Phoneme-Guided Fusion for Environment-Aware Speech and Sound Deepfake Detection
6 224 EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge
7 232 From Signal Separation to Feature Decomposition: A Framework for Synthetic Speech Detection in Complex Environments

4. Panel Discussion

12:00 โ€“ 12:20 PM ยท 20 min
Participants: Xin Wang, Ming Li Host: Xueping Zhang

5. Closing Remarks

12:20 โ€“ 12:30 PM ยท 10 min
Host: Ming Li

Acknowledgement to participants, organizers, reviewers, sponsor, and ICME.

Overall Timeline

Time Event Duration Participants Host
9:30 โ€“ 9:40 Opening Remarks 10 min โ€” Ming Li
9:40 โ€“ 10:40 Invited Talk & Q&A 60 min Xin Wang Ming Li
10:40 โ€“ 11:00 Oral Session 20 min Summary paper Ming Li
11:00 โ€“ 11:10 Coffee Break 10 min โ€” Ming Li
11:10 โ€“ 12:00 Oral Session & Awards 50 min All papers Ming Li
12:00 โ€“ 12:20 Panel Discussion 20 min Ming Li, Xin Wang Xueping Zhang
12:20 โ€“ 12:30 Closing Remarks 10 min โ€” Ming Li