AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

A large-scale benchmark spanning 21 spoofing systems, five emotion conditions, and ~260 hours of data.

Submitted to EMNLP 2026

Why AffectDF?
Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and recent large audio-language model (LALM)-based attacks. Existing emotional spoofing datasets are also limited in scale and attack diversity, typically covering only voice conversion (VC) or text-to-speech (TTS) attacks. We introduce AffectDF, the most comprehensive benchmark for emotionally expressive speech deepfakes, spanning TTS, VC, emotional VC, and LALM-based spoofing attacks across both acted and spontaneous emotional speech. AffectDF contains approximately 260 hours of speech generated using 21 spoofing attacks across five emotional states. We benchmark state-of-the-art SDD systems under conventional and emotional spoofing conditions. Our experiments reveal severe robustness degradation when models trained on conventional benchmarks are evaluated on AffectDF, with several systems approaching near-random performance. Surprisingly, even large-scale emotional training does not consistently improve cross-domain robustness, indicating that current SDD systems fail to learn generalized spoof representations under emotional and prosodic variability. Robustness further varies substantially across emotional states, attack families, and acted vs spontaneous emotional speech conditions. These findings expose fundamental limitations of current SDD systems and establish AffectDF as a benchmark for developing more robust spoof detection models.
Dataset Architecture & Evaluation Framework
Dataset Architecture and Evaluation Framework
AffectDF: Spoofing Attack Configuration

AffectDF spans 21 spoofing attacks across Train, Development, and Test partitions with disjoint speakers and attack systems. Generation types cover TTS, VC, EVC, VC+EVC, and LALM-based systems across five emotional states using both the ESD (acted) and MSP-Podcast (spontaneous) corpora.

Attack ID Generation Type Model Base Data Samples No. Spks Split
TRAIN
A01LALM-EVCQwen 2.5-OmniESD35,0004Train
A02LALM-EVCQwen 2.5-Omni (steered)ESD28,0004Train
A03TTSCosyVoiceESD5,9994Train
A04TTSCosyVoice2ESD6,0004Train
A05TTSQwen3-TTSESD6,0004Train
DEVELOPMENT
A06LALM-EVCMiniCPMESD17,5002Dev
A07TTSCosyVoice3ESD2,8302Dev
TEST
A08LALM-EVCKimi-AudioESD34,9924Test
A09LALM-EVCKimi-Audio (steered)ESD27,9944Test
A10VC+EVCVevo2ESD24,0004Test
A11VCVevo 2ESD6,0004Test
A12EVCGenVCESD24,0004Test
A13VCGenVCESD6,0004Test
A14VCGenVCMSP4,1084Test
A15VCConsistencyVCESD6,0524Test
A16VCTriAAN-VCESD6,0004Test
A17VCDDDMVCESD6,0004Test
A18TTSStyle-TTS2ESD6,0004Test
A19TTSStyle-TTS2MSP4,1074Test
A20TTSF5-TTSESD6,0004Test
A21TTSF5-TTSMSP4,1074Test
Benchmark Evaluation Results

Conventional training fails under emotional spoofing. ASVspoof2019-trained models perform well on ASVspoof2019/2021 but degrade sharply on AffectDF, with RawNet2 and AASIST reaching 59.71% and 56.40% EER. ASVspoof5 training improves AffectDF for some models, but remains inconsistent across architectures and collapses on EmoFake for several systems.

Emotional training does not guarantee generalization. Training on AffectDF improves some emotional-domain results, but sharply hurts performance on ASVspoof2019 and ASVspoof5 for nearly all models. This suggests current SDD systems still learn dataset- and attack-specific cues rather than transferable spoof representations.

Table 2: Cross-domain evaluation of SDD models trained on ASVspoof2019, ASVspoof5, and AffectDF. Results in EER (%).

ModelTrain DataASVspoof-2019ASVspoof-2021ASVspoof5EmoFakeAffectDF
RawNet2ASVspoof20194.608.0840.6721.7159.71
AASISTASVspoof20190.838.1535.5313.6456.40
XLSR-SLSASVspoof20190.563.0425.438.8444.91
XLSR-MambaASVspoof20190.201.6415.540.6929.78
ProSDDASVspoof20190.423.8716.143.7031.04
RawNet2ASVspoof524.7525.5943.6149.4933.02
AASISTASVspoof523.1622.7425.7762.7118.00
XLSR-SLSASVspoof527.0026.5439.6258.5732.28
XLSR-MambaASVspoof513.6513.677.2720.7735.27
ProSDDASVspoof519.0418.087.3825.0612.49
RawNet2AffectDF43.0247.2544.656.5923.17
AASISTAffectDF44.5248.7845.7522.6036.91
XLSR-SLSAffectDF64.8360.5656.5418.4136.20
XLSR-MambaAffectDF46.4753.0947.6728.3941.60
ProSDDAffectDF58.1556.4363.1622.0848.14

Failure is emotion-dependent, but not emotion-specific. The highest EER shifts across neutral, happy, angry, sad, and surprise depending on the model and training data. This shows that no single emotion is universally hardest; failures arise from the interaction between emotional prosody, attack type, and detector architecture.

AffectDF training still leaves large emotion gaps. Even after training on emotionally expressive data, models show wide EER variation across emotions. This indicates that current detectors do not learn emotion-invariant spoof cues, even when exposed to balanced emotional training conditions.

Table 3: Emotion-wise EER (%) results. N = Neutral, H = Happy, A = Angry, S = Sad, Su = Surprise. EmoFake does not include Sad.

ModelTrain Data EmoFake AffectDF
NHASSu NHASSu
RawNet2ASVspoof201919.6616.5723.0028.5456.2160.9363.8248.8864.28
AASISTASVspoof201917.2011.6315.0013.5765.4561.0750.9141.8248.56
XLSR-SLSASVspoof20195.208.099.0310.7759.9451.9231.8732.1239.46
XLSR-MambaASVspoof20190.860.430.940.5428.5231.7826.1321.4223.99
ProSDDASVspoof20192.143.604.572.2931.7033.5626.2024.0833.14
RawNet2ASVspoof552.2948.0742.8645.3042.2533.9925.0030.4420.38
AASISTASVspoof570.4659.5156.3463.0622.4817.4814.0422.0010.98
XLSR-SLSASVspoof559.7155.6759.1658.5039.7330.7427.0538.4920.25
XLSR-MambaASVspoof520.7117.2917.0022.2940.8541.5522.8323.1420.06
ProSDDASVspoof527.4622.1720.9418.7416.2511.959.6013.837.35
RawNet2AffectDF7.515.006.665.1425.6726.7917.1815.4015.05
AASISTAffectDF25.0623.1426.7115.8345.9849.2322.5725.1118.03
XLSR-SLSAffectDF14.8616.5715.6323.3441.8643.3025.0725.4021.15
XLSR-MambaAffectDF29.0031.2630.6622.1457.6052.3124.6527.1522.33
ProSDDAffectDF21.3414.6614.4621.2046.1150.0243.8448.0335.23
Qwen-2.5-OmniInference-only29.3722.9625.5031.4954.1648.7437.8123.7641.61
Qwen-3.0-OmniInference-only35.1231.7339.5839.7441.1437.6942.4740.8544.99
VoxtralInference-only52.2450.8152.7950.6756.4656.4851.9950.5750.56

Prompting LALMs is insufficient for robust SDD. Qwen-2.5-Omni performs reasonably on ASVspoof2019 and EmoFake but degrades on ASVspoof5 and AffectDF, while Voxtral remains poor across nearly all benchmarks. General audio-language understanding does not directly translate to spoof detection.

LALM robustness is highly benchmark-dependent. Qwen-3.0-Omni performs better on ASVspoof5 than on ASVspoof2019 or AffectDF, showing unstable behavior across evaluation domains. This suggests inference-only LALMs remain sensitive to attack diversity and emotional variability.

Table 4: Inference-only evaluation of LALM models. Results in EER (%).

ModelASV19ASV5EmoFakeAffectDF
Qwen-2.5-Omni29.5446.2325.1545.29
Qwen-3.0-Omni42.4227.9134.1839.81
Voxtral46.0650.6251.7654.20

Speaking style alone does not explain difficulty. The acted–spontaneous gap changes by model and attack: ASVspoof2019-trained systems often struggle more on acted attacks, while ASVspoof5-trained systems show partial transfer to both acted and spontaneous conditions.

Emotional training is not style-invariant. Although AffectDF training uses acted emotional speech, several models perform better on spontaneous attacks than acted counterparts. Robustness depends on the source style, generation model, and detector architecture — not simply on whether emotional data was used in training.

Table 5: Attack-wise EER (%) for acted (A13, A18, A20) and spontaneous (A14, A19, A21) emotional speech generation systems. Overall = pooled EER.

ModelTrain Data Acted Spontaneous
A13A18A20Overall A14A19A21Overall
RawNet2ASVspoof201962.3353.0251.1655.5144.2644.2739.7442.74
AASISTASVspoof201962.9049.4860.3357.4743.1162.3148.2351.38
XLSR-SLSASVspoof201942.8145.8074.8548.6434.1044.0140.9341.83
XLSR-MambaASVspoof20196.4327.4253.7031.4322.1827.0525.5725.73
ProSDDASVspoof201914.2218.4715.5716.469.4354.1522.4929.72
RawNet2ASVspoof535.2536.4041.0737.2523.2753.0162.3647.48
AASISTASVspoof521.5714.6738.5625.1712.2422.0439.8125.69
XLSR-SLSASVspoof536.9325.8541.7034.9032.6232.8262.8242.66
XLSR-MambaASVspoof517.2241.3035.0434.112.0235.4137.0331.91
ProSDDASVspoof55.2513.8518.5013.725.8419.9637.1123.11
RawNet2AffectDF34.381.581.3518.974.792.632.693.64
AASISTAffectDF40.9533.5921.4134.3821.2120.6217.8019.71
XLSR-SLSAffectDF30.4528.9523.9728.2817.0626.2527.4424.67
XLSR-MambaAffectDF47.1540.0916.6640.4418.8218.7616.1418.16
ProSDDAffectDF48.069.6013.8726.9251.8621.8927.4734.18
Qwen-2.5-OmniInference-only45.5148.2948.9047.5744.6447.1745.6045.87
Qwen-3.0-OmniInference-only36.6053.3445.9146.2126.8143.6150.6440.81
VoxtralInference-only54.7755.8953.8354.8546.9449.9637.6145.18
AffectDF is More Emotionally Expressive than Conventional Benchmarks

We analyze emotional expressiveness across AffectDF and existing SDD benchmarks using automatic speech emotion recognition (SER) with Emotion2Vec+-large. AffectDF exhibits substantially stronger and more diverse emotional coverage than ASVspoof2019 and ASVspoof5, which are dominated by neutral speech. This confirms that AffectDF introduces genuine prosodic and emotional variability not present in conventional benchmarks — and explains why models trained on neutral-dominated data fail on AffectDF.

Normalized emotion distribution

Figure 2

Normalized emotion distribution from Emotion2Vec+-large predictions across AffectDF, ASVspoof2019, and ASVspoof5. AffectDF shows substantially higher coverage of happy, angry, sad, and surprise compared to conventional benchmarks.

Emotion confusion matrix

Figure 3

Emotion confusion matrix for AffectDF using Emotion2Vec+-large predictions. Predictions are distributed across multiple emotion categories, confirming meaningful emotional variability in the dataset.

More Information

This dataset is designed to improve the performance of speech security systems against state-of-the-art emotional synthetic speech. It is released for research purposes only. Refer to the paper for more details.