Synthetic audio (sine waves and formant-based speech) cannot evaluate ML-based VAD. Silero correctly rejects formant synthesis as non-human -- this demonstrates quality, not limitation. Real speech samples needed for proper ML VAD testing.
Parent: Live
Synthetic audio (both sine waves and formant-based speech) cannot adequately evaluate ML-based VAD systems like Silero. This is a feature, not a bug - it demonstrates that Silero correctly distinguishes real human speech from synthetic/artificial audio.
Approach: Simple sine waves (200Hz fundamental + 400Hz harmonic)
Results:
- RMS VAD: 28.6% accuracy (treats as speech)
- Silero VAD: 42.9% accuracy (confidence ~0.18, below 0.5 threshold)
Conclusion: Too primitive - neither VAD treats it as real speech
Approach: Sophisticated formant synthesis with:
- 3 formants (F1, F2, F3) matching vowel characteristics
- Fundamental frequency + 10 harmonics
- Amplitude modulation for formant resonances
- Natural variation (shimmer/jitter simulation)
- Proper attack-sustain-release envelopes
Audio patterns generated:
- 5 vowels (/A/, /E/, /I/, /O/, /U/) with accurate formant frequencies
- Plosives (bursts of white noise)
- Fricatives (filtered noise at high frequencies)
- Multi-word sentences (CVC structure)
- TV dialogue (mixed voices + music)
- Crowd noise (5+ overlapping voices)
- Factory floor (machinery + random clanks)
Results:
| VAD Type | Accuracy | Key Observation |
|---|---|---|
| RMS | 55.6% | Improved from 28.6% (detects all loud audio as speech) |
| Silero | 33.3% | Max confidence: 0.242 (below 0.5 threshold) |
Specific Silero responses:
- Silence: 0.044 ✓ (correctly rejected)
- White noise: 0.004 ✓ (correctly rejected)
- Formant speech /A/: 0.018 ✗ (rejected as non-human)
- Plosive /P/: 0.014 ✗ (rejected as non-human)
- TV dialogue: 0.016 ✗ (rejected despite containing speech-like patterns)
Approach: 3-word sentence (multiple CVC patterns) processed in 32ms chunks
Results: 0/17 frames detected as speech
Highest confidence: Frame 6: 0.242 (still below 0.5 threshold)
Conclusion: Even with sustained context, Silero rejects formant synthesis
Silero was trained on 6000+ hours of real human speech. It learned to recognize:
- Natural pitch variations (jitter)
- Harmonic structure from vocal cord vibrations
- Articulatory noise (breath, vocal tract turbulence)
- Formant transitions (co-articulation between phonemes)
- Natural prosody (stress, intonation patterns)
Our formant synthesis, while mathematically correct, lacks:
- Irregular glottal pulses (vocal cords don't vibrate perfectly)
- Breathiness (turbulent airflow through glottis)
- Formant transitions (smooth movements between phonemes)
- Micro-variations in pitch and amplitude
- Natural noise from the vocal tract
Silero rejecting synthetic speech means:
- It won't be fooled by audio synthesis attacks
- It's selective about what counts as "human speech"
- It provides high-quality speech detection for real-world use
What synthetic audio CAN test:
- Pure noise rejection (✓ Silero: 100%)
- Energy-based VAD (RMS threshold)
- Relative comparisons (is A louder than B?)
What synthetic audio CANNOT test:
- ML-based VAD accuracy (Silero, WebRTC neural VAD)
- Speech vs non-speech discrimination
- Real-world performance
Pros:
- Ground truth labels
- Realistic evaluation
- Free datasets available (LibriSpeech, Common Voice, VCTK)
Cons:
- Large downloads (multi-GB)
- Need preprocessing (segmentation, labeling)
- Not reproducible (depends on dataset)
Recommended datasets:
- LibriSpeech: 1000 hours, clean read speech
- Common Voice: Multi-language, diverse speakers
- VCTK: 110 speakers, UK accents
Pros:
- Reproducible
- Controllable (generate specific scenarios)
- Compact (10-100MB model)
Cons:
- Requires model download
- Still not perfect human speech
- Adds dependency
Available TTS:
- Piper (ONNX, Home Assistant) - 20MB model
- Kokoro (ONNX, 82M params) - ~80MB model
- Both already have trait-based adapters in
src/tts/
- Synthetic audio for RMS VAD - Tests energy-based detection
- Real speech samples for Silero VAD - Tests ML-based detection
- TTS for edge cases - Generate specific scenarios (background noise, multiple speakers)
Created this document + test cases showing the limitation.
WebRTC VAD is simpler than Silero (rule-based, not neural) and may work better with synthetic audio for testing.
# LibriSpeech test-clean (346MB, 5.4 hours)
wget https://www.openslr.org/resources/12/test-clean.tar.gz
tar -xzf test-clean.tar.gz
# Use for VAD accuracy benchmarkingDownload Piper or Kokoro models and use for generating test scenarios:
let tts = PiperTTS::new();
tts.initialize().await?;
let audio = tts.synthesize("Hello world", "en_US-amy-medium").await?;
let vad_result = silero.detect(&audio.samples).await?;- Formant generator:
src/vad/test_audio.rs - Realistic audio tests:
tests/vad_realistic_audio.rs - Original sine wave tests:
tests/vad_background_noise.rs
Key Takeaway: Silero correctly rejecting formant synthesis demonstrates its quality as a VAD system. It distinguishes real human speech from synthetic/artificial audio.
For comprehensive VAD testing, we need real human speech samples, not synthetic audio.
The formant synthesis work is still valuable for:
- Testing energy-based VAD (RMS threshold)
- Generating background noise patterns
- Understanding speech acoustics
- Placeholder until TTS models are downloaded
But it cannot properly evaluate ML-based VAD like Silero.