Skip to main content

Resource

Voice fraud detection benchmark methodology and metrics

Physics-based anti-spoofing means measuring acoustic signal properties (spectral envelope trajectories and source excitation phase patterns) that synthetic pipelines struggle to reproduce consistently, then documenting those measurements for audit. Sonotheia reports Equal Error Rate (EER) and minimum Detection Cost Function (minDCF) on the ASVspoof5 evaluation partition (spoof attack types A17-A32).

While raw database benchmarks establish baseline detection accuracy, real-world deployment requires blind, in-situ calibration on the partner's actual telephony codec path to control false-positive rates.

What is physics-based anti-spoofing?

Physics-based anti-spoofing analyzes measurable acoustic dynamics rather than opaque embedding scores alone. Sonotheia tracks spectral envelope trajectories and source excitation phase patterns, logs sensor contributions, and ties verdicts to documented thresholds so teams can defend the control under model-risk review.

ASVspoof5 dataset validation scope

We validate our physics-based sensors against the ASVspoof5 evaluation partition. This benchmark includes advanced spoofing methods such as neural codec laundering, state-of-the-art TTS generation, and voice conversion algorithms. Performance sweeps map out-of-domain generalization limits across different noise and codec conditions.

What we do not claim yet

We do not claim universal, out-of-the-box accuracy below a 2% equal error rate across arbitrary networks. Telephony codecs, room acoustics, and microphone profiles shape the signal. Performance is subject to local channel characteristics and in-situ calibration. Real-time claims apply only to evaluated pilot and demo integrations.

Regulatory compliance and model validation expectations

Interagency guidelines and compliance bodies demand that technology verification remains auditable. The FINRA 2026 Annual Regulatory Oversight Report cautions that firms using AI-driven verification systems must implement vendor due-diligence and model performance audits. The FinCEN deepfake media alert (FIN-2024-DEEPFAKEFRAUD) and the NCUA AI resource point to the necessity of structured validation documentation in vendor risk assessments.

Vendor comparison at a glance

Evaluation DimensionValidated Telephony (G.711/AMR-NB)Out-of-Scope Channels (WebRTC/Zoom)
Validation MetricEqual Error Rate (EER) & minDCF trackedNot validated under active regimes
Calibration MethodBlind in-situ calibration sweepsStandard default thresholds only
Generalization BaselineASVspoof5 A17-A32 evaluationOpaque training set assumptions
Supervisory TraceExplainable sensor outputs loggedOpaque score outputs only

Frequently asked questions

What does physics-based anti-spoofing mean in plain terms?
It means measuring acoustic signal properties that live speech produces consistently and that synthetic pipelines often distort, then recording those measurements so a verdict is reviewable rather than a black-box score.
What is Equal Error Rate (EER)?
Equal Error Rate (EER) is the rate at which the false acceptance rate matches the false rejection rate. It is the primary metric used in biometric security and liveness verification evaluations.
Which codecs are included in the validation sweeps?
Our testing includes wideband audio, G.711 (u-law and a-law), and AMR-NB (narrowband telephony).
How does calibration protect against false-positives?
Local sweeps analyze the channel's background noise floor and acoustic characteristics, adjusting sensor thresholds so that normal caller speech does not trigger spoof alerts.

Official regulatory references

Ready to evaluate Sonotheia for your voice channel?

Request a demo