Resource
How to prove a voice call was synthetic without storing biometrics
You do not need a voiceprint to show a call was likely synthetic. You need measurements of the audio, documented thresholds, and a record an examiner can reconstruct. Sonotheia analyzes acoustic behavior in memory and discards the audio after the run, creating no voiceprints and no biometric templates. The output is a forensic risk event with acoustic tags, reason codes, and a decision trace rather than a standalone confidence score.
Zero biometric storage is a product constraint here, not a configuration toggle. That constraint is what makes the evidence portable into a supervisory file: nothing in the record identifies the speaker, so the artifact can be reviewed without inheriting a biometric-retention obligation.
What does proof mean in a regulated channel?
For banks, credit unions, broker-dealers, and RIAs, proof is not a model percentage. It is a package another person can review: what was measured and on which segment of the call, which documented threshold turned that measurement into a flag, what alternatives were ruled out (codec degradation, bandwidth limits, environmental noise), and which frozen baseline and calibration version were live when the decision was made. That is the difference between a detection score and a defensible decision.
How does Sonotheia measure without identity signals?
The served pipeline runs several independent sensors in parallel, each tracking a physical property of speech rather than who is speaking. They include voicing duration measured against human respiratory limits, room response and ambient consistency across the call, multi-scale permutation entropy of pitch and energy movement, pitch discontinuity, and an acoustic ensemble over hand-crafted features. Two further sensors record channel provenance and cast no vote. No sensor builds a speaker template, so nothing in the pipeline needs enrollment or consent workflows.
Are the models inspectable?
The verdict path uses non-neural classifiers over hand-crafted acoustic features, so a reviewer can trace a conclusion back to the measurements that produced it rather than to an opaque embedding. Features target properties of natural speech instead of the fingerprints of any particular synthesizer, which is why the approach is not trained against the specific attack generators it is meant to catch. Audio is processed in memory and discarded when the run ends.
How does this become FinCEN or FINRA documentation?
The same evidence path produces supervisory documentation: a decision trace linking verdict to measurements to thresholds, a calibration and baseline lock so historical decisions stay reproducible, a counter-hypothesis review before escalation, and evidence packages designed to support SAR narratives and examiner review. Detail on constructing those narratives is in the voice fraud SAR evidence guide.
Who is this for, and who is it not for?
It fits compliance, fraud, and risk teams accountable to auditors and regulators, especially where a voice decision authorizes money movement or account control. It does not fit teams that only need a score with no examiner-facing rationale, or programs that require enrolled voice biometrics as the primary control. Those programs need a biometric system, and this is deliberately not one.
What should you ask any vendor, including us?
Do you store voiceprints or raw audio after analysis? Can an investigator follow measurement to threshold to verdict without a data-science escort? Can you show the exact production baseline that was live when the decision was made? Does the evidence package map to the supervisory documentation your counsel already uses? Ask us the same four and hold the answers to the same standard.
Frequently asked questions
- Can you prove a call was synthetic without a voiceprint?
- You can show it was likely synthetic. Acoustic measurements, documented thresholds, and a reconstructable decision trace do not require a speaker template, because they describe the audio rather than the identity of the speaker.
- Do you store voiceprints or raw audio after analysis?
- No. Audio is processed in memory and not retained after analysis. We do not create or store voiceprints or biometric templates.
- Does this replace enrolled voice biometrics?
- No, and it is not meant to. If your program requires enrolled voice biometrics as the primary control, you need a biometric system. This measures the audio instead of the speaker.
- What does an examiner actually receive?
- A decision trace linking the verdict to the measurements and the thresholds that applied, the calibration and baseline version live at the time, and the counter-hypotheses considered before escalation.
Official regulatory references
Ready to evaluate Sonotheia for your voice channel?
Request a demo