A voice that sounds exactly like your kid, your parent, or your boss can now be built from a few seconds of source audio pulled off a voicemail greeting, an old video, or a recorded call. 3 seconds (of clean audio is enough today to clone a voice convincingly, per multiple 2026 security reviews) The FBI's Internet Crime Complaint Center counted $893 million in 2025 fraud losses with confirmed AI involvement, 40% of it — $352 million — from people 60 and older. The instinct most people reach for when a call feels wrong is to listen closely for a robotic tone; that is not the check the FTC actually recommends, and on a live call it mostly can't be: the tools built to catch AI-generated audio check a saved file, not a voice still talking to you. Here's what actually still works, and in what order.
Why listening for the robot voice stopped working
The three-second benchmark isn't marketing copy — it's the plain mechanics of how a modern voice-cloning model works: feed it a short, clean sample and it can hold a live back-and-forth in that voice, with a real person typing responses behind the scenes. What used to give a fake away — flat delivery, mistimed pauses, a slightly wrong cadence — narrows with every new model generation, and what's left gets erased further by ordinary compression: a call routed through a cell network or re-encoded by a messaging app loses exactly the fine detail a detector needs. Independent reviews of today's voice-AI detectors are blunt about it: vendor accuracy claims come from favorable test sets, and real-world accuracy on short, re-encoded, or unfamiliar audio cuts sharply. And the tools that do exist check a file someone hands them after the fact — real-time detection during a live call exists only inside call-center-grade enterprise systems, not on the phone in your hand.
The check that actually works here isn't better hearing. It's a callback.
What to do the moment a call asks you for money or access
- This is the FTC's own guidance for exactly this call, not 'listen closely.' A clone can be flawless; a callback to a number the scammer doesn't control ends the scam in one step, regardless of how good the voice sounds.
- A cloned voice repeats what a scammer types; it does not know a private detail nobody has posted. The FBI and FTC both recommend agreeing on a family code word before an emergency happens, not during one.
- Wire transfers, gift cards, and cryptocurrency get demanded because none of them can be reversed — a real emergency almost never requires exactly one of those three, right now, from you specifically.
- ElevenLabs began embedding an inaudible SynthID watermark in audio generated on its platform in June 2026, readable with its free public detector. A clean result only rules out that one platform's output made after that date — it proves nothing about a live call, an older clip, or any other voice tool.
- The FTC (ReportFraud.ftc.gov) and the FBI's IC3 (ic3.gov) both take reports from calls that didn't succeed — that data is what let investigators total AI-enabled losses at $893 million for 2025, and it's the evidence base future protections get built on.
Which of those matters first depends on what's actually happening — a live call under pressure needs a different first move than a recording someone forwards you after the fact.
Start with what you actually have
What a watermark can and can't tell you
Two checks, and only one of them scales to a live call
| SynthID watermark ElevenLabs audio only | Listening for tells the old advice | |
|---|---|---|
| What it actually checks | Whether a saved file's audio pattern matches ElevenLabs' own generation signature | Whether the voice sounds natural to a human ear |
| Works on a live phone call | No — requires a saved file run through the detector afterward | In theory, but modern clones increasingly defeat it |
| Coverage | Only ElevenLabs-generated audio made after June 2026 — other cloning platforms aren't covered | Applies, unreliably, to any voice from any source |
| What a clean result proves | Nothing about who's on the phone right now — only that no ElevenLabs watermark was found in that file | Nothing reliable — vendor accuracy claims come from favorable test sets and real-world accuracy cuts sharply outside them |
None of this closes the gap — it names it. Google DeepMind built the SynthID watermark ElevenLabs adopted, but only a handful of the platforms capable of cloning a voice have adopted any watermark at all, and none has volunteered a way to check a call in progress. In April 2026, Senator Maggie Hassan sent oversight letters to four voice-cloning companies — ElevenLabs, LOVO, Speechify, and VEED — asking whether any of them verify consent before cloning a voice, watermark their output, or report misuse to law enforcement; as of this writing, the answers aren't public. Until that changes, the defense that actually works predates all of it: call back on a number you already had.
Where this goes wrong
Five ways this check gets skipped when it shouldn't be
This is the third of a kind, after our guide to checking whether an image is AI-generated and our guide to checking whether text was written by AI — voice is the modality where the honest answer runs closest to don't rely on detection at all, because the fastest-growing use is a live call no file-based tool can reach. The dictionary has short entries for terms used here — jailbreak, guardrails — if any of the underlying mechanics need unpacking further. San Francisco's own order forcing Apple and Google to pull deepfake "nudify" apps, covered here, is the same underlying technology aimed at a different harm.
- AI can clone a familiar voice from three seconds of audio, security researchers and the FBI both confirm.
- The FBI counted $893 million in 2025 fraud with confirmed AI involvement, 40% from people over 60.
- ElevenLabs began watermarking its AI audio in June 2026 — but only its own platform's output.
- Detection tools check saved audio files; none can check a live phone call while it's happening.
- Caveat: no detector proves a call is safe — verify independently instead of listening for tells.