Zendoric
← Back to the day · September 4, 2026

Three seconds of audio is all a voice clone needs — and every defense we have arrives too late

🕒 Published on Zendoric: September 4, 2026 · 09:12

✨ AI-generated · how it's made

The FBI has, for the first time in 26 years, given AI-driven fraud its own category: 22,000 complaints and $893 million lost in 2025, with $352 million of it taken from people over 60. The technical barrier collapsed to three seconds of recorded speech. Our thesis: this is not a model-capability problem, it is an identity-verification problem — and the industry's safeguards are built to investigate crimes, not prevent them.

The number that should anchor this story comes from the FBI's Internet Crime Complaint Center. In its 2025 annual online crime report, published in April 2026, the bureau broke out "AI-driven fraud" as a separate category for the first time in the report's 26-year history: more than 22,000 complaints and over $893 million in losses. Victims aged 60 and over accounted for $352 million of that — roughly 39% of the total from a single age bracket. Averaged across complaints, that is about $40,000 stolen per reported case, which is not petty crime. The FBI itself notes the figures only cover fraud that victims recognized and reported.

Engineer Tim Green, writing at SmarterArticles and cited by GIGAZINE, argues that even that caveat understates the problem. Because most people who receive an AI-cloned voice call never learn that AI was involved, he treats the $893 million as a floor, not a ceiling. His illustrative case is Sharon Brightwell, a Florida resident who in July 2025 received a call from what sounded like her daughter crying, claiming to have hit a pregnant woman while driving and to have had her phone seized by police. A second voice, posing as a lawyer, asked for $15,000 in bail and — the operationally clever part — warned her not to state the purpose of the withdrawal at the bank, because it could damage her daughter's reputation. Within an hour she had handed cash to a courier. The daughter had never been in an accident.

The technical core is what makes this different from decades of grandparent scams. Modern voice cloning needs roughly three seconds of audio to produce speech that is hard to distinguish from the original. Three seconds is a voicemail greeting, a podcast clip, a single Instagram story. As Green puts it, one grandchild appearing in one TikTok video supplies everything a scammer needs. The tools are also cheap and plentiful: Consumer Reports has found that most voice-cloning products — it names Descript, ElevenLabs, Lovo, PlayHT, Resemble AI and Speechify — lack effective safeguards against misuse.

ElevenLabs, the most prominent of them, is not doing nothing. It bans impersonation in its usage policy, publishes a classifier that flags audio likely generated by its own system, links generated content back to the account that made it, and blocks cloning of protected individuals during election periods. Green's critique is not that these measures are weak but that they are structurally the wrong shape: almost all of them are reactive. They help investigators trace a voice after the money is gone. None of them stop a three-second clip from becoming a clone in the first place. His explanation for why is uncomfortable and probably correct — a fast-moving, competitive market will not impose friction on itself if that friction costs it customers.

Our read: this is the short-term AI harm we keep saying deserves more attention than distant superintelligence scenarios, and it follows the pattern we have flagged before — generation costs collapsed, verification costs did not. Cloning a voice is now essentially free; confirming a voice is real still costs a phone call, a code word, an hour of doubt. Fraud rushes into exactly that gap. It also explains why the elderly are targeted first: higher average savings make each successful call more profitable, and, as Green notes, telling someone a hundred times that voices can be faked does nothing once a voice that sounds like their grandchild is asking for help. Awareness campaigns are not a defense against a stimulus that bypasses reasoning.

Where we differ from the doom framing is on the remedy. The capability itself is genuinely valuable — the same synthesis restores speech to people who have lost it and makes content accessible across languages, and banning it would be both futile and costly. The fixable layer is identity. Green's proposal, requiring verified identity from anyone who creates a clone, is the right instinct: govern the point of generation, not the existence of the technology, and make cloning traceable to a person before the audio exists rather than after the savings are gone. Banks can add out-of-band confirmation for urgent cash withdrawals; families can agree on a passphrase tonight, for free. Long term, we still think AI ends up expanding what people can do and how long they stay healthy. Getting there honestly means admitting that the transition includes a wave of very effective crime, and building the boring plumbing — verified identity, provenance, friction where money moves — before the next $893 million becomes a rounding error.

🔗 Related on Zendoric

Sources & references