RankShield
RANKSHIELD NETWORK Get started

Deepfake Voice Scams: How Business Owners Beat a CEO Voice Clone

AI can clone your CEO from seconds of audio, and you cannot catch it by ear. Here is the callback-and-authorization control model that stops the wire anyway.

August 28, 2026 · 12 min read · deepfake voice scam protection business
Share

A voice on the phone that sounds exactly like your CEO, calling to confirm an urgent wire transfer, is no longer a hypothetical. AI can now clone a convincing version of anyone’s voice from a short sample of audio, the kind that is trivially available from a conference talk, a podcast, a webinar, or a voicemail greeting. The FBI has started tracking this directly: in its 2025 report it named AI-related fraud as a formal category for the first time, roughly 893 million dollars across more than 22,000 complaints, and it specifically flagged voice cloning being layered into business email compromise, used to place a follow-up call that appears to come from a CFO or CEO to reinforce a written wire instruction (FBI IC3 2025 Annual Report1). The uncomfortable truth this guide is built around is that you cannot reliably tell a good clone from the real person by listening, so the defense cannot be detection. It has to be a process that does not depend on trusting the voice at all. I run a security company and work with business owners on exactly this, and the honest framing up front is that no tool detects every deepfake and anyone selling you one is overpromising. What actually works is a small set of verification controls that make a convincing voice worthless to an attacker, and this post lays them out.

How common are deepfake voice scams against businesses?

Common enough that the FBI created a category for it, and growing. Business email compromise has been the costliest quiet crime in business for years, and it is now being upgraded with AI. In its 2025 report the FBI put business email compromise losses at 3.046 billion dollars for the year, and separately logged AI-related fraud as a new formal descriptor at roughly 893 million dollars across more than 22,000 complaints (FBI IC3 2025 Annual Report1). Inside that, the bureau specifically described voice cloning being used to place follow-up calls impersonating executives to confirm wire transfers, with the confirmed AI-component cases alone accounting for more than 30 million dollars.

The case that put this on every board agenda was Arup, the global engineering firm, which lost 25.6 million dollars in early 2024 when a finance employee in its Hong Kong office was walked through 15 transfers during a video call where every other participant, including a convincing version of the company’s UK chief financial officer, was an AI-generated deepfake (CNN Business2). The money was moved in a single day and has not been recovered. Arup was a full video deepfake, which is still expensive and rare, but it proved the model works, and the cheaper version, a voice-only clone reinforcing an email, is now within reach of ordinary criminals.

That is the part small and mid-sized businesses underestimate. You do not need to be a target worth 25 million dollars, and the attacker does not need a Hollywood video call. A cloned voice on a phone, confirming an email your bookkeeper already half-believes, is enough to move a five-figure or six-figure wire, and five-figure wires out of a small business are exactly the losses the FBI’s numbers are built from. The technology has commoditized; the defense has to as well.

THE PATTERN

What a deepfake voice scam actually looks like

StageWhat the attacker doesWhat they are counting on
SetupClone the exec voice from public audioA podcast, talk, or voicemail is enough
PretextEmail an urgent, confidential paymentYou not wanting to question the boss
ReinforceA cloned-voice call confirms the emailYou trusting the voice you recognize
PressureUrgency plus secrecy, act nowYou skipping your normal checks
PayoutWire to a mule account, then goneNo callback and no dual approval

Every stage assumes you will trust the surface. The controls below break the chain at the reinforce and payout stages.

Why can’t you just detect a deepfake voice by ear?

Because modern voice cloning reproduces the specific things you use to recognize someone, and the tells that used to give it away are disappearing. A few years ago cloned speech had a robotic flatness, odd pacing, or missing emotion. Current tools capture timbre, accent, cadence, and filler words well enough that a short call, especially over a compressed phone connection that already strips detail, does not give your ear enough to work with. The human brain is built to accept a familiar voice as proof of identity, and that instinct is exactly what the attack rents.

Detection technology exists, but it is not something you can rely on in the moment a call comes in. Deepfake-audio detectors work probabilistically, they lag behind the generation tools that are improving monthly, and they are not sitting on your accountant’s desk phone flagging calls in real time. Treating detection as the defense puts you in a race you cannot win, because the attacker only has to beat your ear or your detector once, on the one call that matters, while you have to catch every attempt. Any vendor promising to reliably detect every deepfake is selling confidence, not protection.

This is why the entire defense has to move off the question can I tell if this voice is real. That question is a trap. The right question is can this caller complete the request without passing a check that a voice clone cannot fake, and the answer to that one is entirely in your control. A clone can reproduce your CEO’s voice; it cannot call you back on the number you already have for them, it does not know your internal code word, and it cannot be the second approver on a payment. The defense is built on things the attacker does not have, not on things you hope to notice.

What verification controls stop a CEO voice clone?

Three process controls, layered, make a perfect voice clone worthless. None of them requires you to detect anything. The first is an out-of-band callback: any request to move money or change payment details is verified by calling the person back on a number you already have on file, never a number they give you on the call or in the email. A cloned voice cannot answer your CEO’s real phone. This one control, applied without exception, defeats the entire category, because the attack depends on you acting on the inbound contact rather than reaching the real person independently.

The second is a verbal code word or challenge phrase agreed in advance for financial requests, known only to the people who authorize payments and never written in email or Slack. When a call comes in asking to release a wire, the person asks for the phrase. The real executive knows it; the clone does not. It feels slightly theatrical the first time and then becomes routine, and it is the cheapest high-assurance check a small business can deploy, because it costs nothing and works even if the callback somehow reaches a compromised line. The third is dual authorization: no single person can move money or change vendor bank details alone, above a threshold you set. Two people, verifying independently, means the attacker has to defeat two humans and both of the controls above, which collapses the success rate.

Around those three, add the cultural control that matters most: make it explicitly safe, and expected, to slow down. The scam runs on urgency and secrecy, a confidential deal that must close now, do not tell anyone, just get it done. State plainly, from the top, that no real executive will ever be angry at an employee for verifying a payment request, and that a demand for secrecy or speed is itself the signal to verify harder. As we covered in the business owner’s playbook against SIM swaps, the same principle holds across identity attacks: the control that works is an independent second channel, not a sharper eye. The diagram below shows exactly where each control breaks the attack.

DOWNLOADABLE INFOGRAPHIC

Where each control stops the fake-CEO wire

RANKSHIELD // THE CLONE CANNOT PASS A CHECK IT DOES NOT HAVE CLONED VOICE "release the wire" CALLBACK call the KNOWN number clone cannot answer it CODE WORD ask the agreed phrase clone does not know it DUAL APPROVAL two people, independent must both be fooled Each gate fails the attacker on its own: No callback path The real exec answers, or the number does not connect. Request dies. Wrong / no phrase A voice can be cloned; a shared secret spoken aloud cannot. Request dies. Second approver Two humans verifying independently collapse the odds. Request dies. Voice cloning is now layered into BEC to confirm wires (FBI IC3 2025). Arup lost $25.6M to a deepfake call (2024). Defense is PROCESS, not detection. No tool catches every deepfake; a callback the clone cannot answer does.
The attack needs every step to land. A callback, a code word, or dual approval each break the chain on their own. Free to share with attribution.

How do you roll this out without slowing the business down?

Write it as a one-page payment-verification policy, apply it only where money moves, and make it a point of pride rather than a burden. The controls sound heavy in the abstract, but in practice they touch a narrow slice of activity: outbound wires, changes to vendor or payroll bank details, and any payment request that arrives with urgency or secrecy. Everything else runs normally. Define a dollar threshold above which dual authorization is mandatory, set the rule that bank-detail changes always require an out-of-band callback regardless of amount, and agree the code word with the handful of people who authorize payments. That is the whole program, and it fits on a single page.

The most important move is cultural, and it is free: give your finance and admin staff explicit, standing permission to say no and to verify, even when the request appears to come from the owner. Most successful attacks in this category succeed not because the technology fooled someone, but because an employee did not feel able to slow down a demand that seemed to come from the top. Tell them directly that you would rather have a wire delayed by an hour than lost forever, that you will never penalize a verification, and that a request pushing them to skip the checks is the reddest flag there is. An employee who feels safe pausing is worth more than any detector.

This is also where the RankShield philosophy and this problem meet, honestly stated. No product of ours, and no product anyone sells, will stop a cloned voice from reaching your bookkeeper’s phone, so this defense is process, run by your people. What we build is the same idea applied to software: verifiable AI security is about not trusting that an action is authentic because it looks or sounds right, but requiring it to prove itself against something an impostor cannot forge. A callback the clone cannot answer, a code word it does not know, a second human it cannot become: those are the human version of the same principle, and against a synthetic voice they are the version that actually protects the money.

EXPOSURE CHECK

Would a voice clone get a wire out of your team?

  1. How are urgent wire requests from leadership verified?
  2. Do vendor or payroll bank-detail changes require out-of-band verification?
  3. Is there a verbal code word for financial requests?
  4. Does moving money above a threshold need two approvers?
  5. Would staff feel safe telling the "CEO" to wait while they verify?

What actually protects you from a CEO voice clone?

Not your ear, and not a detector. The technology to clone a convincing version of your CEO or CFO from a few seconds of public audio is here and cheap, the FBI is now tracking it as a formal category, and it is being bolted onto the business email compromise playbook precisely because a familiar voice is what makes a fake wire request feel real. You cannot reliably tell a good clone from the real person by listening, so any defense built on detection is a race you will eventually lose, on the one call that counts.

What works is a small, boring set of controls that do not care whether the voice is real: an out-of-band callback to a number you already have, a verbal code word the clone does not know, dual authorization so no one person can move money alone, and a culture that makes it safe and expected to slow an urgent request down. Those controls fit on one page, touch only the moments where money moves, and defeat the entire category because they require the caller to prove something an impostor cannot forge. That is the same principle we build into software as verifiable AI security: do not trust that something is authentic because it looks or sounds right, make it prove it. Against a synthetic voice, that principle is the difference between a delayed wire and a lost one.

FREQUENTLY ASKED

Questions, answered.

Jamie Kloncz
Jamie KlonczCEO, RankShield · online

Can AI really clone my CEO’s voice, and how much audio does it need?

Jamie Kloncz

Yes, and it needs surprisingly little. Modern voice-cloning tools can produce a convincing version of a specific person’s voice from a short sample, often just seconds to a minute of clear speech, capturing their timbre, accent, pacing, and characteristic phrasing. For most executives that sample already exists in public: a conference talk, a podcast appearance, a webinar recording, a promotional video, an earnings call, or even a voicemail greeting. An attacker does not need to breach anything to get it; they just need to find your CEO or CFO speaking online, which for anyone in a leadership role is usually trivial. The quality is good enough that over a phone call, where the connection already compresses and degrades audio, the clone loses very little of what would otherwise give it away. This is why the old advice to listen for a robotic or unnatural voice no longer holds. The practical takeaway is that you should assume your executives’ voices can be cloned and build your payment controls on that assumption, rather than hoping an attacker will not bother or that the result will sound obviously fake.

How can I tell if a voice on the phone is a deepfake?

Jamie Kloncz

Honestly, you usually cannot, and building your defense around trying to is a mistake. A few years ago cloned speech had tells, flat emotion, odd rhythm, strange pauses, but current tools have largely closed those gaps, and a phone connection strips away the fine detail your ear would need anyway. Detection software exists, but it works probabilistically, lags behind rapidly improving generation tools, and is not running on your staff’s desk phones flagging calls as they come in. So the reliable answer to is this voice real is not something you can produce in the moment. The good news is that you do not need to answer that question at all. Instead of trying to detect the clone, you verify the request through a channel the clone cannot use: call the person back on a number you already have on file, ask for a code word agreed in advance, and require a second approver for money movement. A voice clone can sound perfect and still fail every one of those checks, because none of them depend on how the voice sounds. Shift your effort from detection, which you will lose, to verification, which you control.

What is an out-of-band callback and why does it stop these scams?

Jamie Kloncz

An out-of-band callback means verifying a request by contacting the person through a separate, trusted channel that you initiate, rather than responding on the same channel the request arrived on. In practice: when a call, email, or message asks you to move money or change payment details, you hang up or set it aside and call the executive or vendor back on a phone number you already have on file from before this request, not a number they provided during the interaction. It stops deepfake voice scams because the entire attack depends on you acting on the inbound contact. A cloned voice can call you, but it cannot answer your CEO’s real phone when you call the number you already had. If the caller is genuine, they pick up and confirm in ten seconds; if it was an impostor, your callback reaches the real person, who tells you they never made the request, and the scam collapses. The one rule that makes it work is that the callback number must come from your own records, never from the suspicious message, because attackers will happily supply a number that rings their own phone. Applied without exception to every payment and bank-detail change, this single control defeats the whole category.

What is a verbal code word and is it really worth setting up?

Jamie Kloncz

A verbal code word is a phrase agreed in advance among the people who authorize payments, used to confirm identity on financial requests, and it is one of the highest-value, lowest-cost controls a business can put in place. The idea is simple: you and your finance staff agree on a word or phrase that is never written down in email, chat, or documents and is known only to that small group. When a call comes in asking to release a wire or change bank details, the person handling it asks for the phrase. A real executive knows it; a cloned voice, no matter how convincing, does not, because the secret was never spoken in any channel an attacker could capture. It is worth setting up precisely because it defends against the thing detection cannot: it does not matter how perfect the voice is if the caller cannot produce the shared secret. It costs nothing, takes minutes to establish, and works even in the rare case where an attacker has somehow compromised a phone line. The only discipline it requires is keeping the phrase out of writing and changing it if you suspect it has leaked. It feels slightly awkward the first time someone asks the boss for the code word, and then it becomes a normal, respected part of how your team protects the company’s money.

Are small businesses actually targeted, or only big companies like Arup?

Jamie Kloncz

Small and mid-sized businesses are very much targeted, and in many ways they are the core of the problem rather than the exception. The Arup case, a 25.6 million dollar loss to a deepfake video call, made headlines because of its size and sophistication, but it represents the high end of the spectrum. The everyday version is far cheaper to run and aimed squarely at smaller organizations: a cloned voice on an ordinary phone call, reinforcing an email your bookkeeper already half-believes, to move a five- or six-figure wire. The FBI’s billions in annual business email compromise losses are built largely from exactly these mid-range thefts, not a handful of giant ones. Smaller businesses are attractive because they often lack formal payment-verification controls, concentrate approval authority in one or two people, and have close, trusting cultures where questioning a request that appears to come from the owner feels inappropriate. Those are the conditions the scam exploits. The reassuring flip side is that the defenses do not require an enterprise budget: a callback rule, a code word, and dual authorization are free or nearly free, and they are more effective for a small team than any technology, because they remove the single points of trust that attackers rely on.

Does any software or tool stop deepfake voice fraud?

Jamie Kloncz

No single tool stops it reliably, and you should be skeptical of any product that claims to. Deepfake-audio detectors exist and can help as one signal, but they are probabilistic, they trail the generation tools that improve every month, and they are not practically deployed on the phone lines where these attacks land, so they cannot be your primary defense. The durable protection is process, not a product: an out-of-band callback to a known number, a verbal code word, dual authorization on money movement, and a culture that makes verifying a payment safe and expected. Those controls work regardless of how convincing the voice is, because they never rely on judging the voice at all. Where technology genuinely helps is a layer up, in reducing the account-takeover and email-compromise footholds that make these scams credible in the first place, so keeping strong multi-factor authentication, watching for mailbox rule changes, and protecting executive accounts all matter. But for the specific moment when a voice that sounds like your CEO asks to release a wire, the thing that saves you is a verification step the clone cannot pass, run by a person who feels safe insisting on it. Treat any tool as support for that process, never as a replacement for it.

Try one of the suggested questions above.

References

  1. FBI Internet Crime Complaint Center (IC3). 2025 Internet Crime Report (BEC losses $3.046B in 2025; AI-related fraud a new category, ~$893M across 22,000+ complaints; voice cloning layered into BEC to confirm wire transfers).
  2. CNN Business. Arup revealed as victim of $25 million deepfake scam involving Hong Kong employee (deepfake video call, HK$200M / $25.6M across 15 transfers, funds unrecovered; May 16, 2024).
Jamie Kloncz
WRITTEN BY

Jamie Kloncz

Founder & CEO, RankShield

Jamie Kloncz is the founder and CEO of RankShield, the verifiable AI and quantum security platform. He started the company after two attacks landed in a single week: his phone was cloned, and his business was hit by a click-fraud campaign. One targeted him as a person, the other his livelihood, and no single tool defended both. That experience, together with surviving an AI voice-clone scam, shaped RankShield’s core belief: the threats of the AI age are personal first, and trust should be something you can check, not just extend.

Make every AI action provable.

RankShield is the verifiable, quantum-safe AI security platform — protection you can check, not just trust.