Select Language

Choose your language

株式会社ヤグラ

Select Language

Choose your language

株式会社ヤグラ

Select Language

Choose your language

Insight

Executive Voices Cloned in 3 Seconds: AI Voice Scams and "90-Second OOB Verification"

In early 2024, an incident occurred in Hong Kong where HK$200 million (approx. US$26 million) was transferred in 15 transactions across 5 accounts immediately following a video conference. Many of the meeting participants were deepfakes, and public sources (info.gov.hk) suggest the meeting itself may have been pre-recorded. The assumption that "attending a meeting equals verified identity" is no longer valid. This article analyzes the playbook of multi-participant conference scams driven by voice cloning, breaking it down into 5 stages: Recon → Lure → Meeting → Authorize → Transfer. Based on Yagura's philosophy of "verifying" processes rather than trying to "detect" fakes, we propose a 90-second protocol centered on Out-of-Band (OOB) two-path verification. We also examine why human eyes and ears are unreliable, the technical background of generating voices from just 3 seconds of audio, blind spots in Japanese approval processes, and an implementation checklist. Generative AI-driven attacks will become increasingly sophisticated, and lower attack costs will drive an exponential rise in incidents. Companies failing to leverage generative AI for defense face heightened risks of elimination. Yagura is committed to addressing this challenge head-on.

AI SOCとは? 仕組み・従来型SOCとの違い

1. Why the "Meeting" Became the Decisive Blow

Method: A meeting is called under the CFO's name, making it appear as though multiple people are present. Long monologues, screen sharing, and muting are used to suppress two-way interaction, hiding rough lip-syncing and lag.
Result: 15 transfers split across 5 accounts. Subsequent reports identified the victim company as Arup in the UK.
Lesson: Believing "attending a meeting = identity verified" is a mistake. Meetings without live questions or out-of-band (OOB) dual-channel verification only serve to reinforce authority bias.

2. Why "Intuition-Based Detection" Fails

2-1. Voice Differentiation is Unreliable

Studies show that human voice differentiation accuracy is unstable, and the improvement effect of training is limited. Relying on ears for defense design is inherently fragile.

2-2. Faces are Trusted More Than the Real Thing

Experimental results show that AI-generated faces are rated as more trustworthy than real ones. A screen lined with familiar colleagues maximizes this bias.

2-3. Protect with "Processes," Not "Gut Feelings"

The core defense is to integrate OOB dual-channel verification, cool-down periods, and mutual approval. Ensure security through procedures, not human senses.

3. The Attack Chain (Recon → Lure → Meeting → Authorize → Transfer)

3-1. Recon (Asset Gathering)

Voice and face samples are harvested from speaking presentation videos and webinars. With TTS research (VALL-E series), voice likeness can be cloned from just 3 seconds of audio, requiring minimal source material (Microsoft / arXiv).

3-2. Lure (Inducement)

The attacker creates information asymmetry using the CFO's name to emphasize "confidentiality" and "urgency," with phrases like "see attachment" or "details will be explained in the meeting."

3-3. Meeting (Faked Meeting)

The presence of multiple participants creates artificial "social proof." If pre-recorded video is used with overlaid audio, two-way interaction can be avoided entirely.

3-4. Authorize (Bypassing Approval)

Attackers exploit the false assumption of "verified during the meeting" to bypass internal controls. Pressure tactics like "immediate" and "strictly confidential" prevent cool-down periods.

3-5. Transfer (Remittance)

Funds are split and sent to multiple accounts to scatter the trail (15 transfers across 5 accounts in the aforementioned case). Recovery depends on a minute-by-minute initial response.

4. Technical Reality: Why a Voice Can Be Cloned in 3 Seconds

Neural Codec LM (VALL-E series): This method converts audio into discrete codes and generates speech using a language model. It can replicate voice quality and prosody (intonation) from a 3-second prompt audio.
Cross-lingual: VALL-E X retains voice quality even in other languages, making it natural-sounding even when a foreign executive gives instructions in Japanese.
Implications: Identity verification relying solely on voiceprints is no longer viable. Organizations must switch to short codephrases and verification through separate channels.

5. Three Common Blind Spots in Japanese Corporate Approval Workflows

"Meeting = Identity Verification" Fallacy: A meeting without two-way verification does not constitute identity verification. The Hong Kong government has explicitly highlighted cases involving "pre-recorded meetings."
Incomplete Out-of-Band (OOB) Verification: Establish policy-mandated OOB verification, such as telephone callbacks, as recommended by FBI/CISA.
Lack of Cool-down for Large/Exceptional Transactions: It is critical to institutionalize a minimum 5-minute pause and mutual approval process.

6. The Reality of Victim Scale

IC3 2024: Losses reached $16.6B (+33% YoY), with BEC accounting for $2.77B.
FFKC (Financial Fraud Kill Chain): The success rate is 66%. Here too, a minute-by-minute initial response is key.
Intermediary Destinations: The UK and Hong Kong frequently appear as destination countries.

Declining attack costs are driving up the scale of damage. Defenders must leverage generative AI to automate and accelerate defenses to keep pace.

7. A 90-Second "OOB Dual-Channel" Protocol (For Meetings & Calls)

Policy: Do not rely on intuition. Secure the transaction through "processes," not "likeness."

7-1. Pre-Meeting Prep (~30 seconds)

Complete preparation within 30 seconds. First, update codephrases monthly, anchoring them to company-specific information. Next, restrict callbacks to a whitelist of numbers registered in the internal directory, and formally document a 5-minute cool-down period in policies. Finally, insert short, unannounced live questions at the start of meetings to verify real-time responses.

7-2. During the Meeting (~90 seconds)

Within 90 seconds during the meeting, perform a simultaneous dual-channel verification using the codephrase alongside a separate channel like an internal phone or enterprise chat. Simultaneously state, "We will perform OOB verification per policy. We will pause for 5 minutes," to declare a release of time pressure. Concurrently, check environmental metadata (background, shadows, voice quality, presence of interruptions) and ask high-specificity questions regarding funding sources, transaction IDs, and internal codes. If the response is vague, suspend the meeting immediately and call back using a known number.

7-3. Phone Calls (Assuming Voice Cloning)

For phone calls, never accept incoming calls at face value; always call back using a known number. The baseline premise is to secure the connection through process rather than trying to evaluate "if the voice sounds like the person" by ear.

7-4. Initial Response After Transfer

If a transfer is accidentally made, request FFKC from your financial institution immediately (66% success rate).

8. Addressing Objections (For Internal Alignment)

"Aren't detection tools enough?": Detection is only the final safety net. The primary battlefield lies in processes like OOB, cool-down periods, and mutual approval.
"I can tell by face and voice": Voice differentiation is unstable, and AI-synthesized faces are actually prone to overestimation. Protect with procedures, not intuition.
"I cannot disobey urgent instructions from superiors": The solution is to formalize "Urgent = Yellow Flag" in policies, creating a culture where anyone can say, "I will pause for 5 minutes for OOB verification."

9. Implementation Checklist (Quick Start)

Policy: □ Mandatory OOB / □ Callback rules / □ 5-minute cool-down / □ Mutual approval by two people
Process: □ Live questioning at meeting start / □ Monthly codephrase rotation
People (Training): □ Simulated AI meeting exercises / □ BEC roleplay
Technology: □ OOB button integration in approval workflows / □ Log retention
Response: □ Financial institution contact templates / □ Posted initial response procedures

10. FAQ

Q1. How many seconds of voice audio are needed to pose a threat?
A. Clones can be created with just 3 seconds of audio. Ban voice-only identity verification, and switch to codephrases combined with separate channel verification.

Q2. Is there any evidence of "realistic" meetings being used for fraud?
A. Official cases exist involving AI voice overlays on pre-recorded video. Meetings without two-way verification should be treated as unverified.

Q3. What is the first step if something feels suspicious during a meeting?
A. Declare, "We will pause for 5 minutes for OOB verification," and call back using a number from the internal directory.

Q4. How does this relate to BEC?
A. Both are attacks targeting decision-making processes. IC3 2024 reports $16.6B in losses, with BEC at $2.77B and a 66% FFKC success rate. Response speed and OOB verification are key.

Q5. Which regions require caution for outbound transfers?
A. The UK and Hong Kong frequently appear as intermediary destinations. Note that they are less likely to raise suspicions because they are common business partners.

11. Conclusion: The Challenge is the "Verification System," Not the "Ability to Detect Persuasiveness"

A voice created in 3 seconds, the authority of a meeting, and time pressure. When these three align, human intuition fails. That is why defense priorities must shift to processes: OOB dual-channel verification, cool-down periods, and mutual approval. If an incident occurs, initiate FFKC immediately.
In an era where attacks are highly sophisticated, low-cost, and growing exponentially, defenders must adopt generative AI for automation and operational strengthening. Organizations failing to leverage generative AI for defense risk being phased out of the market. Yagura directly addresses this reality, implementing "90-second OOB verification" in your operations.

Related Services: Click here to learn more about Yagura Awareness, which builds employee and executive resilience through multi-channel training, including vishing (voice) and deepfakes.

ヤグラAIセキュリティ

丸わかり資料を

無料でダウンロード

生成AI時代に求められるサイバー環境の変化や

サービスの概要資料についてお送りいたします。

ヤグラAIセキュリティ

丸わかり資料を

無料でダウンロード

生成AI時代に求められるサイバー環境の変化やサービスの概要資料についてお送りいたします。

ヤグラAIセキュリティ

丸わかり資料を

無料でダウンロード

生成AI時代に求められるサイバー環境の変化やサービスの概要資料についてお送りいたします。