By Anat Goldstein and Farook Sattar
A finance worker at a multinational firm received what appeared to be a legitimate call from someone claiming to be the company’s CFO. The person receiving the call ended up transferring $25 million. The reality: The CFO’s voice was AI generated. The CFO never made that call.
That incident is no longer an outlier. Voice fraud in banking increased by roughly 30 percent in 2025, with AI-powered synthetic voices capable of fooling even experienced professionals. Synthetic voice attacks against financial institutions rose sharply in 2024, increasing by a factor of approximately 20 compared to the prior year, as data from the same period indicated that roughly one in 750 banking calls was flagged as potentially fraudulent.
Commercial AI voice-cloning tools have significantly reduced the barrier to entry for attackers. As synthetic voice technology becomes more affordable and accessible, attacks can scale more easily and adapt in real time. These capabilities support voice cloning, deepfake impersonation and synthetic identity schemes. While such attacks are not yet universal across the industry, reported incidents are increasing rapidl, and many industry observers expect them to become a major concern for banks by 2030.
Traditional methods of detection, rules-based systems, statistical modeling and manual review have struggled to keep pace. Rules-based systems often produce high false-positive rates. Statistical approaches can be time-consuming and resource-intensive. Manual review cannot scale quickly enough. Many existing workflows also lack the capacity for real-time analysis of the large volumes of voice data involved.
Banks have begun warning customers about increasingly sophisticated impersonation attempts and customers themselves now expect stronger fraud protection. The practical question for many institutions is no longer whether they will encounter synthetic-voice attacks, but how effectively they can distinguish genuine callers from generated ones under everyday phone conditions.
Where current defenses are losing ground
A core concern for many banks is the erosion of trust in the voice channel. Surveys indicate that roughly 85 percent of consumers no longer answer calls from unknown numbers, effectively limiting the usefulness of outbound and inbound voice contact. Call-center impersonation, spoofing and deepfake voice fraud are frequently cited as contributing factors.
Branded calling, which displays the caller’s identity and purpose before the phone is answered, improves transparency, yet it addresses only legitimate caller visibility. It does not prevent deepfake impersonation once a call is connected, nor does it stop social-engineering tactics that follow. From a commercial standpoint, the pricing of branded-call services and related analytics reflects that trust in the voice channel has become both a security issue and a revenue issue.
Some carriers and technology providers are shifting emphasis from post-call detection to earlier trust signals, embedding verified identity at the network level so that authentication information is visible before the call is answered. In the United Kingdom this approach has reached nearly full subscriber coverage, illustrating how infrastructure-level solutions can operate at scale. Pre-call validation of this kind offers a practical complement to detection tools that operate only after a conversation has begun.
Data quality remains a practical constraint. The effectiveness of any fraud-prevention system depends on the accuracy and completeness of the underlying information. Clean, verified and well-structured data continues to be essential for voice-related detection as well as for broader identity verification and anti-money-laundering processes.
Resource limitations also vary considerably across the industry. Institutions with smaller fraud teams, more limited technical infrastructure, or tighter technology budgets can face a slower path to evaluating and deploying specialized tools, a dynamic that affects banks of various sizes, not only the largest commercial institutions.
Finally, many real-time voice detection systems that perform well in controlled laboratory settings show reduced accuracy under everyday phone conditions. Background noise, audio compression and channel distortion commonly degrade performance. As a result, institutions are examining approaches that combine voice characteristics with behavioral cues and transactional context rather than relying on voice signals alone.
How banks are approaching detection:
Call center impersonation
Call-center impersonation has become one of the faster-growing forms of identity fraud. Using synthetic audio, attackers pose as trusted organizations, banks, payment processors or even law-enforcement agencies to obtain account details or personal information from customers or employees.
Banks are exploring detection tools that analyze voice characteristics in real time to help distinguish genuine human speech from generated audio. These tools are typically deployed as one layer within broader authentication and fraud-monitoring processes rather than as a stand-alone defense, and are designed to support secure handling of high-risk transactions including wire transfers, digital wallet payments and ACH activity.
Account takeover (ATO) via voice spoofing
Account takeover through voice spoofing occurs when criminals use synthetic or deepfake voices to gain unauthorized access to a customer’s banking, payroll, health-savings or other accounts. The goal is typically to move money or extract personal information (O’Driscoll, 2021; TrustBuilder, 2025).
Both banks and customers feel the effects. Institutions can face reputational harm and increased operational costs from investigating and remediating incidents. Customers often experience direct financial losses, time spent resolving the problem and in some cases secondary damage such as compromised accounts across other platforms.
To address these risks, some providers have introduced multi-layer detection approaches that combine signal-based analysis, behavioral monitoring and AI-driven detection to identify suspicious activity. Banks evaluating such tools typically view them as one component within a broader authentication and monitoring framework rather than a complete solution in themselves.
Social engineering amplified by synthetic media
Social-engineering schemes increasingly rely on synthetic media. Attackers can generate realistic voices and place calls that impersonate bank executives, customer-service staff or IT support personnel. In some cases, generative AI is also used to test or bypass existing voice-detection systems and to support large-scale credential-stuffing campaigns.
Banks are examining detection platforms that analyze multiple signals rather than voice alone, combining audio analysis with broader contextual information to improve performance under everyday phone conditions. Institutions typically treat these tools as one element within a layered fraud-monitoring process.
Ongoing challenges and emerging approaches
Banks evaluating AI-based voice-fraud tools continue to encounter several practical limitations. Detection systems that work well in controlled settings often lose accuracy when processing real-world phone audio that includes background noise, compression or long-duration calls. Many models also remain difficult for front-line staff to interpret, and the underlying training data can lack the variety needed for consistent performance across different conditions and attack methods.
Industry and research efforts are therefore focusing on several practical directions:
- Developing lighter-weight detection methods that can operate efficiently on everyday call center audio
- Improving the ability of systems to generalize across different synthesis techniques and noisy environments
- Making detection results more transparent so that fraud teams and call-center staff can understand and act on the output
- Combining voice analysis with additional signals, behavioral cues and transactional context to create multilayered defenses
- Exploring digital audio watermarking as a way to identify content generated by legitimate tools, or to flag its absence
Two related trends are also drawing attention. First, some fraudsters now coach victims in real time during calls. Behavioral signals such as changes in speaking intensity or response patterns may help distinguish these coached interactions from ordinary conversations. Second, a few organizations are beginning to consult behavioral specialists to better understand human responses during fraud attempts and to test whether those insights can strengthen detection models.
These developments remain works in progress. Institutions that have tracked how detection tools perform under their own call conditions, documented false-positive rates and evaluated multi-signal approaches have found themselves better prepared as the technology and the threat continue to evolve.
Deepfake voice fraud is still viewed by many institutions as an emerging rather than fully operational threat. Banks are nonetheless examining how detection and authentication practices may need to evolve as the technology becomes more accessible.
By the end of the decade, the cost of running large generative models is expected to fall substantially — estimates suggest inference costs could drop by more than 90 percent compared with 2025 levels. Lower costs could make sophisticated synthetic-media attacks more widely available and potentially more adaptive when combined with large language models.
Current detection systems often perform adequately in controlled settings but lose accuracy under ordinary phone conditions that include noise, compression and channel distortion. As a result, attention is shifting toward approaches that combine voice analysis with behavioral cues and transactional context rather than relying on voice signals alone. Digital audio watermarking is another area of interest: legitimate generative-AI providers may embed markers that identify content as synthetic, so that the absence of such a marker becomes an additional signal worth investigating.
Institutions that have tracked the real-world performance of their existing tools, documented false-positive rates and evaluated multi-signal methods have found themselves better positioned as both the threat and the available defenses continue to change.
Institutions that have tracked the real-world performance of their existing tools, documented false-positive rates and evaluated multi-signal methods have found themselves better positioned as both the threat and the available defenses continue to change. Industry exercises such as the UK Home Office Deepfake Detection Challenge illustrate increasing collaboration among government, academic and private-sector participants.
The voice channel remains an important customer-contact method for many institutions. Maintaining its usefulness will depend on clearer authentication signals, more reliable detection under everyday conditions and continued attention to the human elements of social engineering.
Anat Goldstein is the founder and CEO of FinOptima Solutions, an AI-driven fraud intelligence company, where Farook Sattar is senior researcher and advisor.
Go to aba.com’s fraud resource section for updated information and analysis concerning threats to banks and bankers.









