real-time deepfake detection is coming to video calls — is it fast enough?

Business professional on a video conference call, representing real-time deepfake detection in virtual meetings

In January 2025, Singapore recorded 43 cases of live-video deepfake scams. By January 2026, that number had jumped to more than 1,200 cases — a nearly 28x increase in a single year. One of those cases involved a Singaporean CFO who was tricked into wiring $500,000 during a video call where every other participant, including his own boss, was an AI fabrication. The meeting looked and sounded completely normal. It wasn't.

The fake meeting is now a real threat

For years, "seeing is believing" survived deepfakes because live video was assumed to be too technically demanding to fake in real time. That assumption is collapsing fast. Generative models can now clone a face and voice convincingly enough to sit through an entire meeting undetected, and fraud teams are only starting to catch up.

Researchers at the Fraunhofer Institute for Secure Information Technology (SIT) in Darmstadt, Germany, recently unveiled a prototype built specifically to flag manipulated video and audio streams during a live call, not after the fact. It's a proof-of-concept, but it signals where the industry is heading: from static identity checks at onboarding to continuous, in-session verification.

How the Fraunhofer prototype works

The system combines audio and video analysis to score the probability that a call is being manipulated in real time, running locally on a high-performance laptop rather than sending sensitive footage to external servers. When it flags multiple warning signs, it prompts participants to verify identity through a second channel — a phone call, a separate messaging app, a pre-agreed code word.

That last detail matters. The researchers trained their model on thousands of real and spoofed recordings under typical videoconferencing conditions — compressed streams, fluctuating bandwidth, blur filters, background noise — precisely the artifacts that make live-call deepfakes so much harder to catch than a doctored photo or a pre-recorded clip.

Why this problem is scaling faster than the defenses

The market is responding, but not fast enough to outrun the fraud curve. Combined voice and facial deepfake checks are projected to more than double, from over 6 billion in 2026 to more than 12.2 billion by 2028, with annual market revenue climbing from roughly $3 billion to $6.1 billion over the same period, according to Biometric Update's 2026 Deepfake Fraud Detection Market Report. That growth reflects genuine investment — but it also reflects how far behind most organizations still are.

Most video-call fraud today isn't stopped by technology at all. It's stopped by a wire transfer that happens to get a second look, or an employee who gets a nagging feeling and calls back on a known number. That's not a security program — it's luck.

What fraud and security teams should do now

Enterprise-grade, real-time deepfake detection for video and voice is still maturing, but that doesn't mean teams should wait on the sidelines. A few steps close the gap immediately:

  • Treat high-value video calls like high-value transactions. Any call involving a wire transfer, credential reset, or sensitive data request should require out-of-band confirmation before action is taken — never approval within the call itself.
  • Deploy continuous, not one-time, verification. A face or voice match at login says nothing about who's speaking ten minutes into the call. Detection needs to run for the full duration of high-risk interactions.
  • Build a callback protocol that doesn't rely on caller ID or the video window. Pre-agreed verification channels defeat impersonation far more reliably than "does this look/sound right to me."
  • Pressure-test your own executives. The CFO in the Fraunhofer report wasn't careless — the deepfake was simply good enough. Run internal simulations to see how your team actually responds under pressure.

The window to get ahead of this is closing

Fraunhofer's prototype is still working its way toward integration with platforms like Teams and Zoom, and legal questions — like whether call participants must consent to being analyzed — remain unresolved. Enterprises don't have the luxury of waiting for those answers to settle before fraud volumes climb further.

The organizations that get ahead of this will be the ones that build real-time, multimodal deepfake detection into their video and voice channels today, rather than reacting after the next CFO gets a call that looks and sounds exactly right. See how Corsound AI's Deepfake Detect identifies manipulated audio and video in real time, before a fabricated meeting becomes a very real loss.

Photo: Diva Plavalaguna / Pexels

See Corsound AI Voice Intelligence In Action
Thank you.
Your submission has been received.
Oops! Something went wrong while submitting the form.