iPhones were the last holdout against deepfake injection attacks — not anymore

Close-up of a smartphone performing a facial recognition biometric scan

For years, security teams treated iPhones as the safe option. Apple's tightly controlled camera pipeline made it far harder for fraudsters to feed fake video directly into a biometric check. That assumption just collapsed: according to iProov's 2026 Threat Intelligence Report, injection attacks against iOS devices rose 741% year-on-year, with activity surging even faster in the second half of 2025. The device-type shortcut that identity teams leaned on is gone.

The safe zone that stopped being safe

Injection attacks bypass the camera or microphone entirely, feeding fabricated video or audio straight into the verification pipeline instead of capturing a live person. For most of the deepfake era, this technique worked best against Android and web-based flows, while Apple's hardware-level protections made iOS a comparatively hard target. That gap has now closed, and fraud teams that built risk models around device type are exposed on a channel they considered low-risk.

The scale of the broader problem is accelerating just as fast. Data from Surfshark cited by Biometric Update puts global deepfake fraud losses at $3.7 billion to date — and 2025 and 2026 alone account for 89% of that total. This isn't a slow-building risk. It's a curve that bent sharply upward in the last eighteen months.

Why single-signal detection keeps losing

The market's response tells its own story. Vendors are racing to patch specific gaps: voice-native audio models to catch cloned speech, neural watermarking to trace AI-generated audio after the fact, and real-time meeting monitors that flag synthetic participants on video calls. Each of these is a genuine improvement — and each also confirms that a single biometric signal, checked once, is no longer enough.

From ‘liveness’ to ‘genuine presence’

As Dr. Chris Allgrove of Ingenium Biometric Laboratories put it in recent industry commentary, “the conversation itself has shifted from basic liveness to genuine presence. In a world of injection attacks and deepfakes, simply proving that a user is alive isn't really enough anymore.” Proving a pulse is no longer the bar. Proving that the person on the call, in the video, or on the line matches who they claim to be — continuously, across modalities — is the new one.

What multi-signal defense actually looks like

For banks, telecoms, and HR platforms rebuilding their fraud posture, the practical shift looks like this:

  • Cross-modal matching — verifying that a voice and a face genuinely belong to the same person, so a cloned voice alone or a swapped face alone isn't enough to pass.
  • Real-time detection, not post-hoc review — catching synthetic audio and video during the interaction itself, before a wire transfer clears or an account opens.
  • No dependency on a pre-enrolled database — so verification still works for first-time callers, new hires, or first-party onboarding, not just returning customers.
  • Device-agnostic assumptions — treating iOS, Android, and desktop channels as equally exposed, rather than deprioritizing any one platform.

This is precisely the gap between point-solution deepfake detectors and layered biometric verification. A tool that only listens for cloned audio, or only watches for face-swapped video, still leaves the other channel open.

What to do before the next injection attempt

Security and fraud leaders don't need to wait for a formal certification cycle to start closing this gap. Three moves matter most right now:

  1. Audit whether any risk scoring still treats device type (particularly iOS) as a mitigating factor — and remove that assumption.
  2. Stress-test contact center and video onboarding flows against injection-style attacks, not just presentation attacks like printed photos or replayed video.
  3. Move toward verification that checks voice and face together in real time, rather than relying on a single biometric signal at a single point in the journey.

The 741% figure isn't a warning about some future risk — it's a description of what's already happening on the devices security teams trusted most. Corsound AI's Deepfake Detect identifies synthetic audio and video in real time, and pairs with Voice-to-Face AI to confirm that the voice and face on a call genuinely belong to the same person — no database required. If your verification stack still leans on a single signal or a device-type assumption, now is the time to close that gap.

Photo: ShotPot / Pexels

See Corsound AI Voice Intelligence In Action
Thank you.
Your submission has been received.
Oops! Something went wrong while submitting the form.