Highlights
Accept US mobile driver's licenses from every wallet 3 years in a row named a Leader First to achieve iBeta Level 3 on iOS and Android Introducing GovFaceMatch
01/04
01/04
Back to webinars

Webinars

Deepfake detection in production: The research problem it creates

On 14 academic benchmarks, a state-of-the-art public deepfake detector scored 91.2%. On Incode's identity verification data, it fell to just over 60%. Efim Boieru, Senior Manager of Machine Learning at Incode, walks through the experiments his team ran to find out what a production detector actually needs.

On demand · September 22, 2026

On 14 academic benchmarks, GenD, a state-of-the-art public deepfake detector, separated real faces from fakes with 91.2% accuracy. On Incode's identity verification data, that number fell to just over 60%.

That gap raises a tough question: if public benchmarks can't validate a production detector, what can? In this webinar, Efim Boieru, Senior Manager of Machine Learning at Incode, walks through the experiments his team ran to find out, from testing expert human labelers against Incode's own models to feeding a detector samples from a generator that didn't exist when it was trained.

Speakers

  • Efim Boieru, Senior Manager of Machine Learning, Incode

The talk covers how Incode builds training data that looks like real production traffic, including fine-tuning open-source generators on internal data and running internal red-team exercises. It also covers the three approaches Incode relies on to handle generators it has never seen: few-shot adaptation, agent-based monitoring, and generalization.

The next shift is already visible. Agentic fraud, where AI agents research targets, generate deepfakes, and test injection methods continuously, is turning deepfake detection into one piece of a much larger orchestration problem.

Key takeaways

  • Public benchmark scores measure performance on public benchmarks. A detector that scores 91.2% AUROC on academic datasets dropped to just over 60% on production identity verification data.
  • Human labeling doesn't scale. A two-year-old Incode model beat five expert deepfake labelers, even when their answers were combined by majority vote.
  • A benchmark only scores generators that already existed when it was built. Few-shot adaptation cut Incode's error rate on new generators roughly 10x.
  • Strong detectors generalize. Incode's detector correctly grouped samples from Nano Banana Pro, a generator that didn't publicly exist when the model was trained.
  • Agentic fraud is replacing manual attacks, making deepfake detection one layer in a defense that spans the device, the camera, and the image.

Watch this webinar on demand

Tell us who you are and we'll share the full recording.