Voice and deepfake

Deepfake vs voice cloning: what is the difference?

Voice cloning synthesises a person’s speech from audio samples. Deepfake is the broader term for any synthetic media of a real person, including video and still images. Voice cloning is audio-only deepfaking, so every voice clone is a deepfake but not every deepfake involves voice.

The distinction matters operationally because the two need different amounts of raw material. A usable voice clone can be built from a short public recording. A convincing live video deepfake of several people at once is a much higher bar, which is why voice-only attacks are far more common.

Documented cases

  • Voice only: a UK energy firm wired about EUR 220,000 in 2019 after a single cloned phone call.
  • Voice only, refused: LastPass had its CEO cloned over WhatsApp and the employee declined to act.
  • Full video: Arup faced an entire meeting of AI-generated colleagues, and lost about US$25.6 million.

The control that breaks it

  • Use the same rule for both: verify out of band before acting, regardless of medium.
  • Assume any public-facing executive can be voice-cloned. Plan for it rather than trying to prevent it.
  • Never treat a recognised voice as authentication.

Related: vishing vs voice cloning · deepfake vs synthetic media · Voice cloning