Voice cloning synthesises a person’s speech from audio samples. Deepfake is the broader term for any synthetic media of a real person, including video and still images. Voice cloning is audio-only deepfaking, so every voice clone is a deepfake but not every deepfake involves voice.
The distinction matters operationally because the two need different amounts of raw material. A usable voice clone can be built from a short public recording. A convincing live video deepfake of several people at once is a much higher bar, which is why voice-only attacks are far more common.
Related: vishing vs voice cloning · deepfake vs synthetic media · Voice cloning