SPOT FAKE IMAGES & VIDEO
What still gives away AI-generated images, video and cloned voices in 2026 — and why the visual tells are disappearing…
Count the fingers. Check the teeth. Look for garbled text on signs. Look for the weird ear. These were real failures of 2023-era image models and they are largely solved. Advice built on them is worse than useless now, because it hands you a confident all-clear on a fake that simply does not have that flaw. If a guide's main advice is finger-counting, it was written for a world that no longer exists.
Images: look at physics, not anatomy
Generative image models learned what scenes look like without learning the rules that produce them. So the durable signals are about consistency rather than craft.
Find the light source, then check every shadow in the frame points away from it, and that reflections in glass, water, eyes and metal show what should actually be there. Multiple objects lit from subtly different directions is a strong signal — it is a rule the model was never given.
The subject gets the model's attention; the background gets plausibility. Look for architecture that cannot be built, repeated patterns in crowds, objects merging into each other, and text that is confidently shaped but meaningless at a glance.
Real lenses produce depth: some things sharp, others not, in a way governed by distance. Generated images often render everything with similar clarity, and skin, fabric and stone can be too clean — no dust, no wear, no pores where pores should be.
Hands slightly too large for the arm, a chair the wrong size for the table, five posts in one shot and six in another. Models learn appearance, not measurement.
Video: look at time, not frames
The same scene rules apply, and continuity adds new ones — because a video model has to keep a world consistent across time, which is much harder than making one convincing frame.
A logo changes slightly, a necklace moves to the other side, a room gains a window. Cross-shot consistency comes from reference material, and reference material has gaps.
Hair, cloth, water and dust betray generated video, because they should carry inertia — continuing after the body stops, settling, catching. Synthetic motion often moves with the subject and stops when it stops.
Counter-intuitive, and one of the strongest signals. Production pipelines use character reference sheets to hold a face steady across shots, which produces identical bone structure, hair behaviour and wear from every angle. Real filming drifts: hair moves, make-up degrades, fabric creases differently. Uncanny sameness is itself the tell. See how AI films are actually made for why.
Coherence degrades with length, so generated footage tends to arrive in short pieces with cuts falling at suspiciously regular intervals. A long unbroken take of a person doing something complicated is still expensive to fake.
Audio is frequently replaced after the fact. If the voice keeps identical texture and cadence while the visible space changes — indoors to outdoors, small room to large — the audio and video probably did not come from the same event.
Voice: stop listening, start verifying
This is the section where honest advice contradicts instinct. You will probably not hear it. Voice cloning requires only a few seconds of sample audio — Microsoft researchers demonstrated reproduction from roughly three seconds in 2023 — and a short phone call gives your ear very little to work with.Source: Microsoft Research, VALL-E — “Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers”, January 2023. Microsoft's own description: high-quality personalised speech from “only a 3-second enrolled recording of an unseen speaker”. Trained on 60,000 hours of speech from over 7,000 speakers; code not released. A 2024 successor claimed human parity.
Which is why the defence has to be procedural, not perceptual. The habits that work regardless of how convincing the voice sounds: a family code word agreed in advance, and a callback rule — you hang up and dial the number already saved in your phone, never one the caller supplies. Treat urgency and secrecy as the finding rather than the context. The full set, with the documented cases, is on the deepfake defence page.
Where you do have a recording to examine, the useful signals are environmental: breathing that does not fit the sentence rhythm, room tone that stays constant while the speaker supposedly moves, and emotional delivery that is appropriate but never quite spontaneous — grief with correct words and even pacing.
Every signal above will weaken. The one that does not is provenance: where did this come from, and does it exist anywhere else? A real event has more than one witness. Search for the earliest copy, look for a second angle, check whether the person or outlet posted it on their own verified channel. Media with no traceable origin and no corroboration is the actual red flag — regardless of how it looks or sounds.
Detection tools: signals, never proof
Software claiming to identify deepfakes degrades sharply on compressed, re-encoded, screenshotted or re-filmed media — which describes essentially everything that circulates on social platforms. The same limitation applies here as with text detectors: a percentage is not a verdict, and nobody should face consequences on one. Watermarking and provenance standards embedded at the point of creation are a more promising direction than after-the-fact detection, because they carry a claim rather than guessing at one.
Synthetic images and video are not inherently deceptive — the illustrations on this site are generated, and labelled. The harm is unlabelled synthetic media used to make someone appear to do or say something they did not. As the tells disappear, the burden shifts from your eyes to the record: verified channels, corroboration, provenance. Learn the signals, but do not stake anything important on them alone.