ONLINEAGENT_OPS 2026.Q3 HOME ARTICLES CRAFT RECORD BLOG HUBS FAQ SEARCH
HOMEARTICLESWHO OWNS WHAT YOU GENERATE
EXPLAINER · RIGHTS

Why AI Output Looks Fake

The model was tuned to make images people like, not images that look real. What that costs you, and the specific words that undo it.

READ9 min
WORDS1,803
SECTIONS6
TYPEGUIDE
CHECKED31 AUG 26

Your prompt says photorealistic. The output still looks like an advert. The reason is not your wording — it is what the model was trained to prefer, and you have to work against it deliberately.

TL;DR — THE SHORT VERSION
  • Models were tuned toward what people rate as pretty. One widely used dataset was filtered by a model trained to predict the answer to "how much do you like this image, 1 to 10". Pretty is not the same as real.
  • So realism is a fight against the default, not a word you add. "Photorealistic" is a style label; the model already thinks it is being realistic.
  • Ask for imperfection by name — skin texture, asymmetry, one light source, a slightly wrong moment. Perfection is the tell.
  • Delete the quality spam. "Masterpiece, 8K, ultra-detailed, award-winning" pushes output toward the over-processed look you are trying to escape.
  • Video adds physics. Weight, contact, and cloth are where generated motion gives itself away, not resolution.
  • The strongest single move is constraint — one camera, one light, one moment. Realism is what is left when you stop asking for everything.
IN PLAIN ENGLISH

Imagine asking a thousand photographers for "a portrait" and averaging every result. You would get something smooth, centred, well-lit and slightly lifeless — because every specific, awkward, interesting choice cancelled out.

That average is roughly what a generative model gives you first. Realism is the work of putting the specifics back in.

Why the default looks like an advert

01

The model was scored on being liked

Image models are not trained only to be accurate. They are also tuned toward images people rate highly, and that tuning is explicit and documented.

One widely used dataset, LAION-Aesthetics, was built by training a model to predict the answer to a single question: "How much do you like this image on a scale from 1 to 10?" Images scoring highly were emphasised in training for Stable Diffusion's first version.LAION, "LAION-Aesthetics" (laion.ai/blog/laion-aesthetics) — the predictor was trained on human ratings against that question, using CLIP embeddings; the 5+ subset was used for Stable Diffusion V1. This is documented for that model and that dataset. Aesthetic preference tuning is common practice across the field, but do not read this as "every model uses LAION-Aesthetics" — modern pipelines differ and most are not published in this detail. Checked 31 Aug 2026

Read that question again, because it explains the whole problem. Nobody was asked "does this look like a photograph". They were asked whether they liked it. The model learned to please, and pleasing images are lit well, composed cleanly, and smoothed.

02

Averaging removes exactly what reads as real

A generative model produces something like a consensus. Fine, high-frequency detail — pores, stray hairs, dust, small blemishes, uneven texture — varies enormously between real photographs, so it averages toward nothing. Smooth is what is left when detail disagrees.

This is why skin is the first giveaway. Real skin scatters light unevenly at a very small scale. Averaged skin does not, so it reads as rendered, and viewers spot it before they can say why.This paragraph is mechanism, reasoned from how diffusion models are trained and what averaging does to high-frequency detail. It is not a measured claim, and this site has not tested it. The training-objective point above it is documented; this consequence is the standard explanation, not an experiment.

THE COUNTER-ARGUMENT WORTH KNOWING

"Aesthetic" filtering is not a neutral quality measure. An audit of the LAION-Aesthetics predictor found it disproportionately keeps images captioned with women and filters out images captioned with men and LGBTQ+ people — so the filter encodes taste and cultural assumptions, not objective quality."The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor", arXiv 2601.09896. Reported finding; this site has not reproduced the audit. Checked 31 Aug 2026

Worth holding onto for two reasons. It tells you the "default look" is a specific taste rather than a standard — and it is a reminder that what a model finds beautiful was decided by somebody.

The tells, in the order people notice them

Fix them in this order. The first two do more than the rest combined.

1 · Skin and surface. Poreless, evenly lit, faintly glowing. Everything looks recently cleaned.BIGGEST
2 · Light that does not agree. Shadows implying two or three light sources, catchlights in eyes that do not match the key, a subject lit differently from the room.BIGGEST
3 · Too much symmetry. Faces balanced in a way real faces are not; objects centred; horizons level.HIGH
4 · Composed like a stock photo. Nothing cropped awkwardly, nothing half out of frame, no accident anywhere.HIGH
5 · Background incoherence. Repeating patterns, objects that do not resolve, architecture that cannot be built.MEDIUM
6 · Depth of field that is decoration. Blur applied as a look rather than falling off with distance from one focal plane.MEDIUM
TAKEAWAY

Almost every tell is a form of too clean, too even, too arranged. If you remember one thing: real photographs contain accidents. Ask for one.

What to write instead

03

Name the imperfection

The model will not add flaws on its own — flaws are what the tuning removed. So they have to be requested as explicitly as anything else in the frame.

Instead of "photorealistic portrait", write the physical facts: visible skin texture and pores, a few flyaway hairs, uneven skin tone, faint shine on the forehead, slight asymmetry in the expression. These are not aesthetic words. They are descriptions of a real face.

The same logic applies to objects and rooms — scuffs, fingerprints, dust, uneven wear, a cable that was never tidied.

04

Give light one source and a direction

Incoherent lighting is the tell people cannot name but always feel. The fix is to be boringly specific: one key light, stated position, stated quality, and what happens on the other side.

"Single window to camera left, overcast daylight, soft falloff, no fill on the shadow side" gives the model a physical situation to solve. "Cinematic lighting" gives it a mood board.

05

Delete the quality words

This is the change people resist most and it usually helps most. Terms like masterpiece, 8K, ultra-detailed, hyperrealistic, award-winning, trending do not describe anything physical. They push output toward the heavily processed, high-contrast, over-sharpened images that score well on "do you like this" — which is the exact look you are trying to leave.Reasoning from the training objective described above, not a controlled test. Try it both ways on your model: run the same prompt with and without the quality stack and compare. Vendor guidance increasingly favours plain descriptive language over keyword stacking, but the strength of the effect varies by model and by version.

Replace them with camera facts — a specific focal length, an aperture, a film stock or sensor, a shutter speed. These constrain geometry and exposure instead of vibe.

06

Add the thing a camera does

Real images are made by a device with limits, and those limits are signatures: grain or sensor noise, slight motion blur, chromatic aberration at the edges, a highlight that clips, focus that is very slightly off.

Ask for one or two. Asking for all of them produces a different kind of fake — a photograph of a photograph.

Video: the tell is physics, not pixels

07

Weight, contact, and cloth

Generated video usually fails at mass. Feet skim instead of pressing. Objects are set down without settling. Hair and fabric move as if underwater, or move independently of the body carrying them.

Resolution does not fix any of it, and neither does a longer style description. What helps is describing the physical event rather than the appearance: what touches what, what takes weight, what stops moving first.

08

Ask for less motion than you want

The more a model has to invent per second, the more it drifts — faces morph, backgrounds reorganise, details flicker between frames. A held shot with one small movement survives; a shot with a moving camera, a moving subject and a changing background does not.

The practical rule: one significant movement per shot. If you need three things to happen, that is three shots, and cutting between them is also what real filmmaking does.

THE ORDER TO TRY THINGS

Change one thing at a time, or you will not learn which change worked.

1. Strip the quality words. 2. Fix the light to one source. 3. Add surface imperfection. 4. Add one camera artefact. 5. Only then adjust style.

Most people do this backwards — they add style words first and never remove anything. Realism is mostly subtraction.

Where this stops working

Some things are still hard regardless of prompt. Text inside an image, hands doing something specific, crowds where every face must hold up, reflections that have to agree with the scene, and any object the model has seen little of. Prompting improves the odds; it does not remove the ceiling.

And realism is not always the goal. A stylised image that is confidently stylised reads as intentional. The uncanny valley is reserved for work that was aiming at real and missed — which is an argument for either committing to a look or going all the way to photographic, and not stopping in between.

The realism pass

1 — Have I removed every word that describes quality rather than a thing?
2 — Is there exactly one light source, with a stated direction?
3 — Have I named at least two physical imperfections?
4 — Is there a real camera constraint — focal length, aperture, stock?
5 — Is anything in this frame an accident, or is it all arranged?
6 — For video: is there only one significant movement?
SOURCES AND HONESTY ABOUT THEM

LAION, "LAION-Aesthetics", laion.ai — the predictor, its question, and its use in Stable Diffusion V1 · "The Algorithmic Gaze of Image Quality Assessment", arXiv 2601.09896 — the audit of that predictor's biases. Both checked 31 August 2026.

What is documented and what is not. The training objective and the audit are published and cited above. The prompting advice is craft, not measurement — it is reasoned from that objective and from what the tells are, and this site has not run controlled comparisons across models. Techniques that help on one model can do nothing on another, and versions change. Test the pairs on your own model rather than trusting the list.

The through-line: the model is not trying to make a photograph. It is trying to make something you will like. Realism is the work of asking, very specifically, for the parts of reality that nobody likes.