Writing · 30 July 2026
Why your face changes in AI photos, and how to stop it
Reference photographs are not enough to hold a likeness. The fix is to state who you are twice — once as pixels, once as words — and it is the single thing that decides whether a result is usable.
You upload a photograph of yourself. You ask for a different outfit, better light, a wedding backdrop. What comes back is a beautiful picture of somebody who is almost you — a cousin, maybe. The nose is a little different. The jaw is softer. Something in the eyes has gone.
Nearly everyone who has used an AI photo tool has had this experience, and nearly everyone assumes the fix is a better reference photograph. It usually is not.
Why a reference photograph is not enough
An image model does not copy your face across. It reads the reference, forms an impression, and then draws a face from that impression while simultaneously satisfying every other instruction it was given.
That is the problem. The other instructions are strong.
"Golden hour, low sun behind the subject, warm rim through the hair, 50mm at f/1.8" is a specific and forceful description of light. Satisfying it means changing how the face is lit, which means changing the shadows, which means changing the apparent shape of the cheekbones and the depth of the eye sockets. The model is not ignoring your face. It is being pulled between two sets of instructions, and the styling ones are more concrete.
Add a hairstyle change and makeup on top, and the styling instructions now outnumber and out-specify the single image of you. The likeness loses.
Saying it twice
The approach that works is to stop relying on the picture alone, and state your identity a second time — in words.
Before anything is generated, a vision model looks at the reference and writes a description of the face. Not a flattering description; a forensic one. The distance between the eyes. Whether the jaw is square or tapered. The shape of the philtrum. Where the brow ridge sits. Whether the face is broadest at the cheekbones or the jaw.
That description is then pinned into the generation brief as a checklist — a set of conditions the output has to meet, sitting alongside the styling instructions rather than beneath them.
Now the model is not weighing one photograph against ten paragraphs. It is weighing ten paragraphs against ten paragraphs, one set of which is about you.
The part that is counter-intuitive
The description deliberately leaves things out.
It does not describe the hair. It does not describe the makeup. It does not describe the complexion in the reference photograph.
The instinct is to include everything — surely more detail is more likeness? — and it is wrong, for a simple reason: those are the things you are about to change on purpose. If the identity checklist says "shoulder-length dark hair, minimal makeup" and the styling instructions say "a sleek low bun and a matte brick-red lip", the model now has two art directors contradicting each other, and it will split the difference in a way that satisfies neither.
So the description covers only what styling cannot move. Bone structure, proportion, the set of the features. Everything else is left for you to decide.
Angles, and the thing a single photograph cannot tell you
A front-on photograph contains no information about the side of your head. If a pose turns you three-quarters, the model has to invent your profile — and it will invent a generic one.
That is why a reference set is better than a reference photograph. Four angles — front, both three-quarters, and a profile — generated once from your original and then attached to later work, gives the model the sides of a head it would otherwise be guessing at. It matters most in exactly the shots people want: the turned-away, over-the-shoulder ones.
What to do with your own photograph
If you are uploading a face to any tool like this:
- Even, frontal light. Hard side-lighting bakes shadows into the reference that the model reads as bone structure.
- A neutral expression, or a small one. A wide smile changes the shape of the whole lower face, and the model may carry that shape into a photograph you wanted to be serene.
- Nothing across the face. Sunglasses, a hand, a dupatta over the cheek — anything obscuring the structure has to be invented.
- Reasonable resolution, and no beauty filter. A smoothing filter has already removed the fine structure that makes you recognisable. The model cannot restore what the filter took.
A useful test of any tool: generate the same person twice in two very different scenes, and put the results side by side. If they look like siblings rather than one person, the tool is treating your face as a suggestion.
See yourself in the outfit.
Upload one photograph of your face and pick what you are dressing for. 30 free credits, no card.