Writing · 1 August 2026
What AI virtual try-on actually does — and what it cannot
An honest account of how an image model puts you in an outfit you have never worn, why the face is the hard part, and where the technology still falls over.
There is a particular moment in buying occasion wear that nothing online has ever solved. You are looking at a lehenga on a model who is not you, in a light that is not the light at the venue, and you are trying to answer a question the photograph cannot answer: would that look right on me, there?
Virtual try-on is an attempt at that question. It is worth being precise about what it does, because the marketing around it is not.
What actually happens
You give the system two things: a photograph of your face, and a garment — either a photograph of one, or a choice from a catalogue. It gives back a new photograph of you wearing it.
That new photograph is not a composite. Nothing is cut out and pasted. An image model generates the whole frame from scratch, guided by a written brief and the reference images. Which means the fabric falls the way fabric falls, the light on your face matches the light in the room, and the drape follows your posture — because all of it is being drawn at once rather than assembled.
It also means the model is inventing everything it was not told. That is the source of both the quality and every problem below.
The brief matters more than the pictures
Most of the work in a system like this is not the model. It is what you say to it.
A brief that says "wearing a red lehenga" produces a red lehenga, some red lehenga, drawn from the average of every red lehenga on the internet. A brief that names the weave, the border, the fall of the dupatta, the hour of the day, the lens, and the distance to the subject produces something specific. The difference between the two is not a better model — it is about ten paragraphs of instruction.
This is why an occasion matters so much here. "Wedding" is not one thing. A haldi is outdoors, mid-morning, in marigold and turmeric yellow, and everybody expects to get stained. A reception is indoors, artificially lit, and structured. Told only "wedding", a model will average those into something that belongs at neither.
The hard part is your face
Everything above is styling, and styling instructions are strong. Strong enough that they will drag your face along with them — you ask for a warmer light and a different hairline arrives with it.
A reference photograph, on its own, does not hold a likeness. This is the single most common failure in every AI photo product, and the reason so many results get described as "it's nice, but it isn't me."
The fix is to say who you are twice: once as pixels, once as words. The system writes a description of the face before it generates anything — bone structure, the set of the eyes, the shape of the mouth, the things that styling cannot move — and pins that description into the brief as a checklist the model has to satisfy. Deliberately it does not describe the hair or the makeup, because those are things you are about to change on purpose.
There is a whole post about that one, because it is the thing this technology most often gets wrong.
What it cannot do
Being honest about the edges is more useful than another paragraph of enthusiasm.
- It does not know your measurements. It draws a plausible body, guided by what you tell it. It cannot tell you whether a blouse will fit.
- It is not a substitute for seeing the fabric. Colour on a screen is a negotiation between the model, your display and the room you are sitting in. A deep maroon can render as burgundy.
- Fine embroidery is approximated. At the resolutions these models work at, a heavily worked border becomes an impression of that border. Get close enough and it will not survive.
- It generates, so it can be wrong. Occasionally a hand has an extra knuckle or a dupatta goes somewhere physics would not allow. Generating again is usually the answer, which is why it should be cheap to do.
Where it is genuinely useful
Not as a fitting room. As a way of narrowing.
You are choosing between six outfits for an event, and you cannot picture five of them on yourself. Seeing them — in roughly the right light, on roughly you — eliminates four in about a minute, and that is a real thing to have done before you spend an afternoon in a shop or an evening returning a parcel.
It is also, quietly, a way of trying something you would not have. Most people wear a narrow range of colours because they have never seen themselves in the others.
See yourself in the outfit.
Upload one photograph of your face and pick what you are dressing for. 30 free credits, no card.