what sign up
nixm nixm

← learn

the t-shirt and the car — what an ai video model quietly drops

September 20, 2026

AI video models silently drop whatever they cannot carry: a described prop becomes an invented one, a fifth reference goes missing, a wide shot loses a chest print, and text never changes what a reference is wearing. Four limits hold in practice — bind don't describe, one wardrobe line per person, four references total, two camera positions per fifteen seconds. nixm shows the cast and the price before it renders, and refuses when the prompt and the reference list disagree.

We spent a week trying to get three women, one t-shirt and a car into fifteen seconds of video. Here is what the engine did with each ask, in order, and what finally held.

Ask one — the shirt that came back wrong

Three characters, all in white tanks and jeans, one of them in a branded tee, wide shot of a city street cutting to medium. The tee came back with the right shirt and the wrong type — the logo was there, the letters were not ours. The wide shot never showed it anyway.

That failure had a cause we could read in the request. The shirt existed as a reference image. The prompt described it: "a bold black square logo with the word nixm in white." The engine drew a black square and wrote something in it. It had never been shown the shirt. A described thing is invented; a bound thing is reproduced. The fix was one field: send the image, refer to it as a token, never describe it.

Ask two — it held, then it didn't

Same shot, shirt bound as an image reference. It held perfectly — on one character, at medium distance. So the binding works. We then asked for the same shirt on the same character in a wide three-shot with a second camera move. Gone again.

Two more causes, both visible in hindsight. The prompt said "#Element1 is wearing #Image1" and two sentences later "all three in white tanks." The engine was told two things about one person and picked one. And at a wide framing a chest print is maybe forty pixels; there is nothing to reproduce. One instruction per wearer, and a print needs a medium shot or closer.

Ask three — a different car, and a missing woman

A matte black vintage convertible, uploaded as a reference. Three characters posed on it under two spotlights, fifteen seconds, three camera positions. The car came back as a black convertible — a different car. One of the three women did not come back at all.

The car was described, not bound: the prompt said "a matte black grungy cadillac" while the reference sat unused. Same failure as the tee, third time. The missing woman was the new one — three faces plus two objects is five references, and at five the engine dropped one. Three camera positions in fifteen seconds did not help. References share one budget, four is reliable, and fifteen seconds holds two camera positions rather than three.

Ask four — everything held, then the wardrobe didn't

One shot, locked off, ten seconds. Three women on the car, one in the tank, car and tank both bound as images, wardrobe one sentence per name. Everything held: the car was ours, the tank was ours, three faces.

Then we asked for the two others to be in white tees instead of what they had worn in their reference photos. Text again. The tees did not take; they arrived dressed as their references. A reference carries its wardrobe with it. Text does not change clothes.

What this adds up to

None of these were model failures in the sense people mean when they say "AI slop." Every one was the request asking for more than the engine holds, and the engine dropping the part it could not carry without telling anyone. The output looked finished. That is what makes it slop: confident, complete, and wrong about one thing you did not check. Related: why AI videos look bad.

The limits, as we found them:

LimitWhat happens past it
Bind, don't describeDescribe a thing that exists as an image and the engine stops looking at the image
One wardrobe line per personA group line overrides the individual one, for whoever it shouldn't
A print needs a medium shotAt wide framing a chest logo is pixels, and pixels get invented
Four references totalCast and props share the budget; the fifth is the one that goes missing
Two camera positions / 15sA third shot, or a drift across faces, costs a face
Wardrobe is the referenceWords won't redress a character — the reference's clothes come with it

The first habit: the receipt

Before anything renders, show the cast, the props, the length and the price, and do not fire until someone says yes. Half of these failures would have been caught by looking at the receipt — "why is the car not in the props?" — before spending a credit on a render that could not work.

The second habit: make the machine refuse

A rule for whatever is writing the prompt: every reference that is listed is used, every reference that is used is listed, and nothing listed is described. We made our system refuse to fire when the lists and the prompt disagree. The refusal costs nothing. The render would have cost three credits and a stranger's face.

The recast trick

The one we are most pleased with: a character in new clothes is a still — face bound, wardrobe written — and then that still becomes the character's reference for the scene. The wardrobe is now an image, and images hold. It is not something you can do in one prompt anywhere; it is a loop, and it needs a place to keep the cast. More on holding a character across a sequence: how to make a short film with AI, and on getting a first clip out at all, how to make an AI video.

See the one that worked

Three women, one shirt, one car: the showcase signal. Or cast your own character in the cinematic studio — the first HD video with sound is free.

frequently asked

Why did the AI change my product in the video?

Almost always because it was described in words rather than bound as an image. A described thing gets invented from the description; a bound reference gets reproduced. If both happen at once, the description usually wins and the reference sits unused.

How many references can an AI video model hold?

Four is reliable in our testing, and cast and props share that budget. At five, one gets dropped silently — in our case an entire character who simply did not appear in the render.

Why does my character's logo or print disappear?

Framing. At a wide shot a chest print is roughly forty pixels, which is not enough to reproduce, so the model invents something logo-shaped. A print needs a medium shot or closer to survive.

Can I change what a character is wearing by describing it?

No. A reference image carries its wardrobe with it, and text will not override that. Render a still of the character in the new clothes first, then cast that still as the reference.

How many camera positions fit in a 15-second clip?

Two. A third position, or a drift across several faces, reliably cost us a face in the same render.

Why do AI videos look finished but wrong?

Because the engine drops what it cannot carry without reporting it. Nothing errors — you get a complete, confident clip that is wrong about one detail you did not think to check.

What is a render receipt?

A summary shown before anything is generated: the cast, the props, the length and the price. It catches missing references before a credit is spent rather than after.

Does nixm refuse bad prompts?

Yes. If the reference list and the prompt disagree — a prop listed but never used, or a token pointing past what was supplied — the render does not fire and nothing is charged.

try it now — nixm makes cinematic ai video, with sound. no subscription. first video free.

open the studio →

Enter your email and receive a link to sign in. No passwords, just a secure signal to connect.

why connect

  • 20 free credits to start
  • your signals, remembered
  • buy more credits anytime
  • image & video generation features
or
+