AI & Tech AI Tools & Models

AI Fight Scene Test: How to Write a Better Prompt

We rewrote one difficult AI fight prompt and tested it in Seedance, MiniMax and FLUX 3 to see what better structure fixed and what it could not.

Three action figures struggle in a miniature hotel lobby while their hands blur together and a red gift remains sharp beside the elevator.
One suited hero, two agressive dude, and one red gift were enough to show how quickly AI video still loses control of a fight scene.

It is not a difficult equation in AI video: a good prompt gives you a better chance of a good result, but finding the balance between clear logic and a prompt so overloaded that the model starts losing the point is much harder. In our previous AI video generator test, the biggest problem was getting the characters to understand their roles and stop disappearing like people in a Marvel film after Thanos snapped his fingers. In these three retests, the characters remained much more distinct, the first attacker generally attacked first, and the second understood that he had to wait for his turn instead of merging into somebody else's torso or wandering away from the scene. None produced a completely believable fight because, once the role confusion improved, the problems with contact, weight, and human decision-making became much easier to see.

Character merging was mainly a problem in the fighting scene from our original comparison, where three people had to touch, react and change position inside the same continuous shot. We ran one deliberately difficult prompt once per model, so this article shows what clearer prompt structure changed in that particular scene rather than pretending we discovered a universal rule for every AI video.

Quick answer

A clean hierarchy and clearly formulated tasks give an AI video model the best chance of following a complicated action prompt. Use this order: Task → Subjects → Setting → Action Timeline → Camera → Lighting → Style & Audio

Give every character a distinct identity, fixed position and one job at a time, because the model should never have to guess who performs the next action.

We could have used this structure in our original test, but we deliberately did not babysit every movement because we wanted to see how intelligently current models could interpret a normal prompt on their own. They struggled, so this guide has one practical purpose: showing how much the same scene improves when the hierarchy and each character's task are made explicit, while also showing which problems better prompting still cannot fix.

Why we chose a fight scene in the first place

AI video companies love action clips; open Higgsfield or almost any model launch page and you will eventually find somebody throwing a beautiful punch while the camera swings around as if the whole thing came out of a very expensive film. I always wonder how many generations went into those few good seconds, and how much editing joined separate clips together before we saw the finished demo.

So in our five-model AI video comparison, we deliberately chose a scene that would expose that problem: one hero, two attackers, a red gift box and an elevator inside a luxury hotel lobby, all squeezed into ten continuous seconds.

Seedance 2.5 came closest in the original test, but the second attacker more or less stopped existing. MiniMax confused the fight, FLUX turned the gun confrontation into a polite handover, and Gemini rebuilt parts of the hotel as the action continued, giving us something closer to Doctor Strange than a fixed lobby.

The original prompt described “two armed attackers” and then referred to the first and second attacker, but neither man had a persistent appearance, weapon or side of the frame. A person unconsciously fills in who arrives first, where the other man waits and how everyone moves between instructions. The model had to invent all of that while preserving three identities, two weapons, the gift, furniture, elevator and room.

This resembles the gap measured by WorldReasonBench, a 2026 benchmark that evaluates whether generated video preserves physically, socially, logically, and informationally consistent world states. Its researchers found that modern models can create convincing-looking footage while still failing dynamics, causality, or information preservation. Our test was much smaller, but the problem was visible: an individual frame could look like an action movie while the sequence connecting the frames made very little sense.

What we changed in the prompt

Luma's text-to-video guidance tells users to describe the subject, action, camera and style, while Google's Veo guide separates cinematography, subject, action, context and ambiance. Runway recommends simple physical language and clear positional identifiers when several people need different movements. We borrowed the useful parts rather than pretending one company's formula is a magic spell.

The rewritten prompt made five practical changes:

  1. Every man received a different appearance, outfit and weapon.
  2. The hero began in the centre, Attacker A entered from the left and Attacker B entered from the right.
  3. Only one attacker could engage the hero at a time.
  4. Four timestamped blocks defined the order of events.
  5. A locked wide camera kept the elevator, table and people inside one fixed frame.

We also removed some of the obsession with making the fight fast. Ten seconds is already cruel enough without asking the model to behave like Jackie Chan after six espressos.

The exact AI fight prompt we tested

TASK
Create one continuous 10-second cinematic action shot in 16:9 with synchronized native audio.

SUBJECTS
Three adult men must remain visually distinct and retain the same faces, bodies and clothing throughout the shot.

The Hero is a charismatic Chinese man wearing a dark navy suit and white shirt. He begins in the centre carrying a small red gift box.

Attacker A is a bald man wearing a black leather jacket and carrying a handgun. He enters from the left.

Attacker B is a bearded man wearing a grey jacket and carrying a short baton. He enters from the right.

SETTING
A fixed luxury hotel lobby. An open elevator is centred on the back wall. A small side table stands to the left of the elevator. The elevator, table, walls, furniture and lighting remain in exactly the same positions throughout the shot.

ACTION TIMELINE
[00:00–00:02] The Hero places the red gift box on the side table. Attacker A approaches from the left. Attacker B remains visible farther away on the right and does not attack yet.

[00:02–00:05] Attacker A attacks first. The Hero controls Attacker A's gun wrist, removes the handgun and knocks Attacker A onto the floor on the left. Every movement is shown through continuous physical contact. Attacker B remains separate on the right.

[00:05–00:08] Attacker B attacks from the right with the baton. The Hero blocks his baton arm and knocks Attacker B onto the floor on the right. Attacker A remains visible on the floor on the left.

[00:08–00:10] With both attackers still distinct and lying on their separate sides, the Hero steps backward into the centred elevator. The elevator doors close. The red gift box remains visible on the side table.

CAMERA
Use one locked frontal wide shot. Keep the elevator, side table, Hero and both attackers inside the frame. The entire sequence is one unbroken shot with no change of viewpoint.

LIGHTING
Use warm luxury-hotel practical lighting with stable exposure and clearly readable characters.

STYLE AND AUDIO
Precise modern action-thriller choreography with believable weight and timing. Make the action readable rather than excessively fast. Include realistic footsteps, clothing movement, weapon clatter and physical impact sounds.

Each attacker interacts only with the Hero, never with the other attacker. Preserve every character's identity and position. Non-graphic action with no blood, gore, dialogue, subtitles or on-screen text.

How we ran the retest

On 3 September 2026, we ran the same 10-second, 16:9 prompt once in MiniMax H3, FLUX 3 Video and Seedance 2.5, then judged the first completed result. MiniMax was generated at 2K. We did not keep rerolling until everybody suddenly became a stunt professional.

Gemini Omni 1.1 Flash was supposed to be the third comparison instead of FLUX, but it rejected the non-graphic prompt three times. Rewriting it until the safety system accepted a materially different version would have spoiled the comparison, so we used FLUX 3 instead. That rejection is still useful information for anyone choosing a model for fictional action, although it cannot be treated as a quality result because no video was generated.

ModelWhat the clearer prompt fixedWhat still failedTGK verdict
Seedance 2.5Best role separation, action flow, readable attacks, hotel continuity and soundGunman moved too close, transitions were too clean and the baton attack was heavily telegraphedClear winner
MiniMax H3, 2KFollowed the order well and ended with both attackers visibly defeatedHands and arms deformed during contact, people looked mushy in motion and the takedowns felt compressedSecond
FLUX 3 VideoPreserved the elevator, gift and character positions surprisingly wellSlow rehearsal-like choreography, weak force and an attacker waiting unnaturally for his turnThird

Seedance 2.5 was the only result that felt close to a film

0:00
/0:10

First completed 10-second result using our rewritten action prompt, tested 3 September 2026.

Seedance won comfortably because it was the only result that felt close to a proper action scene rather than a model moving figures between required poses. During the gun confrontation, the hero turned his body into the attacker and both men continued moving through the struggle, which created a much clearer sense of resistance. The second attacker carried momentum into his baton swing and the hero visibly reacted.

That is what I mean by the Jackie Chan comparison. Seedance did not recreate his martial arts style, but it understood more of the visual grammar: one threat arrives, the main G reacts, the movement is large enough to read, and then the next threat enters. The footsteps, hotel ambience, and impact sounds also gave imperfect physics more weight. Compared with the other two, it was not close.

It still made the gunman enter convenient grabbing distance, and the hero reset into a suspiciously clean neutral stance between attacks. Real fights contain ugly transitions where people stumble or recover their balance. Seedance preferred action-film logic, especially when the second attacker wound up his baton like a videogame heavy attack and gave the hero enough warning to make tea first.

MiniMax H3 followed the plot but struggled with contact

0:00
/0:10

First completed 10-second result using our rewritten action prompt, tested 3 September 2026.

MiniMax finished second. The improved hierarchy kept both attackers recognisable, followed the order and ended with a clear frame of the hero standing while both attackers were down, but I expected more from the 2K option.

When one hand grabbed another, the shapes started smearing and rebuilding around the contact. This looked more like motion deformation or temporal artifacting than ordinary motion blur, because the anatomy itself stopped feeling solid. Faces, clothing and limbs became waxy while the static hotel stayed sharper. Attacker A also walked closer with his gun extended and presented his wrist to the hero, so the prompt's sequence was followed largely because the attacker supplied the solution. The takedowns had the same compressed logic: contact happened, then somebody was suddenly defeated without enough mechanics between those states.

FLUX 3 preserved the room better than the fight

0:00
/0:10

First completed 10-second result using our rewritten action prompt, tested 3 September 2026.

FLUX kept the elevator recognisable, left the red gift where it belonged, separated the cast and even placed the gun on the floor after the first attacker fell. Considering what happened to the lobby in our original Gemini result, that stability deserves credit.

The choreography was the weakest of the three. It looked like people slowly rehearsing while somebody read instructions from the side. The hero eventually grabbed the gunman's wrist, but there was little acceleration, resistance or leverage, and the man bent backward before the visible force justified his fall. Meanwhile, Attacker B waited for his numbered ticket to be called. Our prompt told him to wait, yet a believable person would still circle, threaten or react. FLUX obeyed the order so literally that the fight looked staged in the least flattering way.

What the better prompt fixed, and what it could not

The rewritten prompt clearly reduced role confusion. Attacker A had one job, Attacker B had another, the hero remained the hero and the actions happened in a legible order, which was a large improvement over the disappearing characters, friendly fire and body merging from our first attempt.

Current models still seem better at understanding which events belong in a sequence than at generating the physical reason one event leads to another. They can show a gunman, a wrist grab and a defeated attacker, yet the gunman may enter grabbing range, the hand may take control without leverage and the body may fall before enough force reaches it.

We could improve Attacker B by telling him to circle and look for an opening without making contact until the first man falls. Even then, I would not claim that prompt engineering has fixed AI fighting. It fixed much of the role confusion here and exposed the harder problem underneath: contact, weight, reaction timing and cause-and-effect between bodies.

How to reduce character merging in your own AI videos

Give every person a visual identity, then attach every action to it. “The bald dude (attacker) on the left raises the handgun” is safer than “the first man attacks,” especially after people cross the frame. State who is active, who waits, where each person finishes and what must remain visible before the next action begins.

Use a wide, stable camera when continuity matters more than spectacle. Positional language helped all three completed models in this test, while timestamps gave them a better chance of preserving the order even when they compressed the movement.

For a clip you genuinely need, split the fight into shorter shots. Establish the hotel and gift, generate each interaction separately, then create the elevator ending from a controlled frame. Reference images can define the cast, clothing and room before movement begins; Google's advanced Veo workflow uses reference “ingredients” and first/last frames for this kind of tighter control.

Editing several useful pieces sounds less magical than a one-prompt demo, but it is probably how many suspiciously perfect launch clips became perfect in the first place. It can also become expensive: in our 10-second AI video cost test, a single generation ranged from $1 to $3.60 before paying for the attempts that failed.

Final verdict

The rewritten prompt stopped most of the character confusion in our hotel fight. Seedance 2.5 was the clear winner because it came closest to showing bodies responding to intention, momentum and resistance; MiniMax understood the plot but became mushy during contact, while FLUX protected the room and props at the cost of a painfully rehearsed fight.

Rewrite the hierarchy before paying for ten more copies of the same vague prompt. If the roles become clear and the physics still look wrong, the prompt has probably done as much as it can. Split the sequence into shots, use reference frames and edit the good moments together.

John Wick can relax for another week.


Sources