I handed Claude and ChatGPT the same 30-second film and asked each to direct it. One continuous shot, nowhere to hide. Same story, same star, same location. Two completely different films came back. Here’s what they chose, what it cost, and why the gap between them is the whole point.
Give two AIs the same creative brief, and they’ll hand you back two different films.
Picture the scene. A man in a battered black leather jacket steps out of a plane, high above a coastline you might just recognise. Strapped to his chest is a hard case, the kind you’d expect to hold nuclear launch codes. His face says the fate of the free world is riding on the next few minutes. He freefalls over the water, snaps open a canopy, lands soft on the sand, and kneels beside a small child with the gravity of a man completing the most important mission of his life.
He opens the case.
Inside, cradled in custom-cut foam, sits one bright yellow plastic beach spade.
Cut to later. The same man is buried up to his neck in sand while the kid cheerfully pats the mound flat. He turns his eyes to the camera and holds the look. No words. He doesn’t need any.
I didn’t shoot a frame of this. I gave the idea to two of the best AI models available right now, Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra, both on their maximum intelligence setting, and asked each one to direct it. Same story. Same star (Max, my AI alter ego, who is basically me if I’d made very different life choices). Same beach, same 30 seconds, one continuous generation with no edit to rescue a weak take. Each model wrote its own script and its own shot list. Then I sat back and watched two different films arrive.
That gap, between two directors handed the identical brief, is the part worth your Saturday.
Why one shot, and why it matters
“Can AI make a video” stopped being an interesting question a while ago. The answer is yes, and it has been for a bit.
The interesting question now is quieter and harder: how good a director is the thing? Directing isn’t generating pixels. It’s choices. What do you show first. Where do you put the camera. When do you cut the music. How long do you hold on a face before the joke lands. Get those right and 30 seconds feels like a film. Get them wrong and it feels like a tech demo.
So I set the rules to make directing the whole test. One story. Thirty seconds. A single continuous generation using the strongest video model I can get my hands on today (SeeDance 2.5). No cutting between takes to quietly fix a bad decision. Whatever the model chooses, it lives or dies in one breath.
Then I gave the exact same brief to both and let each one direct.
The unglamorous part that changes everything
Here’s the bit most people skip, and it’s the bit that actually matters.
I didn’t type a paragraph and pray. Before either model directed anything, I cast the film. Max has a character reference sheet: same face, same jacket, same swept-back hair from every angle. I built the rest of the world in Midjourney as reference images. The coastline from the air. The kid on the beach in the red hat. The little black case. The yellow spade in its foam. Then I handed those images to SeeDance alongside each model’s prompt.

Think of it as casting and set-dressing before you roll a single frame.
Skip that step and you’re leaving everything to the video model’s imagination, which really means its training data. You’ll get a man and a beach and a spade, and they’ll drift shot to shot, because the model is guessing each time. Give it references and you get your man, your beach, your spade, holding steady across a continuous 30-second take.
That’s the difference between a lucky clip you got once and a scene you can actually control. If you take one thing away for your own experiments, take that.
Same brief, two directors
This is where it got genuinely fun.
Claude opened on the child. A small face turned up to the sky, hopeful, waiting for something you can’t see yet. Only then does it cut to the deadly-serious man in the plane. The mystery does the heavy lifting. Your brain is already asking “who is he doing all this for?” a full beat before the punchline arrives.
ChatGPT opened on the action. Straight into the aircraft, goggles on, a committed dive over the skyline. Pure adrenaline, no riddle, then the reveal. It plays like a straight action sequence that happens to end in a joke.
Both stuck the landing on the deadpan finish, and even there they diverged. Claude held a flat, unbroken stare into the lens, the pure resignation of a parent who has been here before. ChatGPT had Max start with his eyes shut, almost peaceful, then slowly open them to camera, followed by a tiny “hmmph” and a face-shrug that really lands the ending. Softer. More surrender-to-fatherhood than sharp irony.
Same story. Same last shot on paper. Two different films in the feeling.
Neither is wrong. That’s the whole point of a director. Personally, I loved how Claude started with the kid, and I loved how ChatGPT chose to end. I’m not going to crown a winner. I’ll let you watch them back to back and pick your own. (I have my leanings. I’d rather hear yours first.)
What it actually cost
Let’s talk money, because “look what AI can do” is cheap and the invoice usually isn’t.
Claude’s planning and final prompt ran about $12.81 in credits. ChatGPT chewed through a five-hour usage window fast, though no extra cash on top of the monthly plan. Midjourney sat inside an existing $8-a-month plan, so no marginal cost for the reference images. The real spend was Runway credits, roughly 2,040 per generation, because SeeDance at 1080p and a full 30 seconds only runs on credits (and Runway is retiring its unlimited tier at the end of November, which stung a little). Total time, including sitting around deciding how to frame the comparison, was one to two hours.
So, not free. Not even close to free if you generate a lot.

Now price it against the alternative. To shoot this for real you’d need a plane or a helicopter, a licensed skydiver who can act, a film crew, a permit to fly over one of the most carefully managed skylines on earth, a child actor with a chaperone, insurance, and a full day of everyone’s time. You’re not comparing AI to zero. You’re comparing it to a five-figure shoot and a week of logistics.
Suddenly thirteen dollars and an afternoon reads very differently.
The bottleneck moved
A year ago, the headline on an experiment like this would have been about physics. Bodies that moved like real bodies. No melting faces, no warping limbs, no hands sprouting a seventh finger. That alone was the win.
A few months later, the frontier moved to consistency: keeping a character recognisably the same person from shot to shot. I showed that was basically a solved problem in my last Max short (see below).
Now both are table stakes. Max looks like Max and moves like an actual human the whole way through, in both films, without me fighting the tool for any of it.
Which is the actual story here. When the tool stops being the bottleneck, the bottleneck becomes you.
Your taste. Your story. Your instinct for opening on the child instead of the plane, for cutting the music the instant the spade appears, for holding on a deadpan face two seconds longer than feels comfortable. The models can direct now. The open question is whether the person feeding them can.
The most valuable thing in that little black case was never the spade. It was the string of small choices around it, what to show, where to point the camera, when to say nothing at all. The tools have quietly caught up on almost everything else. That part is still ours.
So pick your own absurd 30 seconds. Cast it properly. Then go direct something worth watching.
(The geek in me remains ever hopeful you will.)
Discover more from Hotelemarketer by Jitendra Jain (JJ)
Subscribe to get the latest posts sent to your email.

Claude – even with the +\- $5 extra cost…..