A one man production

THE BOXING REEL

30 seconds. One person. No camera, no gym, no crew.
written, directed and generated by SARTHAK
every image, every frame and the voice pipeline documented below
Director's Cut
A frame from the finished boxing reel
A frame from the finished 30 second reel. Fully generated.

00Synopsis

The goal was a 30 second reel where I train in an old Mumbai boxing gym and talk about the work I ship. Everything on screen is generated. The face is mine, the voice is mine, the gym never existed. The build has one rule that made it work: lock one thing at a time. First the face, then the outfit, then the location, then the voice, and only then the video. Every step below shows the input, the tool and the exact prompt, so you can run the same pipeline for yourself.

01The Tools

Higgsfield
HiggsfieldSeedream 5.0 Pro for every still. Seedance 2.5 for the video. Soul Cinematic for the empty location plates. All generation ran here.
Claude
Claude CodeModel: Fable 5. Used as the director and the pipeline: writing and iterating every prompt, reading frames, timing the cuts to the audio, and mixing the final sound with ffmpeg.
Sonilo
Sonilo MusicThe 30 second background track was generated here from a written music brief, then mixed under the voice locally.
ElevenLabs
ElevenLabsThe voice clone engine, used through Higgsfield. Cloned from a 30 second phone recording, used for drafts and alternate reads.
ffmpegCleaning the phone voice recording, and the final music mix with ducking under the voice. Free, local, no credits.
An iPhone voice memoThe 30 second voiceover was recorded on a phone in a quiet room. That one file drives the lip sync and the timing of every cut.

02Lock the face first

Start with 5 or 6 real photos: front, smile, both three quarter angles, one profile. Generate ONE close up master and make the model match it. Do not use many photos as references later, the models average the faces and you stop looking like you. After this step, the master image is the only face reference you ever pass.

Real reference photo, front
INPUT: real photo, front
Real reference photo, three quarter
INPUT: real photo, three quarter
Real reference photo, profile
INPUT: real photo, profile
Generated identity master
OUTPUT: the locked identity master, Seedream 5.0 Pro, 2K
The identity master prompt · Seedream 5.0 Pro, 2K, refs: your real photos
Tight chest-up studio portrait of the identical man from the reference photos, faithful likeness to the reference photos, facing the camera directly with a neutral relaxed expression and direct eye contact, head and shoulders centred, plain medium-grey seamless studio background, professional character reference portrait. Describe your real features here in detail: face shape, cheeks, jawline, nose, lips, eyes, eyebrows, hair colour and style, glasses, jewellery, skin tone. Then close with: visible fine skin texture with natural pores, slight natural sheen rather than glossy retouched finish, no digital smoothing, no beauty filter, no AI-airbrushed look, natural anatomy, high-end but unretouched commercial photography style, soft diffused studio lighting without harsh reflections, sharp focus on skin texture and eyes, single subject only, no props, no text, no watermark.

03Lock the looks

From the master, build the full set in the first outfit: full body, waist up, close up, both profiles. One generation per view, never a grid. Small faces inside grids drift. This collage becomes the single reference that carries face and clothes together.

The five view character collage
OUTPUT: the five view set, assembled after generating each view alone

04Lock the outfit

Generate the wardrobe alone, ghost mannequin style, no person in it. This becomes its own reference so the clothes never change between shots.

Ghost mannequin boxing outfit
OUTPUT: the boxing kit as a floating outfit, grey background
The ghost mannequin outfit prompt · GPT Image 2 or Seedream, 2K
OUTFIT ONLY, ghost-mannequin (invisible body, garments floating and styled as worn, no person, no head, no hands): old-school Mumbai boxing-gym training kit sized for a 24-year-old man, about 180 cm, about 88 kg, broad shoulders, solid build. Hero piece: black cotton ribbed tank vest, relaxed athletic cut, soft worn cotton with a faded texture, narrow straps, plain front, slightly darkened sweat patch at the chest, no graphic, no logo. Bottoms: black woven boxing shorts, mid-thigh length, wide leg, elasticated waistband with a plain black drawstring, side slits, matte nylon with a faint sheen. Hands: black cotton boxing hand wraps wound as worn, a little frayed and chalky at the knuckles, thumb loops visible. Footwear: worn black low-top canvas trainers with gum soles, laces loosely tied. Seamless medium-grey studio backdrop, even soft neutral studio light, no harsh shadows, about 5500K. Clean studio reference sheet, photoreal, fabrics read as real cotton and nylon with weave and wear, no retouching, no text, no watermark, no logos.

05Lock the angles

A video model needs your face from every side the camera will see. Generate the missing angles one by one, each using the master as the first reference. Keep every portrait on the same grey background so they read as one person.

Front portrait in the outfit
OUTPUT: straight on
Three quarter portrait in the outfit
OUTPUT: three quarter
Profile portrait in the outfit
OUTPUT: full profile
The angle prompt, reusable for any angle · Seedream 5.0 Pro, 2K, 3:4, ref 1: master, ref 2: real photo of that side
Chest-up studio portrait of the identical man from the first reference image, exact same face, same 24-year-old, youthful, no aging, same glasses, same beard, same hair, light natural sheen of sweat on the forehead, wearing the same black ribbed cotton tank vest as the first reference. Neutral calm expression, mouth closed. [PICK ONE ANGLE LINE: Full side profile, head turned 90 degrees, the same side of the face as the second reference toward the camera, eyes looking forward not at the lens. OR: Head turned about 40 degrees so the camera sees more of the same side of the face as the second reference, eyes to the lens.] Plain medium-grey seamless studio background, soft diffused studio light, unretouched natural photograph, visible skin texture, no beauty filter, no digital smoothing, single subject, no props, no text, no watermark.

06Register everything as elements

In Higgsfield, save each locked image as an element. In the video prompt you then write the element name instead of describing things again. This is the single biggest consistency win.

ElementWhat it holdsMade with
@sarthak-boxing1My face in the boxing kit, the identitySeedream 5.0 Pro
@front-portraitStraight-on face, used for straight shotsSeedream 5.0 Pro
@right-side and @3/4-righsideProfile and three quarter, used so only my good side is filmedSeedream 5.0 Pro
@outfitThe full boxing kit, ghost mannequinGPT Image 2
@locked-gymThe empty gym, the whole locationSoul Cinematic

07Lock the location

Generate the location empty, no people, like a real location scout. Describe the geometry precisely: where the bag hangs, where the light comes from, where the open floor is. The geometry decides which side of your face the camera sees later, so plan it on purpose.

The gym location in the final film
OUTPUT: the locked location. This gym does not exist.
The location plate prompt · Soul Cinematic, 16:9, batch of 4, pick one
Old Mumbai textile-mill hall converted into a neighbourhood boxing gym, Byculla, 3/4 angle wide shot from the floor, camera at chest height. Cavernous industrial room, gritty, sweat-stained and lived-in, not a fitness showroom. Tall cast-iron columns, exposed steel roof trusses, high arched windows with broken and dusty panes throwing hard shafts of afternoon light across the floor, full of floating dust. Focal point: one worn heavy bag on a rusted chain hanging from a beam, leather cracked, gaffer tape around the middle, with open scuffed concrete floor beside it where a boxer would stand. In the depth a simple rope ring with frayed padded ropes and a single bare tungsten bulb hanging over it. Clutter: skipping ropes on the floor, a steel water bottle and a crumpled towel on a wooden bench, chalk dust, a pedestal fan. Walls of peeling lime plaster and exposed brick. NO modern gym equipment, no machines. Generic unbranded objects, NO logos, NO brand names, NO visible text anywhere. Photorealistic, documentary, warm dust and crushed blacks, 2020s Mumbai. No people. 16:9.

08The voice

Record the full script on your phone in a quiet room, one continuous take, natural pace. Clean it locally, then upload it to the video model as the audio reference. The model lip syncs to it word for word and your own pauses become the edit points. Do not let the model invent a voice for the final film.

ffmpeg, the cleaning chain
ffmpeg -i voice_raw.m4a -af "highpass=f=80,lowpass=f=12000,afftdn=nf=-28,loudnorm=I=-18:TP=-2:LRA=9" -ar 48000 voice_clean.wav

09The master prompt

This is the full Seedance 2.5 prompt that generated the film in one 30 second run, with the cleaned voice uploaded as the audio reference. The structure matters more than any single line: scene, references, location map, first frame, format, optics, camera, action, performance, physics, lighting, grade, audio, style, locks. Cuts land on the ends of sentences in the recording. Punches never overlap words except where one word rides one punch by design.

Final film frame, punching
Final frame: the combination
Final film frame, talking
Final frame: the talk
Final film frame, the bench
Final frame: the bench
The full 30 second Seedance 2.5 prompt · Seedance 2.5, 16:9, 30s, 1080p, audio reference attached
SCENE CONTEXT
Afternoon. @sarthak-boxing1 walks and talks through an old Byculla mill boxing gym, telling the camera how fast he ships: finishing his hand wrap on the way in, fast crisp punches on the bag between phrases, one natural uppercut, a breather in the light, a moment on the bench, and out. Six shots, one environment, one light, cut on the phrase boundaries of his recorded voice. The voice is the provided audio reference, lip-synced word for word; the body, breath and sweat match it. Everything realistic: real boxing rhythm, real hand-speed, no exaggerated or impossible motion.

ACTIVE REFERENCES
@sarthak-boxing1, a 24-year-old man, youthful, fresh-faced, mid-session, sweating, relaxed confidence. HIS FACE IS THE SINGLE MOST IMPORTANT ELEMENT: it matches the reference 100% in every frame of every shot, the same full round-oval face with soft full cheeks, the same rounded jawline under the short full dark beard and connected moustache, the same straight medium nose, the same dark brown eyes, thick dark straight eyebrows, thick black textured quiff with short faded sides, small silver stud earring in each ear, thin black-bead and gold chain, thick black square acetate glasses that never come off. The face never narrows, never lengthens, never ages; identical from every angle and in every shot.
Face-angle references, used per shot as stated below: @front-portrait for any straight-on or near-straight view; @3/4-righside for any three-quarter view; @right-side for any profile view. Only these three angles ever appear; the other side of his face is never shown.
@outfit: black ribbed cotton tank vest, black boxing shorts, black hand wraps, worn black low-top trainers with gum soles. 100% matches the reference, identical in every shot.
@AUDIO: his real recorded voice, one continuous take. Dialogue is lip-synced to it exactly, word for word, with natural relaxed mouth movement, normal jaw opening, no exaggerated lip shapes, small natural head movements while speaking, exactly like a real person talking on camera. Its phrasing drives the cuts and the punches. Layer his natural breathing, sharp exhales on the punches, breath settling by the bench. No generated voice, no subtitles.
(No other people in frame. No props beyond what is already in the room: the heavy bag, the wooden bench, the glove rack, the towel, the bulb.)

LOCATION MAP
@locked-gym: STYLE REFERENCE ONLY, not a fixed keyframe; model extends the world, the subject moves through real space, never pinned 1:1. Foreground right of centre: the heavy bag on its long chain, open scuffed concrete floor to its LEFT where he stands. Left wall: high mill windows throwing hard shafts of dusty afternoon light across the floor from left to right, a cast-iron column mid-frame, a leaning mirror and a speed bag in the depth. Background: the rope ring with a bare tungsten bulb hanging above it, peeling plaster walls, wooden roof. Right wall: gloves and mitts on hooks, a wooden bench with a towel and a steel bottle in the right foreground. Dust hangs in the air. The room, the light direction and the time of day are identical in every shot.

FIRST FRAME / BLOCKING
Non-empty opening frame: medium-wide, @sarthak-boxing1 already walking in from frame-left through a shaft of window light toward the bag, pulling the LAST turn of the black wrap over his left knuckles, towel over his shoulder, his face in the three-quarter angle of @3/4-righside toward the lens, the heavy bag a dark shape ahead of him right of centre. Asymmetric framing, subject already in motion.

FORMAT MODE
Multi-shot. Six shots, hard cuts, no dissolves, no transitions. Each cut lands on a phrase boundary in @AUDIO (the end of a sentence) and on an action (a punch, a lean, a step, a sit). Same room, same light, same wardrobe in every shot.

OPTICS
Rectilinear primes throughout, anamorphic optical flares off the window shafts, prime-lens character, 180 degree shutter motion blur. Shot 1: 35mm medium-wide. Shot 2: 35mm medium. Shot 3: 50mm MEDIUM-CLOSE, chest-up, never tighter than chest-up, the whole head with air above it always in frame. Shot 4: 35mm medium. Shot 5: 50mm medium-close, chest-up on the bench. Shot 6: 35mm medium. NO extreme close-ups of the face anywhere in the film; the tightest framing in any shot is chest-up.

CAMERA
Handheld throughout, live operator breathing, organic micro-drift, short whips that resettle. The camera only ever sees him straight on, in the @3/4-righside three-quarter, or in the @right-side profile; it never crosses to the other side of him.
Shot 1 (line 1): medium-wide, walking backward ahead of him as he walks and talks toward the bag, his face held in the @3/4-righside angle, the light shaft passing over him.
Shot 2 (lines 2-3): medium at the bag, starting in the @right-side profile at chest height as he punches on the words, a small snap of the frame with each hit, then drifting to the @3/4-righside angle as he stops the bag and leans on it for line 3.
Shot 3 (lines 4-5): medium-close, chest-up, his face near straight on as in @front-portrait but the framing loose; a fast two-punch burst at the cut snaps the frame; he speaks; then a half-step pull-back for one natural rising uppercut, and back to the loose chest-up as he steadies the bag and delivers the three-weeks line, face fully visible, never macro.
Shot 4 (lines 6-7): medium, he pushes off the bag and takes slow steps back into the light toward frame-left, camera tracking alongside, face in the @3/4-righside angle; on "the sweat" the camera drops slightly to catch his forearm.
Shot 5 (line 8): medium-close on the bench, chest-up, face straight on as in @front-portrait, static with breath in it, the bag soft behind.
Shot 6 (line 9): medium, he stands, dry grin, last line to the lens, turns to the @3/4-righside angle and walks out of frame; the bag swinging gently behind.

ACTION
Shot 1: he pulls the final turn of the wrap over his left knuckles, presses the velcro strap closed with his right hand, the wrap is FINISHED in one clean motion within the first two seconds and his hands never wrap again, flexes the fingers once, and talks as he walks.
Shot 2: at the bag, profile to camera; on each word, "Websites", "Apps", "Dashboards", "AI videos", "Product shoots", one fast, crisp, realistic punch, real boxing hand-speed, straight lefts and rights snapping the bag, sweat flicking off with the sharpest hits, the words riding the rhythm; after the fifth punch he catches the bag with his forearm and leans on it for "one guy, one laptop, and a lot of AI."
Shot 3: on the phrase boundary a fast one-two, two quick straight punches, real speed, no wind-ups; he catches the bag, speaks the rounds line; then one natural rising uppercut, thrown from the legs the way a boxer actually throws it, the bag lifting slightly on its chain and dropping back; he steadies it with both hands, chest rising, one exhale, and speaks the three-weeks line, still, chest-up framing, face clear.
Shot 4: he shoves the bag away and walks back into the light while speaking; on "the sweat" he looks at his own forearm, a drop running; wipes his forehead with the back of the wrap; "okay, the sweat is real" with the glance back to camera.
Shot 5: sits on the edge of the bench, towel off the shoulder into his hands, still, speaks to the lens.
Shot 6: stands, dry grin, last line, towel in hand, walks out of frame.
Camera: described in CAMERA, six shots, cuts on phrase boundaries, no self-cuts inside a shot.

PERFORMANCE
A real athlete between rounds, not an actor: punches are fast, compact and economical, elbows in, feet pivoting, no flourishes; breath audible, shoulders loose, sweat at the temples. Words during punches ride the rhythm naturally; everywhere else lines are spoken with the head easy and mobile like real speech, small nods, small tilts, mouth unobstructed, lips readable, eyes on the lens. Lip-sync natural and relaxed: normal jaw, no over-articulation. Small wry half-smile in the eyes. A straight sincere look for the three-weeks line and the closing offer. The dry grin only on the last line. Pore-level realism: smooth young skin with real texture, sweat beads and runs, vellus hair, wet living eyes with window catch-lights. Glasses stay on and stay clear.

PHYSICS
The bag has real weight: it dents under the knuckles, swings and recoils on its long chain, the chain creaks; the uppercut lifts it slightly and it drops back, nothing cartoonish; stopping it rocks him back a step; leaning tilts it. Punches travel at real human speed with real motion blur. Sweat leaves his skin on the sharp hits and falls in real arcs. Wraps crease at the knuckles, the strap stays closed, the vest sticks damp to his back, the bench creaks, the towel has weight. Correct contact shadows under him and under the bag. Nothing floats, nothing loops.

LIGHTING
Natural light only, identical in every shot. Hard afternoon window shafts from the high windows on the left wall, raking left to right, full of floating dust; sweat and the bag leather catch the shafts as speculars; the bare tungsten bulb over the ring in the depth as a warm background point. WB locked 5000K.

COLOR GRADE
Warm dust and worn leather neutrals dominant, crushed blacks in the mill shadows, the hard window shafts as the secondary beam, one warm amber accent living only in the sweat highlights, the bulb and one lens-flare glint, color tied to source and surface.

AUDIO
NO MUSIC. SFX ONLY plus @AUDIO, diegetic sound throughout. No score, no soundtrack. Gym ambience, the five fast punches under the service words, the one-two and the single uppercut thud, chain creak, feet on concrete, the bench creak, distant Mumbai traffic and crows.
VOICE: @AUDIO is the voice. Lip-sync exact, word for word, natural and relaxed; its phrasing drives the cuts and the punches. Around it: sharp exhales on every punch, one catch of breath after the uppercut, breath settling on the bench. Voice mixed clearly above the ambience. No subtitles.
The lines in @AUDIO, in order:
1. "Hey. I'm Sarthak. People ask how I ship so fast."
2. "Websites. Apps. Dashboards. AI videos. Product shoots."
3. "One guy, one laptop, and a lot of AI."
4. "I don't do long timelines. I do rounds."
5. "Three weeks of work in three days. That's not a flex. That's just AI, used properly."
6. "And no, nothing here is real. The gym, the bag, the sweat…"
7. "…okay, the sweat is real."
8. "If you want something built, and built fast, you know where to find me."
9. "And if you want the prompt for this, comment BUILD."

STYLE
8K photorealistic, no 3D render, no game engine, no game-cutscene aesthetic. Naturalistic master cinematography, anamorphic optical flares, fine film grain.

OUTPUT SETTINGS
16:9. 30 seconds. Real-time throughout, no slow-motion, no speed ramps.

POSITIVE LOCKS
Six shots, hard cuts only on phrase boundaries of @AUDIO, no transitions. Same gym, same light, same wardrobe in every shot. His face matches the reference 100% in every frame, full round-oval face, full cheeks, rounded jaw, same beard, same glasses on always, 24 years old, no aging, no narrowing of the face. The tightest framing anywhere is chest-up; no extreme close-ups of the face. The hand wrap is completed once in shot 1 and never wraps again. The heavy bag stays on its chain right of centre and is the only thing he hits; he always stands to its left facing it. His face is only ever seen straight on (@front-portrait), three-quarter (@3/4-righside) or profile (@right-side); the other side never appears. Punches are real human speed; one uppercut only. Voice is @AUDIO only, lip-synced naturally. No other people. No bell, no extra props. No text anywhere. Eyes stay natural, no eye glow. Contact shadows read clearly.

10The music and the mix

The clip comes back with voice and gym sound only. The music is a separate 30 second track, mixed underneath locally so it ducks when the voice speaks and breathes back up in the gaps.

Sonilo Music, 30 seconds
Uplifting, fun, swaggering instrumental for a 30-second boxing-gym reel, 100 BPM, no vocals, no pads, clean start, hard clean ending. Dusty funk-breakbeat drums, hand claps on 2 and 4, warm round bass, a bright plucked guitar riff, short punchy brass stabs, a low dhol and tabla layer buried in the mix. Mid-range left open for a spoken voice on top. Sparse intro for 3 seconds, full groove with stabs landing like punches, a lift in the middle, a breather, a playful lick, warm held chords near the end, then the full groove returns and lands one clean final hit with a short tail.
ffmpeg, the ducking mix
ffmpeg -i dialogue.wav -i music.wav -filter_complex "[0:a]loudnorm=I=-16:TP=-1.5:LRA=8,asplit=2[dv][sc];[1:a]volume=-9dB,afade=t=out:st=29.5:d=0.5[m];[m][sc]sidechaincompress=threshold=0.05:ratio=3.5:attack=60:release=500[md];[dv][md]amix=inputs=2:duration=first:normalize=0,alimiter=limit=0.95[out]" -map "[out]" mix.wav
ffmpeg -i film.mp4 -i mix.wav -map 0:v -map 1:a -c:v copy -c:a aac -b:a 320k final.mp4

11What not to do

I burned a lot of credits learning these. Take them for free.

  1. Never put many faces in one image. Grids and split-screens shrink the face and the likeness dies. One portrait per generation.
  2. Never pass many photos as references at once. The model averages them into a stranger. One locked master beats six selfies.
  3. Describe only what changes. The face lives in the reference, never in words. Wardrobe, light and voice live in fixed blocks you paste unchanged into every prompt.
  4. Say what you want, never what you do not want. "A plain pocketless shirt" works. "No pocket" gets you a pocket.
  5. Keyframes are neutral. Mouth closed, calm face. All acting, biting, smiling and talking happens in the video model, not in the still.
  6. Eye level beats low angle for likeness. Wide lenses close to the face distort it. 50mm at eye level is the safe home.
  7. No talking over action. Punches in the pauses, words in the stillness. Lip sync dies under motion blur.
  8. Record your own voice. A real 30 second voice memo gives you your accent, your pace, and free edit points at every sentence end.
  9. Avoid extreme close-ups in AI video. Chest-up is the tightest safe framing for a talking face.
  10. Lock one thing at a time: face, outfit, angles, location, voice, then video. If a step is not right, do not generate the next one.

Want something built properly?

DM me. One person. Every department.

Made in Mumbai. No camera was used.

The character sheet Read next The Character Sheet How to make AI hold your face. Every view, every prompt. The ride into war Read next The Ride Into War A present-day ride that match-cuts into a Rajput battle.