Guides / Filmmaking

Make a video from a script, step by step

You have a script. The mistake is to paste all of it into one prompt box. Split it into shots, give each shot its own camera move, record the words separately, and cut the pieces together. Here is one 66 word script taken all the way to a finished piece.

  • 12 min read
Shot listShot 1VoiceThe cutFinished

What you make, in order. Tap one to jump to its step.

This guide follows one script from the page to an MP4. The script is 66 words, four beats, and every beat became a 6 second clip, three on Kling v3 and the rooftop on Seedance 2.5, each with the camera move from the composer's chip row. The narration is ElevenLabs v3, the music is Lyria 3.5, and the cut was made in Studio and exported from it. Every clip and still here is real output, made for this guide, and every performer is fictional.

One thing to know first: this is the fastest honest shape for a scripted video, and it deliberately gives each shot a different person. A script that needs the same character in every shot is a different job; that is the short film guide.

What goes wrong

MistakeWhy it failsFix
Pasting the whole script into one promptThe model reads it as one long description and returns one generic clip with no cutsOne shot per beat, 4 to 8 seconds, each with its own prompt
Writing feelings instead of pictures"She feels lost" gives the camera nothing to seeWrite what the lens sees: the place, the light, one action, the frame size
Asking the video model to speak the scriptNative speech costs more, drifts from the words, and cannot be re-timedRecord the voice separately in the audio composer and lay it under the cut
Forgetting the cameraEvery shot comes back as a locked wide, and the edit feels like slidesPick a camera move chip per shot: push-in for the hook, crane up for the reveal
Leaving a performer's wardrobe to chanceA jacket you did not describe comes back open, or missingName the garments, how they close, and the colours

Before you start

  • A script of 60 to 90 words. Spoken at a calm pace that is 25 to 35 seconds, which is four to six shots.
  • One line per shot saying where it is, who is in it, one action, and the frame size (wide, medium, close).
  • A NOLGIA account. New accounts start with 350 free credits, and the price of every shot is on the Generate button before you press it.
  • Nothing to upload. Every shot here is Text to video.
The script (66 words)
At two in the morning the city belongs to the ones who are still awake. The dancer in the empty car park, rehearsing for nobody. The chef closing up, lit by the last blue flame. The two friends on a rooftop who missed the last train and did not mind. The taxi driver singing to an empty back seat. Nobody is watching. That is the point.

Turn the script into a shot list

Read the script aloud and mark where the picture changes. In ours it changes four times, one per sentence about a person, so the shot list has four rows. The opening line and the last two lines are narration over the first and last shots, not shots of their own.

ShotFrameCamera moveWhat the lens seesSeconds
1 DancerWidePush-in, mediumEmpty car park, magenta and cyan strip lights, one dancer alone in the middle6
2 ChefMediumTruck left, subtleDark kitchen, one gas burner still lit, a chef wipes down and turns it off6
3 RooftopWide from behindCrane up, mediumTwo friends on a ledge, backs to camera, the skyline behind them6
4 TaxiMedium, from the back seatHandheld, subtleThe driver sings and taps the wheel, rain and neon on the glass6
Four shots, 24 seconds of picture, under 25 seconds of narration. The camera move column is a chip you press, not a sentence you write.

Tip:Describe the wardrobe as a costume department would. Name the garments and how they close, and the model keeps them. The prompt says "black hooded tracksuit, the jacket zipped up to the chin", and the jacket stays zipped for the whole dance.

Generate each shot in the video composer

Open the video composer. Under Start from, keep Text to video. Pick the model with the chip above the prompt: we used Kling v3 (Kling 3.0 in the menu) for the dancer, the chef and the taxi, because it keeps a moving body believable, and Seedance 2.5 for the rooftop. Type the shot's prompt, then press one chip in the Camera move row and pick a strength. The composer shows the exact sentence it will add to your prompt under the chips, and never rewrites what you typed.

  • Set Length to 6 s and the shape to 16:9. Kling v3 renders 3 to 15 seconds and Seedance 2.5 4 to 30; the composer only offers lengths the model can make.
  • The price is on the Generate video button. It is the live quote for this model, length and quality, and it is what you are charged.
  • Press Generate once per shot. Each clip lands in your Library and in the Recent generations row under the form.
Shot 1, as sent (camera chip: Push-in, medium)
A multi-storey car park at night, empty, wet concrete floor reflecting magenta and cyan strip lights, thin haze. A young woman with cropped bleached hair wearing a black hooded tracksuit with a thin white stripe, the jacket zipped up to the chin, black cargo trousers and white trainers, dances alone in the centre of the frame, a slow controlled groove, arms cutting through the haze. Wide shot. Cinematic, anamorphic flare on the lights, fine film grain. No text, no logos.
A dancer in a zipped black hooded tracksuit moving alone in an empty car park under pink and cool strip lights
Shot 1, made with Kling v3. Push-in, medium. Second take: the first walked up to the lens instead of dancing.
A chef in a black apron over a white tee wipes down a steel pass beside a lit gas burner in a dark kitchen, then turns it off
Shot 2, made with Kling v3. Truck left, subtle. One take.
Two friends in a red puffer jacket and a camel coat sit on a rooftop ledge with their backs to the camera, dark rooftops and window lights below, as the camera rises
Shot 3, made with Seedance 2.5. Crane up, medium. Kling v3 put lit signs with lettering on the skyline in three takes; Seedance 2.5 kept it clean once the skyline was written as dark rooftops and window lights.
A taxi driver in a flat cap sings and taps the wheel, seen from the back seat, rain and neon on the windscreen
Shot 4, made with Kling v3. Handheld, subtle. One take.
Shots 2 to 4, as sent
SHOT 2 (Truck left, subtle): A small late-night restaurant kitchen, stainless steel, most lights off, one gas burner still lit with a blue flame. A chef in his forties with a shaved head, a black apron over a white tee, wipes down the steel pass, then reaches over and turns the burner off; the blue flame throws cool light on his face against warm amber pendant lights behind him. Medium shot, shallow depth of field, thin haze, cinematic. No text, no logos.

SHOT 3 (Crane up, medium, on Seedance 2.5): A city rooftop at night, two friends in their twenties sit side by side on a low ledge with their backs to the camera, a man in a red puffer jacket and a woman in a long camel coat; beyond them an anonymous city of dark rooftops and plain distant towers with small warm window lights, no famous landmarks, a hazy sodium glow in the sky, wind in their hair; one laughs and leans against the other. Wide shot from behind, haze, cinematic. No signs, lettering or logos anywhere.

SHOT 4 (Handheld, subtle): Inside a taxi at night, seen from the empty back seat: the driver, a man in his fifties with grey stubble and a flat cap, sings along to the radio and taps the wheel, rain on the windscreen, streaks of neon and traffic lights sliding across the glass and his face. Shallow depth of field, cinematic, film grain. No text, no logos.
The video composer on a computer with Text to video selected, the shot 1 prompt typed, the Push-in camera chip pressed at medium strength, Kling 3.0 chosen and the Generate video button showing the price
On a computer. The camera chip, the sentence it adds, and the price on the button.
The same composer on a phone: Kling 3.0, the prompt, the camera move row with Push-in pressed, and the length and quality chips
On a phone. The same form, one column.

Watch for this:Generate is one click, and the price is the charge. There is no confirmation step. The number on the button is the server's quote for this model, length and quality, and a submit at any other price is refused, so what you see is what you are charged.

Record the narration in the audio composer

Open the audio composer, choose Voiceover, paste the whole script as the prompt and pick a voice. We used ElevenLabs v3 with the voice Rachel at speed 0.95. Text to speech is priced by the length of the script, so the price on the button changes as you type. Our 66 words came back as 25.2 seconds of audio.

For music, the same composer makes a score from a description. We asked Lyria 3.5 for a slow synth pulse for a night-time city film and used the first 30 seconds under the cut. Keep the description about mood and tempo, not about the pictures.

The music prompt
A slow, moody synthwave pulse at 84 bpm for a night-time city film: warm analogue pads, a soft side-chained bass, distant reverb, no drums until the second half, no vocals, cinematic and intimate
The audio composer with Voiceover selected, the 66 word script in the prompt box, ElevenLabs v3 chosen and the price on the Generate button
The narration, ready to generate. The script is the prompt, and the price follows its length.

Cut it together in Studio

Open Timeline and press New timeline; it asks for a project first, so name a new one. In the Media column on the left, open Library, find each shot and press its +: the clip lands at the playhead on a fresh track. Press End to move the playhead to the end of what is there before you add the next one, so the four shots run in sequence. Then press Home and add the narration and the score the same way; they land on audio tracks under the picture.

  • Studio saves as you go. There is no Save button.
  • A 2 minute score is longer than the cut. Put the playhead at the end of the picture, press C with nothing selected to split every clip under it, click the tail of the score and press Delete.
  • Press Export video in the header when the cut plays the way you want. The MP4 lands in your Library, at no charge.
Studio on a computer: four video clips staggered across four tracks in sequence, the narration and the score on two audio tracks beneath, the dancer in the program monitor at one and a half seconds
The cut on a computer. Four shots in sequence, each on the track its + gave it; the narration on A1, the score on A2, trimmed to the picture.
Studio on a phone: a note that the editor is best edited on a desktop, the program monitor, and the NOLGIA Agent chat for edits
On a phone. Studio plays the cut and hands edits to the agent; the timeline itself wants a bigger screen.

Watch for this:The last line lands on black. Four 6 second shots are 24 seconds of picture and the narration runs 24.6. Studio has no fades, so we let the timeline run to 25 seconds: the last words land on black and the score stops with the picture. To avoid it, write one fewer word, or make the last shot 8 seconds.

The finished piece

The finished 2 a.m. piece: the dancer, the chef, the rooftop and the taxi in turn, with narration and music
Made with Kling v3, Seedance 2.5, ElevenLabs v3 and Lyria 3.5, cut in Studio and exported from it. Four shots, one script, one export.

Twenty five seconds, four performers, one voice. Nothing in it needed a face to match from shot to shot, which is exactly why this shape is the fast one.

Which tool for which job

JobWhereWhy
Shots from text, one per beatVideo composer, Text to videoCamera move chips, live price, 16:9 at 6 s
One character's look across shots (wardrobe, props, set)Characters plus the image composer, then Animate an imageSee the short film guide
Narration and musicAudio composerVoice priced by the words, music by the run
The cut and the exportStudio, from TimelineTrim, layer audio, export at no charge
A hands-off first draftNOLGIA AgentPaste the script and ask for shots; you still review each one

Where you still regenerate

  • Signs you did not ask for. Kling v3 put lit signs with lettering on the rooftop skyline in three takes, even with "no signs" in the prompt. Seedance 2.5 rendered it clean once the skyline was described as what is there: dark rooftops and window lights.
  • A walk-on instead of a dance. The first dancer take stepped up to the lens instead of dancing. Say the action again at the end of the prompt, and run it twice.
  • Sound. Kling v3 renders its own sound in the composer; we made these shots without it, so the narration and the score are the whole soundtrack. If a shot's sound fights the score, mute its track in Studio.
  • Length. A beat that needs more than one move is two shots, or the Extend option on the same model.

What it costs

Each shot is priced by model, length and quality, the narration by the number of characters in the script, the music per run. The cut and the export are free. The live rates are below.

  • Kling v3Pro and up

    Per 5s clip (audio on by default)

    720p 35 credits · 1080p 47 credits · 4K 117 credits

  • Seedance 2.5Pro and up

    Per 5s clip

    480p 26 credits · 720p 56 credits · 1080p 137 credits

    1080p renders at 16:9 and 9:16

  • ElevenLabs v3Every plan

    Per 1,000 characters

    Standard 6 credits

  • Lyria 3.5Every plan

    Per generation

    Standard 5 credits

Read live from the catalog when this page loads. Every price is shown before you generate. See every rate.

Checklist

  • Script read aloud and timed; one shot per change of picture.
  • Every shot prompt names the place, the light, one action, the frame size and the wardrobe.
  • A camera move chip pressed on every shot, strength chosen.
  • Narration recorded separately, music described by mood and tempo.
  • Shots trimmed in Studio, voice aligned to the first shot, own sound muted where it fights the score.
  • Exported once, watched once end to end before sharing.

Try it yourself

The four prompts above, the camera chips named beside them, Kling v3 at 6 s and 16:9 (Seedance 2.5 for the rooftop), and the script pasted into the audio composer as Voiceover on ElevenLabs v3 reproduce this piece. The model rolls differently every time; the place and the wardrobe are what the words hold.

The clips

  • Shot 1, the dancer (6 s)MP4 video · 1280x720 · 1.3 MBDownload
  • Shot 2, the chef (6 s)MP4 video · 1280x720 · 1.3 MBDownload
  • Shot 3, the rooftop (6 s)MP4 video · 1280x720 · 630 KBDownload
  • Shot 4, the taxi (6 s)MP4 video · 1280x720 · 1.7 MBDownload
The four source clips, as generated, so you can build the same cut in Studio.

Questions and answers

How long a script fits in one video?
60 to 90 words is 25 to 35 seconds spoken, which is four to six shots of 6 seconds. Longer scripts are the same method with more rows in the shot list. Count the seconds: our 66 words ran 0.6 seconds past four 6 second shots.
Can I paste the whole script into the video composer?
You can, and you will get one clip that tries to show all of it. One shot per beat is what gives you cuts, camera moves and a voice you can move.
Will the same person appear in every shot?
Not with this method; each shot has its own performer, on purpose. For one character across shots, save a character and build the shots from frames, as the short film guide does.
Why not let the video model speak the lines?
It costs more per shot, the words can drift, and you cannot re-time the delivery against the cut. A separate narration is cheaper and stays editable.
Which model should I use for the shots?
Kling v3 for shots where a body moves, and Seedance 2.5 where Kling kept drawing signs. Veo 3.1 is for a line spoken on camera; this method records the voice separately, so it does not need it.
What happens if a shot fails?
The run stops and the credits come back. Press Generate again once the price on the button is what you expect.

Turn your script into shots

Open the video composer

Filmmaking

A short film with AI: the pipeline from script to final cut

A 25 second film, Last Train, from a five line script to a graded cut: a saved character, a saved location, a prop, frames first, then shots, sound, Studio, and the export to Premiere or Resolve.

· 13 min read