Guides / Audio

The same voice in every video

A narrator who sounds different in every episode is not a narrator. Save the voice on a character and stop choosing it by hand.

  • 10 min read
VoiceSpeak asBed and cue

What you make, in order. Tap one to jump to its step.

You keep the same voice across videos by saving it on a character. A character on NOLGIA carries a voice: a speech model and one of that model's named voices. When you make narration in the audio composer, the Speak as button picks the character and sets the model and the voice to the saved pair. You never choose them by hand again, so a line recorded next month is read by the same voice as a line recorded today.

This guide makes one narrator, a fictional campaign voice we named Isla Marr, and has her read a line over three unrelated clips: a runway, a night drive and a rooftop rave. Each clip also gets a music bed and one sound cue, laid under the voice in Studio. Every clip and every line is real output; the model and the voice never changed between them.

What goes wrong, and the fix

MistakeWhy it hurtsFix
Picking the voice by hand each timeOne session on v3, the next on Turbo, and the same name reads differentlySave the model and the voice on a character and use Speak as
Choosing a model with no voice listMiniMax Speech, Inworld, Dia and Orpheus publish no named voices, so there is nothing to pinUse ElevenLabs or Kokoro for a recurring narrator
Giving the character a voice clip and expecting narrationA clip voice is for models that take a reference track; text to speech needs a catalog voiceSet a catalog voice for narration, a clip only when a video model asks for one
Changing the speed per takeSame voice, different pace, reads as two peopleLeave speed alone; fit the script to the clip instead
Rewriting the line to sound more like the voiceThe voice is fixed; the words are the variableWrite for the voice you saved and keep sentence length steady

Before you start

  • A character on the Characters page. Creating one needs a name and at least one photo: press Add reference images, upload the photo, name the character and press Create character. Ours is a generated portrait, so no real person is involved.
  • A speech model with named voices. The Voice dialog lists only those: the ElevenLabs models and Kokoro.
  • The clips you want narrated, in your Library. Ours are three five second clips without sound: the runway and the rooftop on Kling v3, the night drive on Seedance 2.5.

Save the voice on the character

On the Characters page, open the ... menu on the character's card and choose Voice. The dialog offers two kinds: Catalog voice and Voice clip. Keep Catalog voice. Pick the Speech model, then the Catalog voice; a Play preview button plays a recorded sample of that voice for free. Add a language if you want one, then Save. The card now shows the voice under the description.

The Voice dialog for the character Isla Marr: Catalog voice selected, Speech model ElevenLabs v3, Catalog voice Sarah, a Play preview button with its sample line, an optional language field and a Save button
The Voice dialog. Catalog voice, ElevenLabs v3, the named voice Sarah. Play preview costs nothing.
The character card for Isla Marr on the Characters page, with her portrait in the Front slot, the badge Not checked yet, and the line Sarah · ElevenLabs v3 under the description
Saved. The card reads Sarah · ElevenLabs v3. Face check is left off; a voice does not need it.

Narrate with Speak as

Open the audio composer on the Voiceover task and press Speak as. Pick the character. The model picker switches to the saved model and the Voice list to the saved voice, and a line under the button confirms it: Speaking as Isla Marr · Sarah. Type the line and press Generate audio. We did this five times in a row for two guides, and the pair never moved.

The audio composer with Speak as Isla Marr selected, the note Speaking as Isla Marr · Sarah, ElevenLabs v3 in the model picker and the Generate audio button with its live price
Speak as sets both fields. The model and the voice follow the character.
Line one, over the runway
This is the sound of silk at midnight.
Line two, over the night drive
Nothing here was built to stand still.
Line three, over the rooftop
You already know the ending. Wear it anyway.

Watch for this:If the note says the voice is unavailable. Speak as checks that the saved model still publishes the saved voice. If a voice leaves the catalog, the note tells you to pick another on the Characters page rather than reading with a stand-in.

Put a bed and a cue under the voice

A line on its own sounds bare. Make a music bed in the audio composer's Music task and one sound cue in Sound effect, then build each clip's timeline in Studio: Timelines, New timeline, and in the left panel's Library tab the clip with its +, the take at the playhead where the line starts, the bed, and the cue.

Watch for this:Studio has no volume control yet. Every lane plays at full level, with no fade and no ducking. So the bed does not dip under the voice by itself: you choose the passage. Ask the music model for strong dynamics, a few seconds full and a few seconds hushed, then play the track in the Source preview, press Mark In where a full bar leads into a hushed stretch, Mark Out five seconds later, and press + with the playhead at 0.

Studio with the runway clip on V1, the narration on A1 from one second, a five-second window of the music bed across A2 and the camera shutters on A3 for the first second
The runway timeline. Clip on V1, Isla's line on A1, the bed window on A2, the shutters on A3. Every sound is its own item.
Runway bed, Lyria 3.5 (window from 60.1 s)
Sleek fashion-show runway track, 120 bpm, deep kick and a crisp clicking hi-hat, glossy synth stabs, instrumental, no vocals. Strong dynamics: two seconds of the full groove, then four seconds of a hushed, near-silent filtered pulse, then the full groove slams back in, and that shape repeats through the track.
Night drive bed, Lyria 3.5 (window from 57.3 s)
Dark synthwave night-drive track, 100 bpm, driving analog bass arpeggio, gated snare, cold pads, instrumental, no vocals. Strong dynamics: two seconds of the full drive, then four seconds of a sparse, hushed passage with only a low filtered bass, then the full drive returns, and that shape repeats through the track.
Rooftop bed, Stable Audio 3 Medium (window from 3.1 s)
Warehouse rave build for a rooftop party at night, 128 bpm, instrumental, no vocals: a soft muffled kick and a quiet rising filtered riser that builds for five seconds, then a huge drop with open hi-hats and a big supersaw lead, and that quiet build and big drop repeat through the track.

The cues are one short line each on ElevenLabs Sound Effects v2, trimmed with Mark Out so they stay clear of the voice: camera shutters for the first second of the runway, the engine and spray after the drive's line, and a crowd cheer as the rooftop confetti falls. Press play: the voice is the same across all three because the model and the voice were the same request each time, and it sits above the music.

A model in an emerald green long-sleeved gown walking a mirrored runway through strobes and haze, narrated over a runway beat with camera shutters
Runway. Line one over a fashion-show pulse, shutters at the start. The voice sits 5.6 dB above the music while she speaks.
An original white concept coupe with no badges driving past plain teal and magenta neon tubes on a wet street, spray from the tyres, narrated over a dark synth drive with the engine and spray after the line
Night drive. The same voice over a synth drive; the engine and spray come in after the line. 4.7 dB above the music.
A crowd in dark coats and hoods with arms raised under green lasers and confetti on a rooftop, narrated over a quiet build that drops with a crowd cheer
Rooftop. The line rides the quiet build; the drop and the cheer land with the confetti. 11.8 dB above the music.

We measured each export: no gap of silence longer than half a second, and the voice at least 3 dB above everything else while she speaks. Studio mixes every lane at the same level, so these numbers come from the passages we chose, not from a setting.

Three ways to run it

PathHowBest when
FastestAudio composer, Speak as, one line at a timeA handful of lines for a cut you are making now
Hands offBrief the NOLGIA Agent with the character's name; it can list your characters and generate audio with the saved voiceA series of episodes from a script
Most controlThe API: POST /generate/audio with the character's id, or the MCP tool nolgia_text_to_audio from your own codeBatch narration from a document or a pipeline

The API path is documented under Characters and locations: a character's voice is used by audio generation with the character id, and by video generation when you ask for the character's voice. We ran the composer path for this guide, not the other two.

Where you still regenerate

  • Delivery still varies a little. Same voice, same model, but each take is a new read. If a line lands flat, run it again; the voice will not move.
  • No volume or fade in Studio yet. A bed plays at full level. If a passage buries the line, move Mark In to a quieter one or ask the music model for more dynamics.
  • No cloning here. You cannot make a catalog voice from your own recording. A voice clip on a character is a reference for models that take one, not a text to speech voice.
  • Speed is API only. The composer has no speed control today.

What it costs

Saving a voice costs nothing; previewing one costs nothing. Each line bills the model's per character rate with a minimum. Each music bed and each cue is one flat charge. The clips were priced by length, and the Studio exports are not billed.

  • ElevenLabs v3Every plan

    Per 1,000 characters

    Standard 6 credits

  • Lyria 3.5Every plan

    Per generation

    Standard 5 credits

  • Stable Audio 3 MediumEvery plan

    Per generation

    Standard 3 credits

  • ElevenLabs SFXEvery plan

    Per generation

    Standard 4 credits

  • Kling v3Pro and up

    Per 5s clip (audio on by default)

    720p 35 credits · 1080p 47 credits · 4K 117 credits

  • Seedance 2.5Pro and up

    Per 5s clip

    480p 26 credits · 720p 56 credits · 1080p 137 credits

    1080p renders at 16:9 and 9:16

Read live from the catalog when this page loads. Every price is shown before you generate. See every rate.

Try it yourself

The narrator's portrait

  • Isla Marr portrait, 1408x768 JPEGJPEG photo · 1408x768 · 483 KBDownload
Create a character from it on the Characters page, set the catalog voice ElevenLabs v3, Sarah, then read the three lines with Speak as. Make the clips with the prompts below, 16:9, five seconds: the runway and the rooftop on Kling v3, the night drive on Seedance 2.5 with Audio Off (Kling v3 renders its own sound in the composer, so mute its track in Studio), and the beds with the three bed prompts in the step above.
Runway clip
Fashion film, one continuous shot. A model in a sculpted emerald green floor-length gown with long sleeves walks down a mirrored runway toward the camera, camera strobes flashing from the crowd in silhouette on both sides, white haze, slow motion, low tracking shot moving backwards, anamorphic flares, deep blacks. No text, no logo.
Night drive clip
Premium car commercial, one continuous shot. An original concept coupe that resembles no production car: a low, wide, wedge-shaped matte white body with a single thin continuous light bar across the nose, no grille, no badges, a smooth flush glass canopy, enclosed rear wheel arches and a sharp ducktail rear. It slides sideways through a wet night street lined with bare concrete walls and plain vertical neon tubes in teal and magenta, no shop windows, no signs, no lettering anywhere. Rain streaks in the air, neon reflections stretched across the asphalt, low side tracking shot, spray from the tyres, motion blur on the passing lights, cinematic contrast.
Rooftop clip
Music video, one continuous shot. A rooftop rave at night: a dense crowd in silhouette wearing long dark coats, hoodies and bucket hats, arms raised, under sharp green laser beams fanning through a thick haze layer, confetti bursting in slow motion, a city skyline of blurred lights behind them, low handheld camera looking up, backlit, high contrast. No text, no logo.

Your clips and beds will not match ours; a prompt gives the model the scene or the mood, not the take, so find your own hushed passage for the Mark In. The voice will match, because the character carries it.

Checklist

  1. Character exists with a name and a photo.
  2. Voice dialog: Catalog voice, a model with named voices, preview played, saved.
  3. Audio composer: Speak as shows the character and the note names the voice.
  4. One generation per line; check each take before you cut.
  5. A music bed asked for with strong dynamics, and one cue per clip.
  6. Studio: clip, take at the playhead, a five-second bed window from a hushed passage, the cue clear of the voice, Export video.

Questions and answers

Which models can be a character's voice?
The ones that publish named voices: the ElevenLabs speech models with ten voices and Kokoro with twenty. MiniMax Speech, Inworld, Dia and Orpheus have no voice list, so the dialog does not offer them.
Will every take sound identical?
The voice and the model are identical. The read itself is generated fresh each time, so pacing and emphasis can shift a little between takes. Rerun a line you do not like; the voice stays.
Can I use my own voice?
Not as a text to speech voice. A character can hold a voice clip of 3 to 30 seconds with your consent, for models that take a reference track, but narration from text needs a catalog voice.
Can a video model use the character's voice?
The API has a switch for it on video generation. We did not run it for this guide, so we do not describe how it sounds.
What is the language field for?
An optional tag such as en or en-US stored with the voice, what the character speaks. Leave it empty if you are unsure.
Can Studio lower the music under the voice for me?
Not yet. Studio plays every lane at full level with no volume, fade or ducking. Pick a hushed passage of the bed for under the line with Mark In and Mark Out, as in this guide.

Pin a voice to a character

Open Characters

Audio

Voice, music and sound effects for your video

Make a narration line, a music bed and a sound effect in the audio composer, then lay them under your clip in Studio. Which model does which job, and what each one costs.

· 10 min read