One Story · One Cast · A Voice Each

Multi-Voice StorybookOne Narrator Is a Choice, Not a Default

Most story audio asks a single throat to play an entire cast. C2Story casts the parts instead — a voice per character, saved to the character, identical from the first page to the last.

  • A voice descriptor written onto each character's profile
  • The same voice in chapter one and chapter twelve
  • Only the lines you actually wrote get spoken
  • Narration you can mute, recast or hand to a character
Free to cast a story Your story, your rights
Multi-voice storybook interface showing an illustrated spread on the left and a cast panel with a voice profile and waveform beside each character name on the right
The Cast Panel
Story · Cast · Performance

250,000+

Stories Created

80,000+

Characters Saved

45+

Art Styles

4.9/5

Creator Rating

Cast First, Perform Second

How the Multi-Voice Storybook Works

Three steps, and the middle one is the whole craft — deciding who sounds like what, once, for the whole book.

1

Bring the Story

Paste a story, a chapter or a finished manuscript. The cast is read out of the text, each character is designed, and every spoken line is pulled out verbatim and attached to whoever said it.

4,100 words · 12 scenes · 5 speaking characters found

2

Cast the Voices

Each character gets a voice descriptor — gender, age band, pitch, timbre, articulation — inferred from how your text describes them, not from their name. Change any of it, and the change is what gets saved.

Nana → female, older adult, low pitch, warm, unhurried

3

Perform the Pages

Lines are timecoded into each scene, the narrator does the description between them if you want one, and the storybook is rendered with every character speaking in their own voice.

Scene 4 · 3 lines · 11s · narrator off for this scene

A cast performance, or one narrator reading along?

These are different things and it is worth knowing which you want. A read-aloud is a reading tool: one clear narrator says every word of the book while the words highlight in time, so a beginning reader can follow the text with their eyes — the single voice is the point, and the read-aloud video is the page for it. A multi-voice storybook is a performance: the dialogue is cast, each character sounds like themselves, and description is narrated or dropped. Both are built from the same book and the same illustrations, so making one does not stop you making the other.

Six Things That Change

What a Multi-Voice Storybook Changes

The words stay yours. What changes is who says them, how long they take, and whether anyone is describing over the top.

A diagram comparing five storybook characters all sharing one grey narrator waveform against the same five characters each with their own distinctly coloured waveform

The Cast Becomes Audible

With one narrator, a reader has to work out who is speaking from the words around the line. With a cast, they know before the line is finished — which is how children follow a scene with four people in it.

Comprehension

A Voice Becomes an Attribute

It stops being a setting on the export and becomes a property of the character, stored beside their design. That is the only reason chapter twelve can be guaranteed to match chapter one.

Continuity

Pacing Follows the Dialogue

A scene lasts as long as its lines take to say plus a beat, instead of every page getting an identical slot. Short exchanges move quickly and a long speech is given the room it needs.

Pacing

Description Stops Being Compulsory

A single-voice reading has to narrate everything, because it is the only voice there is. Once the cast carries the dialogue, description becomes optional — you can mute the narrator entirely.

Structure

Delivery Enters the Text

Each line carries a note on how it is said — soft, teasing, flat — drawn from your prose. "She said quietly" stops being a phrase the narrator reads out and becomes something the performance does.

Performance

Dialogue and Narration Separate

They become two tracks with two switches instead of one undifferentiated read. That is what lets the same book exist as a full performance, a dialogue-only cut, or a silent export you score yourself.

Mixing

Things that only exist once the parts are cast

A grandmother who sounds eighty, not fortyTwo brothers who do not sound like one boyA whisper that stays a whisperThe same dragon, nine chapters apartA narrator who steps back when the cast talksA cat called both Smoke and the old cat — one voiceA shout that does not clip the next lineThe villain, never mistaken for the heroA line delivered flat on purposeA child voice that is not a pitched-up adultSilence where the text has no dialogueA whole scene with no narrator at allFirst person, told by the character who lostA question and its answer, two seconds apartThe same cast in English and in ChineseA voice that belongs to the character, not the export
Where the Voice Lives

A Voice Belongs on the Character Sheet

Not on the render, not on the session, not on the button you pressed last Tuesday. On the character — which is the only version of this that survives a whole book.

A character sheet with a voice profile card pinned beneath it, showing labelled fields for gender, age band, pitch, timbre and articulation

Five fields

What a Voice Descriptor Says

Gender, age band, pitch, timbre and articulation. "Female voice, older adult, low pitch, warm timbre, unhurried articulation" is enough to pin a grandmother down and keep her there. Short on purpose — it has to be reproducible.

Written once

Saved, Then Reused

The first time a character speaks, the descriptor is worked out and written to their profile. Every scene after that reads the saved one back. Nothing is re-decided, so nothing can be decided differently.

From the text

Read, Not Guessed

Gender and age come from how your story describes the character, not from their name — because names are unreliable and half of children's fiction is animals, invented names and titles anyway.

The same character under two different names

Stories rarely call anyone one thing. A cat is Smoke in chapter one and the old cat by chapter six; a woman is her name to her friends and her title to everyone else. Treated naively, each of those names becomes a separate character with a separate voice, and the cat changes throat halfway through a conversation. When two names clearly point at the same character, C2Story gives them an identical voice descriptor, so the voice follows the character rather than the label the sentence happened to use. It is the audible half of the same problem the character consistency tool solves for faces, and the character bible generator keeps the record of who is who.

Verbatim or Nothing

It Only Speaks the Lines You Wrote

A voice model asked to find dialogue in a dramatic paragraph will happily invent some. So every line is checked back against your text before anyone says it.

A dialogue extraction diagram showing two quoted lines lifted out of a page of story text as speech bubbles, while a character profile block and a synopsis block are crossed out as not speech

Spoken: Gets a Voice

  • Anything inside quotation marks
  • Lines written as “Name: what they say”
  • A shout, a whisper, a single word
  • Copied exactly — same words, same language

Not Speech: Never Voiced as Dialogue

  • ! A character sheet describing who someone is
  • ! A synopsis paragraph summarising the plot
  • ! Scene description and stage-setting prose
  • ! Any line that is not physically in your text

The last one is a hard check, not an instruction. A candidate line has to appear in the source passage character for character — punctuation and spacing ignored — or it is discarded before it can be spoken. Most setup passages contain no dialogue at all, and the correct number of voiced lines for them is zero.

The Fourth Voice

The Narrator Is a Role You Can Recast

Once the cast carries the dialogue, the narrator stops being the whole book and becomes one more part to direct — or to remove.

A narration settings diagram with toggle rows for narration, dialogue and music beside a switch between an off-screen narrator and a named character narrating in first person

Off-Screen, or One of Them

Third person is the default — an outside voice describing the scene from outside it. Switch to first person and a character narrates in their own voice, carrying their own feeling about what is happening. You pick which character, and it need not be the protagonist.

Or No Narrator at All

Turn narration off and nothing describes over the top: the scene plays on the illustrations and the cast alone. It suits dialogue-heavy stories and older children, who do not need to be told what they can already see on the page.

Three Independent Switches

Narration, dialogue and music each toggle on their own. Full performance, dialogue-only, narration-only, or all three off for a silent export you score yourself. The defaults — all on, a warm mid-range narrator, music following each scene's emotion — suit most picture books.

A Language of Its Own

Narration can follow the story's own language automatically or be pinned to a fixed one. That is how the same book ends up read in two languages by the same cast, with the characters and the artwork untouched.

Three mixes of the same storybook

Full

Narration + cast + music

The default. A narrator carries description, the cast carries dialogue, music follows the emotion of each scene. Closest to a bedtime audiobook.

Dialogue only

Cast, no describing voice

Nobody narrates. The illustrations do the description and the characters do the rest — closer to an animated scene than a reading.

Silent

All three muted

The book with no audio at all, for classrooms that read it live, or for scoring it yourself afterwards.

These are the same book, not three exports you have to rebuild. The cast, the lines and the illustrations do not change — only which tracks are playing.

Timing

A Scene Lasts As Long As Its Lines Do

Fixed slide durations are why so much story audio sounds rushed in one place and dead in another. Dialogue should set the clock.

A timing diagram showing timecoded speech blocks along a horizontal timeline, each labelled with a character name, and one overloaded block splitting into two shorter shots

Timecoded

Every Line Has a Start and an End

Lines are placed inside the scene rather than queued up back to back, so a question and its answer are two moments with a real gap between them instead of one run-on sentence.

Measured

Length From Speech, Not a Template

A scene runs for as long as its dialogue takes to say plus a beat of action. Two words get a short scene; a speech gets the room it needs. Nothing is padded to fill a slot.

Split

Too Much Speech Becomes Two Shots

When one scene carries more dialogue than a single shot can hold, it splits rather than speeding up or dropping a line. Your words survive; the shot count goes up.

Performed in the shot, or cast and mixed in

There are two ways a line reaches your ear, and both are used. In the native route the video model performs the line while it generates the shot, so the voice and the mouth come from the same take and stay in sync — the voice descriptor and the timecoded lines are what steer the performance, and there is no separate speech pass to drift out of alignment. In the cast route a voice is chosen for each character from a catalogue, matched on gender, age, personality and the story’s language, and the dialogue is voiced and mixed into the finished video. The native route wins on lip sync; the cast route wins on predictability, because the same voice comes back identically every time. In both, the character keeps their own voice — only the machinery underneath differs.

Art Direction

One Look and One Cast, Held Together

A character is a face and a voice. Both are decided once and both are reused, which is the only way a book stays one book.

45+ art styles including watercolor, cartoon, classic storybook and anime.

Casting Toolkit

What the Multi-Voice Storybook Tool Actually Does

Six things a general text-to-speech tool cannot do, because it has never had to keep a cast together across a whole book.

A Voice Is Part of the Character, Not the Export

Each character gets a voice descriptor — gender, age band, pitch, timbre and articulation — and that descriptor is saved onto their profile, not onto the render you happen to be making. Every later scene, chapter and session reads it back from the same place. This is what stops a voice from being re-chosen, and quietly re-chosen differently, every time you export something.

Only the Lines You Actually Wrote Get Spoken

Spoken lines are pulled out of your text verbatim and then checked back against the source, character for character. Anything that does not physically appear in what you wrote is dropped before it ever reaches a voice. Character sheets, a synopsis paragraph and plain scene description are not speech and are never dramatised into dialogue that you did not write.

The Narrator Is a Role, Not a Fixture

Narration is a switch. Leave it on and an off-screen voice carries the description; turn it off and the cast carries the scene alone. Or hand narration to a character and let them tell their own story in first person, in their own voice. The narrator has a voice, a language and a point of view, all of which you set — including which character it belongs to.

Timing Comes From the Lines

Each scene is timed by how long its dialogue actually takes to say, plus a beat of action — not by a fixed slide duration everything has to fit inside. When a scene carries more speech than one shot can hold, it splits into two rather than cramming or cutting. Lines get timecodes, so a question and its answer land where they should.

Two Names, One Character, One Voice

Stories call people different things. A cat is "Smoke" in chapter one and "the old cat" in chapter six; a woman is her name to some characters and her title to others. When two names clearly point at the same character, they are given the identical voice descriptor, so the voice does not flip halfway through a scene because the text changed what it called someone.

English, Chinese and Japanese

The story language is detected from the text itself and the cast is voiced in it, with voices chosen from the pool that actually matches that language rather than an accented approximation. Narration can follow the story language automatically or be pinned to a fixed one, which is how you get the same storybook read in two languages with the same cast.

Why voice drift is harder to catch than face drift

A character whose face changes between chapters is obvious: you can put the two pictures side by side and see it. A voice that changes is worse, because you cannot see it and you cannot easily compare it — by the time chapter nine sounds wrong, chapter one is twenty minutes behind you and all you have is a feeling that something slipped. Listeners do notice; they just tend to blame the story rather than the audio. The fix has to be structural rather than attentive: save the voice to the character, reuse the saved one, never re-derive it. That is the same argument the character consistency tool makes about faces, applied to the part of the book you cannot look at.

For Everyone With More Than One Character Talking

Who Makes a Multi-Voice Storybook?

Parents at Bedtime

You already do the voices. This is the version that keeps doing them on the nights you have lost your own — and does them the same way, so nobody complains that the bear sounds different tonight.

Teachers and Librarians

A class following a four-hander loses track of who is speaking within a page. A cast fixes that without you reading every part yourself, and the dialogue-only mix works for children who read the text along with it.

Audiobook Makers

A full cast recording is expensive and slow to schedule. Casting the parts from the manuscript gives you the shape of one in an afternoon — and a character's voice that is guaranteed identical across every chapter.

Series and Serial Authors

Ten instalments in, continuity is the whole job. A voice saved to the character means book three does not have to remember what book one decided, because it is reading back the same record.

Pick the Right Door

Is the Casting the Part You Want?

This page is for people whose story has several characters talking. If you pictured something else — one narrator, a video, a printed book — one of these is a better place to start. They all share the same character library.

Yes — cast my story

Start here. Bring a story, meet the cast, read the voice proposed for each character and change any of it before you spend anything.

One narrator, words highlighting

A reading tool for early readers: the whole book read word for word with the text highlighting in time. A single voice, on purpose.

My picture book, as a watch-along

A finished illustrated book turned into a narrated video with gentle motion on every spread. Format first, casting second.

An illustrated book plus a trailer

The story as a book and a short animated trailer for it. Start here if you want the book to exist first.

Something calm for bedtime

Soft motion, a soothing read and gentle music, tuned for winding down rather than for performance.

My characters keep changing faces

The visual half of the same problem: lock each design once and reuse it everywhere, across a whole run.

I keep losing track of who is who

Pin looks, voices and history into one reusable record, so chapter twelve still agrees with chapter one.

I do not have a story yet

Build the illustrated book first from an idea, then come back and cast it. Nothing here needs a finished manuscript.

I want a video, not a storybook

The story as a shareable video with scenes, motion and audio, rather than pages you turn.

Panels in motion with voices

Comic panels performed as short-form drama with voice and music, for video platforms rather than a reading app.

Another language entirely

The same book built in another language from the start, rather than translated at the end.

Show me everything else

The full tool shelf — generators, converters and editors, all working off one saved character library.

FAQ

Frequently Asked Questions

Your Characters Already Talk. Give Them Voices.

Bring one chapter and hear the cast it describes — a voice for each of them, yours to adjust — before you spend anything.

Free to cast a story English, Chinese, Japanese You keep every right

Multi-Voice Storybook — Casting Is the Part That Lasts

Almost every piece of story audio ever made asks one voice to play everybody. A parent doing bedtime does it by pitching up for the mouse and growling for the bear; an audiobook narrator does it with decades of technique; a text-to-speech tool mostly does not do it at all and reads the whole book in one flat register. A multi-voice storybook takes the other route: the parts are cast. The grandmother gets a low, slow, warm voice because that is who she is; the little brother gets a fast bright one; the wolf sounds like neither of them. The immediate benefit is comprehension — a listener knows who is speaking before the line is over, which is exactly what children lose track of in a scene with four people in it — but the reason it holds up over a whole book is less obvious, and it is about where the voice is stored.

A voice can be a setting on an export, or it can be an attribute of a character. Nearly every tool treats it as the first thing: you pick a voice when you render, and the next time you render you pick again. That works perfectly for one clip and fails quietly across a book, because on the twelfth render nobody remembers precisely what the third one chose. C2Story treats it as the second thing. The first time a character speaks, a short voice descriptor is worked out — gender, age band, pitch, timbre, articulation, something like “female voice, older adult, low pitch, warm timbre, unhurried articulation” — and that descriptor is written onto the character’s own profile, beside their design. Every scene after that reads the saved one back rather than deciding again. Chapter twelve is not a close approximation of chapter one; it is the same record.

The descriptor is read out of your story rather than guessed from a name. That distinction matters more than it sounds, because names are terrible evidence. Children’s fiction is full of animals, invented words, nicknames, titles and characters who are described as one thing and called another, and a guess based on a name gets a meaningful fraction of them wrong in a way that is instantly audible. So the character sheets and descriptions in your own text take priority: if the story says she is an old cat, she is voiced as one. And when two names clearly point at the same character — Smoke in chapter one, the old cat by chapter six — they are given an identical descriptor, so the voice follows the character rather than whichever label the sentence happened to use.

The second half of a multi-voice storybook is deciding what actually counts as a spoken line, and this is where a lot of AI narration goes wrong in a way people find unsettling. Ask a language model to find the dialogue in a dramatic paragraph and it will sometimes produce lines nobody wrote — plausible, in character, and completely invented. No amount of instruction reliably prevents it. So C2Story does not rely on instruction: spoken lines are extracted verbatim and then checked back against the source, character for character with punctuation and spacing ignored, and anything that does not physically appear in your text is discarded before it can reach a voice. Character sheets, synopsis paragraphs and scene description are excluded by definition too — a sentence describing a character is not a sentence she says. Most setup passages contain no dialogue at all, and the right number of voiced lines for them is zero.

Once the cast is carrying the dialogue, the narrator stops being the entire book and becomes one more part you can direct. Narration, dialogue and music are three independent switches. Leave narration on and an off-screen voice carries the description between the spoken lines. Turn it off and the scene plays on the artwork and the cast alone, which suits conversation-heavy stories and older children who do not need to be told what they can already see. Or hand narration to a character: switch the point of view to first person, choose who it belongs to, and they tell their own story in their own voice. It does not have to be the protagonist, and giving it to a supporting character — or to the one who loses — changes the entire book without altering a word of the text.

Timing is the last thing that changes, and it is the one people notice without being able to name. Most story audio is built on fixed durations: every page gets the same slot, so short scenes sit in silence and long ones get rushed. Here the dialogue sets the clock. Each line gets a timecode inside its scene, a scene runs for as long as its speech takes to say plus a beat of action, and when a scene carries more dialogue than one shot can hold it splits into two rather than accelerating the delivery or dropping a line. That is why exchanges breathe: a question and its answer are two separate moments with real space between them, instead of two sentences run together to fit a frame someone chose in advance.

Underneath, lines reach you by one of two routes and both are honest about their trade-off. In the native route the video model performs the line while generating the shot, so the voice and the mouth come out of the same take and stay in sync, steered by the voice descriptor and the timecoded lines — there is no separate speech pass that can drift out of alignment. In the cast route a voice is selected for each character from a catalogue, matched on gender, age, personality and the story’s language, and the dialogue is voiced and mixed into the finished video. Native wins on lip sync; cast wins on repeatability, because the same voice returns identically every time. English, Chinese and Japanese are supported, the story’s language is detected from the text rather than from a setting you have to remember, and narration can either follow it automatically or be pinned — which is how the same book gets read in two languages with the artwork and the characters untouched.

Ready to start? Bring one chapter rather than the whole book. A single chapter is enough to hear whether your cast sounds the way you have always heard them in your head, and that is far cheaper to change now than after you have rendered twelve of them. If what you actually wanted was one clear narrator reading word for word with the text highlighting along, the read-aloud video is the right page — the single voice is the feature there. For a finished illustrated book as a narrated watch-along, use picture book to video; for something calmer, AI bedtime story video; for a book plus a short trailer, the AI animated story creator. If you have no story yet, the AI picture book maker builds the book first. And to keep the visual half of your cast as stable as the audible half, the character consistency tool and the character bible generator do for faces what this page does for voices. Or open your character library and start casting.