How to Turn a Web Novel Chapter into a Full-Cast Audio Drama
Paste a chapter, let the characters be found and cast, preview the opening, then render MP3 and SRT. Real timings, costs and the mistakes that break casting.
Listen to this article
AI narration · about 7 min
Short answer: paste the chapter into an audio drama studio that attributes every line to a speaker, check the cast it proposes, preview the opening, then render. In the VoiceSmiths Audio Drama Studio that is four clicks: Find the characters → check the cast → Preview the opening → Render. An 820-character Chinese chapter rendered in about 35 seconds in our test on 4 October 2026 and came back as 3 minutes 18 seconds of audio, with an MP3, an SRT subtitle file and the script.
The rest of this guide is what decides whether the result sounds like a drama or like a narrator doing voices.
What "full cast" actually means
A single-narrator audiobook reads everything in one voice and leans on the listener to keep track of who is talking. A full-cast audio drama gives the narrator one voice and every character their own, and keeps each character on the same voice for the whole chapter. That needs three things a plain text-to-speech tool does not do:
- Speaker attribution — deciding, line by line, who says it.
- Casting — a distinct voice per character, consistent across the chapter.
- Direction — how each line is delivered: whispered, sobbed, shouted.
Here is what that sounds like on a web-novel scene, five voices, straight from the text:
Step 1: Prepare the chapter (two minutes that save twenty)
Casting is only as good as the attribution, and attribution works from the same cues a reader uses:
- Dialogue in quotation marks. Lines outside quotes are treated as narration.
- Speech verbs near the quote — "she said", "he snapped", "Mara whispered". These are the strongest signal.
- Forms of address — "Master, …", "Mr. Shen, …" tell the studio who is being spoken to, which narrows who is speaking.
What breaks it is the long back-and-forth with no tags at all. Those lines are inferred from turn-taking and may need one or two speakers corrected by hand. If you are editing anyway, adding a "he said" every few lines fixes it at the source.
You can paste up to 8,000 characters at a time. A typical web-novel chapter fits; a long one goes in two halves.
Step 2: Find the characters
Paste the text and click Find the characters. The studio splits the chapter into lines, separates narration from dialogue, names the characters it found and proposes a voice for each from the character voices for that language (15 for Chinese, 27 for English: old storyteller, cold queen, young hero, late-night radio host and so on), matched on gender and age.
Your text is not rewritten. The script view shows your words, split into lines, with a speaker and an optional emotion tag on each.
Step 3: Check the cast
This is the step worth slowing down on:
- Tap an avatar to hear the voice; pick another from the list if it is wrong for the part.
- One character, one voice. Changing a character's voice changes it everywhere in the chapter, which is what keeps them recognisable.
- Give the narrator room. A calm storyteller voice under expressive characters is easier to listen to for twenty minutes than an equally dramatic narrator.
Picking voices by role is covered in more depth in character voices for short dramas.
Step 4: Direct the lines
Each line can carry an emotion tag such as [whispers], [sobbing] or [shouting]. Tags direct the performance and are never read aloud. The studio suggests them from the text around each quote; change any that do not fit. The emotion tags cheatsheet lists what works.
If a line is attributed to the wrong person, change its speaker from the dropdown on that line. Nothing else needs redoing.
Step 5: Preview, then render
Preview the opening plays the first lines with the real cast, so you hear the casting before rendering the whole chapter. Finding the characters is free without an account; previews, renders and adaptations use credits.
Then render the chapter. You get:
- MP3 of the whole chapter
- SRT subtitles, timed line by line, ready for CapCut or Premiere (see SRT subtitles from audio)
- The script, as text
Every file carries an embedded AI-generated label, and by default the audio ends with a spoken "this audio was generated by AI" notice, which you can switch off before rendering.
How long, and how much
From our own renders:
| Chinese | English | |
|---|---|---|
| Audio per 1,000 characters | about 4 minutes | about 1 minute 10 seconds |
| Render time (our test) | 820 characters in ~35 s | — |
| Credits per 1,000 characters | 70 | 70 |
Seventy credits per 1,000 characters works out to about $0.55–0.82 on the monthly plans, or $1.86 with the one-time trial pack (see pricing). Chinese packs far more story into each character, so a Chinese chapter costs roughly a quarter of what the same scene costs in English.
Step 6 (optional): the same chapter in other languages
Once the Chinese or English version is right, Adapt (same cast) rewrites the script into English, Spanish, Portuguese, Japanese, Korean, Indonesian, French or German. Idioms, forms of address and genre tropes are adapted rather than glossed, and with "keep the same voices" on, each actor performs the new language. The scene above exists in seven languages this way; the full walkthrough is in Chinese short drama into English.
What it will not do
- No music or sound effects. You get the performance; add a music bed and effects in your editor.
- Chapter at a time, not whole books. For a full book with a fixed narrator voice, long-form audiobook is the better tool.
- Attribution is not perfect on long untagged exchanges. Expect to fix a line or two per chapter.
FAQ
Can I try it without paying?
Finding the characters is free, with no account: you see the voice proposed for every character. Previewing, rendering and adapting into other languages use credits, starting with a one-time pack whose credits never expire.
Can I use the audio commercially?
Audio you generate is yours and may be used commercially under the terms of service. The character voices are AI-designed originals, not clones of real people.
Which languages can the source chapter be in?
The source chapter can be Chinese or English. The adaptation step outputs nine languages: Chinese, English, Spanish, Portuguese, Japanese, Korean, Indonesian, French and German.
Does it change my wording?
No. The original script keeps your text; only the adaptation step writes new text, and only in the target language.
Where can I hear finished examples?
The theater has 76 episodes made this way, from Lu Xun and Journey to the West to Sherlock Holmes and Poe, each with its script to read along.