VoiceGen
A voice-over studio that turns a script into expressive speech. Mark up delivery with inline tags like [whisper], [pause] or [suspense], or let an AI voice director do it, then render a master WAV.
View project (opens in a new tab)Tech Stack
- React 19
- Vite
- TypeScript
- Tailwind
- Vercel Functions
- Gemini 2.5 TTS
- Gemini 2.5 Flash
- Firebase Auth
- Firestore
Problem
Text-to-speech models can now whisper, build suspense and read like a news anchor. Getting them to do it on cue is the hard part. The controls are either a prompt you rewrite by trial and error or markup like SSML that nobody writing a script wants to learn.
The community I volunteer for had moved to video announcements, and producing them took up to six hours every weekend: coordinating schedules, recording, editing. Voice-overs in the community's own accent, generated from a script, removed both the bottleneck and the cost of a voice subscription.
I wanted the direction to live inside the script, the way a director pencils notes onto a page, and I wanted a long script to come out sounding like one person in one take.
Approach
VoiceGen started as a Google AI Studio prototype. When it outgrew that, I wrote a port plan and rebuilt it as its own codebase with Claude Code, carrying the features across one by one.
The studio is a three-step wizard: Prepare, Preview, Finalize. You paste a script, choose one of ten voices and one of seven personas, listen to a short preview, then render the full master.
Expressiveness comes from bracket tags written inline. There are about two dozen, covering timing ([pause], [break], [short]) and delivery ([whisper], [excited], [announcer], [suspense] and more). Each tag maps to a plain-language instruction for the model and applies to the sentence it opens. There are two ways to add them. AI Enhance runs a voice-director pass that splits the script into contexts and inserts tags for you, with an optional copy review that logs every change it made. The manual palette is the power path, and it is collapsed by default so the sidebar is not a wall of buttons. The editor highlights tags as you type.
The engineering problem was length. A serverless function on Vercel is cut off at 60 seconds, and a long script takes far longer than that to render. So the browser orchestrates the job: each call renders exactly one span of text, the client fans out with bounded concurrency, then stitches the 24 kHz PCM into a single WAV. I measured the span sizes instead of guessing. 400 characters renders in about 13 seconds and 600 in 10 to 12, while anything near 1,000 can run past three minutes, so the app works down a ladder of 900, 600 and 400.
Chunking created a second problem. Every span is an independent call, so the voice drifted and the accent changed between paragraphs. The fix was a fixed seed derived from the text, voice and persona, a low temperature, and the persona instruction injected again on every chunk. The commit history keeps the bugs that led there, including the day the voice began reading its own stage directions aloud.
Small decisions carry the rest. Voice auditions all use one fixed short line, so comparing voices is instant, consistent and cheap. A pacing normaliser collapses stacked pauses. The API key stays on the server, and the browser only ever calls my own routes.
Key Learnings
- Inline tags beat a settings panel. Direction written in the script is readable, editable and travels with the text.
- Measure the model before you design around it. The span ladder came from timing real calls, and the numbers were not what I would have guessed.
- When work is split across independent AI calls, consistency has to be engineered. Seed, temperature and repeated instructions did what a single long call gets for free.
- Offer an easy path and a power path, and hide the second one until it is wanted.
- My own design review concluded that the product is right and the compute model is wrong. Browser-orchestrated rendering works, but a job queue is where this should go next.