Spoken v2.2 Release Notes — More Ways to Make It Yours
Spoken v2.2 Release Notes: More Ways to Make It Yours
Add Music, Sound Effects, Recorded Audio & More to Your Spoken Production — Plus Finer Pacing & Production Control
Spoken Studio v2.2 opens up the production workflow in a major new way.
For the Authors, Publishers and Producers who have been asking us to make this possible: it's finally here.
For the first time, Spoken Studio lets you bring audio beyond generated narration directly into your audiobook production.
Add a sound effect between two passages. Drop music between a scene transition. Add an opening theme. Insert an author-recorded note, a producer-created element, a custom performance, or any other audio your Spoken production calls for. Sound design. Music. Commentary. Custom audio.
This vastly expands the creative canvas of Spoken Studio, opening up new possibilities for studio-grade, immersive audiobook experiences tailored to the needs of virtually any project. As these production tools continue to expand, so does the freedom to build richer, more distinctiveaudio experiences.
This release brings finer control over the audio Spoken creates, too. v2.2 expands passage- and character-level speed controls across both manual and Magic Mode workflows, making it easier to dial in the pacing dynamics of individual performances and entire scenes.
Together with a broad round of Studio refinements, v2.2 gives you more freedom over what goes into your audiobook — and more precision over how it all comes together.
Add Your Own Audio to Any Passage
This is a massive expansion of what you can build inside Spoken Studio.
You can now upload an audio file directly to a passage, including passages created during Magic Prep, and place that audio exactly where you want it within your production.
That could mean adding music at the beginning of a chapter. A sound effect between passages. Author commentary. A specially recorded performance. Producer-created audio. An existing publisher asset. Or something entirely unique to the way you want to tell your story.
Once uploaded, the audio becomes the latest MP3 associated with that passage, giving it a natural home alongside the rest of your audiobook.
For Authors, this opens up entirely new creative possibilities.
For Producers, it creates more freedom to incorporate externally produced elements without stepping outside the Spoken workflow.
And for Publishers, existing audio assets can now be incorporated directly into projects alongside Spoken-produced narration.
Spoken can produce the narration. Now you can bring even more to the production.
Precision Pacing — Finer Speed Control
Pacing can change the feeling of a performance completely.
With v2.2, we've expanded and improved the fidelity of speed controls across both manual and Magic Mode production, giving you finer and more reliable increments for shaping narration:
0.8x · 0.9x · 1.0x · 1.1x · 1.2x
Start with your established pacing, then fine-tune it where the story calls for something different.
Speed can be adjusted at the individual passage level, allowing you to shape the dynamics within a scene, or globally at the character level when you want a particular performance style carried consistently throughout the work.
Slow down a moment that needs room to breathe. Tighten the pace of a faster exchange. Dial in a character whose natural delivery needs just a little adjustment.
It's another production control available when you need it, while staying out of the way when you don't.
“Scene Note:” — Scene-Level Intelligence for Magic Mode
Magic Mode now gives Authors, Producers, and Publishers a powerful new way to prepare a manuscript for performance before it ever enters Spoken Studio.
Add a Scene Note beneath a chapter heading to provide direction for the entire scene — shaping its pacing, rhythm, accents, delivery, emotional arc, or the way multiple characters interact.
(Example Scene Note: “This scene should build slowly and tensely. All characters speak with British accents, with increasingly rapid pacing as the argument escalates.”)
When the manuscript is brought into Spoken, Magic Mode uses that scene-wide direction alongside its analysis of the characters and passages, giving it a clearer understanding of how the scene should unfold as a complete performance.
For authors and publishers preparing a manuscript, and producers shaping a production, Scene Note puts the director's intent into the manuscript itself, before narration begins.
Bugs, Fixes & Studio Refinements
v2.2 also includes a substantial round of improvements throughout Spoken Studio, with a focus on responsiveness, reliability, and keeping production moving smoothly.
This includes:
Faster and more responsive padding updates
Improvements to Custom Voice matching
Smoother dragging and reordering of Chapters / Installments
Improved Speak It recording, upload, and update behavior
A fix for voice previews in the Voices tab not playing back as expected
Additional performance, stability, and workflow refinements throughout Studio
More Ways to Make It Yours
Spoken has always been built to give creators serious control over how their audiobooks are produced.
v2.2 expands the palette.
You can bring your own creative audio into the production, place it exactly where the story needs it, and shape Spoken performances with finer pacing control than before.
For an Author producing a deeply personal work, a Producer adding the finishing touches, or a Publisher working across an ambitious catalog, there are now even more ways to make a Spoken production distinctly your own.
Spoken Multi-Cast™ Deep Dive Webinar
When 81 percent of listeners demand to hear distinct voices for each character, it's time to dive deep into how to bring that immersive performance to life.
Join Spoken founder and CEO Phil Marshall and the Spoken team for a hands-on look at Spoken Multi-Cast™ and the Magic Mode tool that turns a single manuscript into a fully-narrated, multi-cast audiobook.
Thursday August 20, 1 p.m. ET (10 a.m. PT)
When 81 percent of listeners demand to hear distinct voices for each character, it's time to dive deep into how to bring that immersive performance to life.
Join Spoken founder and CEO Phil Marshall and the Spoken team for a hands-on look at Spoken Multi-Cast™ and the Magic Mode tool that turns a single manuscript into a fully-narrated, multi-cast audiobook.
What we'll cover:
Managing accents and the vocal nuances that makes each character feel real
Expert tips for Lexicon
Adding custom sounds (external audio) to your project
A first look at our new speed-control feature for dynamic, natural pacing
When-in-doubt tricks for getting the most out of Spoken Multi-Cast™
The research behind the 81% and what it means for your audiobooks
This session will show you exactly how to give your listeners the immersive, character-driven experience they're asking for with Spoken Multi-Cast™.
Spoken Multi-Cast™ Preferred by Listeners of Fiction
In Largest Study of Its Kind, U.S. Fiction Audiobook Consumers Rate Spoken Multi-Cast Higher Than Human Narration
In Largest Study of Its Kind, U.S. Fiction Audiobook Consumers Rate Spoken Multi-Cast Higher Than Human Narration
Results from Edison Research reveal increased willingness to purchase AI-narrated audiobooks, signaling a shift in market readiness
Spoken, the AI Audiobook Company™, released a groundbreaking independent research study conducted by Edison Research at SSRS showing that Spoken’s multi-cast narration outperforms conventional human narration in engagement, favorability, and perceived quality. The study focused on character-driven fiction – the most prevalent category in today’s audiobook market – marking a pivotal moment for the use of AI in the publishing industry.
The study randomized over 1,000 adult U.S. fiction audiobook listeners into two blinded cohorts, each experiencing samples of a new sci-fi thriller as either a professional, single-narrator human production, or an unedited, one-click Spoken Multi-Cast™ narration. As industry stakeholders increasingly explore AI narration to meet the surging demand for immersive storytelling, this research offers the first large-scale look at how consumers respond to AI technology integrating distinct voices for each speaking character, a format historically reserved for full-cast human productions.
“Up to now, we’d only measured consumers’ opinions on the concept of AI narration,” said Megan Lazovick, vice president of Edison Research. “For the first time, we were able to measure real-time reactions to excerpts from an actual audiobook. Would there be a difference in acceptance between AI narration and human? What about listeners’ willingness to listen or purchase? Could they distinguish the AI version from human? What we found was a clear signal that quality matters, no matter how the narration is produced, and listeners are open to whatever improves their experience.”
Spoken CEO Phil Marshall added that this data confirms what he’s believed since founding Spoken: “As an audio-only reader myself, getting lost in the story is what matters most. At the end of the day, what readers want will drive decisions in the industry. High-quality, multi-cast narration is what readers want, and a less expensive, one-click solution is what authors and publishers want. With Spoken’s unique, patent-pending approach to layering multi-character scenes, we are able to deliver on that promise with an immersive audiobook experience that readers will pay for.”
Key Findings
Superior performance for multi-character stories: In direct comparison, Spoken Multi-CastTM achieved higher ratings than human narration in overall favorability (61% vs. 53%), perceived quality of narration (66% vs. 60%), and overall engagement (58% vs. 49%) for character-driven stories. For exposition without multiple characters, the human narrator was rated higher.
Hearing is believing: While only 31% of listeners said they’d be likely to listen to an AI audiobook at the start of the survey, after hearing an excerpt, 65% – more than double – said they’d be likely to listen to an audiobook narrated using Spoken Multi-CastTM.
They can’t tell the difference: 61% of those who heard the Spoken version thought it was human, while 65% thought the human was human.
Listeners are equally likely to buy AI audiobooks: Purchase intent for audiobooks featuring Spoken Multi-CastTM narration (46%) is statistically comparable to that of human-narrated works (49%). This is true after they learned it was AI, suggesting no commercial barrier to adoption.
Frequent listeners gravitate to multi-cast: 81% of frequent audiobook listeners are interested in hearing distinct voices for each character, with 51% being very interested.
Quality and immersiveness are more important than cost and celebrity: The top factors increasing likelihood of listening to AI-narrated audiobooks were improved quality, greater immersion, and using multiple character voices, not hearing a celebrity voice.
These findings arrive at an important time for the audiobook industry. The global audiobook market has experienced double-digit growth for many years and is projected to surpass $35 billion by 2030 (Grand View Research), yet production is constrained by the cost and logistics of studio recording and audio production. Spoken’s multi-cast technology was built to close that gap, moving AI narration beyond the flat, uncanny performances that have limited consumer trust in synthetic voice. This research suggests that listeners not only accept AI narration but prefer it when it delivers a more engaging experience. The combination of expanded production capacity and demonstrated listener preference signals a new benchmark for what audiobook narration can be.
About the Survey Methodology
More than 1,000 U.S. fiction audiobook listeners were randomly assigned audiobook excerpts using human narration or Spoken Multi-CastTM AI narration and were asked a series of follow-up questions. The data was weighted to match the demographics of audiobook consumers as established by The Infinite Dial 2026 by Edison Research at SSRS and SiriusXM Media. Each treatment group was also weighted to match each other by age, gender, ethnicity, and interest in Sci-Fi/Fantasy audiobooks.
For more information and to hear the samples examined in the survey, visit https://spoken.press/edisonsurvey. To try Spoken for yourself, go to https://www.spoken.press/.
About Spoken
Spoken, the AI Audiobook Company™, enables authors and publishers to create immersive, multi-voice audiobooks at a fraction of the cost and time of traditional production. Purpose-built for storytelling, Spoken Multi-CastTM layers proprietary agentic AI with emotionally nuanced voice synthesis to deliver high-quality, character-driven performances. Authors can choose from custom AI-generated voices crafted exclusively for their characters or AI-cloned voices from professional narrators, each of whom gets paid with every use. Spoken’s mission is to help storytellers tell stories in ways never before possible. Learn more at www.spoken.press.
About Edison Research at SSRS
Edison Research at SSRS conducts survey research and provides strategic information to a broad array of clients worldwide, having conducted research in 66 countries. Edison Research’s The Infinite Dial® series has been the survey of record for digital audio, social media, podcasting, smart speakers, and other media-related technologies since 1998. The company’s Share of Ear® survey is the only single-source measure of all audio in the U.S. Edison Research is the leading podcast research company in the world, producing the only survey-based data on podcast listening in the U.S., Edison Podcast Metrics, and has conducted research for many companies in the space.
By Storytellers. For Storytellers.
We believe that giving voice to writing isn’t just for those with resources to create elaborate productions or patience to navigate complex publishing hoops. Spoken was created by a small team of storytellers based in Portland, Oregon who believe in empowering self-publishers.