Spoken Multi-Cast™ Preferred by Listeners of Fiction
In Largest Study of Its Kind, U.S. Fiction Audiobook Consumers Rate Spoken Multi-Cast Higher Than Human Narration
In Largest Study of Its Kind, U.S. Fiction Audiobook Consumers Rate Spoken Multi-Cast Higher Than Human Narration
Results from Edison Research reveal increased willingness to purchase AI-narrated audiobooks, signaling a shift in market readiness
Spoken, the AI Audiobook Company™, released a groundbreaking independent research study conducted by Edison Research at SSRS showing that Spoken’s multi-cast narration outperforms conventional human narration in engagement, favorability, and perceived quality. The study focused on character-driven fiction – the most prevalent category in today’s audiobook market – marking a pivotal moment for the use of AI in the publishing industry.
The study randomized over 1,000 adult U.S. fiction audiobook listeners into two blinded cohorts, each experiencing samples of a new sci-fi thriller as either a professional, single-narrator human production, or an unedited, one-click Spoken Multi-Cast™ narration. As industry stakeholders increasingly explore AI narration to meet the surging demand for immersive storytelling, this research offers the first large-scale look at how consumers respond to AI technology integrating distinct voices for each speaking character, a format historically reserved for full-cast human productions.
“Up to now, we’d only measured consumers’ opinions on the concept of AI narration,” said Megan Lazovick, vice president of Edison Research. “For the first time, we were able to measure real-time reactions to excerpts from an actual audiobook. Would there be a difference in acceptance between AI narration and human? What about listeners’ willingness to listen or purchase? Could they distinguish the AI version from human? What we found was a clear signal that quality matters, no matter how the narration is produced, and listeners are open to whatever improves their experience.”
Spoken CEO Phil Marshall added that this data confirms what he’s believed since founding Spoken: “As an audio-only reader myself, getting lost in the story is what matters most. At the end of the day, what readers want will drive decisions in the industry. High-quality, multi-cast narration is what readers want, and a less expensive, one-click solution is what authors and publishers want. With Spoken’s unique, patent-pending approach to layering multi-character scenes, we are able to deliver on that promise with an immersive audiobook experience that readers will pay for.”
Key Findings
Superior performance for multi-character stories: In direct comparison, Spoken Multi-CastTM achieved higher ratings than human narration in overall favorability (61% vs. 53%), perceived quality of narration (66% vs. 60%), and overall engagement (58% vs. 49%) for character-driven stories. For exposition without multiple characters, the human narrator was rated higher.
Hearing is believing: While only 31% of listeners said they’d be likely to listen to an AI audiobook at the start of the survey, after hearing an excerpt, 65% – more than double – said they’d be likely to listen to an audiobook narrated using Spoken Multi-CastTM.
They can’t tell the difference: 61% of those who heard the Spoken version thought it was human, while 65% thought the human was human.
Listeners are equally likely to buy AI audiobooks: Purchase intent for audiobooks featuring Spoken Multi-CastTM narration (46%) is statistically comparable to that of human-narrated works (49%). This is true after they learned it was AI, suggesting no commercial barrier to adoption.
Frequent listeners gravitate to multi-cast: 81% of frequent audiobook listeners are interested in hearing distinct voices for each character, with 51% being very interested.
Quality and immersiveness are more important than cost and celebrity: The top factors increasing likelihood of listening to AI-narrated audiobooks were improved quality, greater immersion, and using multiple character voices, not hearing a celebrity voice.
These findings arrive at an important time for the audiobook industry. The global audiobook market has experienced double-digit growth for many years and is projected to surpass $35 billion by 2030 (Grand View Research), yet production is constrained by the cost and logistics of studio recording and audio production. Spoken’s multi-cast technology was built to close that gap, moving AI narration beyond the flat, uncanny performances that have limited consumer trust in synthetic voice. This research suggests that listeners not only accept AI narration but prefer it when it delivers a more engaging experience. The combination of expanded production capacity and demonstrated listener preference signals a new benchmark for what audiobook narration can be.
About the Survey Methodology
More than 1,000 U.S. fiction audiobook listeners were randomly assigned audiobook excerpts using human narration or Spoken Multi-CastTM AI narration and were asked a series of follow-up questions. The data was weighted to match the demographics of audiobook consumers as established by The Infinite Dial 2026 by Edison Research at SSRS and SiriusXM Media. Each treatment group was also weighted to match each other by age, gender, ethnicity, and interest in Sci-Fi/Fantasy audiobooks.
For more information and to hear the samples examined in the survey, visit https://spoken.press/edisonsurvey. To try Spoken for yourself, go to https://www.spoken.press/.
About Spoken
Spoken, the AI Audiobook Company™, enables authors and publishers to create immersive, multi-voice audiobooks at a fraction of the cost and time of traditional production. Purpose-built for storytelling, Spoken Multi-CastTM layers proprietary agentic AI with emotionally nuanced voice synthesis to deliver high-quality, character-driven performances. Authors can choose from custom AI-generated voices crafted exclusively for their characters or AI-cloned voices from professional narrators, each of whom gets paid with every use. Spoken’s mission is to help storytellers tell stories in ways never before possible. Learn more at www.spoken.press.
About Edison Research at SSRS
Edison Research at SSRS conducts survey research and provides strategic information to a broad array of clients worldwide, having conducted research in 66 countries. Edison Research’s The Infinite Dial® series has been the survey of record for digital audio, social media, podcasting, smart speakers, and other media-related technologies since 1998. The company’s Share of Ear® survey is the only single-source measure of all audio in the U.S. Edison Research is the leading podcast research company in the world, producing the only survey-based data on podcast listening in the U.S., Edison Podcast Metrics, and has conducted research for many companies in the space.
Spoken v2.1.2 — Spoken Multi-Cast™ - A Major Change in Workflow
This release introduces new pre-production tools for Magic Mode, giving you greater control over reviewing passages and speaker assignments before generating audio. It's a powerful new preparation step that helps streamline your workflow. We've also included general improvements and bug fixes across Text-to-Speech, Speak-Its, and overall platform performance.
Better Preparation Tools, More Accurate Narration
This release introduces new pre-production tools for Magic Mode, giving you greater control over reviewing passages and speaker assignments before generating audio. It's a powerful new preparation step that helps streamline your workflow. We've also included general improvements and bug fixes across Text-to-Speech, Speak-Its, and overall platform performance.
Pre-Production & Magic Prep Features
Creating accurate narration starts with accurate preparation. Version 2.1.2 introduces a dedicated pre-production workflow for Magic Mode, making it easier to review and refine your content before generating a base narration layer.
What's New
A new Manage Passages experience allows you to review and edit passages before narration begins.
Verify and update character attribution, speaker assignments, speaker cues, passage text, accents, and emotion cues in one centralized workflow.
Generate narration for a single chapter or your entire work, providing greater flexibility when creating or updating audio.
Review passage changes and regenerate only the content that needs updating, without reprocessing your entire project.
Enhanced passage editing tools make it easier to fine-tune speaker details and ensure your narration is prepared exactly as intended.
Why It Matters
These new preparation tools help ensure your base narration is as accurate as possible from the very first generation, reducing the need for post-production corrections while preserving the simplicity and speed of one-click narration.
Bracketed Cues Support in Text-to-Speech
Text-to-speech now supports bracketed cues within passage text when editing a passage. Use cues such as [British accent], [whispers], or [excited] to influence the delivery of a specific word or phrase.
For the best results, think of narration control in three levels:
Speaker Cues apply to every passage spoken by that character.
Emotion Cues apply to the entire current passage.
Bracketed Cues apply only to the specific words or phrase enclosed in brackets.
Because these controls can influence one another, avoid using them to accomplish the same effect at multiple levels. For example, if you're using a bracketed cue to create a specific emotional moment within a sentence, it's often best to remove the passage-wide Emotion Cue so the two instructions don't compete.
Bracketed cues are intended as a precision tool for isolated moments—not as a replacement for Speaker or Emotion Cues.
(Note: Bracketed cues are supported on Magic Mode + 11L Voices only; they are not supported for Hume models.)
Improvements & Fixes
General improvements to Text-to-Speech inputs and processing.
Performance and reliability enhancements for Speak-Its.
Improved stability across narration and content generation workflows.
Various bug fixes and behind-the-scenes optimizations for a smoother overall experience.
Spoken Studio v2.1.1 Release Notes: Foundation & Performance
Faster Workflows. Smarter Voice Systems. Stronger Reliability.
Faster Workflows. Smarter Voice Systems. Stronger Reliability.
Following the rollout of Magic Mode and the continuity enhancements introduced in v2.1, Spoken Studio v2.1.1 focuses on strengthening the systems that power audiobook creation behind the scenes.
This release delivers meaningful improvements across voice management, project workflows, editor behavior, system responsiveness, and production reliability.
In short:
Spoken Studio should feel faster, smarter, and more predictable throughout the creation process.
Workflow & Editor Refinements
We've continued refining the authoring experience throughout Studio.
This release includes improvements to editor behavior, project management workflows, voice preview handling, narration preparation tools, and character management—helping reduce friction and keep production moving smoothly.
Notable improvements include:
More accurate character analysis for voice assignments, prompts, and accent detection.
Improved character merging and deduplication to reduce duplicate entries across projects.
Enhanced voice preview behavior with improved end-of-preview timing for a more natural listening experience.
Longer preview endings. Based on author demand—particularly those distributing through Author’s Republic—voice previews now include additional trailing silence, making it easier to meet retail submission requirements.
Collectively, these updates create a significantly smoother production workflow.
Performance & Reliability Improvements
Voice retrieval performance for faster voice loading and assignment.
Project processing workflows.
System reliability.
Audio preparation pipelines.
Internal quality control systems.
Magic Mode precision, reducing audible artifacts at the beginning and end of generated clips.
Improved accent preservation during spoken generation when narrator and character voice settings are correctly configured.
These updates help ensure Studio remains responsive and dependable as projects become larger and more ambitious.
Improved Accent Consistency:
When character voice settings are configured correctly, Spoken Studio now does a significantly better job preserving accents throughout the narration process, resulting in more consistent performances across chapters and scenes.
Best practice: Before clicking Make Spoken, verify that each character's Speaker Cue and Accent are correctly assigned—particularly when a character's accent differs from the chapter narrator. If no accent is displayed, simply click the character's colored voice bar to open the voice settings and add the appropriate Speaker Cue and Accent before generating audio.
Bugs & Fixes
As part of v2.1.1, we've addressed a variety of bugs and edge cases throughout the platform, including:
Text processing improvements
Character handling refinements
Editor behavior fixes
Quality-of-life improvements throughout Studio
General stability and performance enhancements
Looking Ahead
This release makes every author, publisher and producer’s workflow better.
v2.1.1 represents an important investment in the systems underneath Spoken Studio — strengthening the platform as we continue expanding what's possible for AI audiobook narration.
By Storytellers. For Storytellers.
We believe that giving voice to writing isn’t just for those with resources to create elaborate productions or patience to navigate complex publishing hoops. Spoken was created by a small team of storytellers based in Portland, Oregon who believe in empowering self-publishers.