Apple Updates Siri with Voice Customization
Apple in its latest iOS 27 beta release, has launched the next generation of Apple Intelligence, powering Siri AI, an entirely new version of Siri, and capabilities that make the apps and experiences more personal and helpful.
Included in the new Siri AI update is the ability for users to customize the assistant with even more expressive voices, as well as a major boost in accuracy, with systemwide dictation for products that support Apple's most advanced on-device model ever, AFM Core Advanced.
Users can customize the expressiveness, pace, and nationality of Siri's voice. There are seven American voice options and four British voices, all of which can be customized with different Pace and Expressivity settings. Pace changes talking speed, while expressivity changes emphasis.
In addition, an updated dictation engine will capture speech as polished text, automatically handling capitalization, punctuation, and formatting in real time as users speak. Apple says improved speech understanding means users can speak naturally and trust that their words will appear accurately and as intended. And with improved speech understanding, users can speak naturally and trust that their words will appear clearly, accurately, and as intended.
Ruth Zive, chief marketing officer of Voices, a global voice solutions provider, sees the move by Apple as a signal of the end of the pre-recorded AI voice. As consumers are increasingly exposed to voices that reflect personality and greater emotional understanding, natural-sounding AI will quickly become the baseline, not the differentiator, and Apple just raised the bar, she says.
"Apple's move to let users choose the nationality, cadence, and expressivity in this latest iteration of Siri is proof that companies are no longer in the race for realism. The next challenge is nuance," Zive states. "Our own research found that 48 percent of enterprise decision makers already rank tone and emotional expressiveness as the single most important vocal factor. Real people don't speak in perfectly consistent patterns, as micro-inflections, regional cadence, and expressivity are all shaped by context. So if these nuances aren't in the data, it won't be in the voice. You can't model what you haven't captured."