Forum Discussion
Japanese AI Text-to Speech Quality
Has anyone had success using a particular voice for the AI Text-to-Speech for Japanese? I have tried several and continue to get feedback that the text-to-speech quality is very poor, and some voices randomly mix Japanese, Chinese, and Korean accents even within a single paragraph. All of the translated audio text has been thoroughly reviewed and is accurate. Here are the latest voices I have tried and the general feedback received:
- Alva: Good (the voice used in the course)
- Akira: Poor
- Hajime: Poor
- Gojo: Acceptable
- Ken: Good (possible alternative, but does not address the root cause)
- Masa: Acceptable
In the Advanced settings, I have Multilingual v2 selected for each one.
Note that even for the voices marked good or acceptable, the reviewer still indicates that the quality is not at a level that they feel can be used. They believe the root cause is that the program invoking text-to-speech does not consistently retain the selected voice option and that it incorrectly auto-detects the language and switches voice options even within a single sentence. I am not sure how to respond to that.
I would appreciate any insights anyone else has about this topic!
6 Replies
- JimVogel-1ea9baCommunity Member
Thank you for sharing this experience. We are also struggling with the Japanese TTS options. I hope Articulate is improving capabilities for this language.
Hi JimVogel-1ea9ba,
Thanks for chiming in! I'm sorry to hear you've also been having challenges with Japanese text-to-speech voices, and I appreciate you sharing your experience.
Are you mainly encountering pronunciation issues, or are there other challenges with the Japanese voices? How are these affecting your course development workflow?
We appreciate your feedback, as it helps us better understand how these challenges affect different customers.
- LisaDobias-c0c7Community Member
I should note that the voice feedback above reflect ease of listening, not pronunciation accuracy.
Hello LisaDobias-c0c7,
I appreciate you reaching out! While we don't have any news on any additional advanced settings for AI TTS Voices, I've shared your insight with the product team. We'll be sure to share any future updates in this thread, so all are in the loop!
I'm curious if you've tried the following:
- Do you notice any change to the ease of listening if you generate audio with v3 (Beta) instead of Multilingual v2?
- Have you used any supported SSML tags to improve cadence or flow?
I'll open this discussion up to our fabulous community to share their experiences and what has worked well for them!
- LisaDobias-c0c7Community Member
Sorry for the slow response. I did not try the options noted above, but will do that next time to see if we can get better results. I did find that the slides where there were the most issues were with technical terms or when there was a mix of text and numbers (i.e., ISO 13485, for example). It could be that the text-to-speech was also confused by that. For English text-to-speech, it worked better when I spelled the words out but I didn't quite know how to do that with translated content.