Gradium AI has released a new default text-to-speech model across its API and Studio, reporting an 81.0% human-rated pass rate on a 500-sentence hard-case benchmark spanning five languages. The company built and open-sourced a multilingual evaluation set covering complex tokens like acronyms, alphanumerics, dates, numbers, and emails, with strict native-speaker scoring. Existing voices, including custom clones, continue working without migration or configuration changes. On Coval’s benchmark, the model achieves a 216 ms median time to first audio with low latency variability, positioning it competitively against rival systems and encouraging developers to adopt it for customer-support voice applications.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.