Microsoft MAI-Voice-2.1 logo

Microsoft MAI-Voice-2.1

Turn text into expressive, natural-sounding speech in second

Artificial Intelligence Audio

MAI-Voice-2.1 is Microsoft AI's text-to-speech family for expressive, natural-sounding speech. Two models: 2.1 for fidelity (audiobooks, voice-over) at ~550 ms model latency and $22 per 1M characters, and 2.1-Flash for live use like call center agents and IVR at ~45 ms model latency and $15 per 1M characters. Both support 23 languages, granular emotion control, and instant voice matching from a short clip with no fine-tuning.

投票数: 0
← 投稿一覧に戻る