Small text-to-speech and voice-understanding models, trained on 20,000 hours of speech · everything below runs in your browser, no login needed
| Number | What it means | Good value |
|---|---|---|
| WER (word error rate) | Share of words the model gets wrong when speaking or transcribing. 0.25 = one word in four is wrong. | lower is better |
| Genuineness (0–6) | How real and human the voice sounds, rated by a scoring model. | higher is better |
| Blend (0–10) | How much vocal color (laughs, sighs, warmth) the voice carries. | mid-range is natural |
| ECAPA similarity | How close the generated voice is to the reference voice (speaker identity). | higher is better |
| DNSMOS / OVRL (0–5) | Automatic sound-quality rating (noise, clarity). | higher is better |
| Model | Training | Speech quality (held-out WER) |
|---|---|---|
| Old (M2 20k-score, step 1628) | 6.67M clips × 1 job (text-to-speech with voice measurements as extra input) | 0.250 overall (0.240 without reference · 0.245 with reference) |
| New (M2 multitask, step 7319) | Same clips × 5 jobs (speech ± reference, transcription, short + full captions), no extra measurements | Formal eval pending — compare by ear in the demo above |
Built 2026-09-15 · models: Qwen3-0.6B backbone + single-layer talker head · audio: 48 kHz MP3 · contact via the TTS-AGI organization.