M2 voice demos — listen and compare

Small text-to-speech and voice-understanding models, trained on 20,000 hours of speech · everything below runs in your browser, no login needed

What is this? We train small speech AI models called M2. Given written text, they speak it out loud — either in a neutral reading voice or copying a short voice example they are given. The newest model can also do the reverse: listen to audio and write down what was said, describe the voice, and name the emotions it hears. This site collects listening demos, audit reports, and dataset inventories from that project. All pages are in plain language.

Start here

▶ 60-sample demo: new model vs old model
30 German + 30 English clips nobody trained on. For each clip: the real recording, the reference voice, the new model speaking (with and without reference), the old model speaking (same two ways, side by side so you can judge better-or-worse by ear), plus what the new model hears: transcript, short voice caption, full caption with emotions.

How the numbers work (short stats guide)

NumberWhat it meansGood value
WER (word error rate)Share of words the model gets wrong when speaking or transcribing. 0.25 = one word in four is wrong.lower is better
Genuineness (0–6)How real and human the voice sounds, rated by a scoring model.higher is better
Blend (0–10)How much vocal color (laughs, sighs, warmth) the voice carries.mid-range is natural
ECAPA similarityHow close the generated voice is to the reference voice (speaker identity).higher is better
DNSMOS / OVRL (0–5)Automatic sound-quality rating (noise, clarity).higher is better

Model scoreboard

ModelTrainingSpeech quality (held-out WER)
Old (M2 20k-score, step 1628)6.67M clips × 1 job (text-to-speech with voice measurements as extra input)0.250 overall (0.240 without reference · 0.245 with reference)
New (M2 multitask, step 7319)Same clips × 5 jobs (speech ± reference, transcription, short + full captions), no extra measurementsFormal eval pending — compare by ear in the demo above

Reports (audits, data, repairs)

Prompt audit: durations & bursts Dataset inventory + repair plan Why 43% could not be repaired
What the training prompts really look like versus the specification, where all 20k/50k/100k-hour datasets live, and why some clips keep their original prompts (with costed fix options).

Ladder listening pages (per data tier)

Two clips per selection bucket for every ladder tier (50 h → 10,000 h per language), with genuineness / burst-blend / emotion statistics. Pages appear here as they finish building.

Built 2026-09-15 · models: Qwen3-0.6B backbone + single-layer talker head · audio: 48 kHz MP3 · contact via the TTS-AGI organization.