Where this edition stands
This is a pre-launch edition. Documentation has been checked, but paid listening sessions have not been completed. No product has a quality score. The method below is a testing commitment, not a claim that these sessions have already happened.
The test, in order
- Buy the plan. Every scored tool must be tested on a paid plan purchased by the publication. Record the purchase, model version, voice identifier and available controls.
- Use the same words. Run the standard scripts without tool-specific rewrites for the first take. Preserve that take, then make four further attempts.
- Keep the failures. Count mispronunciations, unusable takes and correction attempts. Publish the settings and the actual changes needed.
- Listen blind. Hide tool names, shuffle their order and compare at matched playback levels. Listening order is randomised again for a new blind session.
- Measure cost and latency. Include retries and editing time. Time twenty API requests and report median and 95th percentile time to first audio separately.
- Lock the verdict. Set scores and rankings before checking commission rates. A tool with no affiliate programme remains eligible.
Standard scripts
Everyday narration
The last train leaves at seven thirty. Outside, the city is still waking up. Take a breath, check your ticket, and listen for the announcement. There is just enough time to begin again.
Names and numbers
Dr Nguyen booked a table for three on Wednesday, the twenty third of September. The total was thirty seven dollars and fifty cents. Please send the receipt to the research department.
Character dialogue
You came back. I thought the bridge was gone. Keep your voice down, take the lantern, and follow me. Whatever you hear behind that door, do not answer.
Playback conditions
Every scored session must name the computer, operating system, audio interface if used, headphones, speakers and playback level. The initial equipment log is not filled in because there has not yet been a listening session.
Keep lossless masters. Apply the same documented loudness adjustment to listening copies, with no denoising, EQ or time stretching. Preserve originals for anyone checking whether processing changed a finding. Test the final files on headphones, laptop speakers and a phone.
Eight axes, one consistent scale
Each axis is scored from 0 to 10. A 0 means the measured requirement failed; 5 means usable with material compromises; 8 means reliable with minor limitations; 10 requires consistently excellent evidence across the tested conditions. Untested axes remain blank.
The overall score is the unweighted mean of all eight axes, rounded to one decimal place. A review cannot be marked tested until all axes, dated evidence and recordings are present. A use-case recommendation may favour a particular strength, but must explain why it differs from the overall ranking.
Voice naturalness
Breathing, phrasing, stress and audible synthesis artefacts.
Emotional range
Whether a requested emotion is audible, controlled and repeatable.
Language coverage
Quality in tested languages, reviewed by fluent listeners; supported-language counts alone do not earn points.
Pronunciation control
First-take accuracy and the effort needed to correct names, numbers and specialist terms.
Latency
Median and 95th percentile time to first audio across 20 requests, with region and connection recorded.
API quality
Documentation, streaming, useful errors, retries, limits and reproducible integration.
Licensing clarity
Whether the purchased plan clearly permits the intended output and voice use.
Value
Cost per usable finished minute, including rejected takes and editing time.
Retests and corrections
The planned cadence is every ninety days for ranked tools, and sooner after a material model, pricing or licensing change. Each page keeps the source-check date separate from the listening-test date. Corrections must describe what changed and whether the verdict moved.
Affiliate independence
Commission rates never set the score. Paid placements cannot buy a ranking, and a missing affiliate programme is not a reason to omit a tool. Every monetised page identifies its affiliate links before the first one appears. Read the full policy.
The player demonstration
The preview includes four locally generated CMU Flite speech examples so readers can try the controls. They are explicitly labelled as demonstrations, use the visible scripts, and are not recordings from Fish Audio, ElevenLabs, Murf or Speechify. They are excluded from scoring.