Define the job before choosing a model
Separate interactive speech from batch narration. For a conversation, the wait until useful audio begins matters. For overnight generation, total throughput and correction work may matter more. Write down your own acceptance criteria before seeing vendor claims.
Use short turns, a paragraph and difficult names from your actual domain. Include a permitted sample in each required language. Record the exact model, output format, region and request settings. Our shared scripts and method provide a starting point, not a substitute for your workload.
Distinguish latency from throughput
Measure time from sending the request to receiving the first usable audio chunk, and separately to completion. Keep connection setup in the record. Report median and p95 across at least 20 requests; a small sample is an initial screen, not an SLA or a production guarantee.
Fish Audio documents concurrency limits separately from requests per second. It explains that request duration affects the throughput a concurrency allowance can support. This supports testing realistic request lengths rather than comparing bare rate-limit numbers. Fish Audio pricing and rate-limit documentation, checked 16 September 2026.
Keep an evaluation log
For each request, record request ID, model, input units, language, first-audio time, completion time, output duration, status, retries, billed usage and whether the output was usable. Do not publish API keys, personal voice recordings or private request payloads in the log.
Test timeout handling and cancelled requests in an authorised test environment. Check the vendor’s published retry guidance before generating extra traffic. Keep rejected audio and pronunciation corrections in your total cost.
Interpret third-party listening results correctly
Artificial Analysis Speech Arena offers blind listening comparisons. Such evidence can help build an audition list, but it is not a benchmark of your application, language, region or correction workflow. The Voice Bench has not run API benchmarks for either tool yet.
Decide with evidence, not one number
Write a decision record with an accepted use case, rejected use cases, required rights, measured costs and unresolved risks. Do not combine a vendor’s fastest model latency with another model’s voice-quality claim.
Use the pricing-unit guide to normalise quotes and the Fish Audio vs ElevenLabs comparison for the current research shortlist. A direct API winner must wait for measured requests and retained evidence.