MiMo TTS API: Turn Text into a Playable WAV File

A successful MiMo speech request returns JSON containing audio data; saving the JSON body with a .wav extension will not create an audio file. The useful sequence is to keep the response, decode the audio, validate the WAV container, and then listen to the result.
This guide uses Tokenhot's preset-voice mimo-v2.5-tts route. The preset-voice reference specifies POST https://api.tokenhot.ai/v1/chat/completions, Bearer authentication, and an important message convention: spoken text belongs in an assistant message; an optional user message controls style. This is a speech route, so do not replace that convention with a generic chat example.
No live synthesis or listening test was performed for this guide. The supplied Python helper was checked locally using artificial PCM WAV data, including failure cases. Those tests verify file handling, not the sound or behavior of a real model.
Build a preset-voice request
Here is an authored, non-streaming request using the documented English preset Mia:
{
"model": "mimo-v2.5-tts",
"messages": [
{"role": "user", "content": "Read clearly at a steady conversational pace."},
{"role": "assistant", "content": "Your project is ready. Open the dashboard to review the result."}
],
"audio": {"format": "wav", "voice": "Mia"}
}
Keep the spoken copy separate from delivery directions so that changing the voiceover does not accidentally change its tone instruction. Use a voice explicitly listed for your intended language. The helper chooses an explicit preset rather than relying on a cluster-dependent default.
The preset route is not the voice-design or voice-cloning route. Their separate model IDs appear in the same documentation, but reference-audio submission, voice design and streaming output are outside this tutorial. Start with this small request before designing a more complex voice pipeline.
Prepare the JSON without spending API usage
Save mimo_wav.py beside a UTF-8 text file named voiceover.txt. The helper uses the Python standard library. Its local checks were run with Python 3.13.5; it is not a tested claim about every Python or operating-system version.
Run the local preparation command:
python mimo_wav.py prepare \
--text-file voiceover.txt \
--voice Mia \
--style "Read clearly at a steady conversational pace." \
--output mimo-request.json
This writes the JSON request and stops. It does not require a key, open a connection or contact a synthesis service. Review the actual text and preset before proceeding. The file is intentionally not overwritten by a later prepare command; use a new name for an intentional revision.
For credentials, follow the current Tokenhot setup guide. Set TOKENHOT_API_KEY in your environment rather than in the request JSON or Python source. Do not commit the key, a terminal dump containing it, or an unredacted account response.
Synthesize once and retain the response
After authorizing the actual request and its possible charge, run:
python mimo_wav.py synthesize \
--request mimo-request.json \
--run runs/mimo-001 \
--execute
The --execute flag is an explicit network gate. The command reserves a new run directory before sending one POST. It stores your request, safe request metadata and the raw response locally. When decoding succeeds, the directory also contains speech.wav.
Do not infer permanent free use from a zero price displayed elsewhere. This guide has no verified account charge, billing entitlement or free-usage guarantee. Check the terms and price for your actual route before sending requests, and reconcile its usage afterward.
The response is saved before decoding because synthesis and local file creation are different operations. A full response can arrive even if an audio field is malformed or the output directory is unavailable. Repeating the synthesis is not the right first response to that local failure.
Decode audio bytes, not the surrounding JSON
In the documented response, the Base64 payload is at choices[0].message.audio.data, and the illustrated completion uses finish_reason: "stop". The helper checks that structure before decoding. The documentation abbreviates its audio string, so that sample cannot be used as a real audio test fixture.
The relevant local operations are:
import base64
import io
import wave
def inspect_pcm_wav(encoded):
"""Local container inspection, not a speech-quality check."""
raw = base64.b64decode(encoded, validate=True)
with wave.open(io.BytesIO(raw), "rb") as audio:
return {
"channels": audio.getnchannels(),
"sample_rate_hz": audio.getframerate(),
"sample_width_bytes": audio.getsampwidth(),
"frames": audio.getnframes(),
}
This small excerpt only inspects the header. The complete companion additionally rejects empty audio, checks that the declared PCM frames are present, and refuses to overwrite an existing output. It returns sanitized diagnostics rather than printing arbitrary provider error text.
Python 3.13's wave module supports uncompressed PCM WAV and the applicable PCM extensible headers. An unsupported format should trigger investigation, not an attempt to manufacture a header. In particular, do not add a second WAV header to bytes that already include one. PCM streaming is a different framing problem and is deliberately not handled by this script.
The helper reads the sample rate and channel count from the returned container instead of assuming the settings of a different streaming mode. That makes an unexpected format visible before it becomes a hard-to-debug playback issue.
Recover locally when saving failed
When a complete response already exists, use the local save command:
python mimo_wav.py save \
--response runs/mimo-001/response.json \
--output output/recovered-voiceover.wav
This command has no network path and needs no API key. It reads the saved response, decodes its audio and writes to a fresh destination. Keep the original response until you have accepted the output.
| Observation | Next action |
|---|---|
| HTTP rejection with a saved body | Inspect that body privately and correct the request or account issue; do not blindly retry. |
| Missing, invalid or abbreviated audio data | Check the actual response and route. A documentation placeholder is not downloadable audio. |
| Valid JSON but unsupported audio container | Confirm the returned format and decoder; do not guess the encoding. |
| Existing output file | Inspect it or choose a deliberate new output name. |
| Timeout without a usable response | Treat the synthesis outcome as unknown and reconcile the attempt before sending another POST. |
The HTTP helper has a local 64 MiB response guard. This is an implementation limit for the example, not a documented MiMo maximum. It also rejects redirects. A redirected endpoint or substantially larger response requires an intentional implementation change and another check of credential and memory handling.
Listen before using the voiceover
Open the saved WAV in a player. Confirm the entire text is present, product names and numbers are pronounced acceptably, the pacing works, and no unwanted spoken instruction appears. Container validation cannot answer those questions.
Keep content and delivery decisions separate when revising: changing text, preset, style and punctuation simultaneously makes a later result harder to evaluate. Record the exact request for each intentional version, and count all actual synthesis attempts when measuring cost.
At this point the code gives you a reproducible way to handle the documented response and recover a local save. It does not establish voice quality, latency, entitlement or billing for your account. Those are the next observations to record during an authorized real run.
Use the MiMo preset-voice route, decode its Base64 audio into WAV, validate the saved file, and recover locally without repeating synthesis.


