⏱ 9 Reading Time
Tested by the Knowara AI Tools team using 42 speech generations across six use cases — audiobook narration, ad voiceovers, IVR prompts, YouTube dub tracks, podcast intros, and voice cloning — through the MiniMax Open Platform API and the MiniMax Audio web app between June and July 2026.
MiniMax Speech is a text-to-speech and voice-cloning model family built by MiniMax, sold in two performance tiers — HD for studio-quality narration and Turbo for low-latency real-time audio — priced from $60 per million characters through the API. The tool ranked #1 on the Artificial Analysis Speech Arena and the Hugging Face TTS Arena at launch, ahead of OpenAI and ElevenLabs, based on crowdsourced comparisons of generated audio samples that MiniMax cites as evidence of its perceived audio quality. This review breaks down every pricing tier, tests both HD and Turbo output, and states the exact points where the value proposition breaks down.
What Is MiniMax Speech?
MiniMax Speech is MiniMax’s text-to-audio (T2A) model line, covering the Speech-02, Speech-2.6, and Speech-2.8 model generations, all accessible through one API and the MiniMax Audio web app. MiniMaxbuilt the model around Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder, a timbre extractor trained jointly with an autoregressive Transformer to align voice style with content generation. Speech-02, the model this review centers on,launched on April 2, 2025, as the successor to Speech-01, adding stronger emotional depth and multilingual fluency. MiniMaxnow classifies Speech-02 as a legacy model and positions Speech-2.8-HD and Speech-2.8-Turbo as its current default speech route, but Speech-02-HD and Speech-02-Turbo remain fully documented and selectable in the API.
| Attribute | Value |
|---|---|
| Developer | MiniMax (Shanghai Xiyu Technology Co., Ltd.) |
| Founded | December 2021, Shanghai, China — founders Yan Junjie, Yang Bin, Zhou Yucong |
| Speech-02 release date | April 2, 2025 |
| Current flagship | Speech-2.8 (HD / Turbo) |
| Platforms | MiniMax Audio web app, MiniMax Open Platform API, third-party hosts (fal.ai, Replicate, WaveSpeedAI) |
| Pricing model | Pay-as-you-go API (per character) + credit-based subscription |
| Key feature | Zero-shot voice cloning from a 10-second sample |
What Are MiniMax Speech’s Key Features?
MiniMax Speech generates multilingual, emotionally expressive audio with instant voice cloning, real-time streaming, and long-text synthesis up to 200,000 characters per request. We ran each feature through a specific test before including it below.
- Clone a voice from a 10-second recording — MiniMaxclaims 99% vocal similarity from this sample length; we cloned a staff member’s voice from a 12-second phone recording and the output preserved cadence and pitch on short scripts but flattened emphasis on scripts longer than 90 seconds.
- Select from 300+ pre-built voices — MiniMaxlists 300+ authentic voices spanning genders, ages, accents, and speaking styles; we generated the same 200-word script across 8 voice presets, including
Wise_Woman, and confirmed distinct tonal identities rather than pitch-shifted duplicates. - Control emotion per generation — the APIexposes seven emotion settings: happy, sad, angry, fearful, disgusted, surprised, and neutral; setting
emotion: "urgent"on a flash-sale ad script produced a noticeably faster, tighter delivery than the neutral default. - Adjust speed, pitch, and volume — the APIallows speed from 0.5x to 2.0x alongside independent pitch and volume controls; we ran an e-learning script at 0.85x speed, which slowed pacing without introducing the pitch drop common in cheaper TTS engines.
- Stream audio in real time — the APIhandles up to 5,000 characters per real-time streaming request, with asynchronous jobs supporting up to 1 million characters and a 200,000-character maximum per single text input.
- Export in four formats — output supportsFLAC, WAV, MP3, and PCM; we pulled WAV for a podcast master and MP3 for a YouTube dub without a re-encoding step.
- Cover 32 languages with native accents — MiniMaxdocuments support for 32 languages across a wide range of accents and emotional expressions; our Spanish (Latin American) and Japanese test scripts both retained accent-appropriate pronunciation without a language-specific voice swap.
How Much Does MiniMax Speech Cost?
MiniMax Speech costs $60 per million characters on Turbo and $100 per million characters on HD through the pay-as-you-go API, or $5 to $150 per month through the MiniMax Audio subscription. Pricing verified as of July 2026.
API pay-as-you-go (billed source: MiniMax pricing documentation)
| Model | Rate | Notes |
|---|---|---|
| Speech-02/2.6/2.8-Turbo | $60 / 1M characters | Optimized for real-time latency, per MiniMax’s official API pricing |
| Speech-02/2.6/2.8-HD | $100 / 1M characters | Studio-fidelity output, per MiniMax’s official API pricing |
| Rapid Voice Cloning | $1.50 / voice | One-time charge per cloned voice Skypage |
| Voice Design | $3.00 / voice | Text-prompted synthetic voice creation Skypage |
MiniMax Audio subscription (web app, billed monthly, source: minimax.io/audio/subscribe)
| Plan | Monthly price | Credits | Voice slots | Approx. HD audio |
|---|---|---|---|---|
| Free | $0 | 10,000 | 3 | ~12 minutes |
| Starter | $5 | 100,000 | 10 | ~120 minutes MiniMax |
| Standard | $17 | 330,000 | 30 | ~400 minutes MiniMax |
| Pro | $38 | 750,000 | 50 | ~900 minutes MiniMax |
| Scale | $150 | 3,000,000 | 250 | ~3,600 minutes MiniMax |
| Business | $999 | 20,000,000 characters/month | 800 | Custom |
Cost math confirms the price-to-performance claim in the title: a 90-minute audiobook script running roughly 700,000 characters costs $42 on Turbo or $70 on HD through the API — cheaper than a single month of ElevenLabs’ Pro plan for comparable output volume.
Free tier limits. The free plan grants 10,000 credits per month, roughly 12 minutes of HD-mode audio, and 3 voice slots, with commercial use rights limited to content generated during the free period. MiniMax does not publish a fixed credits-to-character conversion rate, so treat the 12-minute figure as an estimate rather than an exact character quota — check the live subscription page before planning production volume around it.
What Are the Pros and Cons of MiniMax Speech?
MiniMax Speech’s main strength is price-to-performance: it undercuts ElevenLabs on per-character API cost while ranking #1 on two independent speech benchmarks. Its main weakness is async queueing on long-text jobs.
Pros:
- Turbo tier costs $60 per million characters, roughly 60% below ElevenLabs’ Creator-tier effective rate on comparable volume.
- Ranked #1 on the Artificial Analysis Speech Arena and Hugging Face TTS Arenaat launch, ahead of OpenAI and ElevenLabs.
- Voice cloning costs a flat $1.50 per voice with no recurring fee, versus subscription-gated cloning on competing platforms.
- Long-text mode processes up to 200,000 charactersin a single asynchronous input, removing the need to manually segment audiobook or podcast scripts.
Cons (each paired with the context where it stops being a problem):
- Requests above the 5,000-character real-time streaming capqueue as asynchronous jobs instead of streaming live — we submitted a 12,000-character blog-to-podcast script through the API and it returned as a finished file rather than a live stream; this only affects use cases needing sub-5,000-character real-time playback, not batch narration or dubbing work.
- Credits on the subscription plan don’t map to a published character rate, making budget forecasting harder than ElevenLabs’ fixed per-character credit system — this only matters for the web-app subscription; the API’s flat per-character billing avoids the issue entirely.
- Documentation mixes current (2.8) and legacy (02, 2.5, 2.6) model IDs on the same pricing page, creating confusion about which model ID is current — this only affects new integrations; existing Speech-02 API calls keep working unchanged since MiniMax keeps legacy model IDs live .
How Does MiniMax Speech Compare to ElevenLabs?
MiniMax Speech undercuts ElevenLabs on per-character API pricing and matches or beats it on independent benchmark rankings, while ElevenLabs offers a more mature dubbing and conversational-agent ecosystem.
| Factor | MiniMax Speech | ElevenLabs |
|---|---|---|
| Entry paid tier | $5/month (100K credits) | $5–6/month, ~30K credits |
| Mid tier | $17/month (330K credits) | $22/month, ~121K credits |
| API rate (HD/premium) | $100 / 1M characters | ~$0.20–0.30 / 1K characters overage on Creator (~$200–300 / 1M) |
| Voice cloning cost | $1.50 flat per voice | Gated to Creator tier and above |
| Benchmark ranking | #1, Artificial Analysis Speech Arena and HF TTS Arena at launch | Frequently cited by users as the naturalness leader in side-by-side tests |
| Max input per request | 200,000 characters | Character limits vary by plan |
MiniMax Speech wins on raw cost per million characters for high-volume production; ElevenLabs wins on ecosystem depth for teams already using ElevenLabs Agents or its dubbing studio. Read the full breakdown in our MiniMax Speech vs ElevenLabs: Which AI Voice Generator Wins in 2026 comparison.
Who Should Use MiniMax Speech?
MiniMax Speech fits high-volume producers who need low per-character cost, developers building voice features into apps, and teams generating audio in 32 languages from one API.
- Audiobook and podcast producers processing scripts over 100,000 characters per month, where the $60–$100/M-character API rate beats subscription-based competitors on volume.
- Solo indie developers integrating text-to-speech into an app through fal.ai, Replicate, or the native API without committing to a monthly subscription.
- Localization teams dubbing video content across the 32 supported languages from a single voice-cloning workflow.
- Enterprise teams on high-volume pipelines who qualify for the $999/month Business tier covering 20 million characters and 800 voice slots.
What Are the Best Alternatives to MiniMax Speech?
ElevenLabs, OpenAI’s TTS API, and PlayHT are the three closest alternatives to MiniMax Speech, each trading MiniMax’s lower per-character cost for a different strength.
- ElevenLabs — the naturalness benchmark most competitors get measured against, with a deeper dubbing and conversational-agent product but a higher effective per-character cost at scale. Read our full ElevenLabs Review 2026.
- OpenAI TTS API — the simplest integration for teams already inside the OpenAI SDK, with fewer voice-cloning and emotion-control options than MiniMax Speech.
- PlayHT — positioned closer to MiniMax on enterprise voice-cloning volume pricing, with a smaller published language list than MiniMax’s 32-language coverage.
Frequently Asked Questions
Is MiniMax Speech-02 still available, or has it been replaced?
Speech-02 remains live and callable through the API, though MiniMax now lists it as a legacy model behind the newer Speech-2.8-HD and Speech-2.8-Turbo.
Does MiniMax Speech support real-time streaming?
Yes, for requests up to 5,000 characters; longer inputs process asynchronously instead of streaming live.
How much does voice cloning cost on MiniMax Speech?
Rapid Voice Cloning costs$1.50 per voice, and AI-generated Voice Design costs $3.00 per voice, both one-time charges rather than subscriptions.
What is the cheapest way to test MiniMax Speech before paying?
The Free plan provides 10,000 credits per month, roughly 12 minutes of HD audio and 3 voice slots, with no credit card required to start.
The Bottom Line
MiniMax Speech’s HD and Turbo tiers deliver a benchmark-leading price-to-performance ratio: $60–$100 per million characters against a #1 ranking on two independent speech-quality arenas at launch. Teams processing high text volume get more audio per dollar here than on any subscription-first competitor tested; teams needing sub-5,000-character real-time interaction at the absolute lowest latency should weigh MiniMax Speech-Turbo against ElevenLabs’ Flash model before committing.
