10 KiB
10 KiB
name, description
| name | description |
|---|---|
| mlx.audio | Generate speech from text and transcribe audio using mlx-audio. Use when the user wants text-to-speech synthesis, speech-to-text transcription, voice cloning, audio separation, or speech-to-speech processing on Apple Silicon. |
MLX-Audio
A speech processing library built on Apple's MLX framework, providing TTS, STT, speech-to-speech (STS), and audio separation optimized for Apple Silicon.
- Repository: https://github.com/Blaizzy/mlx-audio
- License: MIT
CLI Tools
Text-to-Speech (TTS)
mlx_audio.tts.generate --model <model> --text '<text>' [options]
| Flag | Type | Default | Description |
|---|---|---|---|
--model |
string | required | HuggingFace model ID |
--text |
string | required | Text to synthesize |
--voice |
string | — | Voice preset (model-specific) |
--speed |
float | 1.0 | Speech speed multiplier |
--lang_code |
string | a |
Language code |
--play |
flag | — | Play audio immediately |
--output_path |
string | — | Directory to save audio |
--ref_audio |
string | — | Reference audio for voice cloning (CSM) |
Language Codes
| Code | Language |
|---|---|
a |
American English |
b |
British English |
j |
Japanese |
z |
Mandarin Chinese |
e |
Spanish |
f |
French |
Kokoro Voices
| Voice | Description |
|---|---|
af_heart, af_bella, af_nova, af_sky |
American female |
am_adam, am_echo |
American male |
bf_alice, bf_emma |
British female |
bm_daniel, bm_george |
British male |
jf_alpha, jm_kumo |
Japanese |
zf_xiaobei, zm_yunxi |
Chinese |
Examples
# Basic generation
mlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello, world!' --lang_code a
# With voice and speed
mlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --voice af_heart --speed 1.2 --lang_code a
# Play immediately
mlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --play --lang_code a
# Voice cloning with CSM
mlx_audio.tts.generate --model mlx-community/csm-1b --text "Hello from Sesame." --ref_audio ./reference_voice.wav --play
Speech-to-Text (STT)
python -m mlx_audio.stt.generate --model <model> --audio <file> [options]
| Flag | Type | Default | Description |
|---|---|---|---|
--model |
string | required | HuggingFace model ID |
--audio |
string | required | Input audio file |
--language |
string | — | Language code |
--max-tokens |
int | 1024 | Maximum output tokens |
--temperature |
float | 0.0 | Sampling temperature |
--context |
string | — | Hotwords/metadata for context |
--output-path |
string | — | Output directory |
--format |
string | — | Output format (e.g. json) |
--stream |
flag | — | Enable streaming mode |
--verbose |
flag | — | Detailed logging |
Examples
# Basic transcription
python -m mlx_audio.stt.generate --model mlx-community/whisper-large-v3-turbo-asr-fp16 --audio speech.wav --verbose
# With context for technical terms
python -m mlx_audio.stt.generate --model mlx-community/VibeVoice-ASR-bf16 --audio meeting.wav --context "MLX, Apple Silicon, PyTorch" --max-tokens 8192 --format json --verbose
# Parakeet model
python -m mlx_audio.stt.generate --model mlx-community/parakeet-tdt-0.6b-v3 --audio speech.wav --format json --verbose
Python API
TTS
from mlx_audio.tts.utils import load_model
model = load_model("mlx-community/Kokoro-82M-bf16")
for result in model.generate("Hello from MLX-Audio!", voice="af_heart"):
audio = result.audio # mx.array waveform
STT
from mlx_audio.stt.generate import generate_transcription
result = generate_transcription(
model="mlx-community/whisper-large-v3-turbo-asr-fp16",
audio="audio.wav",
)
print(result.text)
STT with Streaming
from mlx_audio.stt import load
# VibeVoice-ASR streaming
model = load("mlx-community/VibeVoice-ASR-bf16")
for text in model.stream_transcribe(audio="speech.wav", max_tokens=4096):
print(text, end="", flush=True)
# Parakeet streaming
model = load("mlx-community/parakeet-tdt-0.6b-v3")
for chunk in model.generate("long_audio.wav", stream=True):
print(chunk.text, end="", flush=True)
Forced Alignment (Qwen3)
from mlx_audio.stt import load
aligner = load("mlx-community/Qwen3-ForcedAligner-0.6B-8bit")
result = aligner.generate("audio.wav", text="I have a dream", language="English")
for item in result:
print(f"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}")
REST API Server (OpenAI-compatible)
Starting the Server
python -m mlx_audio.server [OPTIONS]
| Flag | Type | Default | Description |
|---|---|---|---|
--host |
string | localhost |
Server host |
--port |
int | 8000 |
Server port |
--allowed-origins |
string | * |
CORS allowed origins |
--workers |
int/float | 2 |
Number of workers |
--reload |
flag | — | Enable auto-reload |
--start-ui |
flag | — | Launch Studio UI alongside API |
--log-dir |
string | logs |
Directory for server logs |
Endpoints
GET /v1/models
List available models.
curl http://localhost:8000/v1/models
POST /v1/models?model_name=<name>
Add a model to the server.
DELETE /v1/models?model_name=<name>
Remove a model from the server.
POST /v1/audio/speech
Generate speech from text.
curl -X POST http://localhost:8000/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{
"model": "mlx-community/Kokoro-82M-bf16",
"input": "Hello, world!",
"voice": "af_heart",
"speed": 1.0,
"lang_code": "a",
"response_format": "mp3"
}' --output speech.mp3
| Parameter | Type | Default | Description |
|---|---|---|---|
model |
string | required | Model ID |
input |
string | required | Text to synthesize |
voice |
string | — | Voice preset |
speed |
float | 1.0 | Speech speed |
lang_code |
string | a |
Language code |
ref_audio |
string | — | Reference audio path (voice cloning) |
ref_text |
string | — | Reference transcript |
response_format |
string | mp3 |
Output format |
stream |
bool | false | Enable streaming |
streaming_interval |
float | 2.0 | Streaming chunk interval |
temperature |
float | 0.7 | Sampling temperature |
top_p |
float | 0.95 | Nucleus sampling |
top_k |
int | 40 | Top-k sampling |
repetition_penalty |
float | 1.0 | Repetition penalty |
max_tokens |
int | 1200 | Maximum tokens |
gender |
string | male |
Gender hint |
pitch |
float | 1.0 | Pitch adjustment |
instruct |
string | — | Instruction text |
POST /v1/audio/transcriptions
Transcribe an audio file (multipart/form-data).
curl -X POST http://localhost:8000/v1/audio/transcriptions \
-F file=@audio.wav \
-F model=mlx-community/whisper-large-v3-turbo-asr-fp16
| Parameter | Type | Default | Description |
|---|---|---|---|
file |
file | required | Audio file |
model |
string | required | Model ID |
language |
string | — | Language code |
max_tokens |
int | 1024 | Maximum tokens |
chunk_duration |
float | 30.0 | Chunk duration (seconds) |
stream |
bool | false | Enable streaming |
context |
string | — | Hotwords/context |
text |
string | — | Reference text |
verbose |
bool | false | Detailed output |
Response (NDJSON stream):
{"text": "chunk text", "accumulated": "full text so far"}
WebSocket /v1/audio/transcriptions/realtime
Real-time transcription via WebSocket. Send initial config as JSON, then stream int16 PCM audio as binary frames.
{
"model": "mlx-community/whisper-large-v3-turbo-asr-fp16",
"sample_rate": 16000,
"streaming": true
}
POST /v1/audio/separations
Separate audio sources (multipart/form-data).
curl -X POST http://localhost:8000/v1/audio/separations \
-F file=@audio.wav \
-F model=mlx-community/sam-audio-large-fp16 \
-F description="speech"
| Parameter | Type | Default | Description |
|---|---|---|---|
file |
file | required | Audio file |
model |
string | mlx-community/sam-audio-large-fp16 |
Model ID |
description |
string | speech |
Target description |
method |
string | midpoint |
ODE method (midpoint or euler) |
steps |
int | 16 | ODE steps (2/4/8/16/32) |
Response:
{
"target": "<base64 WAV>",
"residual": "<base64 WAV>",
"sample_rate": 44100
}
Supported Models
TTS Models
| Model | Languages | Notes |
|---|---|---|
| Kokoro | EN, JA, ZH, FR, ES, IT, PT, HI | Fast, high-quality multilingual |
| Qwen3-TTS | ZH, EN, JA, KO, + more | Voice design via instruction |
| CSM | EN | Voice cloning with reference audio |
| Dia | EN | Dialogue-focused |
| OuteTTS | EN | Efficient |
| Spark | EN, ZH | SparkTTS |
| Chatterbox | EN, ES, FR, DE, IT, PT, PL, TR, RU, NL, CS, AR, ZH, JA, HU, KO | Expressive multilingual |
| Soprano | EN | High-quality |
STT Models
| Model | Languages | Notes |
|---|---|---|
| Whisper | 99+ languages | OpenAI's robust model |
| Qwen3-ASR | ZH, EN, JA, KO, + more | Alibaba multilingual |
| Qwen3-ForcedAligner | ZH, EN, JA, KO, + more | Word-level alignment |
| Parakeet | EN (v2), 25 EU languages (v3) | NVIDIA, high accuracy |
| Voxtral | Multiple | Mistral speech model |
| Voxtral Realtime | Multiple | 4B streaming STT |
| VibeVoice-ASR | Multiple | Microsoft 9B, supports diarization and context |
Other Models
| Model | Type | Description |
|---|---|---|
| Sortformer v1/v2.1 | VAD/Diarization | Speaker diarization (up to 4 speakers) |
| SAM-Audio | Separation | Text-guided source separation |
| Liquid2.5-Audio | STS | Speech/text-to-speech and STT |
| MossFormer2 SE | Enhancement | Speech enhancement / noise removal |
Models are available from mlx-community on HuggingFace with various quantization levels (3-bit through 8-bit and fp16/bf16).
Source
- Repository: https://github.com/Blaizzy/mlx-audio
- HuggingFace: https://huggingface.co/mlx-community