Files
platform/seed/skills/mlxaudio/chat/messages.json
T
2026-02-16 19:34:35 +00:00

1 line
69 KiB
JSON
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
[{"role":"user","text":"Help me create the content for this new skill file: /home/pastilhas/projects/officer.dev/data/skills/mlxaudio/SKILL.md\nfrom https://github.com/Blaizzy/mlx-audio"},{"role":"assistant","text":"\n\nLet me start by reading the current (empty) SKILL.md file and fetching information about the mlx-audio project."},{"role":"tool","toolName":"Read","toolInput":{"file_path":"/home/pastilhas/projects/officer.dev/data/skills/mlxaudio/SKILL.md"},"toolUseId":"toolu_01MXmpTWFrzx5Wm2fUTBEqcQ","output":" 1→---\n 2→name: mlx.audio\n 3→description: \n 4→---\n 5→","isError":false},{"role":"tool","toolName":"Read","toolInput":{"file_path":"/home/pastilhas/projects/officer.dev/data/skills/whisper.cpp/SKILL.md"},"toolUseId":"toolu_01RfdTasfwALeQhE26dkENTD","output":" 1→---\n 2→name: whisper.cpp\n 3→description: Transcribe audio files to text using whisper.cpp. Use when the user wants to transcribe audio, convert speech to text, or extract text from an audio/video file.\n 4→---\n 5→\n 6→# Whisper.cpp\n 7→\n 8→API reference for the whisper.cpp HTTP server running at `http://macmini:8178`.\n 9→\n 10→whisper.cpp is a C/C++ port of OpenAI's Whisper speech recognition model. The server accepts audio files via HTTP and returns transcriptions in various formats.\n 11→\n 12→## Server\n 13→\n 14→- **Base URL:** `http://macmini:8178`\n 15→- **No authentication required**\n 16→\n 17→## Endpoints\n 18→\n 19→### POST /inference\n 20→\n 21→Transcribes an audio file. Accepts `multipart/form-data`.\n 22→\n 23→#### Example\n 24→\n 25→```bash\n 26→curl -s http://macmini:8178/inference \\\n 27→ -F file=\"@/path/to/audio.mp3\" \\\n 28→ -F temperature=\"0.0\" \\\n 29→ -F temperature_inc=\"0.2\" \\\n 30→ -F response_format=\"json\"\n 31→```\n 32→\n 33→#### Parameters\n 34→\n 35→##### File (required)\n 36→\n 37→| Parameter | Type | Description |\n 38→|-----------|------|-------------|\n 39→| `file` | file | Audio file to transcribe. Accepts at least WAV and MP3. |\n 40→\n 41→##### Response Format\n 42→\n 43→| Parameter | Type | Default | Description |\n 44→|-----------|------|---------|-------------|\n 45→| `response_format` | string | `json` | Output format: `json`, `verbose_json` (or `vjson`), `text`, `srt`, `vtt` |\n 46→\n 47→##### Language\n 48→\n 49→| Parameter | Type | Default | Description |\n 50→|-----------|------|---------|-------------|\n 51→| `language` | string | `en` | Spoken language code (e.g. `en`, `pt`, `es`, `fr`). Use `auto` for auto-detection. |\n 52→| `detect_language` | bool | `false` | Exit after detecting the language (no transcription). |\n 53→| `translate` | bool | `false` | Translate from source language to English. |\n 54→\n 55→##### Decoding\n 56→\n 57→| Parameter | Type | Default | Description |\n 58→|-----------|------|---------|-------------|\n 59→| `temperature` | float | `0.0` | Sampling temperature. `0.0` is deterministic. |\n 60→| `temperature_inc` | float | `0.2` | Temperature increment on fallback attempts. |\n 61→| `best_of` | int | `2` | Number of candidate decodings to keep. |\n 62→| `beam_size` | int | `-1` | Beam search size. `-1` disables beam search. |\n 63→| `entropy_thold` | float | `2.40` | Entropy threshold — decoder fails and retries if exceeded. |\n 64→| `logprob_thold` | float | `-1.00` | Log probability threshold for decoder failure. |\n 65→| `no_fallback` | bool | `false` | Disable temperature fallback on decode failure. |\n 66→\n 67→##### Segmentation\n 68→\n 69→| Parameter | Type | Default | Description |\n 70→|-----------|------|---------|-------------|\n 71→| `max_len` | int | `0` | Maximum segment length in characters. `0` for unlimited. |\n 72→| `max_context` | int | `-1` | Maximum text context tokens to store. `-1` for unlimited. |\n 73→| `split_on_word` | bool | `false` | Split segments at word boundaries instead of token boundaries. |\n 74→| `no_timestamps` | bool | `false` | Suppress timestamps in output. |\n 75→| `word_thold` | float | `0.01` | Word timestamp probability threshold. |\n 76→\n 77→##### Audio Processing\n 78→\n 79→| Parameter | Type | Default | Description |\n 80→|-----------|------|---------|-------------|\n 81→| `offset_t` | int | `0` | Time offset in milliseconds — skip this much audio from the start. |\n 82→| `offset_n` | int | `0` | Segment index offset. |\n 83→| `duration` | int | `0` | Duration of audio to process in milliseconds. `0` for all. |\n 84→| `audio_ctx` | int | `0` | Audio context size. `0` for all. |\n 85→\n 86→##### Speaker Diarization\n 87→\n 88→| Parameter | Type | Default | Description |\n 89→|-----------|------|---------|-------------|\n 90→| `diarize` | bool | `false` | Enable speaker diarization (requires stereo audio). |\n 91→| `tinydiarize` | bool | `false` | Enable tinydiarize (requires a tdrz model). |\n 92→\n 93→##### Voice Activity Detection (VAD)\n 94→\n 95→| Parameter | Type | Default | Description |\n 96→|-----------|------|---------|-------------|\n 97→| `vad` | bool | `false` | Enable VAD preprocessing. |\n 98→| `vad_threshold` | float | `0.50` | Speech confidence threshold (0.01.0). |\n 99→| `vad_min_speech_duration_ms` | int | `250` | Minimum speech segment duration in ms. |\n 100→| `vad_min_silence_duration_ms` | int | `100` | Minimum silence duration to split segments. |\n 101→| `vad_max_speech_duration_s` | float | `FLT_MAX` | Auto-split segments longer than this (seconds). |\n 102→| `vad_speech_pad_ms` | int | `30` | Padding added around speech segments (ms). |\n 103→| `vad_samples_overlap` | float | `0.10` | Overlap between segments (seconds). |\n 104→\n 105→##### Other\n 106→\n 107→| Parameter | Type | Default | Description |\n 108→|-----------|------|---------|-------------|\n 109→| `prompt` | string | `\"\"` | Initial prompt to condition the model (e.g. for vocabulary hints). |\n 110→| `suppress_nst` | bool | `false` | Suppress non-speech tokens. |\n 111→| `no_context` | bool | `false` | Do not use previous audio context for subsequent segments. |\n 112→| `debug_mode` | bool | `false` | Enable debug output. |\n 113→\n 114→#### Response Formats\n 115→\n 116→##### `json` (default)\n 117→\n 118→Minimal JSON with just the transcribed text.\n 119→\n 120→```json\n 121→{\"text\": \"The transcribed content goes here.\"}\n 122→```\n 123→\n 124→##### `verbose_json` (or `vjson`)\n 125→\n 126→Extended JSON including task type, language, audio duration, per-segment timestamps, token-level timing, confidence scores, and language probability distribution.\n 127→\n 128→##### `text`\n 129→\n 130→Plain text transcription. Includes speaker labels if diarization is enabled.\n 131→\n 132→##### `srt`\n 133→\n 134→SubRip subtitle format with sequential numbering, `HH:MM:SS,mmm` timestamps, and text content.\n 135→\n 136→```\n 137→1\n 138→00:00:00,000 --> 00:00:03,500\n 139→The transcribed content goes here.\n 140→```\n 141→\n 142→##### `vtt`\n 143→\n 144→WebVTT subtitle format with `WEBVTT` header and `HH:MM:SS.mmm` timestamps.\n 145→\n 146→```\n 147→WEBVTT\n 148→\n 149→00:00:00.000 --> 00:00:03.500\n 150→The transcribed content goes here.\n 151→```\n 152→\n 153→### POST /load\n 154→\n 155→Loads a different model file on the server at runtime.\n 156→\n 157→```bash\n 158→curl -s http://macmini:8178/load \\\n 159→ -F model=\"/path/to/model.bin\"\n 160→```\n 161→\n 162→| Parameter | Type | Description |\n 163→|-----------|------|-------------|\n 164→| `model` | string | Path to the model file on the server. |\n 165→\n 166→## Supported Audio Formats\n 167→\n 168→The server accepts at least WAV (16-bit PCM) and MP3 files directly. If the server was started with `--convert`, it can use ffmpeg to handle additional formats (ogg, flac, m4a, etc.).\n 169→\n 170→## Source\n 171→\n 172→- Repository: https://github.com/ggml-org/whisper.cpp\n 173→- Server docs: https://github.com/ggml-org/whisper.cpp/blob/master/examples/server/README.md\n 174→","isError":false},{"role":"tool","toolName":"Read","toolInput":{"file_path":"/home/pastilhas/projects/officer.dev/data/skills/ffmpeg/SKILL.md"},"toolUseId":"toolu_0164FTK5nf9HT8m3GArB4jgy","output":" 1→---\n 2→name: ffmpeg\n 3→description: Process audio and video files using ffmpeg/ffprobe. Use when the user wants to convert, transcode, trim, merge, extract, resize, compress, or analyze multimedia files.\n 4→---\n 5→\n 6→# FFmpeg\n 7→\n 8→CLI reference for FFmpeg v8.x — a complete, cross-platform solution for recording, converting, and streaming audio and video.\n 9→\n 10→Official docs: https://www.ffmpeg.org/documentation.html\n 11→\n 12→## Tools\n 13→\n 14→| Tool | Purpose |\n 15→|------|---------|\n 16→| `ffmpeg` | Transcode, convert, filter, mux/demux multimedia |\n 17→| `ffprobe` | Analyze and inspect multimedia streams |\n 18→| `ffplay` | Play multimedia files (interactive) |\n 19→\n 20→---\n 21→\n 22→## ffmpeg\n 23→\n 24→### Synopsis\n 25→\n 26→```\n 27→ffmpeg [global_options] {[input_options] -i input_url} ... {[output_options] output_url} ...\n 28→```\n 29→\n 30→Options before `-i` apply to the input; options before the output URL apply to the output.\n 31→\n 32→### Global Options\n 33→\n 34→| Flag | Description |\n 35→|------|-------------|\n 36→| `-y` | Overwrite output files without asking |\n 37→| `-n` | Do not overwrite; exit if output exists |\n 38→| `-hide_banner` | Suppress copyright/build info banner |\n 39→| `-loglevel level` | Set log level: `quiet`, `error`, `warning`, `info` (default), `verbose`, `debug` |\n 40→| `-stats` | Print encoding progress/statistics |\n 41→| `-progress url` | Send machine-readable progress to url |\n 42→| `-report` | Dump full command line and log to a file |\n 43→| `-filter_threads n` | Number of threads for filter processing |\n 44→\n 45→### Input/Output Options\n 46→\n 47→| Flag | Description |\n 48→|------|-------------|\n 49→| `-i url` | Input file URL |\n 50→| `-f fmt` | Force input or output format |\n 51→| `-c[:stream] codec` | Select encoder/decoder; use `copy` for stream copying |\n 52→| `-t duration` | Limit duration (as input: read limit; as output: write limit) |\n 53→| `-to position` | Stop at position (timestamp) |\n 54→| `-ss position` | Seek to position (before `-i`: fast input seek; after: output seek) |\n 55→| `-sseof position` | Seek relative to end of file |\n 56→| `-itsoffset offset` | Set input time offset |\n 57→| `-itsscale scale` | Rescale input timestamps |\n 58→| `-metadata key=value` | Set metadata key/value pair |\n 59→| `-disposition value` | Set stream disposition flags |\n 60→| `-target type` | Specify target type: `vcd`, `svcd`, `dvd`, `dv`, `dv50` |\n 61→| `-stream_loop n` | Loop input stream n times (-1 = infinite) |\n 62→| `-frames[:stream] n` | Stop after n frames |\n 63→| `-fs limit` | Set file size limit in bytes |\n 64→| `-timestamp date` | Set recording timestamp |\n 65→\n 66→### Video Options\n 67→\n 68→| Flag | Description |\n 69→|------|-------------|\n 70→| `-vn` | Disable video |\n 71→| `-vcodec codec` | Set video codec (alias for `-c:v`) |\n 72→| `-r fps` | Set frame rate |\n 73→| `-fpsmax fps` | Set maximum frame rate |\n 74→| `-s WxH` | Set frame size |\n 75→| `-aspect ratio` | Set display aspect ratio (e.g. `16:9`) |\n 76→| `-pix_fmt format` | Set pixel format |\n 77→| `-vf filtergraph` | Apply video filter graph (alias for `-filter:v`) |\n 78→| `-pass n` | Two-pass encoding pass (1 or 2) |\n 79→| `-passlogfile prefix` | Two-pass log file prefix |\n 80→| `-vframes n` | Set number of video frames to output |\n 81→| `-autorotate` | Auto-rotate based on metadata (default on) |\n 82→| `-display_rotation angle` | Set video rotation metadata |\n 83→| `-display_hflip` | Horizontal flip metadata |\n 84→| `-display_vflip` | Vertical flip metadata |\n 85→| `-force_key_frames expr` | Force keyframes at specified times/expression |\n 86→| `-copyinkf` | Copy non-key frames at the beginning during stream copy |\n 87→\n 88→### Audio Options\n 89→\n 90→| Flag | Description |\n 91→|------|-------------|\n 92→| `-an` | Disable audio |\n 93→| `-acodec codec` | Set audio codec (alias for `-c:a`) |\n 94→| `-ar freq` | Set audio sample rate (Hz) |\n 95→| `-ac channels` | Set number of audio channels |\n 96→| `-af filtergraph` | Apply audio filter graph (alias for `-filter:a`) |\n 97→| `-sample_fmt fmt` | Set audio sample format |\n 98→| `-channel_layout layout` | Set audio channel layout |\n 99→| `-aq q` | Set audio quality (codec-specific VBR) |\n 100→| `-aframes n` | Set number of audio frames to output |\n 101→\n 102→### Subtitle Options\n 103→\n 104→| Flag | Description |\n 105→|------|-------------|\n 106→| `-sn` | Disable subtitles |\n 107→| `-scodec codec` | Set subtitle codec (alias for `-c:s`) |\n 108→| `-fix_sub_duration` | Fix subtitle durations to avoid overlap |\n 109→\n 110→### Stream Selection\n 111→\n 112→| Flag | Description |\n 113→|------|-------------|\n 114→| `-map input:stream` | Manually select streams for output |\n 115→| `-dn` | Disable data streams |\n 116→\n 117→Stream specifiers: `v` (video), `V` (video, no images), `a` (audio), `s` (subtitle), `d` (data). Index with `:N` (e.g. `a:0` = first audio).\n 118→\n 119→### Hardware Acceleration\n 120→\n 121→| Flag | Description |\n 122→|------|-------------|\n 123→| `-hwaccel method` | HW accel method: `cuda`, `vaapi`, `qsv`, `vulkan`, `auto` |\n 124→| `-hwaccel_device device` | Select HW device |\n 125→| `-init_hw_device type=name` | Initialize HW device |\n 126→\n 127→---\n 128→\n 129→## ffprobe\n 130→\n 131→### Synopsis\n 132→\n 133→```\n 134→ffprobe [options] input_url\n 135→```\n 136→\n 137→### Main Options\n 138→\n 139→| Flag | Description |\n 140→|------|-------------|\n 141→| `-show_format` | Show container format info |\n 142→| `-show_streams` | Show per-stream info |\n 143→| `-show_packets` | Show per-packet info |\n 144→| `-show_frames` | Show per-frame info |\n 145→| `-show_chapters` | Show chapter info |\n 146→| `-show_programs` | Show program info |\n 147→| `-show_entries section=key1,key2` | Show only specific fields |\n 148→| `-show_error` | Show probe errors |\n 149→| `-select_streams specifier` | Filter to specific streams (e.g. `v:0`, `a`) |\n 150→| `-count_frames` | Count frames per stream |\n 151→| `-count_packets` | Count packets per stream |\n 152→| `-read_intervals intervals` | Analyze specific time ranges |\n 153→\n 154→### Output Formats\n 155→\n 156→Set with `-output_format` (or `-of`, `-print_format`):\n 157→\n 158→| Format | Description |\n 159→|--------|-------------|\n 160→| `default` | `[SECTION] key=value [/SECTION]` |\n 161→| `json` | JSON output (most useful for parsing) |\n 162→| `xml` | XML output |\n 163→| `csv` | Comma-separated values |\n 164→| `flat` | Flat `key=value` per line |\n 165→| `ini` | INI-style sections |\n 166→\n 167→### Display Options\n 168→\n 169→| Flag | Description |\n 170→|------|-------------|\n 171→| `-pretty` | Human-readable units and time formatting |\n 172→| `-unit` | Show value units |\n 173→| `-sexagesimal` | Format times as HH:MM:SS.us |\n 174→| `-hide_banner` | Suppress copyright/build info |\n 175→| `-o output_url` | Write output to file instead of stdout |\n 176→\n 177→---\n 178→\n 179→## Common Codecs\n 180→\n 181→### Video Encoders\n 182→\n 183→#### libx264 (H.264)\n 184→\n 185→| Option | Description |\n 186→|--------|-------------|\n 187→| `-preset` | Speed/quality: `ultrafast`, `superfast`, `veryfast`, `faster`, `fast`, `medium` (default), `slow`, `slower`, `veryslow` |\n 188→| `-crf` | Constant quality: 0 (lossless) to 51 (worst). 18-23 is typical |\n 189→| `-profile:v` | `baseline`, `main`, `high` |\n 190→| `-tune` | `film`, `animation`, `grain`, `stillimage`, `fastdecode`, `zerolatency` |\n 191→| `-b:v` | Target bitrate (e.g. `2M`) |\n 192→\n 193→#### libx265 (H.265/HEVC)\n 194→\n 195→| Option | Description |\n 196→|--------|-------------|\n 197→| `-preset` | Same presets as x264 |\n 198→| `-crf` | 0-51, default 28. Similar quality to x264 at lower bitrate |\n 199→| `-profile:v` | `main`, `main10`, `main12` |\n 200→| `-b:v` | Target bitrate |\n 201→\n 202→#### libvpx-vp9 (VP9)\n 203→\n 204→| Option | Description |\n 205→|--------|-------------|\n 206→| `-crf` | 0-63. 31 is a good default |\n 207→| `-b:v` | Target bitrate (set to `0` for pure CRF mode) |\n 208→| `-cpu-used` | Speed: 0 (slowest/best) to 8 (fastest) |\n 209→| `-deadline` | `best`, `good` (default), `realtime` |\n 210→| `-row-mt 1` | Enable row-based multithreading |\n 211→\n 212→#### libsvtav1 (SVT-AV1)\n 213→\n 214→| Option | Description |\n 215→|--------|-------------|\n 216→| `-crf` | 0-63. 30 is a good default |\n 217→| `-preset` | 0 (slowest/best) to 13 (fastest). 8 is a good default |\n 218→| `-b:v` | Target bitrate |\n 219→\n 220→#### libaom-av1 (AOM AV1)\n 221→\n 222→| Option | Description |\n 223→|--------|-------------|\n 224→| `-crf` | 0-63 |\n 225→| `-cpu-used` | 0 (best) to 8 (fastest) |\n 226→| `-b:v` | Target bitrate (set to `0` for pure CRF mode) |\n 227→| `-tiles` | Tile columns x rows for parallelism |\n 228→\n 229→### Audio Encoders\n 230→\n 231→#### aac (Native AAC)\n 232→\n 233→| Option | Description |\n 234→|--------|-------------|\n 235→| `-b:a` | Bitrate: `128k`, `192k`, `256k` |\n 236→| `-profile:a` | `aac_low` (default), `aac_he`, `aac_he_v2` |\n 237→\n 238→#### libmp3lame (MP3)\n 239→\n 240→| Option | Description |\n 241→|--------|-------------|\n 242→| `-b:a` | CBR bitrate: `128k`, `192k`, `320k` |\n 243→| `-q:a` | VBR quality: 0 (best) to 9 (worst). 2 is a good default |\n 244→\n 245→#### libopus (Opus)\n 246→\n 247→| Option | Description |\n 248→|--------|-------------|\n 249→| `-b:a` | Bitrate: `64k` to `256k`. 128k is a good default |\n 250→| `-vbr` | `on` (default), `off`, `constrained` |\n 251→| `-application` | `audio` (default), `voip`, `lowdelay` |\n 252→\n 253→#### libvorbis (Vorbis)\n 254→\n 255→| Option | Description |\n 256→|--------|-------------|\n 257→| `-q:a` | VBR quality: -1 to 10. 5 is a good default |\n 258→| `-b:a` | ABR bitrate |\n 259→\n 260→#### flac (FLAC)\n 261→\n 262→| Option | Description |\n 263→|--------|-------------|\n 264→| `-compression_level` | 0 (fast) to 12 (best). 5 is default |\n 265→\n 266→---\n 267→\n 268→## Common Container Formats\n 269→\n 270→| Format | Extensions | Notes |\n 271→|--------|-----------|-------|\n 272→| `mp4` | .mp4, .m4a, .m4v | Use `-movflags +faststart` for web streaming |\n 273→| `matroska` | .mkv | Supports virtually all codecs |\n 274→| `webm` | .webm | VP8/VP9/AV1 + Vorbis/Opus for web |\n 275→| `avi` | .avi | Legacy; limited codec support |\n 276→| `mpegts` | .ts | Broadcast transport stream |\n 277→| `ogg` | .ogg, .ogv | Vorbis/Opus/Theora container |\n 278→| `wav` | .wav | Uncompressed PCM audio |\n 279→| `flac` | .flac | Lossless audio |\n 280→| `mp3` | .mp3 | MPEG audio layer 3 |\n 281→| `hls` | .m3u8 | HTTP Live Streaming |\n 282→| `dash` | .mpd | DASH adaptive streaming |\n 283→| `gif` | .gif | Animated GIF |\n 284→| `image2` | various | Image sequence input/output |\n 285→| `concat` | text file | Concatenation demuxer (file list) |\n 286→| `null` | — | Discard output (benchmarking) |\n 287→\n 288→---\n 289→\n 290→## Common Video Filters (`-vf`)\n 291→\n 292→| Filter | Description | Example |\n 293→|--------|-------------|---------|\n 294→| `scale=W:H` | Resize video. Use `-1` or `-2` to auto-calculate | `scale=1280:720`, `scale=-2:480` |\n 295→| `crop=W:H:X:Y` | Crop to WxH starting at X,Y | `crop=640:480:100:50` |\n 296→| `pad=W:H:X:Y:color` | Pad video with borders | `pad=1920:1080:(ow-iw)/2:(oh-ih)/2:black` |\n 297→| `overlay=X:Y` | Composite second input over first | `overlay=10:10` |\n 298→| `transpose=N` | Rotate: 0=90ccw+vflip, 1=90cw, 2=90ccw, 3=90cw+vflip | `transpose=1` |\n 299→| `hflip` / `vflip` | Horizontal / vertical flip | `hflip` |\n 300→| `rotate=angle` | Rotate by arbitrary angle (radians) | `rotate=PI/4` |\n 301→| `fps=N` | Change frame rate | `fps=30` |\n 302→| `setpts=expr` | Modify presentation timestamps | `setpts=0.5*PTS` (2x speed) |\n 303→| `trim=start:end` | Extract time range | `trim=start=10:end=20` |\n 304→| `drawtext=opts` | Overlay text | `drawtext=text='Hello':fontsize=24:x=10:y=10` |\n 305→| `fade=t=type:st=S:d=D` | Fade in/out | `fade=t=in:st=0:d=2` |\n 306→| `eq=opts` | Adjust brightness/contrast/saturation | `eq=brightness=0.1:contrast=1.2` |\n 307→| `format=pix_fmt` | Convert pixel format | `format=yuv420p` |\n 308→| `concat=n:v:a` | Concatenate segments | `concat=n=2:v=1:a=1` |\n 309→| `split` / `select` | Duplicate / select frames | `select='eq(pict_type,I)'` |\n 310→| `deinterlace` / `yadif` | Remove interlacing | `yadif=1` |\n 311→| `boxblur=R` | Apply box blur | `boxblur=5:1` |\n 312→| `subtitles=file` | Burn in subtitles from file | `subtitles=subs.srt` |\n 313→| `palettegen` / `paletteuse` | Generate/apply palette for GIF | Used in two-pass GIF creation |\n 314→| `colorchannelmixer` | Mix color channels | `colorchannelmixer=.3:.4:.3:0:.3:.4:.3:0:.3:.4:.3` (grayscale) |\n 315→\n 316→## Common Audio Filters (`-af`)\n 317→\n 318→| Filter | Description | Example |\n 319→|--------|-------------|---------|\n 320→| `volume=V` | Adjust volume | `volume=1.5`, `volume=-3dB` |\n 321→| `loudnorm` | EBU R128 loudness normalization | `loudnorm=I=-16:TP=-1.5:LRA=11` |\n 322→| `atempo=T` | Change tempo (0.5-100.0) | `atempo=2.0` (2x speed) |\n 323→| `aresample=rate` | Resample audio | `aresample=44100` |\n 324→| `amerge` | Merge audio channels | `amerge=inputs=2` |\n 325→| `afade=t=type:st=S:d=D` | Audio fade in/out | `afade=t=in:st=0:d=3` |\n 326→| `highpass=f=freq` | High-pass filter | `highpass=f=200` |\n 327→| `lowpass=f=freq` | Low-pass filter | `lowpass=f=3000` |\n 328→| `equalizer=f:t:w:g` | Parametric EQ | `equalizer=f=1000:t=q:w=1:g=5` |\n 329→| `acompressor` | Dynamic range compression | `acompressor=threshold=-20dB:ratio=4` |\n 330→| `silenceremove` | Remove silence | `silenceremove=1:0:-50dB` |\n 331→| `silencedetect` | Detect silence | `silencedetect=n=-30dB:d=2` |\n 332→| `adelay=delays` | Delay audio channels | `adelay=1000\\|1000` (ms) |\n 333→| `aecho=id:ig:delays:decays` | Add echo effect | `aecho=0.8:0.88:60:0.4` |\n 334→| `pan=layout:gains` | Remix channels | `pan=mono\\|c0=0.5*c0+0.5*c1` |\n 335→\n 336→---\n 337→\n 338→## Common Recipes\n 339→\n 340→### Convert format\n 341→\n 342→```bash\n 343→ffmpeg -i input.mkv output.mp4\n 344→```\n 345→\n 346→### Transcode with CRF quality\n 347→\n 348→```bash\n 349→ffmpeg -i input.mp4 -c:v libx264 -crf 20 -c:a aac -b:a 192k output.mp4\n 350→```\n 351→\n 352→### Extract audio\n 353→\n 354→```bash\n 355→ffmpeg -i video.mp4 -vn -c:a copy audio.m4a\n 356→```\n 357→\n 358→### Extract video (no audio)\n 359→\n 360→```bash\n 361→ffmpeg -i input.mp4 -an -c:v copy output.mp4\n 362→```\n 363→\n 364→### Trim / cut\n 365→\n 366→```bash\n 367→ffmpeg -ss 00:01:30 -to 00:03:00 -i input.mp4 -c copy output.mp4\n 368→```\n 369→\n 370→### Resize video\n 371→\n 372→```bash\n 373→ffmpeg -i input.mp4 -vf \"scale=1280:720\" -c:a copy output.mp4\n 374→```\n 375→\n 376→### Compress video (lower quality)\n 377→\n 378→```bash\n 379→ffmpeg -i input.mp4 -c:v libx264 -crf 28 -preset fast -c:a aac -b:a 128k output.mp4\n 380→```\n 381→\n 382→### Two-pass encoding\n 383→\n 384→```bash\n 385→ffmpeg -i input.mp4 -c:v libx264 -b:v 2M -pass 1 -f null /dev/null\n 386→ffmpeg -i input.mp4 -c:v libx264 -b:v 2M -pass 2 output.mp4\n 387→```\n 388→\n 389→### Concatenate files (concat demuxer)\n 390→\n 391→```bash\n 392→# files.txt contains:\n 393→# file 'part1.mp4'\n 394→# file 'part2.mp4'\n 395→ffmpeg -f concat -safe 0 -i files.txt -c copy output.mp4\n 396→```\n 397→\n 398→### Add subtitles (burn-in)\n 399→\n 400→```bash\n 401→ffmpeg -i input.mp4 -vf \"subtitles=subs.srt\" output.mp4\n 402→```\n 403→\n 404→### Create GIF\n 405→\n 406→```bash\n 407→ffmpeg -i input.mp4 -vf \"fps=10,scale=320:-1:flags=lanczos,split[s0][s1];[s0]palettegen[p];[s1][p]paletteuse\" output.gif\n 408→```\n 409→\n 410→### Add watermark / overlay\n 411→\n 412→```bash\n 413→ffmpeg -i video.mp4 -i logo.png -filter_complex \"overlay=10:10\" output.mp4\n 414→```\n 415→\n 416→### Change speed (video + audio)\n 417→\n 418→```bash\n 419→ffmpeg -i input.mp4 -vf \"setpts=0.5*PTS\" -af \"atempo=2.0\" output.mp4\n 420→```\n 421→\n 422→### Extract frames as images\n 423→\n 424→```bash\n 425→ffmpeg -i input.mp4 -vf \"fps=1\" frame_%04d.png\n 426→```\n 427→\n 428→### Merge audio and video\n 429→\n 430→```bash\n 431→ffmpeg -i video.mp4 -i audio.m4a -c:v copy -c:a copy -shortest output.mp4\n 432→```\n 433→\n 434→### Normalize audio loudness\n 435→\n 436→```bash\n 437→ffmpeg -i input.mp4 -af \"loudnorm=I=-16:TP=-1.5:LRA=11\" -c:v copy output.mp4\n 438→```\n 439→\n 440→### Convert to web-optimized MP4\n 441→\n 442→```bash\n 443→ffmpeg -i input.mp4 -c:v libx264 -crf 23 -preset medium -c:a aac -b:a 128k -movflags +faststart output.mp4\n 444→```\n 445→\n 446→### Probe file info (JSON)\n 447→\n 448→```bash\n 449→ffprobe -v quiet -print_format json -show_format -show_streams input.mp4\n 450→```\n 451→\n 452→### Get duration only\n 453→\n 454→```bash\n 455→ffprobe -v quiet -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 input.mp4\n 456→```\n 457→\n 458→### Get resolution only\n 459→\n 460→```bash\n 461→ffprobe -v quiet -select_streams v:0 -show_entries stream=width,height -of csv=s=x:p=0 input.mp4\n 462→```\n 463→\n 464→---\n 465→\n 466→## Source\n 467→\n 468→- Website: https://ffmpeg.org/\n 469→- Documentation: https://www.ffmpeg.org/documentation.html\n 470→- CLI reference: https://www.ffmpeg.org/ffmpeg.html\n 471→- Filters reference: https://www.ffmpeg.org/ffmpeg-filters.html\n 472→- Codecs reference: https://www.ffmpeg.org/ffmpeg-codecs.html\n 473→- Formats reference: https://www.ffmpeg.org/ffmpeg-formats.html\n 474→- Wiki: https://trac.ffmpeg.org/wiki\n 475→","isError":false},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio","prompt":"Extract all information about mlx-audio: what it is, installation instructions, CLI usage, Python API usage, all available commands, supported models, supported features (TTS, STT, voice activity detection, etc.), configuration options, and any examples. Be comprehensive."},"toolUseId":"toolu_01BTviqNNX8vHWJ2Ci1yB8xS","output":"# MLX-Audio: Comprehensive Overview\n\n## What It Is\n\nMLX-Audio is a speech processing library built on Apple's MLX framework. It provides \"fast and efficient text-to-speech (TTS), speech-to-text (STT), and speech-to-speech (STS)\" capabilities optimized for Apple Silicon devices.\n\n## Installation\n\n**Via pip:**\n```bash\npip install mlx-audio\n```\n\n**CLI tools with uv (latest from PyPI):**\n```bash\nuv tool install --force mlx-audio --prerelease=allow\n```\n\n**Development setup:**\n```bash\ngit clone https://github.com/Blaizzy/mlx-audio.git\ncd mlx-audio\npip install -e \".[dev]\"\n```\n\n## Key Features\n\nThe library offers \"fast inference optimized for Apple Silicon (M series chips)\" along with multilingual support, voice customization, quantization options (3-bit through 8-bit), and an OpenAI-compatible REST API.\n\n## CLI Commands\n\n### TTS Generation\n```bash\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello, world!' --lang_code a\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --voice af_heart --speed 1.2 --lang_code a\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --play --lang_code a\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --output_path ./my_audio --lang_code a\n```\n\n### STT Transcription\n```bash\npython -m mlx_audio.stt.generate --model mlx-community/VibeVoice-ASR-bf16 --audio meeting.wav --output-path output --format json --max-tokens 8192 --verbose\npython -m mlx_audio.stt.generate --model mlx-community/parakeet-tdt-0.6b-v3 --audio speech.wav --output-path output --format json --verbose\n```\n\n## Python API Usage\n\n### TTS Example\n```python\nfrom mlx_audio.tts.utils import load_model\n\nmodel = load_model(\"mlx-community/Kokoro-82M-bf16\")\nfor result in model.generate(\"Hello from MLX-Audio!\", voice=\"af_heart\"):\n audio = result.audio # mx.array waveform\n```\n\n### STT Example\n```python\nfrom mlx_audio.stt.generate import generate_transcription\n\nresult = generate_transcription(\n model=\"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n audio=\"audio.wav\",\n)\nprint(result.text)\n```\n\n### Qwen3-ASR with Forced Alignment\n```python\nfrom mlx_audio.stt import load\n\nmodel = load(\"mlx-community/Qwen3-ASR-0.6B-8bit\")\nresult = model.generate(\"audio.wav\", language=\"English\")\nprint(result.text)\n\naligner = load(\"mlx-community/Qwen3-ForcedAligner-0.6B-8bit\")\nresult = aligner.generate(\"audio.wav\", text=\"I have a dream\", language=\"English\")\nfor item in result:\n print(f\"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}\")\n```\n\n### Voice Cloning (CSM)\n```bash\nmlx_audio.tts.generate --model mlx-community/csm-1b --text \"Hello from Sesame.\" --ref_audio ./reference_voice.wav --play\n```\n\n### VibeVoice-ASR with Context\n```python\nresult = model.generate(\n audio=\"technical_talk.wav\",\n context=\"MLX, Apple Silicon, PyTorch, Transformer\",\n max_tokens=8192,\n temperature=0.0,\n)\n```\n\n### Streaming Transcription\n```python\n# VibeVoice-ASR streaming\nfor text in model.stream_transcribe(audio=\"speech.wav\", max_tokens=4096):\n print(text, end=\"\", flush=True)\n\n# Parakeet streaming\nfor chunk in model.generate(\"long_audio.wav\", stream=True):\n print(chunk.text, end=\"\", flush=True)\n\n# Voxtral Realtime streaming\nfor chunk in model.generate(\"audio.wav\", stream=True):\n print(chunk, end=\"\", flush=True)\n```\n\n## Supported TTS Models\n\n| Model | Description | Languages |\n|-------|-------------|-----------|\n| **Kokoro** | Fast, high-quality multilingual | EN, JA, ZH, FR, ES, IT, PT, HI |\n| **Qwen3-TTS** | Alibaba's multilingual with voice design | ZH, EN, JA, KO, + more |\n| **CSM** | Voice cloning | EN |\n| **Dia** | Dialogue-focused | EN |\n| **OuteTTS** | Efficient model | EN |\n| **Spark** | SparkTTS | EN, ZH |\n| **Chatterbox** | Expressive multilingual | EN, ES, FR, DE, IT, PT, PL, TR, RU, NL, CS, AR, ZH, JA, HU, KO |\n| **Soprano** | High-quality | EN |\n\n**Kokoro Voice Examples:** af_heart, af_bella, af_nova, af_sky (American female); am_adam, am_echo (American male); bf_alice, bf_emma (British female); bm_daniel, bm_george (British male); jf_alpha, jm_kumo (Japanese); zf_xiaobei, zm_yunxi (Chinese)\n\n**Language Codes:** a (American English), b (British English), j (Japanese), z (Mandarin Chinese), e (Spanish), f (French)\n\n## Supported STT Models\n\n| Model | Description | Languages |\n|-------|-------------|-----------|\n| **Whisper** | OpenAI's robust model | 99+ languages |\n| **Qwen3-ASR** | Alibaba's multilingual | ZH, EN, JA, KO, + more |\n| **Qwen3-ForcedAligner** | Word-level alignment | ZH, EN, JA, KO, + more |\n| **Parakeet** | NVIDIA's accurate STT | EN (v2), 25 EU languages (v3) |\n| **Voxtral** | Mistral's speech model | Multiple |\n| **Voxtral Realtime** | 4B streaming STT | Multiple |\n| **VibeVoice-ASR** | Microsoft's 9B with diarization | Multiple |\n\n**Parakeet v3 Languages:** Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, Spanish, Swedish, Russian, Ukrainian\n\n## VAD/Speaker Diarization Models\n\n| Model | Description | \n|-------|-------------|\n| **Sortformer v1** | NVIDIA's end-to-end (up to 4 speakers) |\n| **Sortformer v2.1** | Streaming diarization with AOSC compression |\n\n## Speech-to-Speech (STS) Models\n\n| Model | Purpose |\n|-------|---------|\n| **SAM-Audio** | Text-guided source separation |\n| **Liquid2.5-Audio** | Speech/text-to-speech and STT |\n| **MossFormer2 SE** | Speech enhancement/noise removal |\n\n## Configuration Options\n\n### TTS Parameters\n- `--text`: Input text to synthesize\n- `--voice`: Voice preset selection\n- `--speed`: Speech speed adjustment (e.g., 1.2)\n- `--lang_code`: Language code\n- `--play`: Play audio immediately\n- `--output_path`: Directory for saving audio\n- `--ref_audio`: Reference audio for voice cloning\n\n### STT Parameters\n- `--audio`: Input audio file\n- `--model`: Model identifier\n- `--language`: Language specification\n- `--max_tokens`: Token limit (e.g., 8192)\n- `--temperature`: Generation temperature (0.0 for deterministic)\n- `--context`: Hotwords/metadata for context\n- `--output_path`: Output directory\n- `--format`: Output format (json, etc.)\n- `--stream`: Enable streaming mode\n- `--transcription_delay_ms`: Latency setting (Voxtral Realtime)\n- `--verbose`: Detailed logging\n\n## License\n\nThe project is licensed under MIT.","isError":false},{"role":"assistant","text":"Now let me fetch the README more thoroughly, particularly for details about the API server and quantization."},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio/blob/main/README.md","prompt":"Extract the complete raw content of the README. Focus on: 1) The REST API / OpenAI-compatible server setup and usage, 2) Quantization details, 3) Any additional CLI flags or Python API details not covered before. Get everything."},"toolUseId":"toolu_019imtWTE2Lakb4ZaZMFMcD9","output":"# MLX-Audio README: REST API, Quantization & Advanced Features\n\n## Web Interface & API Server\n\nThe project supports an OpenAI-compatible REST API server. Starting the server enables both a web interface with 3D audio visualization and programmatic access to all audio processing models.\n\n**Starting the Server:**\nLaunch the API server with a simple command to make all TTS, STT, and other audio models accessible via HTTP endpoints.\n\n**API Endpoints:**\nThe server implements OpenAI-compatible REST endpoints, allowing developers to integrate MLX-Audio into existing applications designed for OpenAI's API. This includes standard endpoints for speech synthesis, transcription, and other audio operations.\n\n## Quantization\n\nThe library supports multiple quantization levels: 3-bit, 4-bit, 6-bit, 8-bit, and additional precision options. Quantization reduces model size and improves inference speed on Apple Silicon while maintaining quality. Users can select pre-quantized model variants from Hugging Face (e.g., \"mlx-community/model-name-4bit\") or quantize models locally.\n\n## Additional CLI & Python API Features\n\n- **Speech speed control**: Adjust playback speed for TTS generation\n- **Output path specification**: Save generated audio to custom directories\n- **Immediate playback**: Use `--play` flag to hear results instantly\n- **Streaming support**: Several models support streaming transcription for real-time processing\n- **Context/hotwords**: VibeVoice-ASR and other models accept context strings to improve technical term recognition\n- **Model variants**: Choose between different precision levels and model sizes for performance/quality tradeoffs\n\nThe library emphasizes ease of use across command-line, Python API, and web interface paradigms for audio processing workflows.","isError":false},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio/tree/main/mlx_audio/tts","prompt":"List all files and directories in the TTS module. Look for any server.py, api.py, or similar files that indicate REST API functionality."},"toolUseId":"toolu_01VNu3dPPG68dArCP19kaXH9","output":"# TTS Module Contents\n\nBased on the GitHub page content, the TTS (Text-to-Speech) module directory structure includes:\n\n**Files:**\n- `__init__.py`\n- `audio_player.py`\n- `generate.py`\n- `utils.py`\n\n**Subdirectories:**\n- `models/`\n- `tests/`\n\n## API Functionality Assessment\n\nThere is **no evidence of REST API files** (such as `server.py` or `api.py`) within the TTS module itself. The module appears focused on core TTS functionality through generation, audio playback, and utility functions rather than exposing HTTP endpoints.\n\n**Note:** The parent `mlx_audio` directory does contain a `server.py` file, which may provide API capabilities at the package level, but this is separate from the TTS module's direct implementation.","isError":false},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio/blob/main/mlx_audio/server.py","prompt":"Extract the full content of this server.py file. I need to understand: 1) How to start the server (command), 2) All available API endpoints and their parameters, 3) Request/response formats, 4) Any configuration options. Get everything."},"toolUseId":"toolu_017WJRxSMxEMSVeMGNLcgC22","output":"# MLX Audio Server - Complete Documentation\n\n## Starting the Server\n\n**Command:**\n```bash\npython -m mlx_audio.server [OPTIONS]\n```\n\n**Key startup options:**\n- `--host`: Server host (default: localhost)\n- `--port`: Server port (default: 8000)\n- `--allowed-origins`: CORS allowed origins (default: *)\n- `--workers`: Number of workers (int or float, default: 2)\n- `--reload`: Enable auto-reload\n- `--start-ui`: Launch Studio UI alongside API\n- `--log-dir`: Directory for server logs (default: logs)\n\n## API Endpoints\n\n### Model Management\n\n**GET `/v1/models`**\n- Returns list of available models in OpenAI-compatible format\n- Response: `{\"object\": \"list\", \"data\": [{\"id\": model_name, \"object\": \"model\", ...}]}`\n\n**POST `/v1/models`**\n- Adds new model to API\n- Parameter: `model_name` (query string)\n- Response: `{\"status\": \"success\", \"message\": \"Model {name} added successfully\"}`\n\n**DELETE `/v1/models`**\n- Removes model from API\n- Parameter: `model_name` (query string)\n- Response: 204 No Content on success, 404 if not found\n\n### Speech Generation (TTS)\n\n**POST `/v1/audio/speech`**\n- Generates speech from text\n- Request body (SpeechRequest):\n```\n{\n \"model\": str (required),\n \"input\": str (required),\n \"voice\": str | None,\n \"speed\": float (default: 1.0),\n \"gender\": str (default: \"male\"),\n \"pitch\": float (default: 1.0),\n \"instruct\": str | None,\n \"lang_code\": str (default: \"a\"),\n \"ref_audio\": str | None,\n \"ref_text\": str | None,\n \"temperature\": float (default: 0.7),\n \"top_p\": float (default: 0.95),\n \"top_k\": int (default: 40),\n \"repetition_penalty\": float (default: 1.0),\n \"response_format\": str (default: \"mp3\"),\n \"stream\": bool (default: False),\n \"streaming_interval\": float (default: 2.0),\n \"max_tokens\": int (default: 1200),\n \"verbose\": bool (default: False)\n}\n```\n- Response: Audio stream with appropriate content-type\n\n### Speech-to-Text (STT)\n\n**POST `/v1/audio/transcriptions`**\n- Transcribes audio file\n- Parameters (form-data):\n - `file`: Audio file (required)\n - `model`: Model name (required)\n - `language`: Optional language code\n - `verbose`: Boolean (default: False)\n - `max_tokens`: Integer (default: 1024)\n - `chunk_duration`: Float (default: 30.0)\n - `frame_threshold`: Integer (default: 25)\n - `stream`: Boolean (default: False)\n - `context`: Optional context string\n - `prefill_step_size`: Integer (default: 2048)\n - `text`: Optional text parameter\n\n- Response: NDJSON stream with transcription chunks\n```\n{\"text\": str, \"accumulated\": str}\nor\n{\"text\": str, \"start\": float, \"end\": float, \"is_final\": bool, \"language\": str}\n```\n\n**WebSocket `/v1/audio/transcriptions/realtime`**\n- Real-time transcription via WebSocket\n- Initial message (JSON):\n```\n{\n \"model\": str,\n \"language\": str | None,\n \"sample_rate\": int (default: 16000),\n \"streaming\": bool (default: True)\n}\n```\n- Audio: Send as binary frames (int16 PCM)\n- Responses:\n - `{\"status\": \"ready\", \"message\": \"...\"}`\n - `{\"type\": \"delta\", \"delta\": str}`\n - `{\"type\": \"complete\", \"text\": str, \"segments\": null, ...}`\n\n### Audio Separation\n\n**POST `/v1/audio/separations`**\n- Separates audio using SAM Audio model\n- Parameters (form-data):\n - `file`: Audio file (required)\n - `model`: Model name (default: mlx-community/sam-audio-large-fp16)\n - `description`: Target description (default: \"speech\")\n - `method`: ODE method \"midpoint\" or \"euler\" (default: midpoint)\n - `steps`: ODE steps 2/4/8/16/32 (default: 16)\n\n- Response (SeparationResponse):\n```\n{\n \"target\": \"base64_encoded_wav\",\n \"residual\": \"base64_encoded_wav\",\n \"sample_rate\": int\n}\n```\n\n### Root\n\n**GET `/`**\n- Welcome message with link to UI (http://localhost:3000)\n\n## Configuration\n\n**Environment Variables:**\n- `MLX_AUDIO_NUM_WORKERS`: Override worker count\n- `MLX_AUDIO_ALLOWED_ORIGINS`: Comma-separated CORS origins\n\n**CORS Middleware:**\nConfigured automatically on startup; can be reconfigured via `setup_cors(app, origins_list)`\n\n## Architecture Notes\n\n- Uses FastAPI with async/await\n- Implements asyncio locks for thread-safe model management\n- VAD (Voice Activity Detection) enabled for real-time transcription\n- Supports streaming and non-streaming inference modes\n- Sanitizes NaN/Infinity values for JSON compatibility","isError":false},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio/tree/main/mlx_audio","prompt":"List all files and directories in the mlx_audio root package directory."},"toolUseId":"toolu_01TfedYLiRZqsWufnn4awqmw","output":"# Files and Directories in mlx_audio Root Package\n\nBased on the repository structure provided, the mlx_audio root package contains:\n\n**Directories:**\n- codec\n- sts\n- stt\n- tests\n- tts\n- ui\n- vad\n\n**Files:**\n- `__init__.py`\n- `audio_io.py`\n- `base.py`\n- `convert.py`\n- `dsp.py`\n- `server.py`\n- `utils.py`\n- `version.py`\n\nThis appears to be an audio processing library with modules for speech-to-text (stt), text-to-speech (tts), voice activity detection (vad), speech translation (sts), and codec handling.","isError":false},{"role":"assistant","text":"Now I have comprehensive information. Let me write the SKILL.md file."},{"role":"tool","toolName":"Write","toolInput":{"file_path":"/home/pastilhas/projects/officer.dev/data/skills/mlxaudio/SKILL.md","content":"---\nname: mlx.audio\ndescription: Generate speech from text and transcribe audio using mlx-audio. Use when the user wants text-to-speech synthesis, speech-to-text transcription, voice cloning, audio separation, or speech-to-speech processing on Apple Silicon.\n---\n\n# MLX-Audio\n\nA speech processing library built on Apple's MLX framework, providing TTS, STT, speech-to-speech (STS), and audio separation optimized for Apple Silicon.\n\n- **Repository:** https://github.com/Blaizzy/mlx-audio\n- **License:** MIT\n\n## CLI Tools\n\n### Text-to-Speech (TTS)\n\n```bash\nmlx_audio.tts.generate --model <model> --text '<text>' [options]\n```\n\n| Flag | Type | Default | Description |\n|------|------|---------|-------------|\n| `--model` | string | required | HuggingFace model ID |\n| `--text` | string | required | Text to synthesize |\n| `--voice` | string | — | Voice preset (model-specific) |\n| `--speed` | float | 1.0 | Speech speed multiplier |\n| `--lang_code` | string | `a` | Language code |\n| `--play` | flag | — | Play audio immediately |\n| `--output_path` | string | — | Directory to save audio |\n| `--ref_audio` | string | — | Reference audio for voice cloning (CSM) |\n\n#### Language Codes\n\n| Code | Language |\n|------|----------|\n| `a` | American English |\n| `b` | British English |\n| `j` | Japanese |\n| `z` | Mandarin Chinese |\n| `e` | Spanish |\n| `f` | French |\n\n#### Kokoro Voices\n\n| Voice | Description |\n|-------|-------------|\n| `af_heart`, `af_bella`, `af_nova`, `af_sky` | American female |\n| `am_adam`, `am_echo` | American male |\n| `bf_alice`, `bf_emma` | British female |\n| `bm_daniel`, `bm_george` | British male |\n| `jf_alpha`, `jm_kumo` | Japanese |\n| `zf_xiaobei`, `zm_yunxi` | Chinese |\n\n#### Examples\n\n```bash\n# Basic generation\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello, world!' --lang_code a\n\n# With voice and speed\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --voice af_heart --speed 1.2 --lang_code a\n\n# Play immediately\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --play --lang_code a\n\n# Voice cloning with CSM\nmlx_audio.tts.generate --model mlx-community/csm-1b --text \"Hello from Sesame.\" --ref_audio ./reference_voice.wav --play\n```\n\n### Speech-to-Text (STT)\n\n```bash\npython -m mlx_audio.stt.generate --model <model> --audio <file> [options]\n```\n\n| Flag | Type | Default | Description |\n|------|------|---------|-------------|\n| `--model` | string | required | HuggingFace model ID |\n| `--audio` | string | required | Input audio file |\n| `--language` | string | — | Language code |\n| `--max-tokens` | int | 1024 | Maximum output tokens |\n| `--temperature` | float | 0.0 | Sampling temperature |\n| `--context` | string | — | Hotwords/metadata for context |\n| `--output-path` | string | — | Output directory |\n| `--format` | string | — | Output format (e.g. `json`) |\n| `--stream` | flag | — | Enable streaming mode |\n| `--verbose` | flag | — | Detailed logging |\n\n#### Examples\n\n```bash\n# Basic transcription\npython -m mlx_audio.stt.generate --model mlx-community/whisper-large-v3-turbo-asr-fp16 --audio speech.wav --verbose\n\n# With context for technical terms\npython -m mlx_audio.stt.generate --model mlx-community/VibeVoice-ASR-bf16 --audio meeting.wav --context \"MLX, Apple Silicon, PyTorch\" --max-tokens 8192 --format json --verbose\n\n# Parakeet model\npython -m mlx_audio.stt.generate --model mlx-community/parakeet-tdt-0.6b-v3 --audio speech.wav --format json --verbose\n```\n\n---\n\n## Python API\n\n### TTS\n\n```python\nfrom mlx_audio.tts.utils import load_model\n\nmodel = load_model(\"mlx-community/Kokoro-82M-bf16\")\nfor result in model.generate(\"Hello from MLX-Audio!\", voice=\"af_heart\"):\n audio = result.audio # mx.array waveform\n```\n\n### STT\n\n```python\nfrom mlx_audio.stt.generate import generate_transcription\n\nresult = generate_transcription(\n model=\"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n audio=\"audio.wav\",\n)\nprint(result.text)\n```\n\n### STT with Streaming\n\n```python\nfrom mlx_audio.stt import load\n\n# VibeVoice-ASR streaming\nmodel = load(\"mlx-community/VibeVoice-ASR-bf16\")\nfor text in model.stream_transcribe(audio=\"speech.wav\", max_tokens=4096):\n print(text, end=\"\", flush=True)\n\n# Parakeet streaming\nmodel = load(\"mlx-community/parakeet-tdt-0.6b-v3\")\nfor chunk in model.generate(\"long_audio.wav\", stream=True):\n print(chunk.text, end=\"\", flush=True)\n```\n\n### Forced Alignment (Qwen3)\n\n```python\nfrom mlx_audio.stt import load\n\naligner = load(\"mlx-community/Qwen3-ForcedAligner-0.6B-8bit\")\nresult = aligner.generate(\"audio.wav\", text=\"I have a dream\", language=\"English\")\nfor item in result:\n print(f\"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}\")\n```\n\n---\n\n## REST API Server (OpenAI-compatible)\n\n### Starting the Server\n\n```bash\npython -m mlx_audio.server [OPTIONS]\n```\n\n| Flag | Type | Default | Description |\n|------|------|---------|-------------|\n| `--host` | string | `localhost` | Server host |\n| `--port` | int | `8000` | Server port |\n| `--allowed-origins` | string | `*` | CORS allowed origins |\n| `--workers` | int/float | `2` | Number of workers |\n| `--reload` | flag | — | Enable auto-reload |\n| `--start-ui` | flag | — | Launch Studio UI alongside API |\n| `--log-dir` | string | `logs` | Directory for server logs |\n\n### Endpoints\n\n#### GET /v1/models\n\nList available models.\n\n```bash\ncurl http://localhost:8000/v1/models\n```\n\n#### POST /v1/models?model_name=\\<name\\>\n\nAdd a model to the server.\n\n#### DELETE /v1/models?model_name=\\<name\\>\n\nRemove a model from the server.\n\n#### POST /v1/audio/speech\n\nGenerate speech from text.\n\n```bash\ncurl -X POST http://localhost:8000/v1/audio/speech \\\n -H \"Content-Type: application/json\" \\\n -d '{\n \"model\": \"mlx-community/Kokoro-82M-bf16\",\n \"input\": \"Hello, world!\",\n \"voice\": \"af_heart\",\n \"speed\": 1.0,\n \"lang_code\": \"a\",\n \"response_format\": \"mp3\"\n }' --output speech.mp3\n```\n\n| Parameter | Type | Default | Description |\n|-----------|------|---------|-------------|\n| `model` | string | required | Model ID |\n| `input` | string | required | Text to synthesize |\n| `voice` | string | — | Voice preset |\n| `speed` | float | 1.0 | Speech speed |\n| `lang_code` | string | `a` | Language code |\n| `ref_audio` | string | — | Reference audio path (voice cloning) |\n| `ref_text` | string | — | Reference transcript |\n| `response_format` | string | `mp3` | Output format |\n| `stream` | bool | false | Enable streaming |\n| `streaming_interval` | float | 2.0 | Streaming chunk interval |\n| `temperature` | float | 0.7 | Sampling temperature |\n| `top_p` | float | 0.95 | Nucleus sampling |\n| `top_k` | int | 40 | Top-k sampling |\n| `repetition_penalty` | float | 1.0 | Repetition penalty |\n| `max_tokens` | int | 1200 | Maximum tokens |\n| `gender` | string | `male` | Gender hint |\n| `pitch` | float | 1.0 | Pitch adjustment |\n| `instruct` | string | — | Instruction text |\n\n#### POST /v1/audio/transcriptions\n\nTranscribe an audio file (multipart/form-data).\n\n```bash\ncurl -X POST http://localhost:8000/v1/audio/transcriptions \\\n -F file=@audio.wav \\\n -F model=mlx-community/whisper-large-v3-turbo-asr-fp16\n```\n\n| Parameter | Type | Default | Description |\n|-----------|------|---------|-------------|\n| `file` | file | required | Audio file |\n| `model` | string | required | Model ID |\n| `language` | string | — | Language code |\n| `max_tokens` | int | 1024 | Maximum tokens |\n| `chunk_duration` | float | 30.0 | Chunk duration (seconds) |\n| `stream` | bool | false | Enable streaming |\n| `context` | string | — | Hotwords/context |\n| `text` | string | — | Reference text |\n| `verbose` | bool | false | Detailed output |\n\nResponse (NDJSON stream):\n\n```json\n{\"text\": \"chunk text\", \"accumulated\": \"full text so far\"}\n```\n\n#### WebSocket /v1/audio/transcriptions/realtime\n\nReal-time transcription via WebSocket. Send initial config as JSON, then stream int16 PCM audio as binary frames.\n\n```json\n{\n \"model\": \"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n \"sample_rate\": 16000,\n \"streaming\": true\n}\n```\n\n#### POST /v1/audio/separations\n\nSeparate audio sources (multipart/form-data).\n\n```bash\ncurl -X POST http://localhost:8000/v1/audio/separations \\\n -F file=@audio.wav \\\n -F model=mlx-community/sam-audio-large-fp16 \\\n -F description=\"speech\"\n```\n\n| Parameter | Type | Default | Description |\n|-----------|------|---------|-------------|\n| `file` | file | required | Audio file |\n| `model` | string | `mlx-community/sam-audio-large-fp16` | Model ID |\n| `description` | string | `speech` | Target description |\n| `method` | string | `midpoint` | ODE method (`midpoint` or `euler`) |\n| `steps` | int | 16 | ODE steps (2/4/8/16/32) |\n\nResponse:\n\n```json\n{\n \"target\": \"<base64 WAV>\",\n \"residual\": \"<base64 WAV>\",\n \"sample_rate\": 44100\n}\n```\n\n---\n\n## Supported Models\n\n### TTS Models\n\n| Model | Languages | Notes |\n|-------|-----------|-------|\n| Kokoro | EN, JA, ZH, FR, ES, IT, PT, HI | Fast, high-quality multilingual |\n| Qwen3-TTS | ZH, EN, JA, KO, + more | Voice design via instruction |\n| CSM | EN | Voice cloning with reference audio |\n| Dia | EN | Dialogue-focused |\n| OuteTTS | EN | Efficient |\n| Spark | EN, ZH | SparkTTS |\n| Chatterbox | EN, ES, FR, DE, IT, PT, PL, TR, RU, NL, CS, AR, ZH, JA, HU, KO | Expressive multilingual |\n| Soprano | EN | High-quality |\n\n### STT Models\n\n| Model | Languages | Notes |\n|-------|-----------|-------|\n| Whisper | 99+ languages | OpenAI's robust model |\n| Qwen3-ASR | ZH, EN, JA, KO, + more | Alibaba multilingual |\n| Qwen3-ForcedAligner | ZH, EN, JA, KO, + more | Word-level alignment |\n| Parakeet | EN (v2), 25 EU languages (v3) | NVIDIA, high accuracy |\n| Voxtral | Multiple | Mistral speech model |\n| Voxtral Realtime | Multiple | 4B streaming STT |\n| VibeVoice-ASR | Multiple | Microsoft 9B, supports diarization and context |\n\n### Other Models\n\n| Model | Type | Description |\n|-------|------|-------------|\n| Sortformer v1/v2.1 | VAD/Diarization | Speaker diarization (up to 4 speakers) |\n| SAM-Audio | Separation | Text-guided source separation |\n| Liquid2.5-Audio | STS | Speech/text-to-speech and STT |\n| MossFormer2 SE | Enhancement | Speech enhancement / noise removal |\n\nModels are available from `mlx-community` on HuggingFace with various quantization levels (3-bit through 8-bit and fp16/bf16).\n\n## Source\n\n- Repository: https://github.com/Blaizzy/mlx-audio\n- HuggingFace: https://huggingface.co/mlx-community\n"},"toolUseId":"toolu_0189M4gCkmdB8zkFsFYPXooY","output":"The file /home/pastilhas/projects/officer.dev/data/skills/mlxaudio/SKILL.md has been updated. Here's the result of running `cat -n` on a snippet of the edited file:\n 1→---\n 2→name: mlx.audio\n 3→description: Generate speech from text and transcribe audio using mlx-audio. Use when the user wants text-to-speech synthesis, speech-to-text transcription, voice cloning, audio separation, or speech-to-speech processing on Apple Silicon.\n 4→---\n 5→\n 6→# MLX-Audio\n 7→\n 8→A speech processing library built on Apple's MLX framework, providing TTS, STT, speech-to-speech (STS), and audio separation optimized for Apple Silicon.\n 9→\n 10→- **Repository:** https://github.com/Blaizzy/mlx-audio\n 11→- **License:** MIT\n 12→\n 13→## CLI Tools\n 14→\n 15→### Text-to-Speech (TTS)\n 16→\n 17→```bash\n 18→mlx_audio.tts.generate --model <model> --text '<text>' [options]\n 19→```\n 20→\n 21→| Flag | Type | Default | Description |\n 22→|------|------|---------|-------------|\n 23→| `--model` | string | required | HuggingFace model ID |\n 24→| `--text` | string | required | Text to synthesize |\n 25→| `--voice` | string | — | Voice preset (model-specific) |\n 26→| `--speed` | float | 1.0 | Speech speed multiplier |\n 27→| `--lang_code` | string | `a` | Language code |\n 28→| `--play` | flag | — | Play audio immediately |\n 29→| `--output_path` | string | — | Directory to save audio |\n 30→| `--ref_audio` | string | — | Reference audio for voice cloning (CSM) |\n 31→\n 32→#### Language Codes\n 33→\n 34→| Code | Language |\n 35→|------|----------|\n 36→| `a` | American English |\n 37→| `b` | British English |\n 38→| `j` | Japanese |\n 39→| `z` | Mandarin Chinese |\n 40→| `e` | Spanish |\n 41→| `f` | French |\n 42→\n 43→#### Kokoro Voices\n 44→\n 45→| Voice | Description |\n 46→|-------|-------------|\n 47→| `af_heart`, `af_bella`, `af_nova`, `af_sky` | American female |\n 48→| `am_adam`, `am_echo` | American male |\n 49→| `bf_alice`, `bf_emma` | British female |\n 50→| `bm_daniel`, `bm_george` | British male |\n 51→| `jf_alpha`, `jm_kumo` | Japanese |\n 52→| `zf_xiaobei`, `zm_yunxi` | Chinese |\n 53→\n 54→#### Examples\n 55→\n 56→```bash\n 57→# Basic generation\n 58→mlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello, world!' --lang_code a\n 59→\n 60→# With voice and speed\n 61→mlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --voice af_heart --speed 1.2 --lang_code a\n 62→\n 63→# Play immediately\n 64→mlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --play --lang_code a\n 65→\n 66→# Voice cloning with CSM\n 67→mlx_audio.tts.generate --model mlx-community/csm-1b --text \"Hello from Sesame.\" --ref_audio ./reference_voice.wav --play\n 68→```\n 69→\n 70→### Speech-to-Text (STT)\n 71→\n 72→```bash\n 73→python -m mlx_audio.stt.generate --model <model> --audio <file> [options]\n 74→```\n 75→\n 76→| Flag | Type | Default | Description |\n 77→|------|------|---------|-------------|\n 78→| `--model` | string | required | HuggingFace model ID |\n 79→| `--audio` | string | required | Input audio file |\n 80→| `--language` | string | — | Language code |\n 81→| `--max-tokens` | int | 1024 | Maximum output tokens |\n 82→| `--temperature` | float | 0.0 | Sampling temperature |\n 83→| `--context` | string | — | Hotwords/metadata for context |\n 84→| `--output-path` | string | — | Output directory |\n 85→| `--format` | string | — | Output format (e.g. `json`) |\n 86→| `--stream` | flag | — | Enable streaming mode |\n 87→| `--verbose` | flag | — | Detailed logging |\n 88→\n 89→#### Examples\n 90→\n 91→```bash\n 92→# Basic transcription\n 93→python -m mlx_audio.stt.generate --model mlx-community/whisper-large-v3-turbo-asr-fp16 --audio speech.wav --verbose\n 94→\n 95→# With context for technical terms\n 96→python -m mlx_audio.stt.generate --model mlx-community/VibeVoice-ASR-bf16 --audio meeting.wav --context \"MLX, Apple Silicon, PyTorch\" --max-tokens 8192 --format json --verbose\n 97→\n 98→# Parakeet model\n 99→python -m mlx_audio.stt.generate --model mlx-community/parakeet-tdt-0.6b-v3 --audio speech.wav --format json --verbose\n 100→```\n 101→\n 102→---\n 103→\n 104→## Python API\n 105→\n 106→### TTS\n 107→\n 108→```python\n 109→from mlx_audio.tts.utils import load_model\n 110→\n 111→model = load_model(\"mlx-community/Kokoro-82M-bf16\")\n 112→for result in model.generate(\"Hello from MLX-Audio!\", voice=\"af_heart\"):\n 113→ audio = result.audio # mx.array waveform\n 114→```\n 115→\n 116→### STT\n 117→\n 118→```python\n 119→from mlx_audio.stt.generate import generate_transcription\n 120→\n 121→result = generate_transcription(\n 122→ model=\"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n 123→ audio=\"audio.wav\",\n 124→)\n 125→print(result.text)\n 126→```\n 127→\n 128→### STT with Streaming\n 129→\n 130→```python\n 131→from mlx_audio.stt import load\n 132→\n 133→# VibeVoice-ASR streaming\n 134→model = load(\"mlx-community/VibeVoice-ASR-bf16\")\n 135→for text in model.stream_transcribe(audio=\"speech.wav\", max_tokens=4096):\n 136→ print(text, end=\"\", flush=True)\n 137→\n 138→# Parakeet streaming\n 139→model = load(\"mlx-community/parakeet-tdt-0.6b-v3\")\n 140→for chunk in model.generate(\"long_audio.wav\", stream=True):\n 141→ print(chunk.text, end=\"\", flush=True)\n 142→```\n 143→\n 144→### Forced Alignment (Qwen3)\n 145→\n 146→```python\n 147→from mlx_audio.stt import load\n 148→\n 149→aligner = load(\"mlx-community/Qwen3-ForcedAligner-0.6B-8bit\")\n 150→result = aligner.generate(\"audio.wav\", text=\"I have a dream\", language=\"English\")\n 151→for item in result:\n 152→ print(f\"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}\")\n 153→```\n 154→\n 155→---\n 156→\n 157→## REST API Server (OpenAI-compatible)\n 158→\n 159→### Starting the Server\n 160→\n 161→```bash\n 162→python -m mlx_audio.server [OPTIONS]\n 163→```\n 164→\n 165→| Flag | Type | Default | Description |\n 166→|------|------|---------|-------------|\n 167→| `--host` | string | `localhost` | Server host |\n 168→| `--port` | int | `8000` | Server port |\n 169→| `--allowed-origins` | string | `*` | CORS allowed origins |\n 170→| `--workers` | int/float | `2` | Number of workers |\n 171→| `--reload` | flag | — | Enable auto-reload |\n 172→| `--start-ui` | flag | — | Launch Studio UI alongside API |\n 173→| `--log-dir` | string | `logs` | Directory for server logs |\n 174→\n 175→### Endpoints\n 176→\n 177→#### GET /v1/models\n 178→\n 179→List available models.\n 180→\n 181→```bash\n 182→curl http://localhost:8000/v1/models\n 183→```\n 184→\n 185→#### POST /v1/models?model_name=\\<name\\>\n 186→\n 187→Add a model to the server.\n 188→\n 189→#### DELETE /v1/models?model_name=\\<name\\>\n 190→\n 191→Remove a model from the server.\n 192→\n 193→#### POST /v1/audio/speech\n 194→\n 195→Generate speech from text.\n 196→\n 197→```bash\n 198→curl -X POST http://localhost:8000/v1/audio/speech \\\n 199→ -H \"Content-Type: application/json\" \\\n 200→ -d '{\n 201→ \"model\": \"mlx-community/Kokoro-82M-bf16\",\n 202→ \"input\": \"Hello, world!\",\n 203→ \"voice\": \"af_heart\",\n 204→ \"speed\": 1.0,\n 205→ \"lang_code\": \"a\",\n 206→ \"response_format\": \"mp3\"\n 207→ }' --output speech.mp3\n 208→```\n 209→\n 210→| Parameter | Type | Default | Description |\n 211→|-----------|------|---------|-------------|\n 212→| `model` | string | required | Model ID |\n 213→| `input` | string | required | Text to synthesize |\n 214→| `voice` | string | — | Voice preset |\n 215→| `speed` | float | 1.0 | Speech speed |\n 216→| `lang_code` | string | `a` | Language code |\n 217→| `ref_audio` | string | — | Reference audio path (voice cloning) |\n 218→| `ref_text` | string | — | Reference transcript |\n 219→| `response_format` | string | `mp3` | Output format |\n 220→| `stream` | bool | false | Enable streaming |\n 221→| `streaming_interval` | float | 2.0 | Streaming chunk interval |\n 222→| `temperature` | float | 0.7 | Sampling temperature |\n 223→| `top_p` | float | 0.95 | Nucleus sampling |\n 224→| `top_k` | int | 40 | Top-k sampling |\n 225→| `repetition_penalty` | float | 1.0 | Repetition penalty |\n 226→| `max_tokens` | int | 1200 | Maximum tokens |\n 227→| `gender` | string | `male` | Gender hint |\n 228→| `pitch` | float | 1.0 | Pitch adjustment |\n 229→| `instruct` | string | — | Instruction text |\n 230→\n 231→#### POST /v1/audio/transcriptions\n 232→\n 233→Transcribe an audio file (multipart/form-data).\n 234→\n 235→```bash\n 236→curl -X POST http://localhost:8000/v1/audio/transcriptions \\\n 237→ -F file=@audio.wav \\\n 238→ -F model=mlx-community/whisper-large-v3-turbo-asr-fp16\n 239→```\n 240→\n 241→| Parameter | Type | Default | Description |\n 242→|-----------|------|---------|-------------|\n 243→| `file` | file | required | Audio file |\n 244→| `model` | string | required | Model ID |\n 245→| `language` | string | — | Language code |\n 246→| `max_tokens` | int | 1024 | Maximum tokens |\n 247→| `chunk_duration` | float | 30.0 | Chunk duration (seconds) |\n 248→| `stream` | bool | false | Enable streaming |\n 249→| `context` | string | — | Hotwords/context |\n 250→| `text` | string | — | Reference text |\n 251→| `verbose` | bool | false | Detailed output |\n 252→\n 253→Response (NDJSON stream):\n 254→\n 255→```json\n 256→{\"text\": \"chunk text\", \"accumulated\": \"full text so far\"}\n 257→```\n 258→\n 259→#### WebSocket /v1/audio/transcriptions/realtime\n 260→\n 261→Real-time transcription via WebSocket. Send initial config as JSON, then stream int16 PCM audio as binary frames.\n 262→\n 263→```json\n 264→{\n 265→ \"model\": \"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n 266→ \"sample_rate\": 16000,\n 267→ \"streaming\": true\n 268→}\n 269→```\n 270→\n 271→#### POST /v1/audio/separations\n 272→\n 273→Separate audio sources (multipart/form-data).\n 274→\n 275→```bash\n 276→curl -X POST http://localhost:8000/v1/audio/separations \\\n 277→ -F file=@audio.wav \\\n 278→ -F model=mlx-community/sam-audio-large-fp16 \\\n 279→ -F description=\"speech\"\n 280→```\n 281→\n 282→| Parameter | Type | Default | Description |\n 283→|-----------|------|---------|-------------|\n 284→| `file` | file | required | Audio file |\n 285→| `model` | string | `mlx-community/sam-audio-large-fp16` | Model ID |\n 286→| `description` | string | `speech` | Target description |\n 287→| `method` | string | `midpoint` | ODE method (`midpoint` or `euler`) |\n 288→| `steps` | int | 16 | ODE steps (2/4/8/16/32) |\n 289→\n 290→Response:\n 291→\n 292→```json\n 293→{\n 294→ \"target\": \"<base64 WAV>\",\n 295→ \"residual\": \"<base64 WAV>\",\n 296→ \"sample_rate\": 44100\n 297→}\n 298→```\n 299→\n 300→---\n 301→\n 302→## Supported Models\n 303→\n 304→### TTS Models\n 305→\n 306→| Model | Languages | Notes |\n 307→|-------|-----------|-------|\n 308→| Kokoro | EN, JA, ZH, FR, ES, IT, PT, HI | Fast, high-quality multilingual |\n 309→| Qwen3-TTS | ZH, EN, JA, KO, + more | Voice design via instruction |\n 310→| CSM | EN | Voice cloning with reference audio |\n 311→| Dia | EN | Dialogue-focused |\n 312→| OuteTTS | EN | Efficient |\n 313→| Spark | EN, ZH | SparkTTS |\n 314→| Chatterbox | EN, ES, FR, DE, IT, PT, PL, TR, RU, NL, CS, AR, ZH, JA, HU, KO | Expressive multilingual |\n 315→| Soprano | EN | High-quality |\n 316→\n 317→### STT Models\n 318→\n 319→| Model | Languages | Notes |\n 320→|-------|-----------|-------|\n 321→| Whisper | 99+ languages | OpenAI's robust model |\n 322→| Qwen3-ASR | ZH, EN, JA, KO, + more | Alibaba multilingual |\n 323→| Qwen3-ForcedAligner | ZH, EN, JA, KO, + more | Word-level alignment |\n 324→| Parakeet | EN (v2), 25 EU languages (v3) | NVIDIA, high accuracy |\n 325→| Voxtral | Multiple | Mistral speech model |\n 326→| Voxtral Realtime | Multiple | 4B streaming STT |\n 327→| VibeVoice-ASR | Multiple | Microsoft 9B, supports diarization and context |\n 328→\n 329→### Other Models\n 330→\n 331→| Model | Type | Description |\n 332→|-------|------|-------------|\n 333→| Sortformer v1/v2.1 | VAD/Diarization | Speaker diarization (up to 4 speakers) |\n 334→| SAM-Audio | Separation | Text-guided source separation |\n 335→| Liquid2.5-Audio | STS | Speech/text-to-speech and STT |\n 336→| MossFormer2 SE | Enhancement | Speech enhancement / noise removal |\n 337→\n 338→Models are available from `mlx-community` on HuggingFace with various quantization levels (3-bit through 8-bit and fp16/bf16).\n 339→\n 340→## Source\n 341→\n 342→- Repository: https://github.com/Blaizzy/mlx-audio\n 343→- HuggingFace: https://huggingface.co/mlx-community\n 344→","isError":false}]