[{"role":"user","text":"Help me create the content for this new skill file: /home/pastilhas/projects/officer.dev/data/skills/mlxaudio/SKILL.md\nfrom https://github.com/Blaizzy/mlx-audio"},{"role":"assistant","text":"\n\nLet me start by reading the current (empty) SKILL.md file and fetching information about the mlx-audio project."},{"role":"tool","toolName":"Read","toolInput":{"file_path":"/home/pastilhas/projects/officer.dev/data/skills/mlxaudio/SKILL.md"},"toolUseId":"toolu_01MXmpTWFrzx5Wm2fUTBEqcQ","output":" 1→---\n 2→name: mlx.audio\n 3→description: \n 4→---\n 5→","isError":false},{"role":"tool","toolName":"Read","toolInput":{"file_path":"/home/pastilhas/projects/officer.dev/data/skills/whisper.cpp/SKILL.md"},"toolUseId":"toolu_01RfdTasfwALeQhE26dkENTD","output":" 1→---\n 2→name: whisper.cpp\n 3→description: Transcribe audio files to text using whisper.cpp. Use when the user wants to transcribe audio, convert speech to text, or extract text from an audio/video file.\n 4→---\n 5→\n 6→# Whisper.cpp\n 7→\n 8→API reference for the whisper.cpp HTTP server running at `http://macmini:8178`.\n 9→\n 10→whisper.cpp is a C/C++ port of OpenAI's Whisper speech recognition model. The server accepts audio files via HTTP and returns transcriptions in various formats.\n 11→\n 12→## Server\n 13→\n 14→- **Base URL:** `http://macmini:8178`\n 15→- **No authentication required**\n 16→\n 17→## Endpoints\n 18→\n 19→### POST /inference\n 20→\n 21→Transcribes an audio file. Accepts `multipart/form-data`.\n 22→\n 23→#### Example\n 24→\n 25→```bash\n 26→curl -s http://macmini:8178/inference \\\n 27→ -F file=\"@/path/to/audio.mp3\" \\\n 28→ -F temperature=\"0.0\" \\\n 29→ -F temperature_inc=\"0.2\" \\\n 30→ -F response_format=\"json\"\n 31→```\n 32→\n 33→#### Parameters\n 34→\n 35→##### File (required)\n 36→\n 37→| Parameter | Type | Description |\n 38→|-----------|------|-------------|\n 39→| `file` | file | Audio file to transcribe. Accepts at least WAV and MP3. |\n 40→\n 41→##### Response Format\n 42→\n 43→| Parameter | Type | Default | Description |\n 44→|-----------|------|---------|-------------|\n 45→| `response_format` | string | `json` | Output format: `json`, `verbose_json` (or `vjson`), `text`, `srt`, `vtt` |\n 46→\n 47→##### Language\n 48→\n 49→| Parameter | Type | Default | Description |\n 50→|-----------|------|---------|-------------|\n 51→| `language` | string | `en` | Spoken language code (e.g. `en`, `pt`, `es`, `fr`). Use `auto` for auto-detection. |\n 52→| `detect_language` | bool | `false` | Exit after detecting the language (no transcription). |\n 53→| `translate` | bool | `false` | Translate from source language to English. |\n 54→\n 55→##### Decoding\n 56→\n 57→| Parameter | Type | Default | Description |\n 58→|-----------|------|---------|-------------|\n 59→| `temperature` | float | `0.0` | Sampling temperature. `0.0` is deterministic. |\n 60→| `temperature_inc` | float | `0.2` | Temperature increment on fallback attempts. |\n 61→| `best_of` | int | `2` | Number of candidate decodings to keep. |\n 62→| `beam_size` | int | `-1` | Beam search size. `-1` disables beam search. |\n 63→| `entropy_thold` | float | `2.40` | Entropy threshold — decoder fails and retries if exceeded. |\n 64→| `logprob_thold` | float | `-1.00` | Log probability threshold for decoder failure. |\n 65→| `no_fallback` | bool | `false` | Disable temperature fallback on decode failure. |\n 66→\n 67→##### Segmentation\n 68→\n 69→| Parameter | Type | Default | Description |\n 70→|-----------|------|---------|-------------|\n 71→| `max_len` | int | `0` | Maximum segment length in characters. `0` for unlimited. |\n 72→| `max_context` | int | `-1` | Maximum text context tokens to store. `-1` for unlimited. |\n 73→| `split_on_word` | bool | `false` | Split segments at word boundaries instead of token boundaries. |\n 74→| `no_timestamps` | bool | `false` | Suppress timestamps in output. |\n 75→| `word_thold` | float | `0.01` | Word timestamp probability threshold. |\n 76→\n 77→##### Audio Processing\n 78→\n 79→| Parameter | Type | Default | Description |\n 80→|-----------|------|---------|-------------|\n 81→| `offset_t` | int | `0` | Time offset in milliseconds — skip this much audio from the start. |\n 82→| `offset_n` | int | `0` | Segment index offset. |\n 83→| `duration` | int | `0` | Duration of audio to process in milliseconds. `0` for all. |\n 84→| `audio_ctx` | int | `0` | Audio context size. `0` for all. |\n 85→\n 86→##### Speaker Diarization\n 87→\n 88→| Parameter | Type | Default | Description |\n 89→|-----------|------|---------|-------------|\n 90→| `diarize` | bool | `false` | Enable speaker diarization (requires stereo audio). |\n 91→| `tinydiarize` | bool | `false` | Enable tinydiarize (requires a tdrz model). |\n 92→\n 93→##### Voice Activity Detection (VAD)\n 94→\n 95→| Parameter | Type | Default | Description |\n 96→|-----------|------|---------|-------------|\n 97→| `vad` | bool | `false` | Enable VAD preprocessing. |\n 98→| `vad_threshold` | float | `0.50` | Speech confidence threshold (0.0–1.0). |\n 99→| `vad_min_speech_duration_ms` | int | `250` | Minimum speech segment duration in ms. |\n 100→| `vad_min_silence_duration_ms` | int | `100` | Minimum silence duration to split segments. |\n 101→| `vad_max_speech_duration_s` | float | `FLT_MAX` | Auto-split segments longer than this (seconds). |\n 102→| `vad_speech_pad_ms` | int | `30` | Padding added around speech segments (ms). |\n 103→| `vad_samples_overlap` | float | `0.10` | Overlap between segments (seconds). |\n 104→\n 105→##### Other\n 106→\n 107→| Parameter | Type | Default | Description |\n 108→|-----------|------|---------|-------------|\n 109→| `prompt` | string | `\"\"` | Initial prompt to condition the model (e.g. for vocabulary hints). |\n 110→| `suppress_nst` | bool | `false` | Suppress non-speech tokens. |\n 111→| `no_context` | bool | `false` | Do not use previous audio context for subsequent segments. |\n 112→| `debug_mode` | bool | `false` | Enable debug output. |\n 113→\n 114→#### Response Formats\n 115→\n 116→##### `json` (default)\n 117→\n 118→Minimal JSON with just the transcribed text.\n 119→\n 120→```json\n 121→{\"text\": \"The transcribed content goes here.\"}\n 122→```\n 123→\n 124→##### `verbose_json` (or `vjson`)\n 125→\n 126→Extended JSON including task type, language, audio duration, per-segment timestamps, token-level timing, confidence scores, and language probability distribution.\n 127→\n 128→##### `text`\n 129→\n 130→Plain text transcription. Includes speaker labels if diarization is enabled.\n 131→\n 132→##### `srt`\n 133→\n 134→SubRip subtitle format with sequential numbering, `HH:MM:SS,mmm` timestamps, and text content.\n 135→\n 136→```\n 137→1\n 138→00:00:00,000 --> 00:00:03,500\n 139→The transcribed content goes here.\n 140→```\n 141→\n 142→##### `vtt`\n 143→\n 144→WebVTT subtitle format with `WEBVTT` header and `HH:MM:SS.mmm` timestamps.\n 145→\n 146→```\n 147→WEBVTT\n 148→\n 149→00:00:00.000 --> 00:00:03.500\n 150→The transcribed content goes here.\n 151→```\n 152→\n 153→### POST /load\n 154→\n 155→Loads a different model file on the server at runtime.\n 156→\n 157→```bash\n 158→curl -s http://macmini:8178/load \\\n 159→ -F model=\"/path/to/model.bin\"\n 160→```\n 161→\n 162→| Parameter | Type | Description |\n 163→|-----------|------|-------------|\n 164→| `model` | string | Path to the model file on the server. |\n 165→\n 166→## Supported Audio Formats\n 167→\n 168→The server accepts at least WAV (16-bit PCM) and MP3 files directly. If the server was started with `--convert`, it can use ffmpeg to handle additional formats (ogg, flac, m4a, etc.).\n 169→\n 170→## Source\n 171→\n 172→- Repository: https://github.com/ggml-org/whisper.cpp\n 173→- Server docs: https://github.com/ggml-org/whisper.cpp/blob/master/examples/server/README.md\n 174→","isError":false},{"role":"tool","toolName":"Read","toolInput":{"file_path":"/home/pastilhas/projects/officer.dev/data/skills/ffmpeg/SKILL.md"},"toolUseId":"toolu_0164FTK5nf9HT8m3GArB4jgy","output":" 1→---\n 2→name: ffmpeg\n 3→description: Process audio and video files using ffmpeg/ffprobe. Use when the user wants to convert, transcode, trim, merge, extract, resize, compress, or analyze multimedia files.\n 4→---\n 5→\n 6→# FFmpeg\n 7→\n 8→CLI reference for FFmpeg v8.x — a complete, cross-platform solution for recording, converting, and streaming audio and video.\n 9→\n 10→Official docs: https://www.ffmpeg.org/documentation.html\n 11→\n 12→## Tools\n 13→\n 14→| Tool | Purpose |\n 15→|------|---------|\n 16→| `ffmpeg` | Transcode, convert, filter, mux/demux multimedia |\n 17→| `ffprobe` | Analyze and inspect multimedia streams |\n 18→| `ffplay` | Play multimedia files (interactive) |\n 19→\n 20→---\n 21→\n 22→## ffmpeg\n 23→\n 24→### Synopsis\n 25→\n 26→```\n 27→ffmpeg [global_options] {[input_options] -i input_url} ... {[output_options] output_url} ...\n 28→```\n 29→\n 30→Options before `-i` apply to the input; options before the output URL apply to the output.\n 31→\n 32→### Global Options\n 33→\n 34→| Flag | Description |\n 35→|------|-------------|\n 36→| `-y` | Overwrite output files without asking |\n 37→| `-n` | Do not overwrite; exit if output exists |\n 38→| `-hide_banner` | Suppress copyright/build info banner |\n 39→| `-loglevel level` | Set log level: `quiet`, `error`, `warning`, `info` (default), `verbose`, `debug` |\n 40→| `-stats` | Print encoding progress/statistics |\n 41→| `-progress url` | Send machine-readable progress to url |\n 42→| `-report` | Dump full command line and log to a file |\n 43→| `-filter_threads n` | Number of threads for filter processing |\n 44→\n 45→### Input/Output Options\n 46→\n 47→| Flag | Description |\n 48→|------|-------------|\n 49→| `-i url` | Input file URL |\n 50→| `-f fmt` | Force input or output format |\n 51→| `-c[:stream] codec` | Select encoder/decoder; use `copy` for stream copying |\n 52→| `-t duration` | Limit duration (as input: read limit; as output: write limit) |\n 53→| `-to position` | Stop at position (timestamp) |\n 54→| `-ss position` | Seek to position (before `-i`: fast input seek; after: output seek) |\n 55→| `-sseof position` | Seek relative to end of file |\n 56→| `-itsoffset offset` | Set input time offset |\n 57→| `-itsscale scale` | Rescale input timestamps |\n 58→| `-metadata key=value` | Set metadata key/value pair |\n 59→| `-disposition value` | Set stream disposition flags |\n 60→| `-target type` | Specify target type: `vcd`, `svcd`, `dvd`, `dv`, `dv50` |\n 61→| `-stream_loop n` | Loop input stream n times (-1 = infinite) |\n 62→| `-frames[:stream] n` | Stop after n frames |\n 63→| `-fs limit` | Set file size limit in bytes |\n 64→| `-timestamp date` | Set recording timestamp |\n 65→\n 66→### Video Options\n 67→\n 68→| Flag | Description |\n 69→|------|-------------|\n 70→| `-vn` | Disable video |\n 71→| `-vcodec codec` | Set video codec (alias for `-c:v`) |\n 72→| `-r fps` | Set frame rate |\n 73→| `-fpsmax fps` | Set maximum frame rate |\n 74→| `-s WxH` | Set frame size |\n 75→| `-aspect ratio` | Set display aspect ratio (e.g. `16:9`) |\n 76→| `-pix_fmt format` | Set pixel format |\n 77→| `-vf filtergraph` | Apply video filter graph (alias for `-filter:v`) |\n 78→| `-pass n` | Two-pass encoding pass (1 or 2) |\n 79→| `-passlogfile prefix` | Two-pass log file prefix |\n 80→| `-vframes n` | Set number of video frames to output |\n 81→| `-autorotate` | Auto-rotate based on metadata (default on) |\n 82→| `-display_rotation angle` | Set video rotation metadata |\n 83→| `-display_hflip` | Horizontal flip metadata |\n 84→| `-display_vflip` | Vertical flip metadata |\n 85→| `-force_key_frames expr` | Force keyframes at specified times/expression |\n 86→| `-copyinkf` | Copy non-key frames at the beginning during stream copy |\n 87→\n 88→### Audio Options\n 89→\n 90→| Flag | Description |\n 91→|------|-------------|\n 92→| `-an` | Disable audio |\n 93→| `-acodec codec` | Set audio codec (alias for `-c:a`) |\n 94→| `-ar freq` | Set audio sample rate (Hz) |\n 95→| `-ac channels` | Set number of audio channels |\n 96→| `-af filtergraph` | Apply audio filter graph (alias for `-filter:a`) |\n 97→| `-sample_fmt fmt` | Set audio sample format |\n 98→| `-channel_layout layout` | Set audio channel layout |\n 99→| `-aq q` | Set audio quality (codec-specific VBR) |\n 100→| `-aframes n` | Set number of audio frames to output |\n 101→\n 102→### Subtitle Options\n 103→\n 104→| Flag | Description |\n 105→|------|-------------|\n 106→| `-sn` | Disable subtitles |\n 107→| `-scodec codec` | Set subtitle codec (alias for `-c:s`) |\n 108→| `-fix_sub_duration` | Fix subtitle durations to avoid overlap |\n 109→\n 110→### Stream Selection\n 111→\n 112→| Flag | Description |\n 113→|------|-------------|\n 114→| `-map input:stream` | Manually select streams for output |\n 115→| `-dn` | Disable data streams |\n 116→\n 117→Stream specifiers: `v` (video), `V` (video, no images), `a` (audio), `s` (subtitle), `d` (data). Index with `:N` (e.g. `a:0` = first audio).\n 118→\n 119→### Hardware Acceleration\n 120→\n 121→| Flag | Description |\n 122→|------|-------------|\n 123→| `-hwaccel method` | HW accel method: `cuda`, `vaapi`, `qsv`, `vulkan`, `auto` |\n 124→| `-hwaccel_device device` | Select HW device |\n 125→| `-init_hw_device type=name` | Initialize HW device |\n 126→\n 127→---\n 128→\n 129→## ffprobe\n 130→\n 131→### Synopsis\n 132→\n 133→```\n 134→ffprobe [options] input_url\n 135→```\n 136→\n 137→### Main Options\n 138→\n 139→| Flag | Description |\n 140→|------|-------------|\n 141→| `-show_format` | Show container format info |\n 142→| `-show_streams` | Show per-stream info |\n 143→| `-show_packets` | Show per-packet info |\n 144→| `-show_frames` | Show per-frame info |\n 145→| `-show_chapters` | Show chapter info |\n 146→| `-show_programs` | Show program info |\n 147→| `-show_entries section=key1,key2` | Show only specific fields |\n 148→| `-show_error` | Show probe errors |\n 149→| `-select_streams specifier` | Filter to specific streams (e.g. `v:0`, `a`) |\n 150→| `-count_frames` | Count frames per stream |\n 151→| `-count_packets` | Count packets per stream |\n 152→| `-read_intervals intervals` | Analyze specific time ranges |\n 153→\n 154→### Output Formats\n 155→\n 156→Set with `-output_format` (or `-of`, `-print_format`):\n 157→\n 158→| Format | Description |\n 159→|--------|-------------|\n 160→| `default` | `[SECTION] key=value [/SECTION]` |\n 161→| `json` | JSON output (most useful for parsing) |\n 162→| `xml` | XML output |\n 163→| `csv` | Comma-separated values |\n 164→| `flat` | Flat `key=value` per line |\n 165→| `ini` | INI-style sections |\n 166→\n 167→### Display Options\n 168→\n 169→| Flag | Description |\n 170→|------|-------------|\n 171→| `-pretty` | Human-readable units and time formatting |\n 172→| `-unit` | Show value units |\n 173→| `-sexagesimal` | Format times as HH:MM:SS.us |\n 174→| `-hide_banner` | Suppress copyright/build info |\n 175→| `-o output_url` | Write output to file instead of stdout |\n 176→\n 177→---\n 178→\n 179→## Common Codecs\n 180→\n 181→### Video Encoders\n 182→\n 183→#### libx264 (H.264)\n 184→\n 185→| Option | Description |\n 186→|--------|-------------|\n 187→| `-preset` | Speed/quality: `ultrafast`, `superfast`, `veryfast`, `faster`, `fast`, `medium` (default), `slow`, `slower`, `veryslow` |\n 188→| `-crf` | Constant quality: 0 (lossless) to 51 (worst). 18-23 is typical |\n 189→| `-profile:v` | `baseline`, `main`, `high` |\n 190→| `-tune` | `film`, `animation`, `grain`, `stillimage`, `fastdecode`, `zerolatency` |\n 191→| `-b:v` | Target bitrate (e.g. `2M`) |\n 192→\n 193→#### libx265 (H.265/HEVC)\n 194→\n 195→| Option | Description |\n 196→|--------|-------------|\n 197→| `-preset` | Same presets as x264 |\n 198→| `-crf` | 0-51, default 28. Similar quality to x264 at lower bitrate |\n 199→| `-profile:v` | `main`, `main10`, `main12` |\n 200→| `-b:v` | Target bitrate |\n 201→\n 202→#### libvpx-vp9 (VP9)\n 203→\n 204→| Option | Description |\n 205→|--------|-------------|\n 206→| `-crf` | 0-63. 31 is a good default |\n 207→| `-b:v` | Target bitrate (set to `0` for pure CRF mode) |\n 208→| `-cpu-used` | Speed: 0 (slowest/best) to 8 (fastest) |\n 209→| `-deadline` | `best`, `good` (default), `realtime` |\n 210→| `-row-mt 1` | Enable row-based multithreading |\n 211→\n 212→#### libsvtav1 (SVT-AV1)\n 213→\n 214→| Option | Description |\n 215→|--------|-------------|\n 216→| `-crf` | 0-63. 30 is a good default |\n 217→| `-preset` | 0 (slowest/best) to 13 (fastest). 8 is a good default |\n 218→| `-b:v` | Target bitrate |\n 219→\n 220→#### libaom-av1 (AOM AV1)\n 221→\n 222→| Option | Description |\n 223→|--------|-------------|\n 224→| `-crf` | 0-63 |\n 225→| `-cpu-used` | 0 (best) to 8 (fastest) |\n 226→| `-b:v` | Target bitrate (set to `0` for pure CRF mode) |\n 227→| `-tiles` | Tile columns x rows for parallelism |\n 228→\n 229→### Audio Encoders\n 230→\n 231→#### aac (Native AAC)\n 232→\n 233→| Option | Description |\n 234→|--------|-------------|\n 235→| `-b:a` | Bitrate: `128k`, `192k`, `256k` |\n 236→| `-profile:a` | `aac_low` (default), `aac_he`, `aac_he_v2` |\n 237→\n 238→#### libmp3lame (MP3)\n 239→\n 240→| Option | Description |\n 241→|--------|-------------|\n 242→| `-b:a` | CBR bitrate: `128k`, `192k`, `320k` |\n 243→| `-q:a` | VBR quality: 0 (best) to 9 (worst). 2 is a good default |\n 244→\n 245→#### libopus (Opus)\n 246→\n 247→| Option | Description |\n 248→|--------|-------------|\n 249→| `-b:a` | Bitrate: `64k` to `256k`. 128k is a good default |\n 250→| `-vbr` | `on` (default), `off`, `constrained` |\n 251→| `-application` | `audio` (default), `voip`, `lowdelay` |\n 252→\n 253→#### libvorbis (Vorbis)\n 254→\n 255→| Option | Description |\n 256→|--------|-------------|\n 257→| `-q:a` | VBR quality: -1 to 10. 5 is a good default |\n 258→| `-b:a` | ABR bitrate |\n 259→\n 260→#### flac (FLAC)\n 261→\n 262→| Option | Description |\n 263→|--------|-------------|\n 264→| `-compression_level` | 0 (fast) to 12 (best). 5 is default |\n 265→\n 266→---\n 267→\n 268→## Common Container Formats\n 269→\n 270→| Format | Extensions | Notes |\n 271→|--------|-----------|-------|\n 272→| `mp4` | .mp4, .m4a, .m4v | Use `-movflags +faststart` for web streaming |\n 273→| `matroska` | .mkv | Supports virtually all codecs |\n 274→| `webm` | .webm | VP8/VP9/AV1 + Vorbis/Opus for web |\n 275→| `avi` | .avi | Legacy; limited codec support |\n 276→| `mpegts` | .ts | Broadcast transport stream |\n 277→| `ogg` | .ogg, .ogv | Vorbis/Opus/Theora container |\n 278→| `wav` | .wav | Uncompressed PCM audio |\n 279→| `flac` | .flac | Lossless audio |\n 280→| `mp3` | .mp3 | MPEG audio layer 3 |\n 281→| `hls` | .m3u8 | HTTP Live Streaming |\n 282→| `dash` | .mpd | DASH adaptive streaming |\n 283→| `gif` | .gif | Animated GIF |\n 284→| `image2` | various | Image sequence input/output |\n 285→| `concat` | text file | Concatenation demuxer (file list) |\n 286→| `null` | — | Discard output (benchmarking) |\n 287→\n 288→---\n 289→\n 290→## Common Video Filters (`-vf`)\n 291→\n 292→| Filter | Description | Example |\n 293→|--------|-------------|---------|\n 294→| `scale=W:H` | Resize video. Use `-1` or `-2` to auto-calculate | `scale=1280:720`, `scale=-2:480` |\n 295→| `crop=W:H:X:Y` | Crop to WxH starting at X,Y | `crop=640:480:100:50` |\n 296→| `pad=W:H:X:Y:color` | Pad video with borders | `pad=1920:1080:(ow-iw)/2:(oh-ih)/2:black` |\n 297→| `overlay=X:Y` | Composite second input over first | `overlay=10:10` |\n 298→| `transpose=N` | Rotate: 0=90ccw+vflip, 1=90cw, 2=90ccw, 3=90cw+vflip | `transpose=1` |\n 299→| `hflip` / `vflip` | Horizontal / vertical flip | `hflip` |\n 300→| `rotate=angle` | Rotate by arbitrary angle (radians) | `rotate=PI/4` |\n 301→| `fps=N` | Change frame rate | `fps=30` |\n 302→| `setpts=expr` | Modify presentation timestamps | `setpts=0.5*PTS` (2x speed) |\n 303→| `trim=start:end` | Extract time range | `trim=start=10:end=20` |\n 304→| `drawtext=opts` | Overlay text | `drawtext=text='Hello':fontsize=24:x=10:y=10` |\n 305→| `fade=t=type:st=S:d=D` | Fade in/out | `fade=t=in:st=0:d=2` |\n 306→| `eq=opts` | Adjust brightness/contrast/saturation | `eq=brightness=0.1:contrast=1.2` |\n 307→| `format=pix_fmt` | Convert pixel format | `format=yuv420p` |\n 308→| `concat=n:v:a` | Concatenate segments | `concat=n=2:v=1:a=1` |\n 309→| `split` / `select` | Duplicate / select frames | `select='eq(pict_type,I)'` |\n 310→| `deinterlace` / `yadif` | Remove interlacing | `yadif=1` |\n 311→| `boxblur=R` | Apply box blur | `boxblur=5:1` |\n 312→| `subtitles=file` | Burn in subtitles from file | `subtitles=subs.srt` |\n 313→| `palettegen` / `paletteuse` | Generate/apply palette for GIF | Used in two-pass GIF creation |\n 314→| `colorchannelmixer` | Mix color channels | `colorchannelmixer=.3:.4:.3:0:.3:.4:.3:0:.3:.4:.3` (grayscale) |\n 315→\n 316→## Common Audio Filters (`-af`)\n 317→\n 318→| Filter | Description | Example |\n 319→|--------|-------------|---------|\n 320→| `volume=V` | Adjust volume | `volume=1.5`, `volume=-3dB` |\n 321→| `loudnorm` | EBU R128 loudness normalization | `loudnorm=I=-16:TP=-1.5:LRA=11` |\n 322→| `atempo=T` | Change tempo (0.5-100.0) | `atempo=2.0` (2x speed) |\n 323→| `aresample=rate` | Resample audio | `aresample=44100` |\n 324→| `amerge` | Merge audio channels | `amerge=inputs=2` |\n 325→| `afade=t=type:st=S:d=D` | Audio fade in/out | `afade=t=in:st=0:d=3` |\n 326→| `highpass=f=freq` | High-pass filter | `highpass=f=200` |\n 327→| `lowpass=f=freq` | Low-pass filter | `lowpass=f=3000` |\n 328→| `equalizer=f:t:w:g` | Parametric EQ | `equalizer=f=1000:t=q:w=1:g=5` |\n 329→| `acompressor` | Dynamic range compression | `acompressor=threshold=-20dB:ratio=4` |\n 330→| `silenceremove` | Remove silence | `silenceremove=1:0:-50dB` |\n 331→| `silencedetect` | Detect silence | `silencedetect=n=-30dB:d=2` |\n 332→| `adelay=delays` | Delay audio channels | `adelay=1000\\|1000` (ms) |\n 333→| `aecho=id:ig:delays:decays` | Add echo effect | `aecho=0.8:0.88:60:0.4` |\n 334→| `pan=layout:gains` | Remix channels | `pan=mono\\|c0=0.5*c0+0.5*c1` |\n 335→\n 336→---\n 337→\n 338→## Common Recipes\n 339→\n 340→### Convert format\n 341→\n 342→```bash\n 343→ffmpeg -i input.mkv output.mp4\n 344→```\n 345→\n 346→### Transcode with CRF quality\n 347→\n 348→```bash\n 349→ffmpeg -i input.mp4 -c:v libx264 -crf 20 -c:a aac -b:a 192k output.mp4\n 350→```\n 351→\n 352→### Extract audio\n 353→\n 354→```bash\n 355→ffmpeg -i video.mp4 -vn -c:a copy audio.m4a\n 356→```\n 357→\n 358→### Extract video (no audio)\n 359→\n 360→```bash\n 361→ffmpeg -i input.mp4 -an -c:v copy output.mp4\n 362→```\n 363→\n 364→### Trim / cut\n 365→\n 366→```bash\n 367→ffmpeg -ss 00:01:30 -to 00:03:00 -i input.mp4 -c copy output.mp4\n 368→```\n 369→\n 370→### Resize video\n 371→\n 372→```bash\n 373→ffmpeg -i input.mp4 -vf \"scale=1280:720\" -c:a copy output.mp4\n 374→```\n 375→\n 376→### Compress video (lower quality)\n 377→\n 378→```bash\n 379→ffmpeg -i input.mp4 -c:v libx264 -crf 28 -preset fast -c:a aac -b:a 128k output.mp4\n 380→```\n 381→\n 382→### Two-pass encoding\n 383→\n 384→```bash\n 385→ffmpeg -i input.mp4 -c:v libx264 -b:v 2M -pass 1 -f null /dev/null\n 386→ffmpeg -i input.mp4 -c:v libx264 -b:v 2M -pass 2 output.mp4\n 387→```\n 388→\n 389→### Concatenate files (concat demuxer)\n 390→\n 391→```bash\n 392→# files.txt contains:\n 393→# file 'part1.mp4'\n 394→# file 'part2.mp4'\n 395→ffmpeg -f concat -safe 0 -i files.txt -c copy output.mp4\n 396→```\n 397→\n 398→### Add subtitles (burn-in)\n 399→\n 400→```bash\n 401→ffmpeg -i input.mp4 -vf \"subtitles=subs.srt\" output.mp4\n 402→```\n 403→\n 404→### Create GIF\n 405→\n 406→```bash\n 407→ffmpeg -i input.mp4 -vf \"fps=10,scale=320:-1:flags=lanczos,split[s0][s1];[s0]palettegen[p];[s1][p]paletteuse\" output.gif\n 408→```\n 409→\n 410→### Add watermark / overlay\n 411→\n 412→```bash\n 413→ffmpeg -i video.mp4 -i logo.png -filter_complex \"overlay=10:10\" output.mp4\n 414→```\n 415→\n 416→### Change speed (video + audio)\n 417→\n 418→```bash\n 419→ffmpeg -i input.mp4 -vf \"setpts=0.5*PTS\" -af \"atempo=2.0\" output.mp4\n 420→```\n 421→\n 422→### Extract frames as images\n 423→\n 424→```bash\n 425→ffmpeg -i input.mp4 -vf \"fps=1\" frame_%04d.png\n 426→```\n 427→\n 428→### Merge audio and video\n 429→\n 430→```bash\n 431→ffmpeg -i video.mp4 -i audio.m4a -c:v copy -c:a copy -shortest output.mp4\n 432→```\n 433→\n 434→### Normalize audio loudness\n 435→\n 436→```bash\n 437→ffmpeg -i input.mp4 -af \"loudnorm=I=-16:TP=-1.5:LRA=11\" -c:v copy output.mp4\n 438→```\n 439→\n 440→### Convert to web-optimized MP4\n 441→\n 442→```bash\n 443→ffmpeg -i input.mp4 -c:v libx264 -crf 23 -preset medium -c:a aac -b:a 128k -movflags +faststart output.mp4\n 444→```\n 445→\n 446→### Probe file info (JSON)\n 447→\n 448→```bash\n 449→ffprobe -v quiet -print_format json -show_format -show_streams input.mp4\n 450→```\n 451→\n 452→### Get duration only\n 453→\n 454→```bash\n 455→ffprobe -v quiet -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 input.mp4\n 456→```\n 457→\n 458→### Get resolution only\n 459→\n 460→```bash\n 461→ffprobe -v quiet -select_streams v:0 -show_entries stream=width,height -of csv=s=x:p=0 input.mp4\n 462→```\n 463→\n 464→---\n 465→\n 466→## Source\n 467→\n 468→- Website: https://ffmpeg.org/\n 469→- Documentation: https://www.ffmpeg.org/documentation.html\n 470→- CLI reference: https://www.ffmpeg.org/ffmpeg.html\n 471→- Filters reference: https://www.ffmpeg.org/ffmpeg-filters.html\n 472→- Codecs reference: https://www.ffmpeg.org/ffmpeg-codecs.html\n 473→- Formats reference: https://www.ffmpeg.org/ffmpeg-formats.html\n 474→- Wiki: https://trac.ffmpeg.org/wiki\n 475→","isError":false},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio","prompt":"Extract all information about mlx-audio: what it is, installation instructions, CLI usage, Python API usage, all available commands, supported models, supported features (TTS, STT, voice activity detection, etc.), configuration options, and any examples. Be comprehensive."},"toolUseId":"toolu_01BTviqNNX8vHWJ2Ci1yB8xS","output":"# MLX-Audio: Comprehensive Overview\n\n## What It Is\n\nMLX-Audio is a speech processing library built on Apple's MLX framework. It provides \"fast and efficient text-to-speech (TTS), speech-to-text (STT), and speech-to-speech (STS)\" capabilities optimized for Apple Silicon devices.\n\n## Installation\n\n**Via pip:**\n```bash\npip install mlx-audio\n```\n\n**CLI tools with uv (latest from PyPI):**\n```bash\nuv tool install --force mlx-audio --prerelease=allow\n```\n\n**Development setup:**\n```bash\ngit clone https://github.com/Blaizzy/mlx-audio.git\ncd mlx-audio\npip install -e \".[dev]\"\n```\n\n## Key Features\n\nThe library offers \"fast inference optimized for Apple Silicon (M series chips)\" along with multilingual support, voice customization, quantization options (3-bit through 8-bit), and an OpenAI-compatible REST API.\n\n## CLI Commands\n\n### TTS Generation\n```bash\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello, world!' --lang_code a\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --voice af_heart --speed 1.2 --lang_code a\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --play --lang_code a\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --output_path ./my_audio --lang_code a\n```\n\n### STT Transcription\n```bash\npython -m mlx_audio.stt.generate --model mlx-community/VibeVoice-ASR-bf16 --audio meeting.wav --output-path output --format json --max-tokens 8192 --verbose\npython -m mlx_audio.stt.generate --model mlx-community/parakeet-tdt-0.6b-v3 --audio speech.wav --output-path output --format json --verbose\n```\n\n## Python API Usage\n\n### TTS Example\n```python\nfrom mlx_audio.tts.utils import load_model\n\nmodel = load_model(\"mlx-community/Kokoro-82M-bf16\")\nfor result in model.generate(\"Hello from MLX-Audio!\", voice=\"af_heart\"):\n audio = result.audio # mx.array waveform\n```\n\n### STT Example\n```python\nfrom mlx_audio.stt.generate import generate_transcription\n\nresult = generate_transcription(\n model=\"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n audio=\"audio.wav\",\n)\nprint(result.text)\n```\n\n### Qwen3-ASR with Forced Alignment\n```python\nfrom mlx_audio.stt import load\n\nmodel = load(\"mlx-community/Qwen3-ASR-0.6B-8bit\")\nresult = model.generate(\"audio.wav\", language=\"English\")\nprint(result.text)\n\naligner = load(\"mlx-community/Qwen3-ForcedAligner-0.6B-8bit\")\nresult = aligner.generate(\"audio.wav\", text=\"I have a dream\", language=\"English\")\nfor item in result:\n print(f\"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}\")\n```\n\n### Voice Cloning (CSM)\n```bash\nmlx_audio.tts.generate --model mlx-community/csm-1b --text \"Hello from Sesame.\" --ref_audio ./reference_voice.wav --play\n```\n\n### VibeVoice-ASR with Context\n```python\nresult = model.generate(\n audio=\"technical_talk.wav\",\n context=\"MLX, Apple Silicon, PyTorch, Transformer\",\n max_tokens=8192,\n temperature=0.0,\n)\n```\n\n### Streaming Transcription\n```python\n# VibeVoice-ASR streaming\nfor text in model.stream_transcribe(audio=\"speech.wav\", max_tokens=4096):\n print(text, end=\"\", flush=True)\n\n# Parakeet streaming\nfor chunk in model.generate(\"long_audio.wav\", stream=True):\n print(chunk.text, end=\"\", flush=True)\n\n# Voxtral Realtime streaming\nfor chunk in model.generate(\"audio.wav\", stream=True):\n print(chunk, end=\"\", flush=True)\n```\n\n## Supported TTS Models\n\n| Model | Description | Languages |\n|-------|-------------|-----------|\n| **Kokoro** | Fast, high-quality multilingual | EN, JA, ZH, FR, ES, IT, PT, HI |\n| **Qwen3-TTS** | Alibaba's multilingual with voice design | ZH, EN, JA, KO, + more |\n| **CSM** | Voice cloning | EN |\n| **Dia** | Dialogue-focused | EN |\n| **OuteTTS** | Efficient model | EN |\n| **Spark** | SparkTTS | EN, ZH |\n| **Chatterbox** | Expressive multilingual | EN, ES, FR, DE, IT, PT, PL, TR, RU, NL, CS, AR, ZH, JA, HU, KO |\n| **Soprano** | High-quality | EN |\n\n**Kokoro Voice Examples:** af_heart, af_bella, af_nova, af_sky (American female); am_adam, am_echo (American male); bf_alice, bf_emma (British female); bm_daniel, bm_george (British male); jf_alpha, jm_kumo (Japanese); zf_xiaobei, zm_yunxi (Chinese)\n\n**Language Codes:** a (American English), b (British English), j (Japanese), z (Mandarin Chinese), e (Spanish), f (French)\n\n## Supported STT Models\n\n| Model | Description | Languages |\n|-------|-------------|-----------|\n| **Whisper** | OpenAI's robust model | 99+ languages |\n| **Qwen3-ASR** | Alibaba's multilingual | ZH, EN, JA, KO, + more |\n| **Qwen3-ForcedAligner** | Word-level alignment | ZH, EN, JA, KO, + more |\n| **Parakeet** | NVIDIA's accurate STT | EN (v2), 25 EU languages (v3) |\n| **Voxtral** | Mistral's speech model | Multiple |\n| **Voxtral Realtime** | 4B streaming STT | Multiple |\n| **VibeVoice-ASR** | Microsoft's 9B with diarization | Multiple |\n\n**Parakeet v3 Languages:** Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, Spanish, Swedish, Russian, Ukrainian\n\n## VAD/Speaker Diarization Models\n\n| Model | Description | \n|-------|-------------|\n| **Sortformer v1** | NVIDIA's end-to-end (up to 4 speakers) |\n| **Sortformer v2.1** | Streaming diarization with AOSC compression |\n\n## Speech-to-Speech (STS) Models\n\n| Model | Purpose |\n|-------|---------|\n| **SAM-Audio** | Text-guided source separation |\n| **Liquid2.5-Audio** | Speech/text-to-speech and STT |\n| **MossFormer2 SE** | Speech enhancement/noise removal |\n\n## Configuration Options\n\n### TTS Parameters\n- `--text`: Input text to synthesize\n- `--voice`: Voice preset selection\n- `--speed`: Speech speed adjustment (e.g., 1.2)\n- `--lang_code`: Language code\n- `--play`: Play audio immediately\n- `--output_path`: Directory for saving audio\n- `--ref_audio`: Reference audio for voice cloning\n\n### STT Parameters\n- `--audio`: Input audio file\n- `--model`: Model identifier\n- `--language`: Language specification\n- `--max_tokens`: Token limit (e.g., 8192)\n- `--temperature`: Generation temperature (0.0 for deterministic)\n- `--context`: Hotwords/metadata for context\n- `--output_path`: Output directory\n- `--format`: Output format (json, etc.)\n- `--stream`: Enable streaming mode\n- `--transcription_delay_ms`: Latency setting (Voxtral Realtime)\n- `--verbose`: Detailed logging\n\n## License\n\nThe project is licensed under MIT.","isError":false},{"role":"assistant","text":"Now let me fetch the README more thoroughly, particularly for details about the API server and quantization."},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio/blob/main/README.md","prompt":"Extract the complete raw content of the README. Focus on: 1) The REST API / OpenAI-compatible server setup and usage, 2) Quantization details, 3) Any additional CLI flags or Python API details not covered before. Get everything."},"toolUseId":"toolu_019imtWTE2Lakb4ZaZMFMcD9","output":"# MLX-Audio README: REST API, Quantization & Advanced Features\n\n## Web Interface & API Server\n\nThe project supports an OpenAI-compatible REST API server. Starting the server enables both a web interface with 3D audio visualization and programmatic access to all audio processing models.\n\n**Starting the Server:**\nLaunch the API server with a simple command to make all TTS, STT, and other audio models accessible via HTTP endpoints.\n\n**API Endpoints:**\nThe server implements OpenAI-compatible REST endpoints, allowing developers to integrate MLX-Audio into existing applications designed for OpenAI's API. This includes standard endpoints for speech synthesis, transcription, and other audio operations.\n\n## Quantization\n\nThe library supports multiple quantization levels: 3-bit, 4-bit, 6-bit, 8-bit, and additional precision options. Quantization reduces model size and improves inference speed on Apple Silicon while maintaining quality. Users can select pre-quantized model variants from Hugging Face (e.g., \"mlx-community/model-name-4bit\") or quantize models locally.\n\n## Additional CLI & Python API Features\n\n- **Speech speed control**: Adjust playback speed for TTS generation\n- **Output path specification**: Save generated audio to custom directories\n- **Immediate playback**: Use `--play` flag to hear results instantly\n- **Streaming support**: Several models support streaming transcription for real-time processing\n- **Context/hotwords**: VibeVoice-ASR and other models accept context strings to improve technical term recognition\n- **Model variants**: Choose between different precision levels and model sizes for performance/quality tradeoffs\n\nThe library emphasizes ease of use across command-line, Python API, and web interface paradigms for audio processing workflows.","isError":false},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio/tree/main/mlx_audio/tts","prompt":"List all files and directories in the TTS module. Look for any server.py, api.py, or similar files that indicate REST API functionality."},"toolUseId":"toolu_01VNu3dPPG68dArCP19kaXH9","output":"# TTS Module Contents\n\nBased on the GitHub page content, the TTS (Text-to-Speech) module directory structure includes:\n\n**Files:**\n- `__init__.py`\n- `audio_player.py`\n- `generate.py`\n- `utils.py`\n\n**Subdirectories:**\n- `models/`\n- `tests/`\n\n## API Functionality Assessment\n\nThere is **no evidence of REST API files** (such as `server.py` or `api.py`) within the TTS module itself. The module appears focused on core TTS functionality through generation, audio playback, and utility functions rather than exposing HTTP endpoints.\n\n**Note:** The parent `mlx_audio` directory does contain a `server.py` file, which may provide API capabilities at the package level, but this is separate from the TTS module's direct implementation.","isError":false},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio/blob/main/mlx_audio/server.py","prompt":"Extract the full content of this server.py file. I need to understand: 1) How to start the server (command), 2) All available API endpoints and their parameters, 3) Request/response formats, 4) Any configuration options. Get everything."},"toolUseId":"toolu_017WJRxSMxEMSVeMGNLcgC22","output":"# MLX Audio Server - Complete Documentation\n\n## Starting the Server\n\n**Command:**\n```bash\npython -m mlx_audio.server [OPTIONS]\n```\n\n**Key startup options:**\n- `--host`: Server host (default: localhost)\n- `--port`: Server port (default: 8000)\n- `--allowed-origins`: CORS allowed origins (default: *)\n- `--workers`: Number of workers (int or float, default: 2)\n- `--reload`: Enable auto-reload\n- `--start-ui`: Launch Studio UI alongside API\n- `--log-dir`: Directory for server logs (default: logs)\n\n## API Endpoints\n\n### Model Management\n\n**GET `/v1/models`**\n- Returns list of available models in OpenAI-compatible format\n- Response: `{\"object\": \"list\", \"data\": [{\"id\": model_name, \"object\": \"model\", ...}]}`\n\n**POST `/v1/models`**\n- Adds new model to API\n- Parameter: `model_name` (query string)\n- Response: `{\"status\": \"success\", \"message\": \"Model {name} added successfully\"}`\n\n**DELETE `/v1/models`**\n- Removes model from API\n- Parameter: `model_name` (query string)\n- Response: 204 No Content on success, 404 if not found\n\n### Speech Generation (TTS)\n\n**POST `/v1/audio/speech`**\n- Generates speech from text\n- Request body (SpeechRequest):\n```\n{\n \"model\": str (required),\n \"input\": str (required),\n \"voice\": str | None,\n \"speed\": float (default: 1.0),\n \"gender\": str (default: \"male\"),\n \"pitch\": float (default: 1.0),\n \"instruct\": str | None,\n \"lang_code\": str (default: \"a\"),\n \"ref_audio\": str | None,\n \"ref_text\": str | None,\n \"temperature\": float (default: 0.7),\n \"top_p\": float (default: 0.95),\n \"top_k\": int (default: 40),\n \"repetition_penalty\": float (default: 1.0),\n \"response_format\": str (default: \"mp3\"),\n \"stream\": bool (default: False),\n \"streaming_interval\": float (default: 2.0),\n \"max_tokens\": int (default: 1200),\n \"verbose\": bool (default: False)\n}\n```\n- Response: Audio stream with appropriate content-type\n\n### Speech-to-Text (STT)\n\n**POST `/v1/audio/transcriptions`**\n- Transcribes audio file\n- Parameters (form-data):\n - `file`: Audio file (required)\n - `model`: Model name (required)\n - `language`: Optional language code\n - `verbose`: Boolean (default: False)\n - `max_tokens`: Integer (default: 1024)\n - `chunk_duration`: Float (default: 30.0)\n - `frame_threshold`: Integer (default: 25)\n - `stream`: Boolean (default: False)\n - `context`: Optional context string\n - `prefill_step_size`: Integer (default: 2048)\n - `text`: Optional text parameter\n\n- Response: NDJSON stream with transcription chunks\n```\n{\"text\": str, \"accumulated\": str}\nor\n{\"text\": str, \"start\": float, \"end\": float, \"is_final\": bool, \"language\": str}\n```\n\n**WebSocket `/v1/audio/transcriptions/realtime`**\n- Real-time transcription via WebSocket\n- Initial message (JSON):\n```\n{\n \"model\": str,\n \"language\": str | None,\n \"sample_rate\": int (default: 16000),\n \"streaming\": bool (default: True)\n}\n```\n- Audio: Send as binary frames (int16 PCM)\n- Responses:\n - `{\"status\": \"ready\", \"message\": \"...\"}`\n - `{\"type\": \"delta\", \"delta\": str}`\n - `{\"type\": \"complete\", \"text\": str, \"segments\": null, ...}`\n\n### Audio Separation\n\n**POST `/v1/audio/separations`**\n- Separates audio using SAM Audio model\n- Parameters (form-data):\n - `file`: Audio file (required)\n - `model`: Model name (default: mlx-community/sam-audio-large-fp16)\n - `description`: Target description (default: \"speech\")\n - `method`: ODE method \"midpoint\" or \"euler\" (default: midpoint)\n - `steps`: ODE steps 2/4/8/16/32 (default: 16)\n\n- Response (SeparationResponse):\n```\n{\n \"target\": \"base64_encoded_wav\",\n \"residual\": \"base64_encoded_wav\",\n \"sample_rate\": int\n}\n```\n\n### Root\n\n**GET `/`**\n- Welcome message with link to UI (http://localhost:3000)\n\n## Configuration\n\n**Environment Variables:**\n- `MLX_AUDIO_NUM_WORKERS`: Override worker count\n- `MLX_AUDIO_ALLOWED_ORIGINS`: Comma-separated CORS origins\n\n**CORS Middleware:**\nConfigured automatically on startup; can be reconfigured via `setup_cors(app, origins_list)`\n\n## Architecture Notes\n\n- Uses FastAPI with async/await\n- Implements asyncio locks for thread-safe model management\n- VAD (Voice Activity Detection) enabled for real-time transcription\n- Supports streaming and non-streaming inference modes\n- Sanitizes NaN/Infinity values for JSON compatibility","isError":false},{"role":"tool","toolName":"WebFetch","toolInput":{"url":"https://github.com/Blaizzy/mlx-audio/tree/main/mlx_audio","prompt":"List all files and directories in the mlx_audio root package directory."},"toolUseId":"toolu_01TfedYLiRZqsWufnn4awqmw","output":"# Files and Directories in mlx_audio Root Package\n\nBased on the repository structure provided, the mlx_audio root package contains:\n\n**Directories:**\n- codec\n- sts\n- stt\n- tests\n- tts\n- ui\n- vad\n\n**Files:**\n- `__init__.py`\n- `audio_io.py`\n- `base.py`\n- `convert.py`\n- `dsp.py`\n- `server.py`\n- `utils.py`\n- `version.py`\n\nThis appears to be an audio processing library with modules for speech-to-text (stt), text-to-speech (tts), voice activity detection (vad), speech translation (sts), and codec handling.","isError":false},{"role":"assistant","text":"Now I have comprehensive information. Let me write the SKILL.md file."},{"role":"tool","toolName":"Write","toolInput":{"file_path":"/home/pastilhas/projects/officer.dev/data/skills/mlxaudio/SKILL.md","content":"---\nname: mlx.audio\ndescription: Generate speech from text and transcribe audio using mlx-audio. Use when the user wants text-to-speech synthesis, speech-to-text transcription, voice cloning, audio separation, or speech-to-speech processing on Apple Silicon.\n---\n\n# MLX-Audio\n\nA speech processing library built on Apple's MLX framework, providing TTS, STT, speech-to-speech (STS), and audio separation optimized for Apple Silicon.\n\n- **Repository:** https://github.com/Blaizzy/mlx-audio\n- **License:** MIT\n\n## CLI Tools\n\n### Text-to-Speech (TTS)\n\n```bash\nmlx_audio.tts.generate --model --text '' [options]\n```\n\n| Flag | Type | Default | Description |\n|------|------|---------|-------------|\n| `--model` | string | required | HuggingFace model ID |\n| `--text` | string | required | Text to synthesize |\n| `--voice` | string | — | Voice preset (model-specific) |\n| `--speed` | float | 1.0 | Speech speed multiplier |\n| `--lang_code` | string | `a` | Language code |\n| `--play` | flag | — | Play audio immediately |\n| `--output_path` | string | — | Directory to save audio |\n| `--ref_audio` | string | — | Reference audio for voice cloning (CSM) |\n\n#### Language Codes\n\n| Code | Language |\n|------|----------|\n| `a` | American English |\n| `b` | British English |\n| `j` | Japanese |\n| `z` | Mandarin Chinese |\n| `e` | Spanish |\n| `f` | French |\n\n#### Kokoro Voices\n\n| Voice | Description |\n|-------|-------------|\n| `af_heart`, `af_bella`, `af_nova`, `af_sky` | American female |\n| `am_adam`, `am_echo` | American male |\n| `bf_alice`, `bf_emma` | British female |\n| `bm_daniel`, `bm_george` | British male |\n| `jf_alpha`, `jm_kumo` | Japanese |\n| `zf_xiaobei`, `zm_yunxi` | Chinese |\n\n#### Examples\n\n```bash\n# Basic generation\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello, world!' --lang_code a\n\n# With voice and speed\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --voice af_heart --speed 1.2 --lang_code a\n\n# Play immediately\nmlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --play --lang_code a\n\n# Voice cloning with CSM\nmlx_audio.tts.generate --model mlx-community/csm-1b --text \"Hello from Sesame.\" --ref_audio ./reference_voice.wav --play\n```\n\n### Speech-to-Text (STT)\n\n```bash\npython -m mlx_audio.stt.generate --model --audio [options]\n```\n\n| Flag | Type | Default | Description |\n|------|------|---------|-------------|\n| `--model` | string | required | HuggingFace model ID |\n| `--audio` | string | required | Input audio file |\n| `--language` | string | — | Language code |\n| `--max-tokens` | int | 1024 | Maximum output tokens |\n| `--temperature` | float | 0.0 | Sampling temperature |\n| `--context` | string | — | Hotwords/metadata for context |\n| `--output-path` | string | — | Output directory |\n| `--format` | string | — | Output format (e.g. `json`) |\n| `--stream` | flag | — | Enable streaming mode |\n| `--verbose` | flag | — | Detailed logging |\n\n#### Examples\n\n```bash\n# Basic transcription\npython -m mlx_audio.stt.generate --model mlx-community/whisper-large-v3-turbo-asr-fp16 --audio speech.wav --verbose\n\n# With context for technical terms\npython -m mlx_audio.stt.generate --model mlx-community/VibeVoice-ASR-bf16 --audio meeting.wav --context \"MLX, Apple Silicon, PyTorch\" --max-tokens 8192 --format json --verbose\n\n# Parakeet model\npython -m mlx_audio.stt.generate --model mlx-community/parakeet-tdt-0.6b-v3 --audio speech.wav --format json --verbose\n```\n\n---\n\n## Python API\n\n### TTS\n\n```python\nfrom mlx_audio.tts.utils import load_model\n\nmodel = load_model(\"mlx-community/Kokoro-82M-bf16\")\nfor result in model.generate(\"Hello from MLX-Audio!\", voice=\"af_heart\"):\n audio = result.audio # mx.array waveform\n```\n\n### STT\n\n```python\nfrom mlx_audio.stt.generate import generate_transcription\n\nresult = generate_transcription(\n model=\"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n audio=\"audio.wav\",\n)\nprint(result.text)\n```\n\n### STT with Streaming\n\n```python\nfrom mlx_audio.stt import load\n\n# VibeVoice-ASR streaming\nmodel = load(\"mlx-community/VibeVoice-ASR-bf16\")\nfor text in model.stream_transcribe(audio=\"speech.wav\", max_tokens=4096):\n print(text, end=\"\", flush=True)\n\n# Parakeet streaming\nmodel = load(\"mlx-community/parakeet-tdt-0.6b-v3\")\nfor chunk in model.generate(\"long_audio.wav\", stream=True):\n print(chunk.text, end=\"\", flush=True)\n```\n\n### Forced Alignment (Qwen3)\n\n```python\nfrom mlx_audio.stt import load\n\naligner = load(\"mlx-community/Qwen3-ForcedAligner-0.6B-8bit\")\nresult = aligner.generate(\"audio.wav\", text=\"I have a dream\", language=\"English\")\nfor item in result:\n print(f\"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}\")\n```\n\n---\n\n## REST API Server (OpenAI-compatible)\n\n### Starting the Server\n\n```bash\npython -m mlx_audio.server [OPTIONS]\n```\n\n| Flag | Type | Default | Description |\n|------|------|---------|-------------|\n| `--host` | string | `localhost` | Server host |\n| `--port` | int | `8000` | Server port |\n| `--allowed-origins` | string | `*` | CORS allowed origins |\n| `--workers` | int/float | `2` | Number of workers |\n| `--reload` | flag | — | Enable auto-reload |\n| `--start-ui` | flag | — | Launch Studio UI alongside API |\n| `--log-dir` | string | `logs` | Directory for server logs |\n\n### Endpoints\n\n#### GET /v1/models\n\nList available models.\n\n```bash\ncurl http://localhost:8000/v1/models\n```\n\n#### POST /v1/models?model_name=\\\n\nAdd a model to the server.\n\n#### DELETE /v1/models?model_name=\\\n\nRemove a model from the server.\n\n#### POST /v1/audio/speech\n\nGenerate speech from text.\n\n```bash\ncurl -X POST http://localhost:8000/v1/audio/speech \\\n -H \"Content-Type: application/json\" \\\n -d '{\n \"model\": \"mlx-community/Kokoro-82M-bf16\",\n \"input\": \"Hello, world!\",\n \"voice\": \"af_heart\",\n \"speed\": 1.0,\n \"lang_code\": \"a\",\n \"response_format\": \"mp3\"\n }' --output speech.mp3\n```\n\n| Parameter | Type | Default | Description |\n|-----------|------|---------|-------------|\n| `model` | string | required | Model ID |\n| `input` | string | required | Text to synthesize |\n| `voice` | string | — | Voice preset |\n| `speed` | float | 1.0 | Speech speed |\n| `lang_code` | string | `a` | Language code |\n| `ref_audio` | string | — | Reference audio path (voice cloning) |\n| `ref_text` | string | — | Reference transcript |\n| `response_format` | string | `mp3` | Output format |\n| `stream` | bool | false | Enable streaming |\n| `streaming_interval` | float | 2.0 | Streaming chunk interval |\n| `temperature` | float | 0.7 | Sampling temperature |\n| `top_p` | float | 0.95 | Nucleus sampling |\n| `top_k` | int | 40 | Top-k sampling |\n| `repetition_penalty` | float | 1.0 | Repetition penalty |\n| `max_tokens` | int | 1200 | Maximum tokens |\n| `gender` | string | `male` | Gender hint |\n| `pitch` | float | 1.0 | Pitch adjustment |\n| `instruct` | string | — | Instruction text |\n\n#### POST /v1/audio/transcriptions\n\nTranscribe an audio file (multipart/form-data).\n\n```bash\ncurl -X POST http://localhost:8000/v1/audio/transcriptions \\\n -F file=@audio.wav \\\n -F model=mlx-community/whisper-large-v3-turbo-asr-fp16\n```\n\n| Parameter | Type | Default | Description |\n|-----------|------|---------|-------------|\n| `file` | file | required | Audio file |\n| `model` | string | required | Model ID |\n| `language` | string | — | Language code |\n| `max_tokens` | int | 1024 | Maximum tokens |\n| `chunk_duration` | float | 30.0 | Chunk duration (seconds) |\n| `stream` | bool | false | Enable streaming |\n| `context` | string | — | Hotwords/context |\n| `text` | string | — | Reference text |\n| `verbose` | bool | false | Detailed output |\n\nResponse (NDJSON stream):\n\n```json\n{\"text\": \"chunk text\", \"accumulated\": \"full text so far\"}\n```\n\n#### WebSocket /v1/audio/transcriptions/realtime\n\nReal-time transcription via WebSocket. Send initial config as JSON, then stream int16 PCM audio as binary frames.\n\n```json\n{\n \"model\": \"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n \"sample_rate\": 16000,\n \"streaming\": true\n}\n```\n\n#### POST /v1/audio/separations\n\nSeparate audio sources (multipart/form-data).\n\n```bash\ncurl -X POST http://localhost:8000/v1/audio/separations \\\n -F file=@audio.wav \\\n -F model=mlx-community/sam-audio-large-fp16 \\\n -F description=\"speech\"\n```\n\n| Parameter | Type | Default | Description |\n|-----------|------|---------|-------------|\n| `file` | file | required | Audio file |\n| `model` | string | `mlx-community/sam-audio-large-fp16` | Model ID |\n| `description` | string | `speech` | Target description |\n| `method` | string | `midpoint` | ODE method (`midpoint` or `euler`) |\n| `steps` | int | 16 | ODE steps (2/4/8/16/32) |\n\nResponse:\n\n```json\n{\n \"target\": \"\",\n \"residual\": \"\",\n \"sample_rate\": 44100\n}\n```\n\n---\n\n## Supported Models\n\n### TTS Models\n\n| Model | Languages | Notes |\n|-------|-----------|-------|\n| Kokoro | EN, JA, ZH, FR, ES, IT, PT, HI | Fast, high-quality multilingual |\n| Qwen3-TTS | ZH, EN, JA, KO, + more | Voice design via instruction |\n| CSM | EN | Voice cloning with reference audio |\n| Dia | EN | Dialogue-focused |\n| OuteTTS | EN | Efficient |\n| Spark | EN, ZH | SparkTTS |\n| Chatterbox | EN, ES, FR, DE, IT, PT, PL, TR, RU, NL, CS, AR, ZH, JA, HU, KO | Expressive multilingual |\n| Soprano | EN | High-quality |\n\n### STT Models\n\n| Model | Languages | Notes |\n|-------|-----------|-------|\n| Whisper | 99+ languages | OpenAI's robust model |\n| Qwen3-ASR | ZH, EN, JA, KO, + more | Alibaba multilingual |\n| Qwen3-ForcedAligner | ZH, EN, JA, KO, + more | Word-level alignment |\n| Parakeet | EN (v2), 25 EU languages (v3) | NVIDIA, high accuracy |\n| Voxtral | Multiple | Mistral speech model |\n| Voxtral Realtime | Multiple | 4B streaming STT |\n| VibeVoice-ASR | Multiple | Microsoft 9B, supports diarization and context |\n\n### Other Models\n\n| Model | Type | Description |\n|-------|------|-------------|\n| Sortformer v1/v2.1 | VAD/Diarization | Speaker diarization (up to 4 speakers) |\n| SAM-Audio | Separation | Text-guided source separation |\n| Liquid2.5-Audio | STS | Speech/text-to-speech and STT |\n| MossFormer2 SE | Enhancement | Speech enhancement / noise removal |\n\nModels are available from `mlx-community` on HuggingFace with various quantization levels (3-bit through 8-bit and fp16/bf16).\n\n## Source\n\n- Repository: https://github.com/Blaizzy/mlx-audio\n- HuggingFace: https://huggingface.co/mlx-community\n"},"toolUseId":"toolu_0189M4gCkmdB8zkFsFYPXooY","output":"The file /home/pastilhas/projects/officer.dev/data/skills/mlxaudio/SKILL.md has been updated. Here's the result of running `cat -n` on a snippet of the edited file:\n 1→---\n 2→name: mlx.audio\n 3→description: Generate speech from text and transcribe audio using mlx-audio. Use when the user wants text-to-speech synthesis, speech-to-text transcription, voice cloning, audio separation, or speech-to-speech processing on Apple Silicon.\n 4→---\n 5→\n 6→# MLX-Audio\n 7→\n 8→A speech processing library built on Apple's MLX framework, providing TTS, STT, speech-to-speech (STS), and audio separation optimized for Apple Silicon.\n 9→\n 10→- **Repository:** https://github.com/Blaizzy/mlx-audio\n 11→- **License:** MIT\n 12→\n 13→## CLI Tools\n 14→\n 15→### Text-to-Speech (TTS)\n 16→\n 17→```bash\n 18→mlx_audio.tts.generate --model --text '' [options]\n 19→```\n 20→\n 21→| Flag | Type | Default | Description |\n 22→|------|------|---------|-------------|\n 23→| `--model` | string | required | HuggingFace model ID |\n 24→| `--text` | string | required | Text to synthesize |\n 25→| `--voice` | string | — | Voice preset (model-specific) |\n 26→| `--speed` | float | 1.0 | Speech speed multiplier |\n 27→| `--lang_code` | string | `a` | Language code |\n 28→| `--play` | flag | — | Play audio immediately |\n 29→| `--output_path` | string | — | Directory to save audio |\n 30→| `--ref_audio` | string | — | Reference audio for voice cloning (CSM) |\n 31→\n 32→#### Language Codes\n 33→\n 34→| Code | Language |\n 35→|------|----------|\n 36→| `a` | American English |\n 37→| `b` | British English |\n 38→| `j` | Japanese |\n 39→| `z` | Mandarin Chinese |\n 40→| `e` | Spanish |\n 41→| `f` | French |\n 42→\n 43→#### Kokoro Voices\n 44→\n 45→| Voice | Description |\n 46→|-------|-------------|\n 47→| `af_heart`, `af_bella`, `af_nova`, `af_sky` | American female |\n 48→| `am_adam`, `am_echo` | American male |\n 49→| `bf_alice`, `bf_emma` | British female |\n 50→| `bm_daniel`, `bm_george` | British male |\n 51→| `jf_alpha`, `jm_kumo` | Japanese |\n 52→| `zf_xiaobei`, `zm_yunxi` | Chinese |\n 53→\n 54→#### Examples\n 55→\n 56→```bash\n 57→# Basic generation\n 58→mlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello, world!' --lang_code a\n 59→\n 60→# With voice and speed\n 61→mlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --voice af_heart --speed 1.2 --lang_code a\n 62→\n 63→# Play immediately\n 64→mlx_audio.tts.generate --model mlx-community/Kokoro-82M-bf16 --text 'Hello!' --play --lang_code a\n 65→\n 66→# Voice cloning with CSM\n 67→mlx_audio.tts.generate --model mlx-community/csm-1b --text \"Hello from Sesame.\" --ref_audio ./reference_voice.wav --play\n 68→```\n 69→\n 70→### Speech-to-Text (STT)\n 71→\n 72→```bash\n 73→python -m mlx_audio.stt.generate --model --audio [options]\n 74→```\n 75→\n 76→| Flag | Type | Default | Description |\n 77→|------|------|---------|-------------|\n 78→| `--model` | string | required | HuggingFace model ID |\n 79→| `--audio` | string | required | Input audio file |\n 80→| `--language` | string | — | Language code |\n 81→| `--max-tokens` | int | 1024 | Maximum output tokens |\n 82→| `--temperature` | float | 0.0 | Sampling temperature |\n 83→| `--context` | string | — | Hotwords/metadata for context |\n 84→| `--output-path` | string | — | Output directory |\n 85→| `--format` | string | — | Output format (e.g. `json`) |\n 86→| `--stream` | flag | — | Enable streaming mode |\n 87→| `--verbose` | flag | — | Detailed logging |\n 88→\n 89→#### Examples\n 90→\n 91→```bash\n 92→# Basic transcription\n 93→python -m mlx_audio.stt.generate --model mlx-community/whisper-large-v3-turbo-asr-fp16 --audio speech.wav --verbose\n 94→\n 95→# With context for technical terms\n 96→python -m mlx_audio.stt.generate --model mlx-community/VibeVoice-ASR-bf16 --audio meeting.wav --context \"MLX, Apple Silicon, PyTorch\" --max-tokens 8192 --format json --verbose\n 97→\n 98→# Parakeet model\n 99→python -m mlx_audio.stt.generate --model mlx-community/parakeet-tdt-0.6b-v3 --audio speech.wav --format json --verbose\n 100→```\n 101→\n 102→---\n 103→\n 104→## Python API\n 105→\n 106→### TTS\n 107→\n 108→```python\n 109→from mlx_audio.tts.utils import load_model\n 110→\n 111→model = load_model(\"mlx-community/Kokoro-82M-bf16\")\n 112→for result in model.generate(\"Hello from MLX-Audio!\", voice=\"af_heart\"):\n 113→ audio = result.audio # mx.array waveform\n 114→```\n 115→\n 116→### STT\n 117→\n 118→```python\n 119→from mlx_audio.stt.generate import generate_transcription\n 120→\n 121→result = generate_transcription(\n 122→ model=\"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n 123→ audio=\"audio.wav\",\n 124→)\n 125→print(result.text)\n 126→```\n 127→\n 128→### STT with Streaming\n 129→\n 130→```python\n 131→from mlx_audio.stt import load\n 132→\n 133→# VibeVoice-ASR streaming\n 134→model = load(\"mlx-community/VibeVoice-ASR-bf16\")\n 135→for text in model.stream_transcribe(audio=\"speech.wav\", max_tokens=4096):\n 136→ print(text, end=\"\", flush=True)\n 137→\n 138→# Parakeet streaming\n 139→model = load(\"mlx-community/parakeet-tdt-0.6b-v3\")\n 140→for chunk in model.generate(\"long_audio.wav\", stream=True):\n 141→ print(chunk.text, end=\"\", flush=True)\n 142→```\n 143→\n 144→### Forced Alignment (Qwen3)\n 145→\n 146→```python\n 147→from mlx_audio.stt import load\n 148→\n 149→aligner = load(\"mlx-community/Qwen3-ForcedAligner-0.6B-8bit\")\n 150→result = aligner.generate(\"audio.wav\", text=\"I have a dream\", language=\"English\")\n 151→for item in result:\n 152→ print(f\"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}\")\n 153→```\n 154→\n 155→---\n 156→\n 157→## REST API Server (OpenAI-compatible)\n 158→\n 159→### Starting the Server\n 160→\n 161→```bash\n 162→python -m mlx_audio.server [OPTIONS]\n 163→```\n 164→\n 165→| Flag | Type | Default | Description |\n 166→|------|------|---------|-------------|\n 167→| `--host` | string | `localhost` | Server host |\n 168→| `--port` | int | `8000` | Server port |\n 169→| `--allowed-origins` | string | `*` | CORS allowed origins |\n 170→| `--workers` | int/float | `2` | Number of workers |\n 171→| `--reload` | flag | — | Enable auto-reload |\n 172→| `--start-ui` | flag | — | Launch Studio UI alongside API |\n 173→| `--log-dir` | string | `logs` | Directory for server logs |\n 174→\n 175→### Endpoints\n 176→\n 177→#### GET /v1/models\n 178→\n 179→List available models.\n 180→\n 181→```bash\n 182→curl http://localhost:8000/v1/models\n 183→```\n 184→\n 185→#### POST /v1/models?model_name=\\\n 186→\n 187→Add a model to the server.\n 188→\n 189→#### DELETE /v1/models?model_name=\\\n 190→\n 191→Remove a model from the server.\n 192→\n 193→#### POST /v1/audio/speech\n 194→\n 195→Generate speech from text.\n 196→\n 197→```bash\n 198→curl -X POST http://localhost:8000/v1/audio/speech \\\n 199→ -H \"Content-Type: application/json\" \\\n 200→ -d '{\n 201→ \"model\": \"mlx-community/Kokoro-82M-bf16\",\n 202→ \"input\": \"Hello, world!\",\n 203→ \"voice\": \"af_heart\",\n 204→ \"speed\": 1.0,\n 205→ \"lang_code\": \"a\",\n 206→ \"response_format\": \"mp3\"\n 207→ }' --output speech.mp3\n 208→```\n 209→\n 210→| Parameter | Type | Default | Description |\n 211→|-----------|------|---------|-------------|\n 212→| `model` | string | required | Model ID |\n 213→| `input` | string | required | Text to synthesize |\n 214→| `voice` | string | — | Voice preset |\n 215→| `speed` | float | 1.0 | Speech speed |\n 216→| `lang_code` | string | `a` | Language code |\n 217→| `ref_audio` | string | — | Reference audio path (voice cloning) |\n 218→| `ref_text` | string | — | Reference transcript |\n 219→| `response_format` | string | `mp3` | Output format |\n 220→| `stream` | bool | false | Enable streaming |\n 221→| `streaming_interval` | float | 2.0 | Streaming chunk interval |\n 222→| `temperature` | float | 0.7 | Sampling temperature |\n 223→| `top_p` | float | 0.95 | Nucleus sampling |\n 224→| `top_k` | int | 40 | Top-k sampling |\n 225→| `repetition_penalty` | float | 1.0 | Repetition penalty |\n 226→| `max_tokens` | int | 1200 | Maximum tokens |\n 227→| `gender` | string | `male` | Gender hint |\n 228→| `pitch` | float | 1.0 | Pitch adjustment |\n 229→| `instruct` | string | — | Instruction text |\n 230→\n 231→#### POST /v1/audio/transcriptions\n 232→\n 233→Transcribe an audio file (multipart/form-data).\n 234→\n 235→```bash\n 236→curl -X POST http://localhost:8000/v1/audio/transcriptions \\\n 237→ -F file=@audio.wav \\\n 238→ -F model=mlx-community/whisper-large-v3-turbo-asr-fp16\n 239→```\n 240→\n 241→| Parameter | Type | Default | Description |\n 242→|-----------|------|---------|-------------|\n 243→| `file` | file | required | Audio file |\n 244→| `model` | string | required | Model ID |\n 245→| `language` | string | — | Language code |\n 246→| `max_tokens` | int | 1024 | Maximum tokens |\n 247→| `chunk_duration` | float | 30.0 | Chunk duration (seconds) |\n 248→| `stream` | bool | false | Enable streaming |\n 249→| `context` | string | — | Hotwords/context |\n 250→| `text` | string | — | Reference text |\n 251→| `verbose` | bool | false | Detailed output |\n 252→\n 253→Response (NDJSON stream):\n 254→\n 255→```json\n 256→{\"text\": \"chunk text\", \"accumulated\": \"full text so far\"}\n 257→```\n 258→\n 259→#### WebSocket /v1/audio/transcriptions/realtime\n 260→\n 261→Real-time transcription via WebSocket. Send initial config as JSON, then stream int16 PCM audio as binary frames.\n 262→\n 263→```json\n 264→{\n 265→ \"model\": \"mlx-community/whisper-large-v3-turbo-asr-fp16\",\n 266→ \"sample_rate\": 16000,\n 267→ \"streaming\": true\n 268→}\n 269→```\n 270→\n 271→#### POST /v1/audio/separations\n 272→\n 273→Separate audio sources (multipart/form-data).\n 274→\n 275→```bash\n 276→curl -X POST http://localhost:8000/v1/audio/separations \\\n 277→ -F file=@audio.wav \\\n 278→ -F model=mlx-community/sam-audio-large-fp16 \\\n 279→ -F description=\"speech\"\n 280→```\n 281→\n 282→| Parameter | Type | Default | Description |\n 283→|-----------|------|---------|-------------|\n 284→| `file` | file | required | Audio file |\n 285→| `model` | string | `mlx-community/sam-audio-large-fp16` | Model ID |\n 286→| `description` | string | `speech` | Target description |\n 287→| `method` | string | `midpoint` | ODE method (`midpoint` or `euler`) |\n 288→| `steps` | int | 16 | ODE steps (2/4/8/16/32) |\n 289→\n 290→Response:\n 291→\n 292→```json\n 293→{\n 294→ \"target\": \"\",\n 295→ \"residual\": \"\",\n 296→ \"sample_rate\": 44100\n 297→}\n 298→```\n 299→\n 300→---\n 301→\n 302→## Supported Models\n 303→\n 304→### TTS Models\n 305→\n 306→| Model | Languages | Notes |\n 307→|-------|-----------|-------|\n 308→| Kokoro | EN, JA, ZH, FR, ES, IT, PT, HI | Fast, high-quality multilingual |\n 309→| Qwen3-TTS | ZH, EN, JA, KO, + more | Voice design via instruction |\n 310→| CSM | EN | Voice cloning with reference audio |\n 311→| Dia | EN | Dialogue-focused |\n 312→| OuteTTS | EN | Efficient |\n 313→| Spark | EN, ZH | SparkTTS |\n 314→| Chatterbox | EN, ES, FR, DE, IT, PT, PL, TR, RU, NL, CS, AR, ZH, JA, HU, KO | Expressive multilingual |\n 315→| Soprano | EN | High-quality |\n 316→\n 317→### STT Models\n 318→\n 319→| Model | Languages | Notes |\n 320→|-------|-----------|-------|\n 321→| Whisper | 99+ languages | OpenAI's robust model |\n 322→| Qwen3-ASR | ZH, EN, JA, KO, + more | Alibaba multilingual |\n 323→| Qwen3-ForcedAligner | ZH, EN, JA, KO, + more | Word-level alignment |\n 324→| Parakeet | EN (v2), 25 EU languages (v3) | NVIDIA, high accuracy |\n 325→| Voxtral | Multiple | Mistral speech model |\n 326→| Voxtral Realtime | Multiple | 4B streaming STT |\n 327→| VibeVoice-ASR | Multiple | Microsoft 9B, supports diarization and context |\n 328→\n 329→### Other Models\n 330→\n 331→| Model | Type | Description |\n 332→|-------|------|-------------|\n 333→| Sortformer v1/v2.1 | VAD/Diarization | Speaker diarization (up to 4 speakers) |\n 334→| SAM-Audio | Separation | Text-guided source separation |\n 335→| Liquid2.5-Audio | STS | Speech/text-to-speech and STT |\n 336→| MossFormer2 SE | Enhancement | Speech enhancement / noise removal |\n 337→\n 338→Models are available from `mlx-community` on HuggingFace with various quantization levels (3-bit through 8-bit and fp16/bf16).\n 339→\n 340→## Source\n 341→\n 342→- Repository: https://github.com/Blaizzy/mlx-audio\n 343→- HuggingFace: https://huggingface.co/mlx-community\n 344→","isError":false}]