This commit is contained in:
2026-02-16 19:34:35 +00:00
commit 9ab0940ca4
784 changed files with 41710 additions and 0 deletions
+173
View File
@@ -0,0 +1,173 @@
---
name: whisper.cpp
description: Transcribe audio files to text using whisper.cpp. Use when the user wants to transcribe audio, convert speech to text, or extract text from an audio/video file.
---
# Whisper.cpp
API reference for the whisper.cpp HTTP server running at `http://macmini:8178`.
whisper.cpp is a C/C++ port of OpenAI's Whisper speech recognition model. The server accepts audio files via HTTP and returns transcriptions in various formats.
## Server
- **Base URL:** `http://macmini:8178`
- **No authentication required**
## Endpoints
### POST /inference
Transcribes an audio file. Accepts `multipart/form-data`.
#### Example
```bash
curl -s http://macmini:8178/inference \
-F file="@/path/to/audio.mp3" \
-F temperature="0.0" \
-F temperature_inc="0.2" \
-F response_format="json"
```
#### Parameters
##### File (required)
| Parameter | Type | Description |
|-----------|------|-------------|
| `file` | file | Audio file to transcribe. Accepts at least WAV and MP3. |
##### Response Format
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `response_format` | string | `json` | Output format: `json`, `verbose_json` (or `vjson`), `text`, `srt`, `vtt` |
##### Language
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `language` | string | `en` | Spoken language code (e.g. `en`, `pt`, `es`, `fr`). Use `auto` for auto-detection. |
| `detect_language` | bool | `false` | Exit after detecting the language (no transcription). |
| `translate` | bool | `false` | Translate from source language to English. |
##### Decoding
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `temperature` | float | `0.0` | Sampling temperature. `0.0` is deterministic. |
| `temperature_inc` | float | `0.2` | Temperature increment on fallback attempts. |
| `best_of` | int | `2` | Number of candidate decodings to keep. |
| `beam_size` | int | `-1` | Beam search size. `-1` disables beam search. |
| `entropy_thold` | float | `2.40` | Entropy threshold — decoder fails and retries if exceeded. |
| `logprob_thold` | float | `-1.00` | Log probability threshold for decoder failure. |
| `no_fallback` | bool | `false` | Disable temperature fallback on decode failure. |
##### Segmentation
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `max_len` | int | `0` | Maximum segment length in characters. `0` for unlimited. |
| `max_context` | int | `-1` | Maximum text context tokens to store. `-1` for unlimited. |
| `split_on_word` | bool | `false` | Split segments at word boundaries instead of token boundaries. |
| `no_timestamps` | bool | `false` | Suppress timestamps in output. |
| `word_thold` | float | `0.01` | Word timestamp probability threshold. |
##### Audio Processing
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `offset_t` | int | `0` | Time offset in milliseconds — skip this much audio from the start. |
| `offset_n` | int | `0` | Segment index offset. |
| `duration` | int | `0` | Duration of audio to process in milliseconds. `0` for all. |
| `audio_ctx` | int | `0` | Audio context size. `0` for all. |
##### Speaker Diarization
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `diarize` | bool | `false` | Enable speaker diarization (requires stereo audio). |
| `tinydiarize` | bool | `false` | Enable tinydiarize (requires a tdrz model). |
##### Voice Activity Detection (VAD)
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `vad` | bool | `false` | Enable VAD preprocessing. |
| `vad_threshold` | float | `0.50` | Speech confidence threshold (0.01.0). |
| `vad_min_speech_duration_ms` | int | `250` | Minimum speech segment duration in ms. |
| `vad_min_silence_duration_ms` | int | `100` | Minimum silence duration to split segments. |
| `vad_max_speech_duration_s` | float | `FLT_MAX` | Auto-split segments longer than this (seconds). |
| `vad_speech_pad_ms` | int | `30` | Padding added around speech segments (ms). |
| `vad_samples_overlap` | float | `0.10` | Overlap between segments (seconds). |
##### Other
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `prompt` | string | `""` | Initial prompt to condition the model (e.g. for vocabulary hints). |
| `suppress_nst` | bool | `false` | Suppress non-speech tokens. |
| `no_context` | bool | `false` | Do not use previous audio context for subsequent segments. |
| `debug_mode` | bool | `false` | Enable debug output. |
#### Response Formats
##### `json` (default)
Minimal JSON with just the transcribed text.
```json
{"text": "The transcribed content goes here."}
```
##### `verbose_json` (or `vjson`)
Extended JSON including task type, language, audio duration, per-segment timestamps, token-level timing, confidence scores, and language probability distribution.
##### `text`
Plain text transcription. Includes speaker labels if diarization is enabled.
##### `srt`
SubRip subtitle format with sequential numbering, `HH:MM:SS,mmm` timestamps, and text content.
```
1
00:00:00,000 --> 00:00:03,500
The transcribed content goes here.
```
##### `vtt`
WebVTT subtitle format with `WEBVTT` header and `HH:MM:SS.mmm` timestamps.
```
WEBVTT
00:00:00.000 --> 00:00:03.500
The transcribed content goes here.
```
### POST /load
Loads a different model file on the server at runtime.
```bash
curl -s http://macmini:8178/load \
-F model="/path/to/model.bin"
```
| Parameter | Type | Description |
|-----------|------|-------------|
| `model` | string | Path to the model file on the server. |
## Supported Audio Formats
The server accepts at least WAV (16-bit PCM) and MP3 files directly. If the server was started with `--convert`, it can use ffmpeg to handle additional formats (ogg, flac, m4a, etc.).
## Source
- Repository: https://github.com/ggml-org/whisper.cpp
- Server docs: https://github.com/ggml-org/whisper.cpp/blob/master/examples/server/README.md