first
This commit is contained in:
@@ -0,0 +1,173 @@
|
||||
---
|
||||
name: whisper.cpp
|
||||
description: Transcribe audio files to text using whisper.cpp. Use when the user wants to transcribe audio, convert speech to text, or extract text from an audio/video file.
|
||||
---
|
||||
|
||||
# Whisper.cpp
|
||||
|
||||
API reference for the whisper.cpp HTTP server running at `http://macmini:8178`.
|
||||
|
||||
whisper.cpp is a C/C++ port of OpenAI's Whisper speech recognition model. The server accepts audio files via HTTP and returns transcriptions in various formats.
|
||||
|
||||
## Server
|
||||
|
||||
- **Base URL:** `http://macmini:8178`
|
||||
- **No authentication required**
|
||||
|
||||
## Endpoints
|
||||
|
||||
### POST /inference
|
||||
|
||||
Transcribes an audio file. Accepts `multipart/form-data`.
|
||||
|
||||
#### Example
|
||||
|
||||
```bash
|
||||
curl -s http://macmini:8178/inference \
|
||||
-F file="@/path/to/audio.mp3" \
|
||||
-F temperature="0.0" \
|
||||
-F temperature_inc="0.2" \
|
||||
-F response_format="json"
|
||||
```
|
||||
|
||||
#### Parameters
|
||||
|
||||
##### File (required)
|
||||
|
||||
| Parameter | Type | Description |
|
||||
|-----------|------|-------------|
|
||||
| `file` | file | Audio file to transcribe. Accepts at least WAV and MP3. |
|
||||
|
||||
##### Response Format
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `response_format` | string | `json` | Output format: `json`, `verbose_json` (or `vjson`), `text`, `srt`, `vtt` |
|
||||
|
||||
##### Language
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `language` | string | `en` | Spoken language code (e.g. `en`, `pt`, `es`, `fr`). Use `auto` for auto-detection. |
|
||||
| `detect_language` | bool | `false` | Exit after detecting the language (no transcription). |
|
||||
| `translate` | bool | `false` | Translate from source language to English. |
|
||||
|
||||
##### Decoding
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `temperature` | float | `0.0` | Sampling temperature. `0.0` is deterministic. |
|
||||
| `temperature_inc` | float | `0.2` | Temperature increment on fallback attempts. |
|
||||
| `best_of` | int | `2` | Number of candidate decodings to keep. |
|
||||
| `beam_size` | int | `-1` | Beam search size. `-1` disables beam search. |
|
||||
| `entropy_thold` | float | `2.40` | Entropy threshold — decoder fails and retries if exceeded. |
|
||||
| `logprob_thold` | float | `-1.00` | Log probability threshold for decoder failure. |
|
||||
| `no_fallback` | bool | `false` | Disable temperature fallback on decode failure. |
|
||||
|
||||
##### Segmentation
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `max_len` | int | `0` | Maximum segment length in characters. `0` for unlimited. |
|
||||
| `max_context` | int | `-1` | Maximum text context tokens to store. `-1` for unlimited. |
|
||||
| `split_on_word` | bool | `false` | Split segments at word boundaries instead of token boundaries. |
|
||||
| `no_timestamps` | bool | `false` | Suppress timestamps in output. |
|
||||
| `word_thold` | float | `0.01` | Word timestamp probability threshold. |
|
||||
|
||||
##### Audio Processing
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `offset_t` | int | `0` | Time offset in milliseconds — skip this much audio from the start. |
|
||||
| `offset_n` | int | `0` | Segment index offset. |
|
||||
| `duration` | int | `0` | Duration of audio to process in milliseconds. `0` for all. |
|
||||
| `audio_ctx` | int | `0` | Audio context size. `0` for all. |
|
||||
|
||||
##### Speaker Diarization
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `diarize` | bool | `false` | Enable speaker diarization (requires stereo audio). |
|
||||
| `tinydiarize` | bool | `false` | Enable tinydiarize (requires a tdrz model). |
|
||||
|
||||
##### Voice Activity Detection (VAD)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `vad` | bool | `false` | Enable VAD preprocessing. |
|
||||
| `vad_threshold` | float | `0.50` | Speech confidence threshold (0.0–1.0). |
|
||||
| `vad_min_speech_duration_ms` | int | `250` | Minimum speech segment duration in ms. |
|
||||
| `vad_min_silence_duration_ms` | int | `100` | Minimum silence duration to split segments. |
|
||||
| `vad_max_speech_duration_s` | float | `FLT_MAX` | Auto-split segments longer than this (seconds). |
|
||||
| `vad_speech_pad_ms` | int | `30` | Padding added around speech segments (ms). |
|
||||
| `vad_samples_overlap` | float | `0.10` | Overlap between segments (seconds). |
|
||||
|
||||
##### Other
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `prompt` | string | `""` | Initial prompt to condition the model (e.g. for vocabulary hints). |
|
||||
| `suppress_nst` | bool | `false` | Suppress non-speech tokens. |
|
||||
| `no_context` | bool | `false` | Do not use previous audio context for subsequent segments. |
|
||||
| `debug_mode` | bool | `false` | Enable debug output. |
|
||||
|
||||
#### Response Formats
|
||||
|
||||
##### `json` (default)
|
||||
|
||||
Minimal JSON with just the transcribed text.
|
||||
|
||||
```json
|
||||
{"text": "The transcribed content goes here."}
|
||||
```
|
||||
|
||||
##### `verbose_json` (or `vjson`)
|
||||
|
||||
Extended JSON including task type, language, audio duration, per-segment timestamps, token-level timing, confidence scores, and language probability distribution.
|
||||
|
||||
##### `text`
|
||||
|
||||
Plain text transcription. Includes speaker labels if diarization is enabled.
|
||||
|
||||
##### `srt`
|
||||
|
||||
SubRip subtitle format with sequential numbering, `HH:MM:SS,mmm` timestamps, and text content.
|
||||
|
||||
```
|
||||
1
|
||||
00:00:00,000 --> 00:00:03,500
|
||||
The transcribed content goes here.
|
||||
```
|
||||
|
||||
##### `vtt`
|
||||
|
||||
WebVTT subtitle format with `WEBVTT` header and `HH:MM:SS.mmm` timestamps.
|
||||
|
||||
```
|
||||
WEBVTT
|
||||
|
||||
00:00:00.000 --> 00:00:03.500
|
||||
The transcribed content goes here.
|
||||
```
|
||||
|
||||
### POST /load
|
||||
|
||||
Loads a different model file on the server at runtime.
|
||||
|
||||
```bash
|
||||
curl -s http://macmini:8178/load \
|
||||
-F model="/path/to/model.bin"
|
||||
```
|
||||
|
||||
| Parameter | Type | Description |
|
||||
|-----------|------|-------------|
|
||||
| `model` | string | Path to the model file on the server. |
|
||||
|
||||
## Supported Audio Formats
|
||||
|
||||
The server accepts at least WAV (16-bit PCM) and MP3 files directly. If the server was started with `--convert`, it can use ffmpeg to handle additional formats (ogg, flac, m4a, etc.).
|
||||
|
||||
## Source
|
||||
|
||||
- Repository: https://github.com/ggml-org/whisper.cpp
|
||||
- Server docs: https://github.com/ggml-org/whisper.cpp/blob/master/examples/server/README.md
|
||||
Reference in New Issue
Block a user