File size: 3,661 Bytes
7c726de 3610c39 7c726de 2738107 7c726de ea17953 7c726de 39e5d04 7c726de 39e5d04 7c726de 39e5d04 7c726de ea17953 7c726de ea17953 7c726de ea17953 7c726de ea17953 7c726de ea17953 49a46f0 ea17953 7c726de | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 | ---
license: mit
language:
- multilingual
- en
- ru
tags:
- whisper
- gguf
- quantized
- speech-recognition
- rust
- candle
base_model:
- openai/whisper-tiny
pipeline_tag: automatic-speech-recognition
---
# WHISPER-TINY - GGUF Quantized Models
Quantized versions of [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny) in GGUF format.
## Directory Structure
```
tiny/
βββ whisper-tiny-q*.gguf # Candle-compatible GGUF models (root)
βββ model-tiny-q80.gguf # Candle-compatible legacy naming (q8_0 format)
βββ config-tiny.json # Model configuration for Candle
βββ tokenizer-tiny.json # Tokenizer for Candle
βββ whisper.cpp/ # whisper.cpp-compatible models
βββ whisper-tiny-q*.gguf
```
### Format Compatibility
- **Root directory** (`whisper-tiny-*.gguf`): Use with **Candle** (Rust ML framework)
- Tensor names include `model.` prefix (e.g., `model.encoder.conv1.weight`)
- Requires `config-tiny.json` and `tokenizer-tiny.json`
- **whisper.cpp/** directory: Use with **whisper.cpp** (C++ implementation)
- Tensor names without `model.` prefix (e.g., `encoder.conv1.weight`)
- Compatible with whisper.cpp CLI tools
- Both directories contain `.gguf` files, not `.bin` files
## Available Formats
**Q5_0** is the "last good" option. Lower quantizations (Q4, Q3, Q2) lead to a sharp loss of quality and meaningless output.
| Format | Quality | Use Case |
|--------| ---------|----------|
| q2_k | Smallest | Extreme compression |
| q3_k | Small | Mobile devices |
| q4_0 | Small | Legacy compatibility |
| q4_k | Small | - |
| q4_1 | Small | Legacy with bias |
| q5_0 | Good | Legacy compatibility |
| q5_k | Good | Good quality |
| q5_1 | Very Good | Legacy with bias |
| q6_k | Excellent | Near-lossless |
| q8_0 | Excellent | **Recommended for production** Minimal loss, benchmarking |
## Usage
### With Candle (Rust)
**Command line example:**
```bash
# Run Candle Whisper with local quantized model
cargo run --example whisper --release -- \
--features symphonia \
--quantized \
--model tiny \
--model-id oxide-lab/whisper-tiny-GGUF
```
### With whisper.cpp (C++)
```bash
# Use models from whisper.cpp/ subdirectory
./whisper.cpp/build/bin/whisper-cli \
--model models/openai/tiny/whisper.cpp/whisper-tiny-q4_k.gguf \
--file audio.wav
```
### Recommended Format
For most use cases, we recommend **q4_k** format as it provides the best balance of:
- Size reduction (~65% smaller)
- Quality (minimal degradation)
- Speed (faster inference than higher quantizations)
## Quantization Details
- **Source Model**: [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny)
- **Quantization Methods**:
- **Candle GGUF** (root directory): Python-based quantization. Directly PyTorch β GGUF
- Adds `model.` prefix to tensor names for Candle compatibility
- **whisper.cpp GGML** (whisper.cpp/ subdirectory): whisper-quantize tool
- Uses original tensor names without prefix
- **Format**: GGUF (GGML Universal Format) for both directories
- **Total Formats**: 10 quantization levels (q2_k through q8_0)
## License
Same as the original Whisper model (MIT License).
## Citation
```bibtex
@misc{radford2022whisper,
doi = {10.48550/ARXIV.2212.04356},
url = {https://arxiv.org/abs/2212.04356},
author = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
title = {Robust Speech Recognition via Large-Scale Weak Supervision},
publisher = {arXiv},
year = {2022},
copyright = {arXiv.org perpetual, non-exclusive license}
}
``` |