File size: 3,661 Bytes
7c726de
 
 
 
 
 
 
 
 
 
 
 
3610c39
7c726de
 
2738107
7c726de
 
 
 
ea17953
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7c726de
 
 
39e5d04
 
7c726de
 
 
 
39e5d04
 
 
 
 
 
7c726de
39e5d04
7c726de
 
 
ea17953
7c726de
ea17953
7c726de
ea17953
 
 
 
 
 
 
7c726de
ea17953
 
 
 
 
 
 
7c726de
 
 
 
 
 
 
 
 
 
 
 
ea17953
49a46f0
ea17953
 
 
 
 
7c726de
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
---
license: mit
language:
- multilingual
- en
- ru
tags:
- whisper
- gguf
- quantized
- speech-recognition
- rust
- candle
base_model:
- openai/whisper-tiny
pipeline_tag: automatic-speech-recognition
---

# WHISPER-TINY - GGUF Quantized Models

Quantized versions of [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny) in GGUF format.

## Directory Structure

```
tiny/
β”œβ”€β”€ whisper-tiny-q*.gguf       # Candle-compatible GGUF models (root)
β”œβ”€β”€ model-tiny-q80.gguf        # Candle-compatible legacy naming (q8_0 format)
β”œβ”€β”€ config-tiny.json           # Model configuration for Candle
β”œβ”€β”€ tokenizer-tiny.json        # Tokenizer for Candle
└── whisper.cpp/               # whisper.cpp-compatible models
    └── whisper-tiny-q*.gguf

```

### Format Compatibility

- **Root directory** (`whisper-tiny-*.gguf`): Use with **Candle** (Rust ML framework)
  - Tensor names include `model.` prefix (e.g., `model.encoder.conv1.weight`)
  - Requires `config-tiny.json` and `tokenizer-tiny.json`
  
- **whisper.cpp/** directory: Use with **whisper.cpp** (C++ implementation)
  - Tensor names without `model.` prefix (e.g., `encoder.conv1.weight`)
  - Compatible with whisper.cpp CLI tools
  - Both directories contain `.gguf` files, not `.bin` files

## Available Formats

**Q5_0** is the "last good" option. Lower quantizations (Q4, Q3, Q2) lead to a sharp loss of quality and meaningless output.

| Format |  Quality | Use Case |
|--------| ---------|----------|
| q2_k |   Smallest | Extreme compression |
| q3_k |  Small | Mobile devices |
| q4_0 |   Small | Legacy compatibility |
| q4_k |   Small | - |
| q4_1 |  Small | Legacy with bias |
| q5_0 |  Good | Legacy compatibility |
| q5_k |  Good | Good quality |
| q5_1 |  Very Good | Legacy with bias |
| q6_k |   Excellent | Near-lossless |
| q8_0 |   Excellent | **Recommended for production** Minimal loss, benchmarking |

## Usage

### With Candle (Rust)

**Command line example:**
```bash
# Run Candle Whisper with local quantized model
cargo run --example whisper --release -- \
  --features symphonia \
  --quantized \
  --model tiny \
  --model-id oxide-lab/whisper-tiny-GGUF 
```

### With whisper.cpp (C++)

```bash
# Use models from whisper.cpp/ subdirectory
./whisper.cpp/build/bin/whisper-cli \
  --model models/openai/tiny/whisper.cpp/whisper-tiny-q4_k.gguf \
  --file audio.wav
```

### Recommended Format

For most use cases, we recommend **q4_k** format as it provides the best balance of:
- Size reduction (~65% smaller)
- Quality (minimal degradation)
- Speed (faster inference than higher quantizations)

## Quantization Details

- **Source Model**: [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny)
- **Quantization Methods**:
  - **Candle GGUF** (root directory): Python-based quantization. Directly PyTorch β†’ GGUF
    - Adds `model.` prefix to tensor names for Candle compatibility
  - **whisper.cpp GGML** (whisper.cpp/ subdirectory): whisper-quantize tool
    - Uses original tensor names without prefix
- **Format**: GGUF (GGML Universal Format) for both directories
- **Total Formats**: 10 quantization levels (q2_k through q8_0)

## License

Same as the original Whisper model (MIT License).

## Citation

```bibtex
@misc{radford2022whisper,
  doi = {10.48550/ARXIV.2212.04356},
  url = {https://arxiv.org/abs/2212.04356},
  author = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},
  title = {Robust Speech Recognition via Large-Scale Weak Supervision},
  publisher = {arXiv},
  year = {2022},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
```