banner

KIKOCIS // 8-BIT GGUF // MEASURED AGAINST BF16
 BF16 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  token_embd โ”€โ”€ bf16 โ”‚   KLD 0.000616
  blk.0..41  โ”€โ”€ q8_0 โ”œโ”€โ”€โ–ถ top-1 98.67 %
  output     โ”€โ”€ bf16 โ”‚   PPL +0.0109
 โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
MiniCPM5-2B ยท Q8_0-XL
llama ยท 2B ยท 128K context ยท โˆ’11 % KLD vs plain Q8_0
FORMAT
GGUF
SIZE
3.18 / 2.68 GB
ARCH
llama, 42 layers
CONTEXT
131,072
QUANTS
Q8_0-XL ยท Q8_0
VALIDATION
KLD / PPL / top-1
RUNS ON
llama.cpp ยท Ollama
LICENSE
Apache-2.0

MiniCPM5-2B โ€” Q8_0-XL + Q8_0 GGUF

8-bit GGUFs of OpenBMB's MiniCPM5-2B, each one measured against the BF16 original. Q8_0-XL keeps the token embeddings and the output head in BF16 and puts every transformer block in Q8_0: 11 % lower KLD and 26 % less perplexity drift than a plain Q8_0, for +0.5 GB. Runs in ~3โ€“4 GB of RAM at 8K context.

This is OpenBMB's model, our quantization. OpenBMB also publishes their own GGUFs; we measured their Q8_0 alongside ours (table below) and it is equivalent to our plain Q8_0. What this repo adds is the XL variant and the numbers.

โœ… Recommended files

Use case File Notes
Best fidelity at 8 bits MiniCPM5-2B-Q8_0-XL.gguf Lowest KLD and PPL drift here. Recommended.
Smallest 8-bit MiniCPM5-2B-Q8_0.gguf Standard Q8_0, 0.5 GB lighter.

๐Ÿ“ฆ Files

Quant Bits (blocks / embed + head) File size
Q8_0-XL 8 / 16 3.18 GB
Q8_0 8 / 8 2.68 GB

The embedding table and output head are unusually large for a 2B model (130,560-token vocabulary ร— 2,048 = 1 GB in BF16 between the two), which is why keeping them in BF16 costs 0.5 GB and is where most of the remaining 8-bit error lives.

๐Ÿ“Š Metrics โ€” fidelity vs the BF16 reference

KLD (Kullbackโ€“Leibler divergence, nats): how far the quant's next-token distribution drifts from BF16. Top-1 = how often the quant's most likely token is the same as BF16's. PPL ฮ” = perplexity above BF16 (11.196).

File Size GB PPL PPL ฮ” KLD mean KLD median KLD p95 KLD p99 Top-1 match
BF16 reference 5.04 11.196 0 0 0 0 0 100 %
Q8_0-XL 3.18 11.207 +0.0109 0.000616 0.000375 0.00193 0.00444 98.67 %
Q8_0 2.68 11.211 +0.0147 0.000692 0.000451 0.00209 0.00463 98.57 %
Q8_0 โ€” OpenBMB's official file, for reference 2.68 11.211 +0.0147 0.000692 0.000450 0.00209 0.00464 98.56 %

Standard errors: ยฑ0.000005 on mean KLD, ยฑ0.003 on PPL ฮ”, ยฑ0.05 points on top-1. The XL gap is well outside them.

Machine-readable: metrics/quant-summary-with-kld.json / .csv. Per-file detail in reports/.

๐Ÿ“ˆ Chart

fidelity vs size

๐Ÿงฎ Will it fit?

Memory โ‰ˆ file size + KV cache. MiniCPM5-2B uses grouped-query attention (2 KV heads), so long context is cheap:

Context KV cache (F16) Q8_0-XL total Q8_0 total
8K 0.35 GB ~3.6 GB ~3.1 GB
32K 1.4 GB ~4.6 GB ~4.1 GB
128K (native) 5.6 GB ~8.9 GB ~8.4 GB

๐Ÿš€ How to run it

# llama.cpp โ€” the GGUF carries OpenBMB's chat template (tools + reasoning)
llama-server -m MiniCPM5-2B-Q8_0-XL.gguf -c 32768 --jinja --temp 1.0 --top-p 0.95

# Ollama โ€” Modelfiles for 8K / 32K / 128K are included
ollama run hf.co/KikoCis/MiniCPM5-2B-GGUF:Q8_0-XL
ollama create minicpm5 -f Modelfile-32k

Sampling: temperature 1.0, top_p 0.95 โ€” OpenBMB's own defaults (generation_config.json). Reasoning: the model thinks before answering; Ollama returns that in a separate thinking field, llama-server in reasoning_content with --jinja. Tools: native function calling through the embedded chat template โ€” use the tools parameter of the OpenAI-compatible API.

โš ๏ธ Good to know

  • At 8 bits both files are already very close to BF16; the XL gain is real but small in absolute terms. Pick XL if you have 0.5 GB to spare, plain Q8_0 otherwise.
  • No weights were changed: this is a faithful quantization of OpenBMB's release.
  • We have not run agentic or task benchmarks on these files โ€” the numbers above measure fidelity to the original model, not its capability. For capability, see OpenBMB's card.

๐Ÿ“ Evaluation methodology

  • What: KLD, perplexity and top-1 agreement of each GGUF against the BF16 GGUF converted from the same weights.
  • Data: wikitext-2 test split (wiki.test.raw), first 64 chunks of 2,048 tokens.
  • Tool: llama-perplexity --kl-divergence-base (BF16 logits) then --kl-divergence per file.
  • Source: openbmb/MiniCPM5-2B at revision f97400052a43d642bbc6e9975e2397e3ae6a6b52.
  • Load test: both files loaded and answered correctly in llama.cpp and Ollama (chat template, reasoning split).
  • Date: 2026-10-10.

๐Ÿ” Provenance & reproducibility

๐Ÿ—’๏ธ Changelog

  • 2026-10-10 โ€” v1: Q8_0-XL and Q8_0, with KLD/PPL/top-1 vs BF16 and OpenBMB's Q8_0 as a reference row.

๐Ÿ“š Credit & license

Model, weights and training: ยฉ OpenBMB โ€” openbmb/MiniCPM5-2B. Quantization, measurements and Modelfiles: KikoCis. Apache-2.0, same as upstream.

Downloads last month
-
GGUF
Model size
3B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for KikoCis/MiniCPM5-2B-GGUF

Quantized
(104)
this model