- ThinkingCap-Qwen3.8-27B — NInfer Blackwell RC1
- Primary artifact
- Quantization profile
- Quality validation
- Target performance
- Native ThinkingCap MTP
- MTP0 overhead
- Context / KV
- Chat template
- Reproducibility
- RC1 status
- Compact vocabulary variants
- Vocabulary-pruning safety procedure
- 147k runtime results
- NInfer runtime requirement
- Chat template
- Scope and limitations
- License
- Credits
- Primary artifact
ThinkingCap-Qwen3.8-27B — NInfer Blackwell RC1
Text-only NInfer export of bottlecapai/ThinkingCap-Qwen3.8-27B, tuned for NVIDIA Blackwell GPUs.
RC1 status: the target quantization profile and native ThinkingCap MTP integration are validated. Proposal-head row tuning, NVFP4 KV-cache validation, final R5 MTP confirmation, real long coding/agentic traces, and untouched holdout/task evaluation are still pending.
Primary artifact
thinkingcap-d48-o44-i53-g64-mtp-nativeq8.ninfer
| Field | Value |
|---|---|
| File size | 17,501,292,292 bytes |
| SHA256 | b1b31d42a12ef848e69f82c8ef5af1c3fedb2ca462df19cbd4b7340c8517c178 |
| NInfer commit | 9e163eee4b8acec21ab0ac765107b6a3f287b217 |
| Text NVFP4 parents | 209 |
| Native MTP layers | 1 |
| Native MTP projection format | Q8_G32_FP16 |
| Proposal head | Q4_G64_FP16, [131072, 5120] |
| Chat template | qwen3_8_sharp.jinja |
| Chat-template SHA256 | cdff39fb26b60dc90faa292e726655c6b21f62db497846e02e4c4bbab942a84a |
The artifact contains text + native MTP. Vision is intentionally not included in this RC.
Quantization profile
Frozen target: D48 + O44 + I53 + G64
D48 — MLP down NVFP4
0,1,4,5,6,7,9,13,14,15,16,17,19,20,22,23,24,28,29,30,31,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,53,54,55,56,58,59,60,61,62
O44 — output NVFP4
GDN:
0,5,9,12,13,14,16,17,18,20,22,28,32,33,34,36,37,40,41,42,44,45,46,48,49,50,52,53,54,56,57,58,60,61,62
Attention:
11,19,27,31,39,51,55,59,63
I53 — input-parent NVFP4
GDN:
0,1,2,4,5,6,8,9,12,14,16,17,21,22,25,28,29,30,32,33,36,37,38,40,42,44,45,46,48,49,50,52,53,54,56,57,58,60,61,62
Attention:
7,11,15,19,23,35,39,43,47,51,55,59,63
G64 — all MLP gate/up parents
All 64 MLP gate/up parents use NVFP4/W4A4.
Total text NVFP4 parent count: 209.
Quality validation
Development/calibration corpus:
ninfer-ppl-en-code-v1- 12 streams
- 782,535 scored tokens
- context 4096
- stride 2048
- KV BF16
This is a development/calibration corpus, not an untouched holdout.
| Model | Mean NLL | PPL |
|---|---|---|
| Original NInfer baseline | 1.4704195156673645 |
4.3510600961173225 |
| D48+O44+I53 | 1.466704419741171 |
4.334925479860586 |
| D48+O44+I53+G64 | 1.469385010577333 |
4.346561229752157 |
G64 vs original baseline: -0.103397% PPL on this development corpus.
Do not interpret the negative delta as proof that the quantized model is intrinsically better; it only means no aggregate PPL regression was measured here.
G64 domain results
| Domain | Mean NLL | PPL |
|---|---|---|
| english_long_form | 2.082259891 |
8.022579 |
| english_reference | 1.809620891 |
6.108131 |
| ninfer_code | 0.511001014 |
1.666959 |
| overall | 1.469385011 |
4.346561 |
Target performance
NVIDIA RTX PRO 4000 Blackwell 24 GB, physical GPU0, ~70 W, FP8 KV, prefill chunk 1024.
G64 full prefill curve (R3)
| Prompt | tok/s |
|---|---|
| 512 | 2040.68 |
| 2,048 | 2289.70 |
| 4,096 | 2228.75 |
| 7,680 | 2119.80 |
| 8,192 | 2097.74 |
| 16,384 | 1940.63 |
| 32,768 | 1694.77 |
| 65,536 | 1397.20 |
G64 standard R3
| Workload | Result |
|---|---|
| pp512 | 2049.11 tok/s |
| pp2048 | 2292.72 tok/s |
| pp8192 | 2155.03 tok/s |
| tg256 | 21.09 tok/s |
| tg1024 | 19.98 tok/s |
| pp2048+tg512 | 2067.23 / 19.92 tok/s |
| pp8192+tg512 | 1978.31 / 19.68 tok/s |
Versus D48+O44+I53:
- mean prefill: +116.709%
- mean decode: +0.487%
Native ThinkingCap MTP
This RC uses ThinkingCap's own native one-layer MTP from the source checkpoint.
Validated Q8 projection objects:
- attention input parent
[14336, 5120] - MLP gate/up
[34816, 5120] - MTP input projection
[5120, 10240] - attention output
[5120, 6144] - MLP down
[5120, 17408]
Proposal head:
Q4_G64_FP16[131072, 5120]
MTP K=0..5 sweep — GPU1 145 W, FP8 KV, optimized head, R3
| K | mean decode tok/s | vs K0 | mean acceptance | mean accepted length |
|---|---|---|---|---|
| 0 | 33.33 |
+0.00% |
— | — |
| 1 | 54.80 |
+64.57% |
81.88% |
1.819 |
| 2 | 69.04 |
+107.43% |
79.36% |
2.587 |
| 3 | 73.86 |
+122.14% |
70.08% |
3.099 |
| 4 | 87.84 |
+164.31% |
68.63% |
3.738 |
| 5 | 91.96 |
+176.85% |
65.39% |
4.258 |
K5 is the current best mean R3 result, but not universally optimal. For tg256, K2/K3 are faster.
| Workload | K0 | K2 | K4 | K5 |
|---|---|---|---|---|
| tg256 | 33.68 |
54.31 |
50.40 |
45.64 |
| tg1024 | 33.57 |
66.28 |
78.22 |
79.96 |
| pp2048+tg512 | 33.28 |
77.14 |
108.35 |
116.45 |
| pp8192+tg512 | 32.78 |
78.41 |
114.40 |
125.79 |
These benchmark-corpus acceptance rates should not be assumed to represent all real chat/coding workloads.
MTP0 overhead
At identical 16K settings, target-only MTP0 vs the text-only G64 artifact measured:
- mean prefill delta:
-2.173% - mean decode delta:
-0.942% - mixed prefill/decode: approximately neutral
- loaded target weight bytes: identical (
16,679,684,352) - sequence/workspace reservations: identical
Pure-PP R3 rows showed about -3.6% in some isolated cases; final R5 confirmation remains pending.
Context / KV
Current reported performance uses FP8 KV.
The RC has run target-only through 65,536 context. NVFP4 KV-cache validation, including 130,048-context fitting, is in progress and is not yet claimed as finalized.
Chat template
Embedded:
tools/chat_templates/qwen3_8_sharp.jinja
SHA256:
cdff39fb26b60dc90faa292e726655c6b21f62db497846e02e4c4bbab942a84a
Reproducibility
NInfer commit:
9e163eee4b8acec21ab0ac765107b6a3f287b217
All cumulative candidates were rebuilt from the original source checkpoint. Existing .ninfer artifacts were not requantized.
RC1 status
Accepted/frozen:
- D48
- O44
- I53
- G64
- native one-layer ThinkingCap MTP integration
- Q8 native-MTP baseline
- optimized Q4 proposal-head integration
- Sharp chat template
- 209 text NVFP4 parents
Pending:
- final R5 MTP-window selection
- proposal-head row-count sweep
- NVFP4 KV-cache validation
- real long coding/agentic speculative traces
- untouched holdout/task evaluation
Compact vocabulary variants
This repository also contains experimental compact-vocabulary NInfer builds for English/coding/agent workloads.
The original Qwen3.8 tokenizer exposes 248,077 public tokens with an aligned NInfer embedding/output-head shape of 248,320 physical rows.
These compact variants are not frequency-only pruning. Tokens are removed only after Unicode/script classification, BPE dependency closure, workload protection, exact tokenizer parity checks, teacher-output audits, and runtime A/B validation.
Available / validated variants
| Variant | Public vocab | Physical rows | Removed vs. full | Status |
|---|---|---|---|---|
| Full vocabulary | 248,077 | 248,320 | 0 | Reference |
| Script cut | 211,375 | 211,456 | 36,702 (14.79%) | Production-canary tested |
| Script + CJK cut | 147,590 | 147,712 | 100,487 (40.51%) | Validated; production-canary ready |
211k script cut
The first compact build removes script-only tokens outside the target English/coding/agent workload while protecting mixed Latin/ASCII tokens, all special and added tokens, workload-observed tokens, and every BPE dependency needed by retained tokens.
Public vocabulary 211,375
Physical vocabulary 211,456
Removed tokens 36,702
Padding rows 81
The 211k artifact has been used as a production canary for the target workload without observed functional issues.
Repository path:
scriptcut-211375/
147k combined Script + CJK cut
The second stage applies the same conservative procedure to CJK-only vocabulary (Han, Hiragana, Katakana and Hangul).
The raw CJK scan found 65,691 candidate tokens. After workload protection and BPE dependency closure, 63,785 CJK tokens remained removable.
Existing script-cut removals 36,702
CJK removals 63,785
-------
Total removed 100,487
Original public vocabulary 248,077
Combined public vocabulary 147,590
Physical vocabulary 147,712
Padding rows 122
Removed 40.51%
The compact artifact keeps the same 131,072-row indexed MTP proposal head. Vocabulary pruning reduces the token embedding and full target LM head; it does not shrink the MTP proposal vocabulary.
147k artifact
- Filename:
thinkingcap-mtp-nvfp4-gateup-combined-147590-phys147712.ninfer - Size: 16,308,912,900 bytes (15.19 GiB)
- SHA-256:
dd37ea4c576a7c25065389cc4e5f371037c1c1c86898bff6163b85990935ab84 - Stored objects: 1,170
- Public vocabulary: 147,590
- Physical vocabulary: 147,712
- Proposal rows: 131,072
Compared with the full artifact:
Full artifact 17,412,164,100 bytes
147k artifact 16,308,912,900 bytes
Reduction 1,103,251,200 bytes (~1.03 GiB / 6.34%)
Suggested repository path:
combined-147590/
Vocabulary-pruning safety procedure
A compact vocabulary is accepted only after all of the following checks.
Conservative Unicode classification
Only tokens belonging exclusively to explicitly targeted writing systems are considered removable. Mixed Latin/ASCII tokens, special tokens and added tokens remain protected.BPE dependency closure
Every retained token is traced through the BPE merge graph. If a removal candidate is needed as an ancestor of a retained token, it is restored.Workload protection
The coding/agent corpus contains 2,772,957 original tokens across 26 files. For CJK, 1,397 directly observed token IDs were protected; including BPE dependencies the protected CJK set grew to 1,906 IDs.Exact tokenizer parity
Retained tokens are compacted together with a deterministicnew_id -> original_idmapping.Teacher-output audit
Full-model outputs are checked against the complete removal set.Runtime A/B validation
Full-vocabulary and compact-vocabulary artifacts are compared with identical prompts, templates and decode settings.
Fidelity results
| Test | Result |
|---|---|
| Coding/agent corpus files | 26 / 26 exact |
| Corpus tokens compared | 2,772,957 |
| Corpus token-sequence mismatches | 0 |
| Real Sharp agent prompt | 24 messages / 28 tools |
| Full prompt tokens | 26,035 |
| Compact prompt tokens | 26,035 |
| Prompt sequence after ID remapping | Exact |
| Tool-result follow-up | 26 messages / 28 tools |
| Follow-up sequence after ID remapping | Exact |
| Teacher tool-call audit | 202 tokens / 0 removed-token hits |
| Teacher follow-up audit | 732 tokens / 0 removed-token hits |
| Full vs. 147k greedy NoSpec response | EXACT MATCH |
An earlier frequency/ranking-based ~180k experiment was rejected because retained text could tokenize differently after pruning. Exact token-sequence preservation for the protected workload is therefore a hard requirement for the current method.
147k runtime results
The 147k build was tested with the external Sharp chat template, MTP speculative
decoding, --draft-tokens 5, --lm-head-draft, NVFP4 KV and a 32,768-token KV
capacity.
Five-run MTP benchmark
Same request, same runtime build, same GPU and decode configuration:
| Metric | Full 248k | Combined 147k | Difference |
|---|---|---|---|
| Mean decode | 85.14 tok/s | 93.70 tok/s | +10.05% |
| Minimum decode | 83.47 tok/s | 92.60 tok/s | |
| Maximum decode | 86.04 tok/s | 94.59 tok/s | |
| Mean MTP acceptance | 61.11% | 67.24% | +6.12 pp |
| Prompt tokens | 26,215 | 26,215 | identical |
The slowest 147k run (92.60 tok/s) was still faster than the fastest Full run (86.04 tok/s) in this five-run comparison.
These measurements are workload- and hardware-specific and are not universal performance guarantees.
VRAM / weight loading
Observed engine startup:
211k script cut ~15.8 GiB weights
147k combined cut ~15.2 GiB weights
The 147k server reported approximately 6.93 GiB free VRAM after startup in the tested 32k/NVFP4/MTP configuration.
NInfer runtime requirement
Compact artifacts require explicit NInfer Q8 and LinearTopK support for their physical vocabulary dimensions.
248,320 physical / 248,077 valid Full vocabulary
211,456 physical / 211,375 valid Script cut
147,712 physical / 147,590 valid Script + CJK cut
For the 147k artifact the runtime must support:
Q8LinearGeometry<147712, 5120>
and LinearTopK must use:
physical rows = 147712
valid rows = 147590
A stock NInfer runtime without these compact shapes will reject the compact LM head.
Chat template
Validation and production-canary testing use the external Qwen Sharp chat template.
For matching agent/tool behavior, use the same Sharp template with NInfer.
Scope and limitations
The compact vocabularies are optimized for the tested English / programming / shell / configuration / agent-tool workload.
They are not intended as universally lossless replacements for the original multilingual vocabulary. Input outside that protected scope can tokenize differently and should be evaluated separately before use.
The safety claim is deliberately narrow: for the protected and tested workload, the compact tokenizer preserves the original token sequences, and the validated non-speculative runtime test produced an exact Full-vs-Compact response match.
License
This is a quantized/converted derivative of bottlecapai/ThinkingCap-Qwen3.8-27B.
ThinkingCap is published under PolyForm Small Business 1.0.0 plus BottleCap's personal-use grant. BottleCap describes upstream Qwen materials as Apache-2.0.
Before redistributing this artifact, verify that your redistribution is permitted by the current source-model license/access terms. Preserve the source LICENSE and NOTICE files in this repository.
Credits
Base model: bottlecapai/ThinkingCap-Qwen3.8-27B
This repository does not claim ownership of the original model weights.
- Downloads last month
- 322