ThinkingCap-Qwen3.8-27B — NInfer Blackwell RC1

Text-only NInfer export of bottlecapai/ThinkingCap-Qwen3.8-27B, tuned for NVIDIA Blackwell GPUs.

RC1 status: the target quantization profile and native ThinkingCap MTP integration are validated. Proposal-head row tuning, NVFP4 KV-cache validation, final R5 MTP confirmation, real long coding/agentic traces, and untouched holdout/task evaluation are still pending.

Primary artifact

thinkingcap-d48-o44-i53-g64-mtp-nativeq8.ninfer

Field Value
File size 17,501,292,292 bytes
SHA256 b1b31d42a12ef848e69f82c8ef5af1c3fedb2ca462df19cbd4b7340c8517c178
NInfer commit 9e163eee4b8acec21ab0ac765107b6a3f287b217
Text NVFP4 parents 209
Native MTP layers 1
Native MTP projection format Q8_G32_FP16
Proposal head Q4_G64_FP16, [131072, 5120]
Chat template qwen3_8_sharp.jinja
Chat-template SHA256 cdff39fb26b60dc90faa292e726655c6b21f62db497846e02e4c4bbab942a84a

The artifact contains text + native MTP. Vision is intentionally not included in this RC.

Quantization profile

Frozen target: D48 + O44 + I53 + G64

D48 — MLP down NVFP4

0,1,4,5,6,7,9,13,14,15,16,17,19,20,22,23,24,28,29,30,31,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,53,54,55,56,58,59,60,61,62

O44 — output NVFP4

GDN: 0,5,9,12,13,14,16,17,18,20,22,28,32,33,34,36,37,40,41,42,44,45,46,48,49,50,52,53,54,56,57,58,60,61,62

Attention: 11,19,27,31,39,51,55,59,63

I53 — input-parent NVFP4

GDN: 0,1,2,4,5,6,8,9,12,14,16,17,21,22,25,28,29,30,32,33,36,37,38,40,42,44,45,46,48,49,50,52,53,54,56,57,58,60,61,62

Attention: 7,11,15,19,23,35,39,43,47,51,55,59,63

G64 — all MLP gate/up parents

All 64 MLP gate/up parents use NVFP4/W4A4.

Total text NVFP4 parent count: 209.

Quality validation

Development/calibration corpus:

  • ninfer-ppl-en-code-v1
  • 12 streams
  • 782,535 scored tokens
  • context 4096
  • stride 2048
  • KV BF16

This is a development/calibration corpus, not an untouched holdout.

Model Mean NLL PPL
Original NInfer baseline 1.4704195156673645 4.3510600961173225
D48+O44+I53 1.466704419741171 4.334925479860586
D48+O44+I53+G64 1.469385010577333 4.346561229752157

G64 vs original baseline: -0.103397% PPL on this development corpus.

Do not interpret the negative delta as proof that the quantized model is intrinsically better; it only means no aggregate PPL regression was measured here.

G64 domain results

Domain Mean NLL PPL
english_long_form 2.082259891 8.022579
english_reference 1.809620891 6.108131
ninfer_code 0.511001014 1.666959
overall 1.469385011 4.346561

Target performance

NVIDIA RTX PRO 4000 Blackwell 24 GB, physical GPU0, ~70 W, FP8 KV, prefill chunk 1024.

G64 full prefill curve (R3)

Prompt tok/s
512 2040.68
2,048 2289.70
4,096 2228.75
7,680 2119.80
8,192 2097.74
16,384 1940.63
32,768 1694.77
65,536 1397.20

G64 standard R3

Workload Result
pp512 2049.11 tok/s
pp2048 2292.72 tok/s
pp8192 2155.03 tok/s
tg256 21.09 tok/s
tg1024 19.98 tok/s
pp2048+tg512 2067.23 / 19.92 tok/s
pp8192+tg512 1978.31 / 19.68 tok/s

Versus D48+O44+I53:

  • mean prefill: +116.709%
  • mean decode: +0.487%

Native ThinkingCap MTP

This RC uses ThinkingCap's own native one-layer MTP from the source checkpoint.

Validated Q8 projection objects:

  • attention input parent [14336, 5120]
  • MLP gate/up [34816, 5120]
  • MTP input projection [5120, 10240]
  • attention output [5120, 6144]
  • MLP down [5120, 17408]

Proposal head:

  • Q4_G64_FP16
  • [131072, 5120]

MTP K=0..5 sweep — GPU1 145 W, FP8 KV, optimized head, R3

K mean decode tok/s vs K0 mean acceptance mean accepted length
0 33.33 +0.00% — —
1 54.80 +64.57% 81.88% 1.819
2 69.04 +107.43% 79.36% 2.587
3 73.86 +122.14% 70.08% 3.099
4 87.84 +164.31% 68.63% 3.738
5 91.96 +176.85% 65.39% 4.258

K5 is the current best mean R3 result, but not universally optimal. For tg256, K2/K3 are faster.

Workload K0 K2 K4 K5
tg256 33.68 54.31 50.40 45.64
tg1024 33.57 66.28 78.22 79.96
pp2048+tg512 33.28 77.14 108.35 116.45
pp8192+tg512 32.78 78.41 114.40 125.79

These benchmark-corpus acceptance rates should not be assumed to represent all real chat/coding workloads.

MTP0 overhead

At identical 16K settings, target-only MTP0 vs the text-only G64 artifact measured:

  • mean prefill delta: -2.173%
  • mean decode delta: -0.942%
  • mixed prefill/decode: approximately neutral
  • loaded target weight bytes: identical (16,679,684,352)
  • sequence/workspace reservations: identical

Pure-PP R3 rows showed about -3.6% in some isolated cases; final R5 confirmation remains pending.

Context / KV

Current reported performance uses FP8 KV.

The RC has run target-only through 65,536 context. NVFP4 KV-cache validation, including 130,048-context fitting, is in progress and is not yet claimed as finalized.

Chat template

Embedded: tools/chat_templates/qwen3_8_sharp.jinja

SHA256: cdff39fb26b60dc90faa292e726655c6b21f62db497846e02e4c4bbab942a84a

Reproducibility

NInfer commit: 9e163eee4b8acec21ab0ac765107b6a3f287b217

All cumulative candidates were rebuilt from the original source checkpoint. Existing .ninfer artifacts were not requantized.

RC1 status

Accepted/frozen:

  • D48
  • O44
  • I53
  • G64
  • native one-layer ThinkingCap MTP integration
  • Q8 native-MTP baseline
  • optimized Q4 proposal-head integration
  • Sharp chat template
  • 209 text NVFP4 parents

Pending:

  • final R5 MTP-window selection
  • proposal-head row-count sweep
  • NVFP4 KV-cache validation
  • real long coding/agentic speculative traces
  • untouched holdout/task evaluation

Compact vocabulary variants

This repository also contains experimental compact-vocabulary NInfer builds for English/coding/agent workloads.

The original Qwen3.8 tokenizer exposes 248,077 public tokens with an aligned NInfer embedding/output-head shape of 248,320 physical rows.

These compact variants are not frequency-only pruning. Tokens are removed only after Unicode/script classification, BPE dependency closure, workload protection, exact tokenizer parity checks, teacher-output audits, and runtime A/B validation.

Available / validated variants

Variant Public vocab Physical rows Removed vs. full Status
Full vocabulary 248,077 248,320 0 Reference
Script cut 211,375 211,456 36,702 (14.79%) Production-canary tested
Script + CJK cut 147,590 147,712 100,487 (40.51%) Validated; production-canary ready

211k script cut

The first compact build removes script-only tokens outside the target English/coding/agent workload while protecting mixed Latin/ASCII tokens, all special and added tokens, workload-observed tokens, and every BPE dependency needed by retained tokens.

Public vocabulary        211,375
Physical vocabulary      211,456
Removed tokens            36,702
Padding rows                  81

The 211k artifact has been used as a production canary for the target workload without observed functional issues.

Repository path:

scriptcut-211375/

147k combined Script + CJK cut

The second stage applies the same conservative procedure to CJK-only vocabulary (Han, Hiragana, Katakana and Hangul).

The raw CJK scan found 65,691 candidate tokens. After workload protection and BPE dependency closure, 63,785 CJK tokens remained removable.

Existing script-cut removals            36,702
CJK removals                            63,785
                                       -------
Total removed                          100,487

Original public vocabulary             248,077
Combined public vocabulary             147,590
Physical vocabulary                    147,712
Padding rows                                122
Removed                                40.51%

The compact artifact keeps the same 131,072-row indexed MTP proposal head. Vocabulary pruning reduces the token embedding and full target LM head; it does not shrink the MTP proposal vocabulary.

147k artifact

  • Filename: thinkingcap-mtp-nvfp4-gateup-combined-147590-phys147712.ninfer
  • Size: 16,308,912,900 bytes (15.19 GiB)
  • SHA-256: dd37ea4c576a7c25065389cc4e5f371037c1c1c86898bff6163b85990935ab84
  • Stored objects: 1,170
  • Public vocabulary: 147,590
  • Physical vocabulary: 147,712
  • Proposal rows: 131,072

Compared with the full artifact:

Full artifact      17,412,164,100 bytes
147k artifact      16,308,912,900 bytes
Reduction           1,103,251,200 bytes (~1.03 GiB / 6.34%)

Suggested repository path:

combined-147590/

Vocabulary-pruning safety procedure

A compact vocabulary is accepted only after all of the following checks.

  1. Conservative Unicode classification
    Only tokens belonging exclusively to explicitly targeted writing systems are considered removable. Mixed Latin/ASCII tokens, special tokens and added tokens remain protected.

  2. BPE dependency closure
    Every retained token is traced through the BPE merge graph. If a removal candidate is needed as an ancestor of a retained token, it is restored.

  3. Workload protection
    The coding/agent corpus contains 2,772,957 original tokens across 26 files. For CJK, 1,397 directly observed token IDs were protected; including BPE dependencies the protected CJK set grew to 1,906 IDs.

  4. Exact tokenizer parity
    Retained tokens are compacted together with a deterministic new_id -> original_id mapping.

  5. Teacher-output audit
    Full-model outputs are checked against the complete removal set.

  6. Runtime A/B validation
    Full-vocabulary and compact-vocabulary artifacts are compared with identical prompts, templates and decode settings.

Fidelity results

Test Result
Coding/agent corpus files 26 / 26 exact
Corpus tokens compared 2,772,957
Corpus token-sequence mismatches 0
Real Sharp agent prompt 24 messages / 28 tools
Full prompt tokens 26,035
Compact prompt tokens 26,035
Prompt sequence after ID remapping Exact
Tool-result follow-up 26 messages / 28 tools
Follow-up sequence after ID remapping Exact
Teacher tool-call audit 202 tokens / 0 removed-token hits
Teacher follow-up audit 732 tokens / 0 removed-token hits
Full vs. 147k greedy NoSpec response EXACT MATCH

An earlier frequency/ranking-based ~180k experiment was rejected because retained text could tokenize differently after pruning. Exact token-sequence preservation for the protected workload is therefore a hard requirement for the current method.

147k runtime results

The 147k build was tested with the external Sharp chat template, MTP speculative decoding, --draft-tokens 5, --lm-head-draft, NVFP4 KV and a 32,768-token KV capacity.

Five-run MTP benchmark

Same request, same runtime build, same GPU and decode configuration:

Metric Full 248k Combined 147k Difference
Mean decode 85.14 tok/s 93.70 tok/s +10.05%
Minimum decode 83.47 tok/s 92.60 tok/s
Maximum decode 86.04 tok/s 94.59 tok/s
Mean MTP acceptance 61.11% 67.24% +6.12 pp
Prompt tokens 26,215 26,215 identical

The slowest 147k run (92.60 tok/s) was still faster than the fastest Full run (86.04 tok/s) in this five-run comparison.

These measurements are workload- and hardware-specific and are not universal performance guarantees.

VRAM / weight loading

Observed engine startup:

211k script cut    ~15.8 GiB weights
147k combined cut  ~15.2 GiB weights

The 147k server reported approximately 6.93 GiB free VRAM after startup in the tested 32k/NVFP4/MTP configuration.

NInfer runtime requirement

Compact artifacts require explicit NInfer Q8 and LinearTopK support for their physical vocabulary dimensions.

248,320 physical / 248,077 valid   Full vocabulary
211,456 physical / 211,375 valid   Script cut
147,712 physical / 147,590 valid   Script + CJK cut

For the 147k artifact the runtime must support:

Q8LinearGeometry<147712, 5120>

and LinearTopK must use:

physical rows = 147712
valid rows    = 147590

A stock NInfer runtime without these compact shapes will reject the compact LM head.

Chat template

Validation and production-canary testing use the external Qwen Sharp chat template.

For matching agent/tool behavior, use the same Sharp template with NInfer.

Scope and limitations

The compact vocabularies are optimized for the tested English / programming / shell / configuration / agent-tool workload.

They are not intended as universally lossless replacements for the original multilingual vocabulary. Input outside that protected scope can tokenize differently and should be evaluated separately before use.

The safety claim is deliberately narrow: for the protected and tested workload, the compact tokenizer preserves the original token sequences, and the validated non-speculative runtime test produced an exact Full-vs-Compact response match.

License

This is a quantized/converted derivative of bottlecapai/ThinkingCap-Qwen3.8-27B.

ThinkingCap is published under PolyForm Small Business 1.0.0 plus BottleCap's personal-use grant. BottleCap describes upstream Qwen materials as Apache-2.0.

Before redistributing this artifact, verify that your redistribution is permitted by the current source-model license/access terms. Preserve the source LICENSE and NOTICE files in this repository.

Credits

Base model: bottlecapai/ThinkingCap-Qwen3.8-27B

This repository does not claim ownership of the original model weights.

Downloads last month
322
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Schestex/ThinkingCap-Qwen3.8-27B-NInfer

Base model

Qwen/Qwen3.8-27B
Quantized
(22)
this model