GLM-5.3-Flash-NVFP4-FP8-Hybrid
This checkpoint combines NVFP4 routed experts with selected FP8 attention projections and dense/shared MLP weights for GLM-5.3-Flash. It was prepared for text serving with SGLang on NVIDIA Blackwell GPUs.
The sources are NVIDIA GLM-5.3-Flash-NVFP4, revision 09b04e5e74bca08ca8549fc736d4cdd8624bfde3, and Z.ai GLM-5.3-Flash, revision eb9eb208eb0d988989d07a6a12d0fdeb5f52574a.
Quantization
- Routed expert weights retain the NVIDIA checkpoint's NVFP4 representation.
- 442 BF16 attention projection and shared MLP modules were converted to E4M3FN FP8 using 128 脳 128 blocks and FP32 dequantization scales computed as block absmax / 448.
- Nine dense MLP modules in the first three layers use FP8 weights and scales from the fixed Z.ai revision above.
- Other tensors retain their source representation, including the KDA core and small
b_proj, indexer, MLAkv_b_proj, routers, mHC, normalization, embeddings, language head, vision, and MTP tensors.
The checkpoint contains 33 safetensors shards, approximately 197.8 GB in total. Download all shards and the accompanying configuration, tokenizer, and processor files.
Serving
The configured CUDA 13.0 environment is available as yilzhang/glm-5.3-flash-nvfp4-fp8-hybrid:cu130. It includes PyTorch 2.13.0+cu130, the tested SGLang GLM hybrid runtime, FlashInfer 0.7.0 (commit b5804cfc), and sglang-kernel 0.4.7. The native hybrid loader and its FlashInfer CUTLASS argument compatibility patch are included; no external source-code mount is needed.
A conservative four-GPU serving example. Replace /absolute/path/to/checkpoint with the downloaded checkpoint directory. The checkpoint is not included in the image.
docker run --rm --gpus '"device=0,1,2,3"' --ipc=host \
-p 30000:30000 \
-v /absolute/path/to/checkpoint:/model:ro \
-e SGLANG_PP_SKIP_PURE_CHUNKED_OUTPUT_COMM=1 \
yilzhang/glm-5.3-flash-nvfp4-fp8-hybrid:cu130 \
python3 -m sglang.launch_server \
--model-path /model --served-model-name glm-5.3-flash \
--host 0.0.0.0 --port 30000 --tp-size 1 --pp-size 4 \
--quantization fp8 --moe-runner-backend flashinfer_cutlass \
--disable-shared-experts-fusion --disable-overlap-schedule \
--attention-backend dsa --dsa-prefill-backend flashinfer_sparse_mla \
--dsa-decode-backend flashinfer_sparse_mla --linear-attn-backend triton \
--kv-cache-dtype fp8_e4m3 --mamba-ssm-dtype bfloat16 \
--pp-max-micro-batch-size 2 --pp-async-batch-depth 1 \
--context-length 16384 --chunked-prefill-size 4096 --max-prefill-tokens 4096 \
--max-running-requests 8 --max-total-tokens 65536 --max-mamba-cache-size 32 \
--mem-fraction-static 0.88 --cuda-graph-backend-prefill disabled \
--cuda-graph-backend-decode full --cuda-graph-max-bs-decode 8 \
--reasoning-parser glm45 --tool-call-parser glm47
First startup compiles kernels and creates runtime caches. This example leaves MTP/speculative decoding disabled.
The fp8 argument selects the hybrid loader using the checkpoint metadata; routed experts still execute with NVFP4. Other serving options depend on hardware and workload. Loading and text generation were tested on four RTX 6000D GPUs with TP1/PP4 and MTP disabled. MTP/speculative decoding compatibility has not been verified.
Example request using the tested low-reasoning setting:
curl http://localhost:30000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"glm-5.3-flash","messages":[{"role":"user","content":"17 + 25 = ?"}],"temperature":0,"max_tokens":1024,"reasoning_effort":"low","skip_special_tokens":false}'
Keep skip_special_tokens=false so the configured reasoning parser receives the delimiters needed to separate reasoning from the final answer. The checkpoint template supports reasoning_effort=low or high; it does not use enable_thinking.
Evaluation scope
In a paired full GSM8K test, this checkpoint scored 1267/1319 (96.06%), versus 1253/1319 (95.00%) for the source NVIDIA NVFP4 checkpoint. Both used the same SGLang runtime, five training examples for few-shot prompting, temperature 0, maximum output length 8192, and reasoning_effort=low. Seven questions that the source answered correctly regressed.
Additional partial evaluations did not establish overall accuracy preservation: IFBench prompt-loose accuracy fell from 80.00% to 77.62% on 126 jointly completed prompts across five seeds; MMMU-Pro accuracy was 22.70% for both checkpoints on 467 jointly completed samples. These incomplete, nonrandom subsets are not full benchmark results, and multimodal quality remains unverified. These are local comparisons, not reproductions of NVIDIA's published scores.
License and attribution
The upstream models are released under MIT. Retain the included LICENSE, including the original Z.AI copyright and permission notice. Credit for the base model belongs to Z.ai and for the source NVFP4 quantization to NVIDIA. This repository distributes a community modification.
- Downloads last month
- 46