MTP works after adding layer 45 to the quantization exclusion lists
#15 opened 1 day ago
by
soonh1618
Why don't the commits in this repo have the "verified" badge?
#14 opened 1 day ago
by
jesang
MTP layer is unquantized but missing from exclude_modules, so vLLM cannot load it for speculative decoding
#13 opened 1 day ago
by
brokenlander
Error in function 'TllmGenFmhaRunner' at .sglang/lib/python3.12/site-packages/flashinfer/data/include/flashinfer/trtllm/fmha/fmhaRunner.cuh:37: Unsupported architecture
#12 opened 8 days ago
by
shakhizat
fix: declare NextN (MTP) layer 45 as BF16 in quantization ignore lists
#11 opened 11 days ago
by
lucifer1004
Serving Recipe: GLM-5.3-Flash-NVFP4 with vLLM at 1M Context on Dual DGX Spark
👍 5
#8 opened 16 days ago
by
Pilcothink
[FIX] Reasoning Loop on Trivial Task same as FP8
🔥 1
#7 opened 17 days ago
by
voves
ignore / exclude_modules omits the MTP layer (45): checkpoint fails to load with speculative decoding
❤️ 4
1
#6 opened 18 days ago
by
jon1012
Where's sglang? Do you have a plan for adding Sglang support?
4
#2 opened 19 days ago
by
IlyaTers