Naive-N0.5-Flash — EXL3 3.5 bpw
Original model: NaiveAI/Naive-N0.5-Flash by the NaiveAI Team, MIT License. This artifact is quantized from the public FP8 release NaiveAI/Naive-N0.5-Flash-FP8. It contains modified weights (EXL3 trellis quantization); the original model is © its authors. Tokenizer and chat template are unchanged from the original.
An EXL3 (ExLlamaV3 trellis) 3.5 bpw quantization of Naive-N0.5-Flash — 309B parameters,
15.5B active, native 1M-token context. The pack is 137 GB (128 GiB) against ~315 GB for the FP8
original, which is what lets it fit across two NVIDIA DGX Spark (GB10) machines.
Quantized with the exllamav3 tooling, adapted for this architecture, run on 8×H100 for speed. It is served on two Sparks with the recipe in sf-stav/Naive-N0.5-Flash-EXL3-2x-DGX-Spark — see Serving below.
Serving
Serve it with the recipe: sf-stav/Naive-N0.5-Flash-EXL3-2x-DGX-Spark.
EXL3 weights do not load in stock vLLM or transformers, and this is a uniform fractional-bitrate
(3.5 bpw) pack, so it needs the exact pinned stack the recipe installs on both nodes: vLLM 0.29.0
- vcruz305/vllm-exl3 + a
vcruz305/exllamav3 build with fractional-K kernels + the
recipe's
vllm_naive_n05model plugin (attention TP2, experts EP2, Ray over RoCE).
cp config/cluster.env.example config/cluster.env && $EDITOR config/cluster.env
bash scripts/install.sh # on BOTH nodes
bash scripts/download.sh # on BOTH nodes — fetches this pack + the DSpark drafter
bash scripts/start.sh # Ray head + worker, then the OpenAI-compatible API on port 8000
download.sh pulls this repository and
NaiveAI/Naive-N0.5-Flash-FP8-Draft for
speculative decoding. The server is OpenAI-compatible at http://<head-ip>:8000/v1, model id
naive-n0.5-flash; loading takes about 7.5 minutes. Reasoning (<think>) and tool calls parse
correctly with the recipe's chat template and parsers — see its README for the details.
Specifications
| Base model | NaiveAI/Naive-N0.5-Flash (FP8 release as the direct source) |
| Architecture | MoE, 309B total / 15.5B active, 48 layers, hidden 4096, 64 heads / 4 KV heads, head_dim 192 |
| Attention | hybrid 39 SWA layers (window 128) + 9 DSA layers (top-2048 selection), GQA4, no full attention |
| MoE | 256 routed experts, 8 per token, expert intermediate 2048 |
| Context / vocab | 1,048,576 tokens / 152,576 |
| Quantization | EXL3 trellis, mul1 codebook — uniform 3.5 bpw, 6-bit lm_head, embeddings unquantized, calibration 250×2048 |
| Format / size | exl3 (--quantization exl3) — |
| License | MIT (inherited from the original model) |
Files
| File | Notes |
|---|---|
model-00001..00024-of-00024.safetensors |
the quantized weights |
model.safetensors.index.json, quantization_config.json |
tensor index and the full EXL3 metadata |
config.json |
carries the EXL3 quantization_config incl. the non_routed_exl3 map (the vLLM-rewritten pack config) |
config.json.native, model.safetensors.index.json.native |
the equivalent native ExLlamaV3 pack metadata |
modeling_naive_n05_flash.py, configuration_naive_n05_flash.py |
model code, loaded via trust_remote_code |
tokenizer + chat_template.jinja, generation_config.json |
unchanged from the original |
Notes
- Beyond 2,048 tokens the DSA layers are approximate. The model's 9 DSA layers attend to the top-2,048 tokens picked by a learned indexer; the recipe runs them as dense attention, which is exact up to 2,048 tokens but past that sees more than the reference model does. Long-context answers may differ from the original; the 1M native context needs a sparse attention backend that is not implemented yet.
- This is an early recipe, not a tuned one. Measured decode on 2× DGX Spark:
34.9 tok/s on pure code, ~23.9 mixed, ~19.5 reasoning (23.5 without speculation). Performance tables, the profile, and the limitations are in the recipe's README. - One request at a time, bf16 KV only, and keep the Sparks otherwise idle (
GPU_MEM_UTIL=0.80).
License and credits
- Original model: NaiveAI/Naive-N0.5-Flash —
© the NaiveAI Team, MIT; a copy
ships here as
LICENSE. Built on Xiaomi's MiMo-V2.5 and DeepSeek Sparse Attention. - This artifact: EXL3 3.5 bpw quantization produced with exllamav3 via the vcruz305 fork. Weights are modified; tokenizer and chat template are not.
- Recipe: sf-stav/Naive-N0.5-Flash-EXL3-2x-DGX-Spark
(MIT). Note that
vllm-exl3is AGPL-3.0 — relevant if you expose the server to others.
- Downloads last month
- 4,005
Model tree for doth4580/Naive-N0.5-Flash-EXL3-3.5bpw
Base model
NaiveAI/Naive-N0.5-Flash-FP8