Naive-N0.5-Flash — EXL3 3.5 bpw

Original model: NaiveAI/Naive-N0.5-Flash by the NaiveAI Team, MIT License. This artifact is quantized from the public FP8 release NaiveAI/Naive-N0.5-Flash-FP8. It contains modified weights (EXL3 trellis quantization); the original model is © its authors. Tokenizer and chat template are unchanged from the original.

An EXL3 (ExLlamaV3 trellis) 3.5 bpw quantization of Naive-N0.5-Flash — 309B parameters, 15.5B active, native 1M-token context. The pack is 137 GB (128 GiB) against ~315 GB for the FP8 original, which is what lets it fit across two NVIDIA DGX Spark (GB10) machines.

Quantized with the exllamav3 tooling, adapted for this architecture, run on 8×H100 for speed. It is served on two Sparks with the recipe in sf-stav/Naive-N0.5-Flash-EXL3-2x-DGX-Spark — see Serving below.

Serving

Serve it with the recipe: sf-stav/Naive-N0.5-Flash-EXL3-2x-DGX-Spark.

EXL3 weights do not load in stock vLLM or transformers, and this is a uniform fractional-bitrate (3.5 bpw) pack, so it needs the exact pinned stack the recipe installs on both nodes: vLLM 0.29.0

cp config/cluster.env.example config/cluster.env && $EDITOR config/cluster.env
bash scripts/install.sh      # on BOTH nodes
bash scripts/download.sh     # on BOTH nodes — fetches this pack + the DSpark drafter
bash scripts/start.sh        # Ray head + worker, then the OpenAI-compatible API on port 8000

download.sh pulls this repository and NaiveAI/Naive-N0.5-Flash-FP8-Draft for speculative decoding. The server is OpenAI-compatible at http://<head-ip>:8000/v1, model id naive-n0.5-flash; loading takes about 7.5 minutes. Reasoning (<think>) and tool calls parse correctly with the recipe's chat template and parsers — see its README for the details.

Specifications

Base model NaiveAI/Naive-N0.5-Flash (FP8 release as the direct source)
Architecture MoE, 309B total / 15.5B active, 48 layers, hidden 4096, 64 heads / 4 KV heads, head_dim 192
Attention hybrid 39 SWA layers (window 128) + 9 DSA layers (top-2048 selection), GQA4, no full attention
MoE 256 routed experts, 8 per token, expert intermediate 2048
Context / vocab 1,048,576 tokens / 152,576
Quantization EXL3 trellis, mul1 codebook — uniform 3.5 bpw, 6-bit lm_head, embeddings unquantized, calibration 250×2048
Format / size exl3 (--quantization exl3) — 137 GB (128 GiB), 24 shards
License MIT (inherited from the original model)

Files

File Notes
model-00001..00024-of-00024.safetensors the quantized weights
model.safetensors.index.json, quantization_config.json tensor index and the full EXL3 metadata
config.json carries the EXL3 quantization_config incl. the non_routed_exl3 map (the vLLM-rewritten pack config)
config.json.native, model.safetensors.index.json.native the equivalent native ExLlamaV3 pack metadata
modeling_naive_n05_flash.py, configuration_naive_n05_flash.py model code, loaded via trust_remote_code
tokenizer + chat_template.jinja, generation_config.json unchanged from the original

Notes

  • Beyond 2,048 tokens the DSA layers are approximate. The model's 9 DSA layers attend to the top-2,048 tokens picked by a learned indexer; the recipe runs them as dense attention, which is exact up to 2,048 tokens but past that sees more than the reference model does. Long-context answers may differ from the original; the 1M native context needs a sparse attention backend that is not implemented yet.
  • This is an early recipe, not a tuned one. Measured decode on 2× DGX Spark: 34.9 tok/s on pure code, ~23.9 mixed, ~19.5 reasoning (23.5 without speculation). Performance tables, the profile, and the limitations are in the recipe's README.
  • One request at a time, bf16 KV only, and keep the Sparks otherwise idle (GPU_MEM_UTIL=0.80).

License and credits

Downloads last month
4,005
Safetensors
Model size
68B params
Tensor type
F16
·
I16
·
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for doth4580/Naive-N0.5-Flash-EXL3-3.5bpw

Quantized
(1)
this model

Space using doth4580/Naive-N0.5-Flash-EXL3-3.5bpw 1