Trellis
Safetensors
naive_n05_flash
exl3
exllamav3
3.5bpw
vllm
Mixture of Experts
long-context
dgx-spark
naive-n0.5-flash
custom_code
3-bit
Instructions to use doth4580/Naive-N0.5-Flash-EXL3-3.5bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Trellis
How to use doth4580/Naive-N0.5-Flash-EXL3-3.5bpw with Trellis:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Model card: link the 2x DGX Spark serving recipe
Browse files
README.md
CHANGED
|
@@ -28,7 +28,9 @@ An **EXL3 (ExLlamaV3 trellis) 3.5 bpw** quantization of **Naive-N0.5-Flash** —
|
|
| 28 |
original**, which is what lets it fit across **two NVIDIA DGX Spark (GB10)** machines.
|
| 29 |
|
| 30 |
Quantized with the [exllamav3](https://github.com/turboderp-org/exllamav3) tooling, adapted for this
|
| 31 |
-
architecture
|
|
|
|
|
|
|
| 32 |
|
| 33 |
## Serving
|
| 34 |
|
|
|
|
| 28 |
original**, which is what lets it fit across **two NVIDIA DGX Spark (GB10)** machines.
|
| 29 |
|
| 30 |
Quantized with the [exllamav3](https://github.com/turboderp-org/exllamav3) tooling, adapted for this
|
| 31 |
+
architecture, run on **8×H100** for speed. It is served on two Sparks with the recipe in
|
| 32 |
+
**[sf-stav/Naive-N0.5-Flash-EXL3-2x-DGX-Spark](https://github.com/sf-stav/Naive-N0.5-Flash-EXL3-2x-DGX-Spark)**
|
| 33 |
+
— see [Serving](#serving) below.
|
| 34 |
|
| 35 |
## Serving
|
| 36 |
|