doth4580 commited on
Commit
bf772c7
·
verified ·
1 Parent(s): 8067fca

Model card: link the 2x DGX Spark serving recipe

Browse files
Files changed (1) hide show
  1. README.md +3 -1
README.md CHANGED
@@ -28,7 +28,9 @@ An **EXL3 (ExLlamaV3 trellis) 3.5 bpw** quantization of **Naive-N0.5-Flash** —
28
  original**, which is what lets it fit across **two NVIDIA DGX Spark (GB10)** machines.
29
 
30
  Quantized with the [exllamav3](https://github.com/turboderp-org/exllamav3) tooling, adapted for this
31
- architecture (the patch that adds it ships with the recipe), run on **8×H100** for speed.
 
 
32
 
33
  ## Serving
34
 
 
28
  original**, which is what lets it fit across **two NVIDIA DGX Spark (GB10)** machines.
29
 
30
  Quantized with the [exllamav3](https://github.com/turboderp-org/exllamav3) tooling, adapted for this
31
+ architecture, run on **8×H100** for speed. It is served on two Sparks with the recipe in
32
+ **[sf-stav/Naive-N0.5-Flash-EXL3-2x-DGX-Spark](https://github.com/sf-stav/Naive-N0.5-Flash-EXL3-2x-DGX-Spark)**
33
+ — see [Serving](#serving) below.
34
 
35
  ## Serving
36