pipenetwork commited on
Commit
2e901d4
·
verified ·
1 Parent(s): 8660ae2

Add quantization cross-links + collection

Browse files
Files changed (1) hide show
  1. README.md +12 -0
README.md CHANGED
@@ -28,6 +28,18 @@ projector and multi-token-prediction heads are not included). The model is a
28
  3 layers dense), with per-head QK-norm, partial RoPE, Gemma-style RMSNorm and the
29
  SwiGLU-OAI activation.
30
 
 
 
 
 
 
 
 
 
 
 
 
 
31
  ## Attention / context note
32
 
33
  MiniMax Sparse Attention (MSA) is implemented here as **full causal attention**.
 
28
  3 layers dense), with per-head QK-norm, partial RoPE, Gemma-style RMSNorm and the
29
  SwiGLU-OAI activation.
30
 
31
+ ## Quantizations
32
+
33
+ Part of the [**MiniMax-M3 MLX** collection](https://huggingface.co/collections/pipenetwork/minimax-m3-mlx-6a2d7776d4e2a69aad841516).
34
+
35
+ | Variant | Size | Notes |
36
+ |---|---|---|
37
+ | [8-bit](https://huggingface.co/pipenetwork/MiniMax-M3-MLX-8bit) | ~453 GB | near-lossless |
38
+ | [6-bit](https://huggingface.co/pipenetwork/MiniMax-M3-MLX-6bit) | ~346 GB | high quality |
39
+ | [4-bit](https://huggingface.co/pipenetwork/MiniMax-M3-MLX-4bit) | ~240 GB | balanced default |
40
+ | **3-bit** (this repo) | ~186 GB | smallest |
41
+ | [mixed-3_6bit](https://huggingface.co/pipenetwork/MiniMax-M3-MLX-mixed-3_6bit) | ~191 GB | experts@3-bit, attn/embeds/router@6-8-bit · best quality-per-GB |
42
+
43
  ## Attention / context note
44
 
45
  MiniMax Sparse Attention (MSA) is implemented here as **full causal attention**.