FLUX.2-klein-4B GGUF

Self-quantized GGUF builds of black-forest-labs/FLUX.2-klein-4B for low-VRAM / edge inference. Runs on an 8GB GPU, or CPU-only.

These weights were produced with stable-diffusion.cpp (--mode convert) and are intended for use with stable-diffusion.cpp / GGUF-aware runtimes. FLUX.2 is by Black Forest Labs; klein 4B is Apache-2.0.

Files & precision

Only the heavy MLP / attention / linear weights are quantized (via --tensor-type-rules); norms and embeddings are kept at higher precision, which preserves quality at low bit-depth.

File Precision (quant) Bits (weights) Size
flux-2-klein-4b-Q2_K.gguf Q2_K ~2-bit K-quant 1.49 GB
flux-2-klein-4b-Q4_K.gguf Q4_K ~4-bit K-quant 2.29 GB
flux-2-klein-4b-Q5_K.gguf Q5_K ~5-bit K-quant 2.72 GB

Quantization rule used:

^.*(_mlp\.(0|2)|_attn\.(proj|qkv)|\.linear(1|2))\.weight=<qtype>

Which to pick: Q4_K is the recommended balance for 8GB. Q5_K for best fidelity if you have headroom. Q2_K for the tightest RAM / smallest storage (softer detail).

Required companion files (not included here)

Usage (stable-diffusion.cpp)

sd-cli \
  --diffusion-model flux-2-klein-4b-Q4_K.gguf \
  --vae flux2-vae.safetensors \
  --llm Qwen3-4B-Q4_K_M.gguf \
  -p "a red fox in an autumn forest, golden hour" \
  --cfg-scale 1.0 --steps 4 --sampling-method euler \
  -W 768 -H 768 --offload-to-cpu --diffusion-fa --vae-tiling -o out.png

klein is step-distilled: use 4 steps and CFG 1.0 (euler). A minimal pure C++ front-end is available as sd-flux2.cpp.

Reference latency (RTX A2000 8GB, 768Γ—768, 4 steps, end-to-end)

Quant Gen time
Q4_K ~14.5 s
Q5_K ~15.8 s
Q2_K ~16.0 s

Acknowledgements

  • Base model: FLUX.2 [klein] 4B β€” Black Forest Labs (Apache-2.0)
  • Quantization + inference engine: stable-diffusion.cpp (leejet), built on ggml
Downloads last month
2,371
GGUF
Model size
4B params
Architecture
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for adithya-balaji/FLUX.2-klein-4B-GGUF

Quantized
(55)
this model

Space using adithya-balaji/FLUX.2-klein-4B-GGUF 1