clef-oQ4e

This model was quantized using oQ mixed-precision quantization.

A mixed-precision MLX checkpoint of Cloudflare/clef, a decision model that answers typed questions (noul, choice, score) about text and images with a probability for each option. It includes the Qwen3.8-27B language backbone, the vision encoder and the joint schema head.

The joint schema head (joint_head.safetensors, joint_head_config.json) is copied unchanged from the source. oMLX serves this checkpoint through POST /v1/systemone once decision model support (jundot/omlx#4315) is available.

The effective average is 5.264 bits/weight for the complete checkpoint.

Quantization and Bit Distribution

  • 4-bit affine quantization, group size 64, for most language-backbone projections and the token embeddings.
  • 5-bit affine quantization, group size 64, for projections boosted by measured layer sensitivity.
  • 6-bit affine quantization, group size 64, for the output projection (lm_head).
  • BF16 for the vision encoder, the joint schema head, norms and the remaining small tensors.

The oQ4e allocation uses measured layer sensitivity and importance-matrix calibration (128 samples x 512 tokens, built-in code and multilingual set).

Storage format Logical weights Share of total Tensor storage Effective bits/weight
Affine 4-bit, group 64 22.986B 81.152% 12.042 GiB 4.50
Affine 5-bit, group 64 2.636B 9.305% 1.687 GiB 5.50
Affine 6-bit, group 64 1.271B 4.489% 0.962 GiB 6.50
BF16 1.432B 5.055% 2.667 GiB 16.00
Total 28.325B 100% 17.358 GiB 5.264

Storage Breakdown

Component Logical weights Storage
Language backbone 27.736B 17.461 GB / 16.262 GiB
Vision encoder 0.461B 0.921 GB / 0.858 GiB
Joint schema head 0.128B 0.256 GB / 0.239 GiB
Total 28.325B 18.638 GB / 17.358 GiB

License

Apache-2.0, following the source model. See LICENSE.

Downloads last month
-
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jundot/clef-oQ4e

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(20)
this model