Noema 1.5 2B MLX 4-bit

4-bit MLX release of Noema 1.5 2B, converted with MLX 0.31.2 and MLX-LM 0.31.3 using affine groupwise quantization with group size 64.

This is an open-weight release, not an open-source release. The weights are publicly downloadable, but no open-source license is granted with this repository at this time.

The measured effective storage rate is 4.503 bits per weight. Non-quantized tensors remain BF16. MTP is disabled, and the tokenizer uses <|im_end|> as EOS and <|endoftext|> as PAD.

Usage

pip install "mlx-lm==0.31.3"

mlx_lm.generate \
  --model NoemaAI-labs/Noema-1.5-2B-MLX-4bit \
  --prompt "Write a Python function that merges two sorted lists." \
  --temp 0 \
  --chat-template-config '{"enable_thinking":false}'

Non-thinking mode with greedy decoding is the recommended default. For harder reasoning tasks, use enable_thinking=true, temperature 1.0, top-p 0.95, and top-k 20, while enforcing an output limit.

The configuration advertises the Qwen3.5 backbone's native 262,144-token limit; Noema independently validated contexts only up to 24,576 tokens. Select context length according to available memory.

Benchmark results, training details, intended uses, and limitations are in the native model card. GGUF builds are available at Noema-1.5-2B-GGUF.

Downloads last month
13
Safetensors
Model size
2B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NoemaAI-labs/Noema-1.5-2B-MLX-4bit

Finetuned
Qwen/Qwen3.5-2B
Quantized
(4)
this model