Instructions to use NoemaAI-labs/Noema-1.5-2B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use NoemaAI-labs/Noema-1.5-2B-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("NoemaAI-labs/Noema-1.5-2B-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use NoemaAI-labs/Noema-1.5-2B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "NoemaAI-labs/Noema-1.5-2B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "NoemaAI-labs/Noema-1.5-2B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use NoemaAI-labs/Noema-1.5-2B-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "NoemaAI-labs/Noema-1.5-2B-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "NoemaAI-labs/Noema-1.5-2B-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "NoemaAI-labs/Noema-1.5-2B-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use NoemaAI-labs/Noema-1.5-2B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "NoemaAI-labs/Noema-1.5-2B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default NoemaAI-labs/Noema-1.5-2B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use NoemaAI-labs/Noema-1.5-2B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "NoemaAI-labs/Noema-1.5-2B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "NoemaAI-labs/Noema-1.5-2B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Noema 1.5 2B MLX 4-bit
4-bit MLX release of Noema 1.5 2B, converted with MLX 0.31.2 and MLX-LM 0.31.3 using affine groupwise quantization with group size 64.
This is an open-weight release, not an open-source release. The weights are publicly downloadable, but no open-source license is granted with this repository at this time.
The measured effective storage rate is 4.503 bits per weight. Non-quantized tensors remain BF16. MTP is disabled, and the tokenizer uses <|im_end|> as EOS and <|endoftext|> as PAD.
Usage
pip install "mlx-lm==0.31.3"
mlx_lm.generate \
--model NoemaAI-labs/Noema-1.5-2B-MLX-4bit \
--prompt "Write a Python function that merges two sorted lists." \
--temp 0 \
--chat-template-config '{"enable_thinking":false}'
Non-thinking mode with greedy decoding is the recommended default. For harder reasoning tasks, use enable_thinking=true, temperature 1.0, top-p 0.95, and top-k 20, while enforcing an output limit.
The configuration advertises the Qwen3.5 backbone's native 262,144-token limit; Noema independently validated contexts only up to 24,576 tokens. Select context length according to available memory.
Benchmark results, training details, intended uses, and limitations are in the native model card. GGUF builds are available at Noema-1.5-2B-GGUF.
- Downloads last month
- 13
4-bit
Model tree for NoemaAI-labs/Noema-1.5-2B-MLX-4bit
Base model
Qwen/Qwen3.5-2B-Base