KrynexAI 25M Instruct

KrynexAI 25M is a compact, hybrid autoregressive chat language model. It combines the efficiency of state-space models with the proven performance of attention layers, optimized using a custom Muon + AdamW optimizer split. ## Model Details - Architecture: Hybrid Mamba2 / Transformer - Block Pattern: 3 Mamba2 blocks : 1 Attention block (repeating) - Parameters: 25,000,000 (25M) - Hidden Dimension: 608 - Layers: 8 (6 Mamba2, 2 Attention) - Vocab Size: 2,048 (Custom Byte-Level BPE) - Context Length: 2048 - Pretraining Tokens: ~25,000,000,000 (25 Billion) - SFT Tokens: 250,000,000 (250 Million) - Optimizer: Muon (for 2D hidden weights) + AdamW (for embeddings, norms, and scalars) - Precision: fp32 master weights with bf16 autocast

Dataset Sources

The base model was pretrained on a 25B token subset of the following datasets:

Dataset Token Allocation Share
FineWeb-Edu 7.50 billion 30%
DCLM 5.00 billion 20%
Cosmopedia-v2 3.75 billion 15%
FineMath-4+ 3.75 billion 15%
FinePhrase 3.00 billion 12%
NPset 2.00 billion 8%

Evaluation Notes

  • PIQA, ARC-Easy, ARC-Challenge, and HellaSwag were evaluated on their respective test splits.
  • ArithMark-2.0 and ArithMark-3.0 were evaluated on their train splits.
  • Results were obtained using zero-shot multiple-choice evaluation.
  • The model was additionally fine-tuned using supervised fine-tuning (SFT).
  • Original author: Pebble (25M), It's fine-tuned version.

Usage

To run the model for text generation, you will need to install the required dependencies. The included Mamba2 implementation relies on CUDA/Triton kernels and is intended to run on a CUDA-enabled GPU. Ampere-class GPUs or newer are recommended.

Note: The model uses custom architecture code, so you must pass trust_remote_code=True when loading both the tokenizer and the model.

License & Attribution

Licensed under Apache License 2.0. Based on original work by Pebble developers.

Downloads last month
430
Safetensors
Model size
24.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support