Lythri

Lythri-7B-A4B-Q4_0_8-GGUF

This repository contains the Q4_0_8 quantization of Lythri-7B-A4B. Q4_0_8 is a mixed-precision quantization proposed in the Lythri technical report: critical tensors (attention projections, ffn_down, PLE gates) are kept at Q8_0 while the remaining tensors use Q4_0. Both formats support native SIMD dot-product acceleration, delivering Q6_K-level quality with 1.16–1.76× faster prefill on models ≥4B.

For model details, evaluation results and limitations, see the original model card.

Why Q4_0_8

Lythri models are highly sensitive to quantization. Standard 4-bit formats (Q4_0, Q4_K_M) cause noticeable quality loss, while Q6_K preserves quality but is slower due to its multi-level scale structure. Q4_0_8 provides a middle ground:

Q4_0 Q6_K Q4_0_8
Quality ✗ Degraded ✓ Good ✓ Near Q6_K
HW Accel ✓ Yes ✗ No ✓ Yes
Speed Fast Slow Fast

Usage

llama.cpp

llama-cli -hf Lythri/Lythri-7B-A4B-Q4_0_8-GGUF -m <file>.gguf --temp 0.7 --top-p 0.9 --top-k 64 --repeat-penalty 1.05

Ollama

ollama run hf.co/Lythri/Lythri-7B-A4B-Q4_0_8-GGUF

LM Studio

Search for Lythri/Lythri-7B-A4B-Q4_0_8-GGUF in the model browser and download.

Q4_0_8 Quantization Guide

Q4_0_8 uses Q4_0 as the base type and overrides critical tensors to Q8_0, achieving Q6_K-level quality with 1.16–1.76× faster prefill via native SIMD dot-product acceleration.

Build llama.cpp

git clone --depth 1 https://github.com/ggerganov/llama.cpp
cd llama.cpp
pip install gguf
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc) --target llama-quantize

Convert & Quantize

# Step 1: Convert HF model to F16 GGUF
python3 convert_hf_to_gguf.py /path/to/model \
    --outfile model-F16.gguf --outtype f16

# Step 2: Quantize with Q4_0_8 rules
# Standard Transformer (all architectures):
build/bin/llama-quantize \
    --tensor-type "attn_q=q8_0" \
    --tensor-type "attn_k=q8_0" \
    --tensor-type "attn_v=q8_0" \
    --tensor-type "attn_output=q8_0" \
    --tensor-type "ffn_down=q8_0" \
    model-F16.gguf model-Q4_0_8.gguf q4_0

# For FFN+PLE architectures (e.g. Gemma 4 E2B/E4B), add:
#   --tensor-type "inp_gate=q8_0"
#   --tensor-type "proj=q8_0"

Tensor Assignment Rules

Tensor Quant Rationale
attn_q, attn_k Q8_0 Softmax sensitivity
attn_v, attn_output Q8_0 Residual stream writer
ffn_down Q8_0 Residual stream writer
ffn_gate, ffn_up Q4_0 Layer-internal, SiLU-attenuated
inp_gate* Q8_0 Sub-network activation control
proj* Q8_0 AltUp residual projection

*Extended rules for FFN+PLE architectures only.

Citation

@techreport{li2026lythri,
  title       = {Lythri Technical Report},
  author      = {Li, Jiawen},
  year        = {2026},
  institution = {Zenodo},
  doi         = {10.5281/zenodo.23179311},
  url         = {https://doi.org/10.5281/zenodo.23179311}
}

License

Lythri is built on Gemma 4 and is released under the Apache License 2.0.

Downloads last month
247
GGUF
Model size
7B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lythri/Lythri-7B-A4B-Q4_0_8-GGUF

Quantized
(4)
this model

Collection including Lythri/Lythri-7B-A4B-Q4_0_8-GGUF