Instructions to use EschaLabs/escha-runtime-qwen3moe with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use EschaLabs/escha-runtime-qwen3moe with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir escha-runtime-qwen3moe EschaLabs/escha-runtime-qwen3moe
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
RE writeup + one clarifying question on escham_reconstruct's per-slot LUT
Hi Escha team,
We've been working on a native Apple Silicon (MLX) port ofQwen3.6-35B-A3B-Escha-W2. The model loader, MoE routing, RoPE, attention,
Q3-Next GatedDeltaNet, and the fp8/int8 residual paths all pass numerical
parity against a Modal-hosted reference. The blocker is the escham_reconstruct
codebook — we've done enough reverse-engineering to be confident in some
structural properties, but have hit a wall on one specific question. Writing
up what we found so far, in case it's useful and to see if you're willing to
sanity-check one hypothesis.
What we've verified empirically (Modal A10G probes, escha==1.0.2+qwen3moe,
CUDA 12.8, ~$0.42 total spend):
- Op is EXACTLY linear across (bi, bj) tile blocks — superposition
|diff|=0 across all tile positions. A single op call on a(bi_max, bj_max, k_max)code tensor is a free test bed forbi_max × bj_max
distinct codebook queries. - K-layers are additive — across (k in 0..15) vs (k in 16..31) the
pair residual is exactly zero. - Within one K-layer, slots interact ONLY pairwise, at overlapping spatial
support — 3-way MöbiusR_3 = δ_ABC - Σδ_pairs + Σδ_solos = 0for all
sampled triples. 15 interacting slot-pairs per K-layer, structure derivable
from(row_class = k%4, col_class = k//4)with thek%4==2col-expansion
rule (details in our notes doc, linked below). - The per-slot LUT
δ_k(v)for a singleint16 vis a full 65536-entry
function that is nonlinear in v's bit representation at every order.
Point (4) is what's blocking us. We probed:
- Additive-bit hypothesis (v = Σbit_i · 2^i → δ_k(v) = Σ_{bit_i(v)==1} P[k,i]):
extracted P[k, i] cleanly via single-bit isolation; but 256 random-v tests
show max_abs_diff ≈ 34.83, rel_l2 ≈ 460%. Predicted magnitude grows linearly
in bit-count while actual saturates at 3-5. - 2-way bit Möbius: 120/120 pairs non-zero (max |cross| = 9.70).
- 3-way bit Möbius: 16/16 triples non-zero (max |R_3| = 7.51).
- Symmetry: δ(-v) is neither δ(v) nor -δ(v).
So the LUT really is 65536 distinct outputs per (k, K-layer). The extracted
compact form (compact.pkl, 120 MB fp16) empirically reproduces single-slot
op calls to bf16 rounding. Full-density accuracy needs the per-pair(v_i, v_j) → cross term, which our rank-4 factorization did NOT capture
(rel_err 180-300%).
One question — no pressure to answer:
Is the per-slot LUT structurally something like:
(a) A scaled Hadamard / DCT of a small learned inner codebook (dim ≪ 65536)?
(b) A lattice quantizer with per-position sign flips (E8 / Leech / product)?
(c) A pair-decomposed structure (v_hi, v_lo) → LUT_hi[v_hi] + LUT_lo[v_lo]
or similar split?
(d) Something else — the runtime just carries the full 65536-entry table?
If (d), we'll ship a Modal-extracted .safetensors bundle of the two LUTs
(~2 GB) alongside the port and call it done. If (a)-(c), a one-line hint
would let us compress properly. Happy to open a PR to your repo with an
accessor function once we know the shape.
What we're NOT asking for: any proprietary training or optimization
detail, or the CUDA kernel itself. Just the analytic form of the codebook
so downstream non-CUDA ports can compress it correctly.
Cost log, extraction scripts, and the full RE diagnosis:
- Notes:
docs/ESCHA_LAYOUT_NOTES.md§3, §3.1, §3.3, §3.4 in
https://github.com/Blaizzy/mlx-video/tree/escha-mlx-port - Probe scripts:
mlx_video/models/qwen3_5_moe_escha/codebooks/modal_*.py - Modal artifacts: on the
escha-codebooksModal Volume (dense per-slot LUTs
cb_K{2,3}.npy, compact.pkl, bit_decomp_probe_v1.pkl, bit_mobius_v1.pkl)
Thanks for the runtime — the linearity + slot-pair-only structure is genuinely
elegant, and MLX Mac users are eager for this model.
— KaedeTai
[link to PR #46]
Hi Kaede,
Thank you for the help. Very impressive analysis. Actually after some internal discussion, we decide to add an official support for MLX runtime and open source it. Stay tuned. Will update here once we release the code.
Hi @yzhwang ,
That's fantastic news — thank you for taking the time to consider this internally, and for the decision to ship official MLX support open source. That's the best possible outcome for the Apple Silicon community.
Really appreciate you engaging so constructively. Happy to help test, benchmark, or provide any Mac-side signal you need during the port. Feel free to ping me here or in a fresh discussion once early code is ready — I have a 128 GB M-series set up specifically for this.
Looking forward to the release!
Best,
Kaede
@yzhwang Thank you for shipping this open source — it's exactly the runtime path Apple Silicon users needed, and doing it in-house with the codec-authors' guarantees (bit-exact reference vectors, Metal kernels gated to goldens) is much cleaner than any community RE could have been. Good job!