TensorFold/Ling-3.0-flash-MLX-8bit
Text Generation • 124B • Updated • 244
Measured peak <192 GB; estimated 256GB fit.
Note Published M3 Studio peak: 132.36 GB, leaving about 124 GB nominal headroom on 256GB. See the card for measured speed and setup; longer contexts and concurrent requests need more memory.
Note Experimental 2-bit MLX with included standalone runtime. Tested on a 256 GiB M3 Ultra Studio: 166.69 GiB peak process RSS in short tests, with Engram tables on SSD; filesystem cache is additional. Up to 9.5 tok/s with MTP off. About 239 GB download. Not stock oMLX compatible; long context unvalidated.