How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf prithivMLmods/CogEvol-4B-GGUF:
# Run inference directly in the terminal:
llama cli -hf prithivMLmods/CogEvol-4B-GGUF:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf prithivMLmods/CogEvol-4B-GGUF:
# Run inference directly in the terminal:
llama cli -hf prithivMLmods/CogEvol-4B-GGUF:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf prithivMLmods/CogEvol-4B-GGUF:
# Run inference directly in the terminal:
./llama-cli -hf prithivMLmods/CogEvol-4B-GGUF:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf prithivMLmods/CogEvol-4B-GGUF:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf prithivMLmods/CogEvol-4B-GGUF:
Use Docker
docker model run hf.co/prithivMLmods/CogEvol-4B-GGUF:
Quick Links

CogEvol-4B-GGUF

CogEvol-4B is an open, post-trained model built on Qwen3.5-4B for Learning Environment Generation (LEG) — given a natural-language course brief, it generates complete, usable learning artifacts in a single pass, either as a structured-JSON slide page or a self-contained interactive HTML page (simulations, visualizations, interactive exercises) that runs directly in the browser. As the small open member of the CogEvol family, it was trained through a three-stage recipe — mixed SFT, followed by Slide RL, then interactive-HTML RL — on 53,687 verified samples, using a hybrid rule-and-VLM reward system hardened against reward hacking so that interactivity is measured via automated probes rather than judged superficially, as detailed in the accompanying paper "CogEvol: Towards Efficient and Reliable Learning Environment Generation" (arXiv:2608.30968). It's deployable via SGLang, vLLM, or llama.cpp (with a quantized Q4_K_M GGUF at ~2.4GB also available), requires thinking mode disabled at inference and the shipped system prompt template for slide generation, and integrates with the OpenMAIC app for fully offline use, with a complete deployment guide, serving scripts, and evaluation tools available on the project's GitHub repository; it's released under the Apache 2.0 license.

Model Files

File Name Quant Type File Size File Link
CogEvol-4B.BF16.gguf BF16 8.42 GB Download
CogEvol-4B.Q3_K_L.gguf Q3_K_L 2.42 GB Download
CogEvol-4B.Q3_K_M.gguf Q3_K_M 2.26 GB Download
CogEvol-4B.Q3_K_S.gguf Q3_K_S 2.07 GB Download
CogEvol-4B.Q4_0.gguf Q4_0 2.54 GB Download
CogEvol-4B.Q4_K_M.gguf Q4_K_M 2.71 GB Download
CogEvol-4B.Q4_K_S.gguf Q4_K_S 2.56 GB Download
CogEvol-4B.Q5_0.gguf Q5_0 2.99 GB Download
CogEvol-4B.Q5_K_M.gguf Q5_K_M 3.07 GB Download
CogEvol-4B.Q5_K_S.gguf Q5_K_S 2.99 GB Download
CogEvol-4B.mmproj-bf16.gguf mmproj-bf16 676 MB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
682
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/CogEvol-4B-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(4)
this model

Collection including prithivMLmods/CogEvol-4B-GGUF

Paper for prithivMLmods/CogEvol-4B-GGUF