Text Generation
GGUF
llama.cpp
quantized
llama-cpp
imatrix
lfm2
hybrid
edge
on-device
tool-calling
reasoning
long-context
multilingual
conversational
Instructions to use KikoCis/LFM2.5-2.6B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use KikoCis/LFM2.5-2.6B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use KikoCis/LFM2.5-2.6B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "KikoCis/LFM2.5-2.6B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KikoCis/LFM2.5-2.6B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
- Ollama
How to use KikoCis/LFM2.5-2.6B-GGUF with Ollama:
ollama run hf.co/KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use KikoCis/LFM2.5-2.6B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use KikoCis/LFM2.5-2.6B-GGUF with Docker Model Runner:
docker model run hf.co/KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
- Lemonade
How to use KikoCis/LFM2.5-2.6B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.LFM2.5-2.6B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use KikoCis/LFM2.5-2.6B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use KikoCis/LFM2.5-2.6B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "KikoCis/LFM2.5-2.6B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download scripts/04_charts.py from KikoCis/LFM2.5-2.6B-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 3.62 kB
-
https://huggingface.co/KikoCis/LFM2.5-2.6B-GGUF/resolve/main/scripts/04_charts.py
- Command line
-
hf download hf://KikoCis/LFM2.5-2.6B-GGUF/scripts/04_charts.py
-
curl -L -o 04_charts.py https://huggingface.co/KikoCis/LFM2.5-2.6B-GGUF/resolve/main/scripts/04_charts.py
3.62 kB
| #!/usr/bin/env python3 | |
| """LFM2.5-2.6B-GGUF — charts from metrics/quant-summary-with-kld.json (one metric each).""" | |
| import json, os | |
| import matplotlib | |
| matplotlib.use("Agg") | |
| import matplotlib.pyplot as plt | |
| REPO = os.environ.get("REPO", "/Volumes/Lexar/aq_lfm25/repo") | |
| M = f"{REPO}/metrics" | |
| d = json.load(open(f"{M}/quant-summary-with-kld.json")) | |
| rows = sorted(d["quants"], key=lambda r: r["size_gb_decimal"]) | |
| lab = [r["quant"] for r in rows] | |
| gb = [r["size_gb_decimal"] for r in rows] | |
| BG, FG, ACC, ACC2 = "#0d1117", "#e6edf3", "#58a6ff", "#f78166" | |
| def style(ax, title, xl, yl): | |
| ax.set_facecolor(BG) | |
| ax.set_title(title, color=FG, fontsize=13, fontweight="bold", pad=12) | |
| ax.set_xlabel(xl, color=FG, fontsize=10) | |
| ax.set_ylabel(yl, color=FG, fontsize=10) | |
| ax.tick_params(colors=FG, labelsize=9) | |
| for s in ax.spines.values(): | |
| s.set_color("#30363d") | |
| ax.grid(alpha=0.18, color=FG, linewidth=0.6) | |
| def fig(): | |
| f, ax = plt.subplots(figsize=(8.2, 4.6)) | |
| f.patch.set_facecolor(BG) | |
| return f, ax | |
| def save(f, name): | |
| f.tight_layout() | |
| f.savefig(f"{M}/{name}", dpi=160, facecolor=BG) | |
| plt.close(f) | |
| print("wrote", name) | |
| # 1) the money chart: fidelity (KLD) vs file size | |
| f, ax = fig() | |
| k = [r["kld_nats_mean"] for r in rows] | |
| ax.plot(gb, k, "o-", color=ACC, lw=2.2, ms=8) | |
| for i, (x, y, t) in enumerate(zip(gb, k, lab)): | |
| dy = 13 if i % 2 == 0 else -20 | |
| ax.annotate(t, (x, y), textcoords="offset points", xytext=(6, dy), | |
| ha="left", color=FG, fontsize=9, fontweight="bold") | |
| ax.set_xlim(min(gb) - 0.25, max(gb) + 0.55) | |
| ax.set_yscale("symlog", linthresh=1e-4) | |
| style(ax, "Quality vs size — KL divergence from the F16 reference (lower = closer)", | |
| "file size (GB)", "mean KLD (nats, log scale)") | |
| save(f, "chart-quality-vs-size.png") | |
| # 2) KLD mean, per quant | |
| f, ax = fig() | |
| ax.bar(lab, [r["kld_nats_mean"] for r in rows], color=ACC, edgecolor="#30363d") | |
| for i, r in enumerate(rows): | |
| ax.text(i, r["kld_nats_mean"], f'{r["kld_nats_mean"]:.4f}', ha="center", | |
| va="bottom", color=FG, fontsize=8.5) | |
| style(ax, "Mean KL divergence vs F16 (lower = more faithful)", "quant", "KLD (nats)") | |
| save(f, "chart-kld-mean.png") | |
| # 3) PPL delta | |
| f, ax = fig() | |
| dl = [r["ppl_delta_vs_f16"] for r in rows] | |
| ax.bar(lab, dl, color=[ACC2 if v > 0 else ACC for v in dl], edgecolor="#30363d") | |
| for i, v in enumerate(dl): | |
| ax.text(i, v, f"{v:+.3f}", ha="center", va="bottom" if v >= 0 else "top", | |
| color=FG, fontsize=8.5) | |
| ax.axhline(0, color=FG, lw=0.8) | |
| style(ax, "Perplexity delta vs F16 (wikitext-2 test, ctx 2048)", "quant", "ΔPPL") | |
| save(f, "chart-ppl-delta.png") | |
| # 4) throughput | |
| f, ax = fig() | |
| x = range(len(lab)) | |
| w = 0.38 | |
| ax.bar([i - w / 2 for i in x], [r["prompt_tps_p512"] for r in rows], w, | |
| label="prompt tok/s (pp512)", color=ACC, edgecolor="#30363d") | |
| ax.bar([i + w / 2 for i in x], [r["gen_tps_n128"] for r in rows], w, | |
| label="generation tok/s (tg128)", color=ACC2, edgecolor="#30363d") | |
| ax.set_xticks(list(x)) | |
| ax.set_xticklabels(lab) | |
| lg = ax.legend(facecolor=BG, edgecolor="#30363d", labelcolor=FG, fontsize=9) | |
| style(ax, "Throughput per quant", "quant", "tokens / s") | |
| save(f, "chart-throughput.png") | |
| # 5) top-1 agreement | |
| f, ax = fig() | |
| t = [100 * r["top1_match_rate"] for r in rows] | |
| ax.bar(lab, t, color=ACC, edgecolor="#30363d") | |
| for i, v in enumerate(t): | |
| ax.text(i, v, f"{v:.2f}%", ha="center", va="bottom", color=FG, fontsize=8.5) | |
| ax.set_ylim(min(t) - 2, 101) | |
| style(ax, "Top-1 token agreement with F16 (higher = same next token more often)", | |
| "quant", "% of tokens where argmax matches") | |
| save(f, "chart-top1-match.png") | |