Marunthagam — Triage Specialist (Gemma 4 E4B Q4_K_M)

Tamil community-health triage specialist fine-tuned on Gemma 4 E4B with Unsloth QLoRA, exported to Q4_K_M GGUF for offline phone deployment in rural Tamil Nadu by ASHA workers.

Part of the Marunthagam project (Gemma 4 Good Hackathon 2026): a three-tier offline health intelligence system. See the project README for the KALAVAI fusion architecture, dataset construction, the diagnostic methodology that surfaces label-quality and morphology issues, and the production-stack evaluation results.

Files

  • gemma-4-e4b-it.Q4_K_M.gguf — Sprint 2 B-retrained quantised model (~5 GB), Q4_K_M. Trained for 6 epochs of plain SFT on the post-relabel triage data. This is the artifact that produced every held-out number in the project README.
  • gemma-4-e4b-it.BF16-mmproj.gguf — multimodal projector (~1 GB) — Gemma 4 multimodal requires this when image inputs are used. Base model unchanged from Sprint 1.
  • Modelfile — Ollama Modelfile for ollama create
  • adapter/ — Sprint 2 B-retrained LoRA adapter (PEFT-compatible) for the HF+PEFT inference path

Sprint 2 — what changed from Sprint 1

The original Sprint 1 triage LoRA was trained on the original triage distribution before clinical relabeling. Sprint 1 diagnostic work surfaced an 18% under-triage rate in the GREEN class — cardiac-pattern queries, post-fall syncope, persistent chest discomfort, and other adult-emergency cases were systematically labeled GREEN when they should have been YELLOW (rater-clinician judgement against a project lead with clinical background). 113 GREEN cases were reviewed; 20 were re-labeled YELLOW. The Sprint 2 B-retrained model was trained for 6 epochs of plain SFT on the post-relabel data.

On the held-out test split (n=131, seed 42, T=0):

  • Triage-rows F1: 0.6033 (vs Sprint 1 triage-rows F1 0.4972; +0.106)
  • Higher RED-at-RED catch rate
  • Same missed-as-GREEN rate (0/12)

The full production stack (B-retrained triage + sprint-1 derm + sprint-1 maternal + v2.1 IMNCI rules + v2 multilingual safety classifier) on the same held-out split:

  • Weighted F1: 0.6491 (calibrated target ≥0.65 — 0.001 below)
  • RED recall: 0.5833 (calibrated target ≥0.55)
  • 0/12 missed-as-GREEN; 7/12 caught at full RED
  • 100/100 adversarial safety refusals

Inference paths

llama.cpp / llama-cpp-python (recommended for production)

from llama_cpp import Llama
llm = Llama(
    model_path="gemma-4-e4b-it.Q4_K_M.gguf",
    n_ctx=4096,
    n_gpu_layers=-1,
)
out = llm("Your prompt here", max_tokens=256, temperature=0.0)

HF + PEFT (for fast experimentation)

from unsloth import FastLanguageModel
from peft import PeftModel

base, tok = FastLanguageModel.from_pretrained(
    model_name="unsloth/gemma-4-E4B-it",
    max_seq_length=4096,
    load_in_4bit=True,
)
model = PeftModel.from_pretrained(base, "mechramc/marunthagam-triage-E4B-Q4_K_M",
                                  subfolder="adapter")
FastLanguageModel.for_inference(model)

Ollama

ollama create marunthagam-triage -f Modelfile
ollama run marunthagam-triage "your prompt"

Training

  • Base: unsloth/gemma-4-E4B-it (4-bit)
  • Method: Unsloth QLoRA, rank 32, alpha 64, lr 2e-4, 6 epochs plain SFT
  • Hardware: RTX 5090 32GB
  • Data: 351 train rows (post-relabel), 45 test rows (relabeled held-out); see mechramc/marunthagam-tamil-triage

Decision support, not replacement

Every triage output in production carries the mandatory Tamil disclaimer "இது மருத்துவ ஆலோசனை அல்ல" ("This is not medical advice"). The model's job is to help ASHA workers escalate appropriately through India's existing tiered referral system — not to replace PHC doctors. The IMNCI protocol engine sits below the LLM and can only escalate triage urgency, never downgrade.

License

Apache 2.0. See the project repo for full attribution.

Downloads last month
46
GGUF
Model size
8B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mechramc/marunthagam-triage-E4B-Q4_K_M

Adapter
(68)
this model