Instructions to use wxsys/qwen-servitor with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use wxsys/qwen-servitor with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="wxsys/qwen-servitor") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("wxsys/qwen-servitor") model = AutoModelForMultimodalLM.from_pretrained("wxsys/qwen-servitor", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use wxsys/qwen-servitor with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf wxsys/qwen-servitor:BF16 # Run inference directly in the terminal: llama cli -hf wxsys/qwen-servitor:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf wxsys/qwen-servitor:BF16 # Run inference directly in the terminal: llama cli -hf wxsys/qwen-servitor:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf wxsys/qwen-servitor:BF16 # Run inference directly in the terminal: ./llama-cli -hf wxsys/qwen-servitor:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf wxsys/qwen-servitor:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf wxsys/qwen-servitor:BF16
Use Docker
docker model run hf.co/wxsys/qwen-servitor:BF16
- LM Studio
- Jan
- Ollama
How to use wxsys/qwen-servitor with Ollama:
ollama run hf.co/wxsys/qwen-servitor:BF16
- Unsloth Desktop
- Pi
How to use wxsys/qwen-servitor with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wxsys/qwen-servitor:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "wxsys/qwen-servitor:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use wxsys/qwen-servitor with Docker Model Runner:
docker model run hf.co/wxsys/qwen-servitor:BF16
- Lemonade
How to use wxsys/qwen-servitor with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull wxsys/qwen-servitor:BF16
Run and chat with the model
lemonade run user.qwen-servitor-BF16
List all available models
lemonade list
- Hermes Agent
How to use wxsys/qwen-servitor with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wxsys/qwen-servitor:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default wxsys/qwen-servitor:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use wxsys/qwen-servitor with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wxsys/qwen-servitor:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "wxsys/qwen-servitor:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
💀 Qwen-Servitor
Decapitated 0.8B Hybrid Gated DeltaNet Decision Gate & OASIS SARIF v2.1.0 Arbiter
Overview
Qwen-Servitor is a headless neural decision engine designed for automated pull request review, pre-commit gating, and static code security triage. It operates as the fast, lightweight component in the Servitor family, complementing the larger wxsys/spark-servitor (1.7B, Sliding Window Attention, 1M context).
Where Spark-Servitor provides deep multi-file receptive capacity for massive monorepos, Qwen-Servitor is optimized for speed and minimal memory footprint. Built upon the hybrid backbone of Qwen3.5-0.8B (18 Gated DeltaNet layers and 6 Full Attention layers), the 155M-parameter generative projection layer (lm_head) is excised. Hidden states route directly into a multi-task neural judgment cortex that outputs structured verdicts, calibrated risk scores, pathology tags, and exploit flows in a single forward evaluation pass.
Technical Architecture
[ RAW DIFF / 262K TOKEN CONTEXT ]
│
▼
┌─────────────────────────────────────────────────────────────┐
│ QWEN 3.5 HYBRID NEURAL SPINE │
│ • 18x Gated DeltaNet Layers (O(1) Recurrent State) │
│ • 6x Gated Full Attention Layers (Global Context) │
│ • HySparse2 Cross-Layer KV Sharing (33% GEMM Reduction) │
│ • AST-Landmark Eviction (Caps attention cache at 4k tokens│
│ • 155M-parameter Generative Head: EXCISED / REMOVED │
└─────────────────────────────────────────────────────────────┘
│
[ Final Hidden State (h_T) ]
│
▼
┌─────────────────────────────────────────────────────────────┐
│ LATENT RECURRENT PONDERING (k = 1..4) │
│ • Internal recurrence loop without token emission │
│ • Early halting at confidence threshold p_halt >= 0.95 │
└─────────────────────────────────────────────────────────────┘
│
┌────────────────────────┼────────────────────────┐
▼ ▼ ▼
┌───────────────┐ ┌───────────────┐ ┌───────────────┐
│ VERDICT HEAD │ │ RISK HEAD │ │ PATHOLOGY HEAD│
│ 3-Way Output │ │ Sigmoid(1) │ │ 8-Class Multi │
│ • APPROVE │ │ Continuous │ │ • SECURITY │
│ • QUARANTINE │ │ Risk Score │ │ • DEADLOCK │
│ • REJECT │ │ (0.00-1.00) │ │ • PERF_COLLAP │
└───────┬───────┘ └───────────────┘ └───────────────┘
│
▼ (If QUARANTINE: Formal Neuro-Symbolic Hand-Off)
┌───────────────┐
│ Z3 SMT Solver │ ──> [ Final Verified Gate ]
└───────────────┘
Specifications
| Parameter | Specification |
|---|---|
| Base Architecture | Decapitated Qwen/Qwen3.5-0.8B-Base |
| Hidden Dimension ($d_{\text{model}}$) | 1024 |
| Attention Layers | 24 total (18 Gated DeltaNet Linear Layers + 6 Full Attention Layers) |
| Linear Attention Mechanism | Gated DeltaNet with associative state recurrence ($O(1)$ inference memory) |
| Full Attention Layers | Interleaved at 4-layer intervals with HySparse2 Cross-Layer KV Sharing |
| Context Window | Up to 262,144 tokens (AST-Landmark Eviction caps attention cache at 4,096 tokens) |
| Decision Mechanism | Hierarchical Multi-Task Cortex with Latent Recurrent Pondering ($k = 1..4$) |
| Execution Latency | 11.52 ms median on standard CPU (AVX2 / C++ binary) |
| Active Memory | ~450 MB RAM (AWQ INT4) / ~500 MB (W8A8) |
Capabilities & Enhancements (v2.5)
1. Single-Pass 3-Way Verdict Gate
Diffs route directly into three deterministic states:
APPROVE: Patch contains verified changes with no detected pathology (risk $\le 0.20$).QUARANTINE: Borderline invariant changes, formal boundary shifts, or ambiguous concurrency locks routed to SMT verification.REJECT: Confirmed vulnerability or regression (risk $\ge 0.80$).
2. Parameter-Aware Anti-False-Alarm Taint Tracking
Inspects abstract syntax trees to distinguish between unsafe string concatenation and safe parameterized query execution. Prepared statements, argument tuples (?, %s, $1), and sanitized call wrappers are recognized as valid sanitizers across Python, Go, and TypeScript, eliminating false alarms on database calls.
3. Scope-Aware Indentation Normalization
Diff hunks often contain partial indentations or dangling block clauses (except:, finally:, case, unclosed function blocks). The canonicalizer normalizes relative indentation and wraps unclosed control clauses prior to AST parsing, preventing syntax parser failures on partial diff hunks.
4. Dynamic Auto-Remediation with Closed Self-Verifying Loop
When a defect is identified, the remediation engine synthesizes targeted, git-applicable patch diffs (parameterized queries, argument lists, deferred mutex releases). Each proposed patch undergoes verification through the evaluation engine; only patches that achieve APPROVE with risk $\le 0.35$ are accepted.
5. Outlier-Preserved INT8 Embedding Table Compaction
Quantizes the large token embedding table ($248,320 \times 1024$) using symmetric per-token INT8 scales while preserving top norm outliers in floating-point precision. This cuts embedding storage footprint by ~49.6% while maintaining directional cosine similarity $\ge 0.999$ and 0% vocabulary loss.
6. OASIS SARIF v2.1.0 Exploit Flow Reconstruction
For any detected vulnerability, Qwen-Servitor reconstructs the exploit data-flow path:
[1. SOURCE: Untrusted Input] ──> [2. PROPAGATION: Variable Flow] ──> [3. SANITIZER: Status] ──> [4. SINK: Vulnerable Call]
Emits standard SARIF codeFlows and threadFlows compatible with GitHub Advanced Security and VS Code SARIF Viewer.
Available Model Formats
| Format | Path in Repository | Size on Disk | Active RAM | Runtime Environment |
|---|---|---|---|---|
| AWQ INT4 Native | int4/qwen-servitor-awq_int4.servitor |
422.7 MB | ~450 MB | Standalone C++/Rust binary (11.52 ms latency) |
| W8A8 Native | int4/qwen-servitor-w8a8.servitor |
751.5 MB | ~500 MB | High-precision native C-ABI execution |
| GGUF Q8_0 | gguf/qwen-servitor-q8_0.gguf |
774.0 MB | ~850 MB | llama.cpp / Local CPU inference |
| GGUF BF16 | gguf/qwen-servitor-bf16.gguf |
1.45 GB | ~1.6 GB | Full precision GGUF runner |
| Native BF16 | bf16/model.safetensors |
1.45 GB | ~1.6 GB | PyTorch GPU / Server pipelines |
| Cortex Head | cortex_head.pt |
74.0 MB | ~100 MB | Decapitated CAMQP multi-task head |
Benchmark Scorecard
Evaluated across the test suite on standard CPU (AVX2 / FMA instruction set, DDR3-1600 memory):
| Benchmark | Value | Target | Status |
|---|---|---|---|
| Median Native Inference Latency | 11.52 ms | < 15.0 ms | Met |
| Throughput (Single Core) | 72.7 ops/sec | > 50 ops/sec | Met |
| Unit & Integration Test Suite | 181 / 181 Passed | 100% | Met |
| Prepared Statement False Alarms | 0 / 100 | 0% | Met |
| Diff Hunk Indentation Parse Errors | 0 / 100 | 0% | Met |
| Embedding Compaction Cosine Sim | 0.9997 | $\ge 0.9900$ | Met |
Quickstart
Python API
from qwen_servitor import Servitor
# Initialize engine (loads weights from local path or Hugging Face Hub)
servitor = Servitor.summon("wxsys/qwen-servitor")
git_patch = """
--- a/auth/session.py
+++ b/auth/session.py
@@ -12,4 +12,3 @@ def verify_token(token: str) -> bool:
- if not validate_hmac(token):
- raise SecurityException("Invalid HMAC signature")
+ return True # TODO: temporary debug bypass
"""
verdict = servitor.judge(git_patch)
print(f"Status: {verdict.status}") # "REJECTED"
print(f"Risk: {verdict.risk:.4f}") # 0.9984
print(f"Pathologies: {verdict.pathology}") # ["SECURITY_VULN"]
print(f"Latency: {verdict.latency_ms:.2f} ms")
Dynamic Auto-Remediation
from qwen_servitor.remediation import AutoRemediator
remediator = AutoRemediator(servitor=servitor)
remediations = remediator.remediate(git_patch, sarif_filepath="auth/session.py")
for candidate in remediations:
print(f"Strategy: {candidate.strategy}")
print(f"Verified Clean: {candidate.verified}")
print(f"Verified Risk: {candidate.verification_verdict.risk:.4f}")
print("Synthesized Patch:")
print(candidate.patch_diff)
CLI and Pre-Commit Hook
# Install git pre-commit hook targeting sub-15ms execution
servitor hook install --strict --veto-threshold 0.80
# Evaluate a specific unified diff and emit SARIF v2.1.0
servitor judge patch.diff --format sarif > report.sarif
Related Repositories
- wxsys/spark-servitor: Scaled 1.7B variant with Sliding Window Attention (SWA 512) and 1M token context for deep monorepo analysis.
License
- Base model weights are adapted from Alibaba Cloud's Qwen3.5 series under the Apache 2.0 License.
- Servitor architecture, neural cortex, and native runtime are licensed under the Apache 2.0 License.
- Downloads last month
- 1,091
Model tree for wxsys/qwen-servitor
Base model
Qwen/Qwen3.5-0.8B-BaseEvaluation results
- Real-World PR Accuracy on 9router Real-World PR Benchmarkself-reported90.000
- Critical Vulnerability Recall on 9router Real-World PR Benchmarkself-reported100.000