Text Generation
Transformers
Safetensors
English
spike_whale
feature-extraction
small-models
chat
mla
kv-cache-compression
triattention
long-context
experimental
custom_code
Instructions to use Quazim0t0/Byrne-TriAtn-86M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Quazim0t0/Byrne-TriAtn-86M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Quazim0t0/Byrne-TriAtn-86M", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Quazim0t0/Byrne-TriAtn-86M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Quazim0t0/Byrne-TriAtn-86M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Quazim0t0/Byrne-TriAtn-86M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Byrne-TriAtn-86M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Quazim0t0/Byrne-TriAtn-86M
- SGLang
How to use Quazim0t0/Byrne-TriAtn-86M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Quazim0t0/Byrne-TriAtn-86M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Byrne-TriAtn-86M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Quazim0t0/Byrne-TriAtn-86M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Byrne-TriAtn-86M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Quazim0t0/Byrne-TriAtn-86M with Docker Model Runner:
docker model run hf.co/Quazim0t0/Byrne-TriAtn-86M
Harness: pass engram_context_ids on cached steps (n-gram memory matches full compute; exercised in real-weight A/B) + no-repeat-ngram(3)
Browse files- generate_triattention.py +8 -1
generate_triattention.py
CHANGED
|
@@ -77,6 +77,12 @@ def generate(model, tok, prompt, max_new_tokens, temperature, top_k, device,
|
|
| 77 |
t0 = time.time()
|
| 78 |
for _step in range(max_new_tokens):
|
| 79 |
logits = _rep_penalty(logits, generated[0, P:], repetition_penalty)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 80 |
if temperature and temperature > 0:
|
| 81 |
probs = torch.softmax(logits / temperature, dim=-1)
|
| 82 |
if top_k:
|
|
@@ -91,8 +97,9 @@ def generate(model, tok, prompt, max_new_tokens, temperature, top_k, device,
|
|
| 91 |
if tok.eos_token_id is not None and nxt.item() == tok.eos_token_id:
|
| 92 |
break
|
| 93 |
cur_pos = torch.tensor([[P + _step]], device=device)
|
|
|
|
| 94 |
out = model(input_ids=nxt, position_ids=cur_pos,
|
| 95 |
-
past_key_values=pkv, use_cache=True)
|
| 96 |
pkv = out.past_key_values
|
| 97 |
logits = out.logits[:, -1, :]
|
| 98 |
cache_len_series.append(pkv[0][0].shape[2])
|
|
|
|
| 77 |
t0 = time.time()
|
| 78 |
for _step in range(max_new_tokens):
|
| 79 |
logits = _rep_penalty(logits, generated[0, P:], repetition_penalty)
|
| 80 |
+
gen = generated[0, P:].tolist()
|
| 81 |
+
if len(gen) >= 3: # no-repeat-ngram(3): A/B'd on real
|
| 82 |
+
t2 = (gen[-2], gen[-1]) # weights, rep-4 -> 0.000
|
| 83 |
+
for i in range(len(gen) - 2):
|
| 84 |
+
if (gen[i], gen[i + 1]) == t2:
|
| 85 |
+
logits[0, gen[i + 2]] = float("-inf")
|
| 86 |
if temperature and temperature > 0:
|
| 87 |
probs = torch.softmax(logits / temperature, dim=-1)
|
| 88 |
if top_k:
|
|
|
|
| 97 |
if tok.eos_token_id is not None and nxt.item() == tok.eos_token_id:
|
| 98 |
break
|
| 99 |
cur_pos = torch.tensor([[P + _step]], device=device)
|
| 100 |
+
ctx = generated[:, -3:-1] # engram trigram context (2 tokens before nxt)
|
| 101 |
out = model(input_ids=nxt, position_ids=cur_pos,
|
| 102 |
+
past_key_values=pkv, use_cache=True, engram_context_ids=ctx)
|
| 103 |
pkv = out.past_key_values
|
| 104 |
logits = out.logits[:, -1, :]
|
| 105 |
cache_len_series.append(pkv[0][0].shape[2])
|