flashback2k commited on
Commit
1eb5124
·
verified ·
1 Parent(s): 21175cf

Model card: usage, training details, limitations

Browse files
Files changed (1) hide show
  1. README.md +124 -15
README.md CHANGED
@@ -1,22 +1,131 @@
1
  ---
2
- language: [en]
3
- base_model: Qwen/Qwen3.5-9B
4
- tags: [sft, reasoning, coding, tool-calling, distillation]
 
5
  license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6
  ---
7
 
8
- # FlashModel — Qwen3.5-9B
9
 
10
- Text model fine-tuned with LoRA r=128 / RSLoRA, merged into the base weights.
11
- Training used 9638 examples and 600 optimizer steps, distilled from open-weight teachers
12
- whose licenses permit training on outputs: DeepSeek-V4-Pro (math, verified answers), DeepSeek-R1-0528
13
- (competitive programming), GLM-4.6 / DeepSeek-V3.2 (tool calling), GPT-OSS-120B (instruction following),
14
- GPT-OSS / Kimi-K2 / DeepSeek-V3.2 (science), via NVIDIA Nemotron SFT datasets (CC BY 4.0 / CC BY-SA 4.0).
15
- Exact runtime versions and the base revision are recorded in `training_metadata.json`.
16
 
17
- The training data contains reasoning-budget tags (`off`, `low`, `medium`, `high`,
18
- `xhigh`, `max`, `adaptive`). Their effect and model quality have not yet been benchmarked.
19
 
20
- Native MTP export is a separate artifact: the base checkpoint's original MTP head
21
- is paired with this target model. No DSpark heads were trained or exported.
22
- Speculative-decoding acceptance and speed must be measured on the target runtime.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ base_model:
3
+ - Qwen/Qwen3.5-9B
4
+ base_model_relation: finetune
5
+ library_name: transformers
6
  license: apache-2.0
7
+ license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
8
+ pipeline_tag: text-generation
9
+ language:
10
+ - en
11
+ tags:
12
+ - qwen3.5
13
+ - reasoning
14
+ - tool-calling
15
+ - distillation
16
+ - sft
17
+ - lora
18
+ datasets:
19
+ - nvidia/Nemotron-SFT-Math-v4
20
+ - nvidia/Nemotron-SFT-Competitive-Programming-v2
21
+ - nvidia/Nemotron-SFT-Agentic-v2
22
+ - nvidia/Nemotron-SFT-Instruction-Following-Chat-v3
23
+ - nvidia/Nemotron-SFT-Science-v2
24
  ---
25
 
26
+ # ⚡ FlashModel-Qwen3.5-9B
27
 
28
+ A reasoning / coding / tool-calling fine-tune of [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B),
29
+ distilled from open-weight frontier teachers (DeepSeek-V4-Pro, DeepSeek-R1-0528, GLM-4.6, DeepSeek-V3.2, GPT-OSS-120B, Kimi-K2).
 
 
 
 
30
 
31
+ **GGUF quants (llama.cpp, Ollama, LM Studio): [flashback2k/FlashModel-Qwen3.5-9B-GGUF](https://huggingface.co/flashback2k/FlashModel-Qwen3.5-9B-GGUF)**
 
32
 
33
+ Merged BF16 weights, text-only (`Qwen3_5ForCausalLM`, no vision tower, no MTP head).
34
+
35
+ ## Usage
36
+
37
+ ### Transformers
38
+
39
+ ```python
40
+ from transformers import AutoModelForCausalLM, AutoTokenizer
41
+
42
+ model_id = "flashback2k/FlashModel-Qwen3.5-9B"
43
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
44
+ model = AutoModelForCausalLM.from_pretrained(model_id, dtype="auto", device_map="auto")
45
+
46
+ messages = [{"role": "user", "content": "Write a Python function that checks whether a number is prime."}]
47
+ inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
48
+ output = model.generate(inputs, max_new_tokens=4096, do_sample=True, temperature=0.6, top_p=0.95, top_k=20)
49
+ print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))
50
+ ```
51
+
52
+ Qwen3.5 support requires a recent `transformers` release.
53
+
54
+ ### vLLM / SGLang
55
+
56
+ ```bash
57
+ vllm serve flashback2k/FlashModel-Qwen3.5-9B --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder
58
+ ```
59
+
60
+ ### Recommended sampling
61
+
62
+ These are the base model's official recommendations ([Qwen3.5 model card](https://huggingface.co/Qwen/Qwen3.5-9B)); they were not re-tuned for the fine-tune.
63
+
64
+ | Mode | temperature | top_p | top_k | min_p | presence_penalty |
65
+ |:--|--:|--:|--:|--:|--:|
66
+ | Thinking, general | 1.0 | 0.95 | 20 | 0.0 | 1.5 |
67
+ | Thinking, precise coding | 0.6 | 0.95 | 20 | 0.0 | 0.0 |
68
+ | Non-thinking, general | 0.7 | 0.8 | 20 | 0.0 | 1.5 |
69
+
70
+ Always pass `--jinja` so the embedded Qwen3.5 chat template (thinking blocks, tool calls) is used.
71
+
72
+ ### Reasoning-budget tag
73
+
74
+ Training system prompts started with a reasoning-budget tag chosen from the length of the teacher's reasoning:
75
+
76
+ ```
77
+ <|reasoning_budget|>medium<|/reasoning_budget|>
78
+ ```
79
+
80
+ Values: `off`, `low`, `medium`, `high`, `xhigh`, `max`. The tag's effect on output length **has not been measured yet**; treat it as experimental.
81
+
82
+ ### Reasoning-budget tag
83
+
84
+ Training system prompts started with a reasoning-budget tag chosen from the length of the teacher's reasoning:
85
+
86
+ ```
87
+ <|reasoning_budget|>medium<|/reasoning_budget|>
88
+ ```
89
+
90
+ Values: `off`, `low`, `medium`, `high`, `xhigh`, `max`. The tag's effect on output length **has not been measured yet**; treat it as experimental.
91
+
92
+ ## About the fine-tune
93
+
94
+ | | |
95
+ |:--|:--|
96
+ | Base | Qwen/Qwen3.5-9B |
97
+ | Method | LoRA r=128 (RSLoRA, α=32) on attention + MLP projections, merged |
98
+ | Data | 9,638 examples / 45M tokens, loss on assistant turns only |
99
+ | Context in training | up to 16,384 tokens (longer examples dropped, never truncated) |
100
+ | Schedule | 1 epoch, 600 steps, lr 5e-5 cosine |
101
+ | Held-out eval loss | 0.5745 (step 100) → 0.5586 (step 600) |
102
+
103
+ ### Training data and teachers
104
+
105
+ Only open-weight teachers whose licenses allow training on their outputs:
106
+
107
+ | Share | Domain | Dataset | Teacher |
108
+ |--:|:--|:--|:--|
109
+ | 40% | Math (answers verified against references) | [nvidia/Nemotron-SFT-Math-v4](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Math-v4) | DeepSeek-V4-Pro |
110
+ | 22% | Competitive programming (Python) | [nvidia/Nemotron-SFT-Competitive-Programming-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Competitive-Programming-v2) | DeepSeek-R1-0528 |
111
+ | 16% | Multi-turn tool calling (judge-filtered) | [nvidia/Nemotron-SFT-Agentic-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Agentic-v2) | GLM-4.6 / DeepSeek-V3.2 |
112
+ | 11% | Instruction following | [nvidia/Nemotron-SFT-Instruction-Following-Chat-v3](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3) | GPT-OSS-120B |
113
+ | 11% | Science reasoning | [nvidia/Nemotron-SFT-Science-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2) | GPT-OSS / Kimi-K2 / DeepSeek-V3.2 |
114
+
115
+ Datasets © NVIDIA, CC BY 4.0 (some Math StackExchange-derived samples CC BY-SA 4.0).
116
+
117
+ The training run's step count, sample count and eval-loss history are in [`training_metadata.json`](training_metadata.json).
118
+
119
+ ## ⚠️ Status and known limitations
120
+
121
+ - **No benchmarks yet.** Lower held-out loss means the model imitates the teachers more closely; it does not by itself prove
122
+ it beats the stock Qwen3.5-9B. Comparative evals are planned and will be added here.
123
+ - **English-centric.** All training data was English. On non-English prompts (e.g. Russian) the model often *reasons* in English.
124
+ - Math and code examples longer than 16k tokens were excluded, which skews those domains toward shorter problems.
125
+ - Inherits the base model's limitations and biases.
126
+
127
+ ## Credits
128
+
129
+ [Qwen team](https://huggingface.co/Qwen) for Qwen3.5 · [NVIDIA](https://huggingface.co/nvidia) for the Nemotron SFT datasets ·
130
+ DeepSeek, Zhipu AI (GLM), OpenAI (GPT-OSS) and Moonshot AI (Kimi) for open-weight teachers ·
131
+ [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp).