alexzhang0118 commited on
Commit
4a7552c
·
verified ·
1 Parent(s): 8a02f13

Upload folder using huggingface_hub

Browse files
Files changed (6) hide show
  1. README.md +206 -0
  2. config.json +44 -0
  3. model.safetensors +3 -0
  4. tokenizer.json +0 -0
  5. tokenizer_config.json +239 -0
  6. vocab.json +0 -0
README.md ADDED
@@ -0,0 +1,206 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ base_model: Qwen/Qwen3-1.7B-Base
6
+ tags:
7
+ - decision-model
8
+ - candidate-scorer
9
+ - calibration
10
+ - selective-prediction
11
+ - on-device
12
+ - mlx
13
+ - onnx
14
+ ---
15
+
16
+ # Decily — calibrated candidate-scorer decision model
17
+
18
+ **Decily** scores a runtime-provided set of candidates against a `state` and a
19
+ `question`, and returns a **calibrated probability distribution** over those
20
+ candidates. It does not generate text. Decily is built for decisions where you
21
+ want a probability you can threshold, not a sentence you have to parse.
22
+
23
+ ```python
24
+ Decily.decide(
25
+ state="Our app crashes on startup after the latest update.",
26
+ question="What is the customer's intent?",
27
+ options=["technical", "billing", "shipping", "returns"],
28
+ ) # -> {"technical": 0.9999, "billing": 0.0001, "returns": 0.0, "shipping": 0.0}
29
+ # measured with the released weights (bf16, T=0.45)
30
+ ```
31
+
32
+ ## This repository
33
+
34
+ **Apple MLX bf16 bundle** (3.47 GB) — backbone and scoring head in one
35
+ safetensors file, the layout the standalone reference implementation
36
+ (`mlx/decision_mlx.py` in [arczhi/decily](https://github.com/arczhi/decily))
37
+ expects.
38
+
39
+ ```bash
40
+ python mlx/decision_mlx.py --model-dir <this repo> \
41
+ --state "Our app crashes on startup after the latest update." \
42
+ --question "What is the customer's intent?" \
43
+ --options "technical,billing,shipping,returns" --temperature 0.45
44
+ ```
45
+
46
+ ## Model family
47
+
48
+ This card covers all four released artifacts; they are the same model in
49
+ different containers.
50
+
51
+ | Variant | Format | Size | Target |
52
+ |---|---|---|---|
53
+ | **Decily-1.7B** | PyTorch bf16 safetensors + `config.json` | ~3.4 GB | server / GPU |
54
+ | **Decily-MLX** | MLX bf16 safetensors | ~3.4 GB | Apple Silicon |
55
+ | **Decily-MLX-4bit** | MLX 4-bit (group-64 affine) | 968 MB | on-device / low memory |
56
+ | **Decily-ONNX-int8** | ONNX int8 (single file) | 1.66 GB | CPU / Windows / edge |
57
+
58
+ ## Model details
59
+
60
+ - **Model type:** cross-encoder candidate scorer. The backbone encodes
61
+ `state + question + candidate` jointly; an attention-pooling layer over all
62
+ tokens feeds a LayerNorm/GELU MLP that outputs one logit per candidate.
63
+ Probabilities come from the softmax over the candidates you pass in.
64
+ - **Backbone:** `Qwen/Qwen3-1.7B-Base` (1.72B parameters, Apache-2.0), with an
65
+ attention-pooling scoring head (`decision_head` recorded in `config.json`;
66
+ pool/head weights live in the checkpoint).
67
+ - **Input limits (training configuration):** `state` ≤256 tokens, `question`
68
+ ≤96 tokens, each candidate ≤64 tokens, 2–16 candidates per call. Longer
69
+ inputs are truncated — for long documents, chunk the state and score chunks
70
+ separately (the cross-encoder logits are candidate-set independent, so
71
+ chunked shortlisting then re-ranking is lossless; see the repository's
72
+ two-stage inference utility).
73
+ - **Precision:** bf16 (PyTorch/MLX), 4-bit affine group-64 (on-device),
74
+ int8 (ONNX).
75
+ - **Temperature:** raw logits are over-confident out of distribution. Fitted
76
+ temperatures observed: ≈1.1 (in-domain), ≈0.45 on unseen label sets, ≈1.35 on
77
+ the zero-shot fair suite. Always report/calibrate the temperature for your
78
+ data (see Evaluation).
79
+
80
+ ## Training summary
81
+
82
+ Full, reproducible recipe: [TRAINING.md](https://github.com/arczhi/decily/blob/main/TRAINING.md)
83
+ in the training repository [arczhi/decily](https://github.com/arczhi/decily)
84
+ (design rationale in `DESIGN.md`, all experiments and ablations in
85
+ `TRAIN-REPORT.md`).
86
+
87
+ | Stage | What | Notes |
88
+ |---|---|---|
89
+ | Teacher | 45-task Qwen3.5-2B decision model (LoRA, frozen backbone) | reference labeler |
90
+ | SFT | Qwen3-1.7B LoRA on 24 task families (hard labels) | 3000 steps ≈1 h |
91
+ | RLCD members | v1 / v2 LoRA + one full-FT member | belief calibration, abstention |
92
+ | Distillation | 4-model probability-space ensemble → **single full-FT student** | KD T=2, α=0.5, +15% belief rows |
93
+ | v5 (this model) | full fine-tune, 3000 steps, effective batch 16, lr 1e-5, 8-bit AdamW | ≈1.5 h on one RTX 5090 32 GB |
94
+
95
+ Key recipe properties:
96
+
97
+ - **Route B (explicit candidate scorer)** beats the LM-head-letter-logits route
98
+ (Route A) by **+10.6 pt** held-out accuracy in a controlled same-backbone test.
99
+ - **Ensemble distillation:** the 4-model ensemble reaches NLL 1.455; the
100
+ distilled single model reaches **1.451** at 1× inference cost.
101
+ - **Consumer hardware:** the whole pipeline trains on a single RTX 5090 32 GB
102
+ in well under a day.
103
+
104
+ ## Evaluation
105
+
106
+ Protocol: accuracy / NLL / ECE after **temperature fitting** (raw T=1 numbers
107
+ are over-confident). "In-task" = the 24 trained task families; "held-out" =
108
+ unseen label sets; "fair suite" = 8 tasks unseen by both this model and the
109
+ external baseline.
110
+
111
+ | Evaluation | Decily (24 tasks) | decider-2b (95 tasks) |
112
+ |---|---|---|
113
+ | In-task, 24 task families (shared training tasks) | **0.863 / 0.390 / 0.034** | 0.811 / 0.453 / 0.032 |
114
+ | Fair suite, zero-shot (8 tasks × 300) | 0.654 / 0.86 / 0.092 | **0.700 / 0.71 / 0.047** |
115
+ | Held-out label sets (60/77-class intents, 1200) | 0.581 / 1.451 / 0.068 | — (scores include tasks it was trained on) |
116
+
117
+ Selective prediction (unseen label sets): taking only the most-confident **5%**
118
+ of predictions gives **93.3%** accuracy (an SFT baseline with the same protocol
119
+ gives 79%); abstention thresholds can be calibrated to a target error rate
120
+ (empirically 4.7% achieved at a 5% target).
121
+
122
+ Honest reading of the numbers:
123
+
124
+ - Decily leads the larger, 95-task decider-2b by **+5.2 pt** on tasks both were
125
+ trained on, at 1.7B vs 2B parameters and ~1/4 of the task coverage.
126
+ - On tasks **neither** model has seen, decider-2b leads by 4.6 pt; the gap is
127
+ concentrated in knowledge-heavy tasks (sciq / pubmedqa / quality).
128
+ - The zero-shot gap is an **efficiency** result, not a ceiling: with this
129
+ recipe, task coverage — not parameter count — is the documented main lever.
130
+ Scaling the 24-task mixture toward 60–95 tasks with the same pipeline is
131
+ expected to close and exceed the baseline (see `TRAINING.md`).
132
+
133
+ ## Uses
134
+
135
+ **Direct use**
136
+
137
+ - Classification / ranking / selection when the label set is known at runtime
138
+ and may change per call (intents, topics, sentiment, NLI-style relations,
139
+ multiple-choice answers, routing, moderation labels, tool selection).
140
+ - Calibration-sensitive automation: threshold the probability, auto-handle the
141
+ confident head and escalate the rest to a human or a larger model.
142
+ - On-device / private inference: 968 MB 4-bit build, no GPU required.
143
+ - Synthetic decision data generation and reward/verifier scoring for LLM
144
+ pipelines (it returns probabilities, not text).
145
+
146
+ **Out of scope**
147
+
148
+ - Text generation, chat, summarization.
149
+ - Open-ended label spaces (the candidate set must be provided at call time).
150
+ - High-stakes decisions without a calibrated threshold and human review.
151
+ - Very long inputs beyond the 256-token state window without chunking.
152
+
153
+ ## Quick start
154
+
155
+ Apple Silicon (MLX) — the standalone reference implementation ships with the
156
+ training repository (`mlx/decision_mlx.py`, pure `mlx.core`, no torch):
157
+
158
+ ```bash
159
+ python mlx/decision_mlx.py --model-dir <Decily-MLX dir> \
160
+ --state "The invoice was paid twice..." \
161
+ --question "What is the customer's intent?" \
162
+ --options "billing,technical,shipping,returns" \
163
+ --temperature 0.45
164
+ ```
165
+
166
+ 4-bit on-device: load the quantized tower with `mlx_lm.load` and apply the
167
+ attention-pooling head from `decision_head.safetensors` (both files are in the
168
+ repository; `config.json` records the head spec).
169
+
170
+ PyTorch: load `model.safetensors` + `config.json` with the training repository's
171
+ model class; the checkpoint keeps backbone and head under their training key
172
+ names (`model.*`, `pool.*`, `head.*`).
173
+
174
+ ## Limitations
175
+
176
+ - **Task coverage:** trained on 24 task families; zero-shot behavior on unseen
177
+ domains trails a 95-task baseline (see Evaluation). Validate on your domain
178
+ before trusting the probabilities.
179
+ - **Calibration drift:** the model needs temperature scaling; the fitted value
180
+ depends on the data distribution (≈0.45–1.35 in our measurements).
181
+ - **Input truncation:** `state` is truncated at 256 tokens; long documents must
182
+ be chunked (importance within a chunk, then combine).
183
+ - **Language:** training data is predominantly English (some multilingual
184
+ intent data); other languages are untested.
185
+ - **No safety layer:** outputs are probabilities over the candidates you
186
+ provide; the model does not refuse or filter inputs. Do not expose it as an
187
+ autonomous decision-maker in safety-critical settings.
188
+ - **Inherited biases:** as a fine-tune of Qwen3-1.7B-Base on public datasets,
189
+ it can reproduce biases present in those datasets (toxicity, sentiment and
190
+ bias-classification tasks were part of the mixture, which reduces but does
191
+ not eliminate this).
192
+
193
+ ## Environmental impact
194
+
195
+ Training used a single RTX 5090 32 GB: teacher ≈4 h, SFT ≈1 h, members ≈1 h,
196
+ distillation ≈1.5 h (plus data conversion). Total well under 10 GPU-hours; no
197
+ model was trained more than once per stage.
198
+
199
+ ## Citation and acknowledgements
200
+
201
+ - Backbone: [Qwen3-1.7B-Base](https://huggingface.co/Qwen/Qwen3-1.7B-Base) (Apache-2.0).
202
+ - The Route B formulation and the calibration/abstention pipeline build on the
203
+ Decily line of work; the external baseline in the tables is
204
+ [Mapika/decider](https://github.com/Mapika/decider).
205
+ - If you use Decily, please cite this model card and link the training
206
+ repository (`TRAINING.md`).
config.json ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen3ForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 151643,
8
+ "eos_token_id": 151643,
9
+ "head_dim": 128,
10
+ "hidden_act": "silu",
11
+ "hidden_size": 2048,
12
+ "initializer_range": 0.02,
13
+ "intermediate_size": 6144,
14
+ "max_position_embeddings": 32768,
15
+ "max_window_layers": 28,
16
+ "model_type": "qwen3",
17
+ "num_attention_heads": 16,
18
+ "num_hidden_layers": 28,
19
+ "num_key_value_heads": 8,
20
+ "rms_norm_eps": 1e-06,
21
+ "rope_scaling": null,
22
+ "rope_theta": 1000000,
23
+ "sliding_window": null,
24
+ "tie_word_embeddings": true,
25
+ "torch_dtype": "bfloat16",
26
+ "transformers_version": "4.51.0",
27
+ "use_cache": true,
28
+ "use_sliding_window": false,
29
+ "vocab_size": 151936,
30
+ "decision_head": {
31
+ "pool": "attention_full_sequence",
32
+ "head_hidden_mult": 4,
33
+ "roles": [
34
+ "state",
35
+ "question",
36
+ "candidate"
37
+ ],
38
+ "concat_order": [
39
+ "state",
40
+ "question",
41
+ "candidate"
42
+ ]
43
+ }
44
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:89aea0a2520db6c10194f1439d1c48d09b93a44f8c2d4afe3715270e9c4225ee
3
+ size 3474788164
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,239 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": false,
3
+ "add_prefix_space": false,
4
+ "added_tokens_decoder": {
5
+ "151643": {
6
+ "content": "<|endoftext|>",
7
+ "lstrip": false,
8
+ "normalized": false,
9
+ "rstrip": false,
10
+ "single_word": false,
11
+ "special": true
12
+ },
13
+ "151644": {
14
+ "content": "<|im_start|>",
15
+ "lstrip": false,
16
+ "normalized": false,
17
+ "rstrip": false,
18
+ "single_word": false,
19
+ "special": true
20
+ },
21
+ "151645": {
22
+ "content": "<|im_end|>",
23
+ "lstrip": false,
24
+ "normalized": false,
25
+ "rstrip": false,
26
+ "single_word": false,
27
+ "special": true
28
+ },
29
+ "151646": {
30
+ "content": "<|object_ref_start|>",
31
+ "lstrip": false,
32
+ "normalized": false,
33
+ "rstrip": false,
34
+ "single_word": false,
35
+ "special": true
36
+ },
37
+ "151647": {
38
+ "content": "<|object_ref_end|>",
39
+ "lstrip": false,
40
+ "normalized": false,
41
+ "rstrip": false,
42
+ "single_word": false,
43
+ "special": true
44
+ },
45
+ "151648": {
46
+ "content": "<|box_start|>",
47
+ "lstrip": false,
48
+ "normalized": false,
49
+ "rstrip": false,
50
+ "single_word": false,
51
+ "special": true
52
+ },
53
+ "151649": {
54
+ "content": "<|box_end|>",
55
+ "lstrip": false,
56
+ "normalized": false,
57
+ "rstrip": false,
58
+ "single_word": false,
59
+ "special": true
60
+ },
61
+ "151650": {
62
+ "content": "<|quad_start|>",
63
+ "lstrip": false,
64
+ "normalized": false,
65
+ "rstrip": false,
66
+ "single_word": false,
67
+ "special": true
68
+ },
69
+ "151651": {
70
+ "content": "<|quad_end|>",
71
+ "lstrip": false,
72
+ "normalized": false,
73
+ "rstrip": false,
74
+ "single_word": false,
75
+ "special": true
76
+ },
77
+ "151652": {
78
+ "content": "<|vision_start|>",
79
+ "lstrip": false,
80
+ "normalized": false,
81
+ "rstrip": false,
82
+ "single_word": false,
83
+ "special": true
84
+ },
85
+ "151653": {
86
+ "content": "<|vision_end|>",
87
+ "lstrip": false,
88
+ "normalized": false,
89
+ "rstrip": false,
90
+ "single_word": false,
91
+ "special": true
92
+ },
93
+ "151654": {
94
+ "content": "<|vision_pad|>",
95
+ "lstrip": false,
96
+ "normalized": false,
97
+ "rstrip": false,
98
+ "single_word": false,
99
+ "special": true
100
+ },
101
+ "151655": {
102
+ "content": "<|image_pad|>",
103
+ "lstrip": false,
104
+ "normalized": false,
105
+ "rstrip": false,
106
+ "single_word": false,
107
+ "special": true
108
+ },
109
+ "151656": {
110
+ "content": "<|video_pad|>",
111
+ "lstrip": false,
112
+ "normalized": false,
113
+ "rstrip": false,
114
+ "single_word": false,
115
+ "special": true
116
+ },
117
+ "151657": {
118
+ "content": "<tool_call>",
119
+ "lstrip": false,
120
+ "normalized": false,
121
+ "rstrip": false,
122
+ "single_word": false,
123
+ "special": false
124
+ },
125
+ "151658": {
126
+ "content": "</tool_call>",
127
+ "lstrip": false,
128
+ "normalized": false,
129
+ "rstrip": false,
130
+ "single_word": false,
131
+ "special": false
132
+ },
133
+ "151659": {
134
+ "content": "<|fim_prefix|>",
135
+ "lstrip": false,
136
+ "normalized": false,
137
+ "rstrip": false,
138
+ "single_word": false,
139
+ "special": false
140
+ },
141
+ "151660": {
142
+ "content": "<|fim_middle|>",
143
+ "lstrip": false,
144
+ "normalized": false,
145
+ "rstrip": false,
146
+ "single_word": false,
147
+ "special": false
148
+ },
149
+ "151661": {
150
+ "content": "<|fim_suffix|>",
151
+ "lstrip": false,
152
+ "normalized": false,
153
+ "rstrip": false,
154
+ "single_word": false,
155
+ "special": false
156
+ },
157
+ "151662": {
158
+ "content": "<|fim_pad|>",
159
+ "lstrip": false,
160
+ "normalized": false,
161
+ "rstrip": false,
162
+ "single_word": false,
163
+ "special": false
164
+ },
165
+ "151663": {
166
+ "content": "<|repo_name|>",
167
+ "lstrip": false,
168
+ "normalized": false,
169
+ "rstrip": false,
170
+ "single_word": false,
171
+ "special": false
172
+ },
173
+ "151664": {
174
+ "content": "<|file_sep|>",
175
+ "lstrip": false,
176
+ "normalized": false,
177
+ "rstrip": false,
178
+ "single_word": false,
179
+ "special": false
180
+ },
181
+ "151665": {
182
+ "content": "<tool_response>",
183
+ "lstrip": false,
184
+ "normalized": false,
185
+ "rstrip": false,
186
+ "single_word": false,
187
+ "special": false
188
+ },
189
+ "151666": {
190
+ "content": "</tool_response>",
191
+ "lstrip": false,
192
+ "normalized": false,
193
+ "rstrip": false,
194
+ "single_word": false,
195
+ "special": false
196
+ },
197
+ "151667": {
198
+ "content": "<think>",
199
+ "lstrip": false,
200
+ "normalized": false,
201
+ "rstrip": false,
202
+ "single_word": false,
203
+ "special": false
204
+ },
205
+ "151668": {
206
+ "content": "</think>",
207
+ "lstrip": false,
208
+ "normalized": false,
209
+ "rstrip": false,
210
+ "single_word": false,
211
+ "special": false
212
+ }
213
+ },
214
+ "additional_special_tokens": [
215
+ "<|im_start|>",
216
+ "<|im_end|>",
217
+ "<|object_ref_start|>",
218
+ "<|object_ref_end|>",
219
+ "<|box_start|>",
220
+ "<|box_end|>",
221
+ "<|quad_start|>",
222
+ "<|quad_end|>",
223
+ "<|vision_start|>",
224
+ "<|vision_end|>",
225
+ "<|vision_pad|>",
226
+ "<|image_pad|>",
227
+ "<|video_pad|>"
228
+ ],
229
+ "bos_token": null,
230
+ "chat_template": "{%- if tools %}\n {{- '<|im_start|>system\\n' }}\n {%- if messages[0].role == 'system' %}\n {{- messages[0].content + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n {%- if messages[0].role == 'system' %}\n {{- '<|im_start|>system\\n' + messages[0].content + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if (message.role == \"user\") or (message.role == \"system\" and not loop.first) %}\n {{- '<|im_start|>' + message.role + '\\n' + message.content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set content = message.content %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is defined and message.reasoning_content is not none %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- else %}\n {%- if '</think>' in message.content %}\n {%- set content = message.content.split('</think>')[-1].lstrip('\\n') %}\n {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\\n').split('<think>')[-1].lstrip('\\n') %}\n {%- endif %}\n {%- endif %}\n {%- if loop.index0 > ns.last_query_index %}\n {%- if loop.last or (not loop.last and reasoning_content) %}\n {{- '<|im_start|>' + message.role + '\\n<think>\\n' + reasoning_content.strip('\\n') + '\\n</think>\\n\\n' + content.lstrip('\\n') }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls %}\n {%- for tool_call in message.tool_calls %}\n {%- if (loop.first and content) or (not loop.first) %}\n {{- '\\n' }}\n {%- endif %}\n {%- if tool_call.function %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {{- '<tool_call>\\n{\"name\": \"' }}\n {{- tool_call.name }}\n {{- '\", \"arguments\": ' }}\n {%- if tool_call.arguments is string %}\n {{- tool_call.arguments }}\n {%- else %}\n {{- tool_call.arguments | tojson }}\n {%- endif %}\n {{- '}\\n</tool_call>' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.first or (messages[loop.index0 - 1].role != \"tool\") %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n<tool_response>\\n' }}\n {{- message.content }}\n {{- '\\n</tool_response>' }}\n {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '<think>\\n\\n</think>\\n\\n' }}\n {%- endif %}\n{%- endif %}",
231
+ "clean_up_tokenization_spaces": false,
232
+ "eos_token": "<|endoftext|>",
233
+ "errors": "replace",
234
+ "model_max_length": 131072,
235
+ "pad_token": "<|endoftext|>",
236
+ "split_special_tokens": false,
237
+ "tokenizer_class": "Qwen2Tokenizer",
238
+ "unk_token": null
239
+ }
vocab.json ADDED
The diff for this file is too large to render. See raw diff