MLX
Safetensors
ONNX
English
qwen3
decision-model
candidate-scorer
calibration
selective-prediction
on-device
Instructions to use alexzhang0118/Decily-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use alexzhang0118/Decily-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download alexzhang0118/Decily-MLX --local-dir Decily-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Upload folder using huggingface_hub
Browse files- README.md +206 -0
- config.json +44 -0
- model.safetensors +3 -0
- tokenizer.json +0 -0
- tokenizer_config.json +239 -0
- vocab.json +0 -0
README.md
ADDED
|
@@ -0,0 +1,206 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
base_model: Qwen/Qwen3-1.7B-Base
|
| 6 |
+
tags:
|
| 7 |
+
- decision-model
|
| 8 |
+
- candidate-scorer
|
| 9 |
+
- calibration
|
| 10 |
+
- selective-prediction
|
| 11 |
+
- on-device
|
| 12 |
+
- mlx
|
| 13 |
+
- onnx
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# Decily — calibrated candidate-scorer decision model
|
| 17 |
+
|
| 18 |
+
**Decily** scores a runtime-provided set of candidates against a `state` and a
|
| 19 |
+
`question`, and returns a **calibrated probability distribution** over those
|
| 20 |
+
candidates. It does not generate text. Decily is built for decisions where you
|
| 21 |
+
want a probability you can threshold, not a sentence you have to parse.
|
| 22 |
+
|
| 23 |
+
```python
|
| 24 |
+
Decily.decide(
|
| 25 |
+
state="Our app crashes on startup after the latest update.",
|
| 26 |
+
question="What is the customer's intent?",
|
| 27 |
+
options=["technical", "billing", "shipping", "returns"],
|
| 28 |
+
) # -> {"technical": 0.9999, "billing": 0.0001, "returns": 0.0, "shipping": 0.0}
|
| 29 |
+
# measured with the released weights (bf16, T=0.45)
|
| 30 |
+
```
|
| 31 |
+
|
| 32 |
+
## This repository
|
| 33 |
+
|
| 34 |
+
**Apple MLX bf16 bundle** (3.47 GB) — backbone and scoring head in one
|
| 35 |
+
safetensors file, the layout the standalone reference implementation
|
| 36 |
+
(`mlx/decision_mlx.py` in [arczhi/decily](https://github.com/arczhi/decily))
|
| 37 |
+
expects.
|
| 38 |
+
|
| 39 |
+
```bash
|
| 40 |
+
python mlx/decision_mlx.py --model-dir <this repo> \
|
| 41 |
+
--state "Our app crashes on startup after the latest update." \
|
| 42 |
+
--question "What is the customer's intent?" \
|
| 43 |
+
--options "technical,billing,shipping,returns" --temperature 0.45
|
| 44 |
+
```
|
| 45 |
+
|
| 46 |
+
## Model family
|
| 47 |
+
|
| 48 |
+
This card covers all four released artifacts; they are the same model in
|
| 49 |
+
different containers.
|
| 50 |
+
|
| 51 |
+
| Variant | Format | Size | Target |
|
| 52 |
+
|---|---|---|---|
|
| 53 |
+
| **Decily-1.7B** | PyTorch bf16 safetensors + `config.json` | ~3.4 GB | server / GPU |
|
| 54 |
+
| **Decily-MLX** | MLX bf16 safetensors | ~3.4 GB | Apple Silicon |
|
| 55 |
+
| **Decily-MLX-4bit** | MLX 4-bit (group-64 affine) | 968 MB | on-device / low memory |
|
| 56 |
+
| **Decily-ONNX-int8** | ONNX int8 (single file) | 1.66 GB | CPU / Windows / edge |
|
| 57 |
+
|
| 58 |
+
## Model details
|
| 59 |
+
|
| 60 |
+
- **Model type:** cross-encoder candidate scorer. The backbone encodes
|
| 61 |
+
`state + question + candidate` jointly; an attention-pooling layer over all
|
| 62 |
+
tokens feeds a LayerNorm/GELU MLP that outputs one logit per candidate.
|
| 63 |
+
Probabilities come from the softmax over the candidates you pass in.
|
| 64 |
+
- **Backbone:** `Qwen/Qwen3-1.7B-Base` (1.72B parameters, Apache-2.0), with an
|
| 65 |
+
attention-pooling scoring head (`decision_head` recorded in `config.json`;
|
| 66 |
+
pool/head weights live in the checkpoint).
|
| 67 |
+
- **Input limits (training configuration):** `state` ≤256 tokens, `question`
|
| 68 |
+
≤96 tokens, each candidate ≤64 tokens, 2–16 candidates per call. Longer
|
| 69 |
+
inputs are truncated — for long documents, chunk the state and score chunks
|
| 70 |
+
separately (the cross-encoder logits are candidate-set independent, so
|
| 71 |
+
chunked shortlisting then re-ranking is lossless; see the repository's
|
| 72 |
+
two-stage inference utility).
|
| 73 |
+
- **Precision:** bf16 (PyTorch/MLX), 4-bit affine group-64 (on-device),
|
| 74 |
+
int8 (ONNX).
|
| 75 |
+
- **Temperature:** raw logits are over-confident out of distribution. Fitted
|
| 76 |
+
temperatures observed: ≈1.1 (in-domain), ≈0.45 on unseen label sets, ≈1.35 on
|
| 77 |
+
the zero-shot fair suite. Always report/calibrate the temperature for your
|
| 78 |
+
data (see Evaluation).
|
| 79 |
+
|
| 80 |
+
## Training summary
|
| 81 |
+
|
| 82 |
+
Full, reproducible recipe: [TRAINING.md](https://github.com/arczhi/decily/blob/main/TRAINING.md)
|
| 83 |
+
in the training repository [arczhi/decily](https://github.com/arczhi/decily)
|
| 84 |
+
(design rationale in `DESIGN.md`, all experiments and ablations in
|
| 85 |
+
`TRAIN-REPORT.md`).
|
| 86 |
+
|
| 87 |
+
| Stage | What | Notes |
|
| 88 |
+
|---|---|---|
|
| 89 |
+
| Teacher | 45-task Qwen3.5-2B decision model (LoRA, frozen backbone) | reference labeler |
|
| 90 |
+
| SFT | Qwen3-1.7B LoRA on 24 task families (hard labels) | 3000 steps ≈1 h |
|
| 91 |
+
| RLCD members | v1 / v2 LoRA + one full-FT member | belief calibration, abstention |
|
| 92 |
+
| Distillation | 4-model probability-space ensemble → **single full-FT student** | KD T=2, α=0.5, +15% belief rows |
|
| 93 |
+
| v5 (this model) | full fine-tune, 3000 steps, effective batch 16, lr 1e-5, 8-bit AdamW | ≈1.5 h on one RTX 5090 32 GB |
|
| 94 |
+
|
| 95 |
+
Key recipe properties:
|
| 96 |
+
|
| 97 |
+
- **Route B (explicit candidate scorer)** beats the LM-head-letter-logits route
|
| 98 |
+
(Route A) by **+10.6 pt** held-out accuracy in a controlled same-backbone test.
|
| 99 |
+
- **Ensemble distillation:** the 4-model ensemble reaches NLL 1.455; the
|
| 100 |
+
distilled single model reaches **1.451** at 1× inference cost.
|
| 101 |
+
- **Consumer hardware:** the whole pipeline trains on a single RTX 5090 32 GB
|
| 102 |
+
in well under a day.
|
| 103 |
+
|
| 104 |
+
## Evaluation
|
| 105 |
+
|
| 106 |
+
Protocol: accuracy / NLL / ECE after **temperature fitting** (raw T=1 numbers
|
| 107 |
+
are over-confident). "In-task" = the 24 trained task families; "held-out" =
|
| 108 |
+
unseen label sets; "fair suite" = 8 tasks unseen by both this model and the
|
| 109 |
+
external baseline.
|
| 110 |
+
|
| 111 |
+
| Evaluation | Decily (24 tasks) | decider-2b (95 tasks) |
|
| 112 |
+
|---|---|---|
|
| 113 |
+
| In-task, 24 task families (shared training tasks) | **0.863 / 0.390 / 0.034** | 0.811 / 0.453 / 0.032 |
|
| 114 |
+
| Fair suite, zero-shot (8 tasks × 300) | 0.654 / 0.86 / 0.092 | **0.700 / 0.71 / 0.047** |
|
| 115 |
+
| Held-out label sets (60/77-class intents, 1200) | 0.581 / 1.451 / 0.068 | — (scores include tasks it was trained on) |
|
| 116 |
+
|
| 117 |
+
Selective prediction (unseen label sets): taking only the most-confident **5%**
|
| 118 |
+
of predictions gives **93.3%** accuracy (an SFT baseline with the same protocol
|
| 119 |
+
gives 79%); abstention thresholds can be calibrated to a target error rate
|
| 120 |
+
(empirically 4.7% achieved at a 5% target).
|
| 121 |
+
|
| 122 |
+
Honest reading of the numbers:
|
| 123 |
+
|
| 124 |
+
- Decily leads the larger, 95-task decider-2b by **+5.2 pt** on tasks both were
|
| 125 |
+
trained on, at 1.7B vs 2B parameters and ~1/4 of the task coverage.
|
| 126 |
+
- On tasks **neither** model has seen, decider-2b leads by 4.6 pt; the gap is
|
| 127 |
+
concentrated in knowledge-heavy tasks (sciq / pubmedqa / quality).
|
| 128 |
+
- The zero-shot gap is an **efficiency** result, not a ceiling: with this
|
| 129 |
+
recipe, task coverage — not parameter count — is the documented main lever.
|
| 130 |
+
Scaling the 24-task mixture toward 60–95 tasks with the same pipeline is
|
| 131 |
+
expected to close and exceed the baseline (see `TRAINING.md`).
|
| 132 |
+
|
| 133 |
+
## Uses
|
| 134 |
+
|
| 135 |
+
**Direct use**
|
| 136 |
+
|
| 137 |
+
- Classification / ranking / selection when the label set is known at runtime
|
| 138 |
+
and may change per call (intents, topics, sentiment, NLI-style relations,
|
| 139 |
+
multiple-choice answers, routing, moderation labels, tool selection).
|
| 140 |
+
- Calibration-sensitive automation: threshold the probability, auto-handle the
|
| 141 |
+
confident head and escalate the rest to a human or a larger model.
|
| 142 |
+
- On-device / private inference: 968 MB 4-bit build, no GPU required.
|
| 143 |
+
- Synthetic decision data generation and reward/verifier scoring for LLM
|
| 144 |
+
pipelines (it returns probabilities, not text).
|
| 145 |
+
|
| 146 |
+
**Out of scope**
|
| 147 |
+
|
| 148 |
+
- Text generation, chat, summarization.
|
| 149 |
+
- Open-ended label spaces (the candidate set must be provided at call time).
|
| 150 |
+
- High-stakes decisions without a calibrated threshold and human review.
|
| 151 |
+
- Very long inputs beyond the 256-token state window without chunking.
|
| 152 |
+
|
| 153 |
+
## Quick start
|
| 154 |
+
|
| 155 |
+
Apple Silicon (MLX) — the standalone reference implementation ships with the
|
| 156 |
+
training repository (`mlx/decision_mlx.py`, pure `mlx.core`, no torch):
|
| 157 |
+
|
| 158 |
+
```bash
|
| 159 |
+
python mlx/decision_mlx.py --model-dir <Decily-MLX dir> \
|
| 160 |
+
--state "The invoice was paid twice..." \
|
| 161 |
+
--question "What is the customer's intent?" \
|
| 162 |
+
--options "billing,technical,shipping,returns" \
|
| 163 |
+
--temperature 0.45
|
| 164 |
+
```
|
| 165 |
+
|
| 166 |
+
4-bit on-device: load the quantized tower with `mlx_lm.load` and apply the
|
| 167 |
+
attention-pooling head from `decision_head.safetensors` (both files are in the
|
| 168 |
+
repository; `config.json` records the head spec).
|
| 169 |
+
|
| 170 |
+
PyTorch: load `model.safetensors` + `config.json` with the training repository's
|
| 171 |
+
model class; the checkpoint keeps backbone and head under their training key
|
| 172 |
+
names (`model.*`, `pool.*`, `head.*`).
|
| 173 |
+
|
| 174 |
+
## Limitations
|
| 175 |
+
|
| 176 |
+
- **Task coverage:** trained on 24 task families; zero-shot behavior on unseen
|
| 177 |
+
domains trails a 95-task baseline (see Evaluation). Validate on your domain
|
| 178 |
+
before trusting the probabilities.
|
| 179 |
+
- **Calibration drift:** the model needs temperature scaling; the fitted value
|
| 180 |
+
depends on the data distribution (≈0.45–1.35 in our measurements).
|
| 181 |
+
- **Input truncation:** `state` is truncated at 256 tokens; long documents must
|
| 182 |
+
be chunked (importance within a chunk, then combine).
|
| 183 |
+
- **Language:** training data is predominantly English (some multilingual
|
| 184 |
+
intent data); other languages are untested.
|
| 185 |
+
- **No safety layer:** outputs are probabilities over the candidates you
|
| 186 |
+
provide; the model does not refuse or filter inputs. Do not expose it as an
|
| 187 |
+
autonomous decision-maker in safety-critical settings.
|
| 188 |
+
- **Inherited biases:** as a fine-tune of Qwen3-1.7B-Base on public datasets,
|
| 189 |
+
it can reproduce biases present in those datasets (toxicity, sentiment and
|
| 190 |
+
bias-classification tasks were part of the mixture, which reduces but does
|
| 191 |
+
not eliminate this).
|
| 192 |
+
|
| 193 |
+
## Environmental impact
|
| 194 |
+
|
| 195 |
+
Training used a single RTX 5090 32 GB: teacher ≈4 h, SFT ≈1 h, members ≈1 h,
|
| 196 |
+
distillation ≈1.5 h (plus data conversion). Total well under 10 GPU-hours; no
|
| 197 |
+
model was trained more than once per stage.
|
| 198 |
+
|
| 199 |
+
## Citation and acknowledgements
|
| 200 |
+
|
| 201 |
+
- Backbone: [Qwen3-1.7B-Base](https://huggingface.co/Qwen/Qwen3-1.7B-Base) (Apache-2.0).
|
| 202 |
+
- The Route B formulation and the calibration/abstention pipeline build on the
|
| 203 |
+
Decily line of work; the external baseline in the tables is
|
| 204 |
+
[Mapika/decider](https://github.com/Mapika/decider).
|
| 205 |
+
- If you use Decily, please cite this model card and link the training
|
| 206 |
+
repository (`TRAINING.md`).
|
config.json
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"Qwen3ForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"attention_bias": false,
|
| 6 |
+
"attention_dropout": 0.0,
|
| 7 |
+
"bos_token_id": 151643,
|
| 8 |
+
"eos_token_id": 151643,
|
| 9 |
+
"head_dim": 128,
|
| 10 |
+
"hidden_act": "silu",
|
| 11 |
+
"hidden_size": 2048,
|
| 12 |
+
"initializer_range": 0.02,
|
| 13 |
+
"intermediate_size": 6144,
|
| 14 |
+
"max_position_embeddings": 32768,
|
| 15 |
+
"max_window_layers": 28,
|
| 16 |
+
"model_type": "qwen3",
|
| 17 |
+
"num_attention_heads": 16,
|
| 18 |
+
"num_hidden_layers": 28,
|
| 19 |
+
"num_key_value_heads": 8,
|
| 20 |
+
"rms_norm_eps": 1e-06,
|
| 21 |
+
"rope_scaling": null,
|
| 22 |
+
"rope_theta": 1000000,
|
| 23 |
+
"sliding_window": null,
|
| 24 |
+
"tie_word_embeddings": true,
|
| 25 |
+
"torch_dtype": "bfloat16",
|
| 26 |
+
"transformers_version": "4.51.0",
|
| 27 |
+
"use_cache": true,
|
| 28 |
+
"use_sliding_window": false,
|
| 29 |
+
"vocab_size": 151936,
|
| 30 |
+
"decision_head": {
|
| 31 |
+
"pool": "attention_full_sequence",
|
| 32 |
+
"head_hidden_mult": 4,
|
| 33 |
+
"roles": [
|
| 34 |
+
"state",
|
| 35 |
+
"question",
|
| 36 |
+
"candidate"
|
| 37 |
+
],
|
| 38 |
+
"concat_order": [
|
| 39 |
+
"state",
|
| 40 |
+
"question",
|
| 41 |
+
"candidate"
|
| 42 |
+
]
|
| 43 |
+
}
|
| 44 |
+
}
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:89aea0a2520db6c10194f1439d1c48d09b93a44f8c2d4afe3715270e9c4225ee
|
| 3 |
+
size 3474788164
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,239 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_bos_token": false,
|
| 3 |
+
"add_prefix_space": false,
|
| 4 |
+
"added_tokens_decoder": {
|
| 5 |
+
"151643": {
|
| 6 |
+
"content": "<|endoftext|>",
|
| 7 |
+
"lstrip": false,
|
| 8 |
+
"normalized": false,
|
| 9 |
+
"rstrip": false,
|
| 10 |
+
"single_word": false,
|
| 11 |
+
"special": true
|
| 12 |
+
},
|
| 13 |
+
"151644": {
|
| 14 |
+
"content": "<|im_start|>",
|
| 15 |
+
"lstrip": false,
|
| 16 |
+
"normalized": false,
|
| 17 |
+
"rstrip": false,
|
| 18 |
+
"single_word": false,
|
| 19 |
+
"special": true
|
| 20 |
+
},
|
| 21 |
+
"151645": {
|
| 22 |
+
"content": "<|im_end|>",
|
| 23 |
+
"lstrip": false,
|
| 24 |
+
"normalized": false,
|
| 25 |
+
"rstrip": false,
|
| 26 |
+
"single_word": false,
|
| 27 |
+
"special": true
|
| 28 |
+
},
|
| 29 |
+
"151646": {
|
| 30 |
+
"content": "<|object_ref_start|>",
|
| 31 |
+
"lstrip": false,
|
| 32 |
+
"normalized": false,
|
| 33 |
+
"rstrip": false,
|
| 34 |
+
"single_word": false,
|
| 35 |
+
"special": true
|
| 36 |
+
},
|
| 37 |
+
"151647": {
|
| 38 |
+
"content": "<|object_ref_end|>",
|
| 39 |
+
"lstrip": false,
|
| 40 |
+
"normalized": false,
|
| 41 |
+
"rstrip": false,
|
| 42 |
+
"single_word": false,
|
| 43 |
+
"special": true
|
| 44 |
+
},
|
| 45 |
+
"151648": {
|
| 46 |
+
"content": "<|box_start|>",
|
| 47 |
+
"lstrip": false,
|
| 48 |
+
"normalized": false,
|
| 49 |
+
"rstrip": false,
|
| 50 |
+
"single_word": false,
|
| 51 |
+
"special": true
|
| 52 |
+
},
|
| 53 |
+
"151649": {
|
| 54 |
+
"content": "<|box_end|>",
|
| 55 |
+
"lstrip": false,
|
| 56 |
+
"normalized": false,
|
| 57 |
+
"rstrip": false,
|
| 58 |
+
"single_word": false,
|
| 59 |
+
"special": true
|
| 60 |
+
},
|
| 61 |
+
"151650": {
|
| 62 |
+
"content": "<|quad_start|>",
|
| 63 |
+
"lstrip": false,
|
| 64 |
+
"normalized": false,
|
| 65 |
+
"rstrip": false,
|
| 66 |
+
"single_word": false,
|
| 67 |
+
"special": true
|
| 68 |
+
},
|
| 69 |
+
"151651": {
|
| 70 |
+
"content": "<|quad_end|>",
|
| 71 |
+
"lstrip": false,
|
| 72 |
+
"normalized": false,
|
| 73 |
+
"rstrip": false,
|
| 74 |
+
"single_word": false,
|
| 75 |
+
"special": true
|
| 76 |
+
},
|
| 77 |
+
"151652": {
|
| 78 |
+
"content": "<|vision_start|>",
|
| 79 |
+
"lstrip": false,
|
| 80 |
+
"normalized": false,
|
| 81 |
+
"rstrip": false,
|
| 82 |
+
"single_word": false,
|
| 83 |
+
"special": true
|
| 84 |
+
},
|
| 85 |
+
"151653": {
|
| 86 |
+
"content": "<|vision_end|>",
|
| 87 |
+
"lstrip": false,
|
| 88 |
+
"normalized": false,
|
| 89 |
+
"rstrip": false,
|
| 90 |
+
"single_word": false,
|
| 91 |
+
"special": true
|
| 92 |
+
},
|
| 93 |
+
"151654": {
|
| 94 |
+
"content": "<|vision_pad|>",
|
| 95 |
+
"lstrip": false,
|
| 96 |
+
"normalized": false,
|
| 97 |
+
"rstrip": false,
|
| 98 |
+
"single_word": false,
|
| 99 |
+
"special": true
|
| 100 |
+
},
|
| 101 |
+
"151655": {
|
| 102 |
+
"content": "<|image_pad|>",
|
| 103 |
+
"lstrip": false,
|
| 104 |
+
"normalized": false,
|
| 105 |
+
"rstrip": false,
|
| 106 |
+
"single_word": false,
|
| 107 |
+
"special": true
|
| 108 |
+
},
|
| 109 |
+
"151656": {
|
| 110 |
+
"content": "<|video_pad|>",
|
| 111 |
+
"lstrip": false,
|
| 112 |
+
"normalized": false,
|
| 113 |
+
"rstrip": false,
|
| 114 |
+
"single_word": false,
|
| 115 |
+
"special": true
|
| 116 |
+
},
|
| 117 |
+
"151657": {
|
| 118 |
+
"content": "<tool_call>",
|
| 119 |
+
"lstrip": false,
|
| 120 |
+
"normalized": false,
|
| 121 |
+
"rstrip": false,
|
| 122 |
+
"single_word": false,
|
| 123 |
+
"special": false
|
| 124 |
+
},
|
| 125 |
+
"151658": {
|
| 126 |
+
"content": "</tool_call>",
|
| 127 |
+
"lstrip": false,
|
| 128 |
+
"normalized": false,
|
| 129 |
+
"rstrip": false,
|
| 130 |
+
"single_word": false,
|
| 131 |
+
"special": false
|
| 132 |
+
},
|
| 133 |
+
"151659": {
|
| 134 |
+
"content": "<|fim_prefix|>",
|
| 135 |
+
"lstrip": false,
|
| 136 |
+
"normalized": false,
|
| 137 |
+
"rstrip": false,
|
| 138 |
+
"single_word": false,
|
| 139 |
+
"special": false
|
| 140 |
+
},
|
| 141 |
+
"151660": {
|
| 142 |
+
"content": "<|fim_middle|>",
|
| 143 |
+
"lstrip": false,
|
| 144 |
+
"normalized": false,
|
| 145 |
+
"rstrip": false,
|
| 146 |
+
"single_word": false,
|
| 147 |
+
"special": false
|
| 148 |
+
},
|
| 149 |
+
"151661": {
|
| 150 |
+
"content": "<|fim_suffix|>",
|
| 151 |
+
"lstrip": false,
|
| 152 |
+
"normalized": false,
|
| 153 |
+
"rstrip": false,
|
| 154 |
+
"single_word": false,
|
| 155 |
+
"special": false
|
| 156 |
+
},
|
| 157 |
+
"151662": {
|
| 158 |
+
"content": "<|fim_pad|>",
|
| 159 |
+
"lstrip": false,
|
| 160 |
+
"normalized": false,
|
| 161 |
+
"rstrip": false,
|
| 162 |
+
"single_word": false,
|
| 163 |
+
"special": false
|
| 164 |
+
},
|
| 165 |
+
"151663": {
|
| 166 |
+
"content": "<|repo_name|>",
|
| 167 |
+
"lstrip": false,
|
| 168 |
+
"normalized": false,
|
| 169 |
+
"rstrip": false,
|
| 170 |
+
"single_word": false,
|
| 171 |
+
"special": false
|
| 172 |
+
},
|
| 173 |
+
"151664": {
|
| 174 |
+
"content": "<|file_sep|>",
|
| 175 |
+
"lstrip": false,
|
| 176 |
+
"normalized": false,
|
| 177 |
+
"rstrip": false,
|
| 178 |
+
"single_word": false,
|
| 179 |
+
"special": false
|
| 180 |
+
},
|
| 181 |
+
"151665": {
|
| 182 |
+
"content": "<tool_response>",
|
| 183 |
+
"lstrip": false,
|
| 184 |
+
"normalized": false,
|
| 185 |
+
"rstrip": false,
|
| 186 |
+
"single_word": false,
|
| 187 |
+
"special": false
|
| 188 |
+
},
|
| 189 |
+
"151666": {
|
| 190 |
+
"content": "</tool_response>",
|
| 191 |
+
"lstrip": false,
|
| 192 |
+
"normalized": false,
|
| 193 |
+
"rstrip": false,
|
| 194 |
+
"single_word": false,
|
| 195 |
+
"special": false
|
| 196 |
+
},
|
| 197 |
+
"151667": {
|
| 198 |
+
"content": "<think>",
|
| 199 |
+
"lstrip": false,
|
| 200 |
+
"normalized": false,
|
| 201 |
+
"rstrip": false,
|
| 202 |
+
"single_word": false,
|
| 203 |
+
"special": false
|
| 204 |
+
},
|
| 205 |
+
"151668": {
|
| 206 |
+
"content": "</think>",
|
| 207 |
+
"lstrip": false,
|
| 208 |
+
"normalized": false,
|
| 209 |
+
"rstrip": false,
|
| 210 |
+
"single_word": false,
|
| 211 |
+
"special": false
|
| 212 |
+
}
|
| 213 |
+
},
|
| 214 |
+
"additional_special_tokens": [
|
| 215 |
+
"<|im_start|>",
|
| 216 |
+
"<|im_end|>",
|
| 217 |
+
"<|object_ref_start|>",
|
| 218 |
+
"<|object_ref_end|>",
|
| 219 |
+
"<|box_start|>",
|
| 220 |
+
"<|box_end|>",
|
| 221 |
+
"<|quad_start|>",
|
| 222 |
+
"<|quad_end|>",
|
| 223 |
+
"<|vision_start|>",
|
| 224 |
+
"<|vision_end|>",
|
| 225 |
+
"<|vision_pad|>",
|
| 226 |
+
"<|image_pad|>",
|
| 227 |
+
"<|video_pad|>"
|
| 228 |
+
],
|
| 229 |
+
"bos_token": null,
|
| 230 |
+
"chat_template": "{%- if tools %}\n {{- '<|im_start|>system\\n' }}\n {%- if messages[0].role == 'system' %}\n {{- messages[0].content + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou may call one or more functions to assist with the user query.\\n\\nYou are provided with function signatures within <tools></tools> XML tags:\\n<tools>\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n</tools>\\n\\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\\n<tool_call>\\n{\\\"name\\\": <function-name>, \\\"arguments\\\": <args-json-object>}\\n</tool_call><|im_end|>\\n\" }}\n{%- else %}\n {%- if messages[0].role == 'system' %}\n {{- '<|im_start|>system\\n' + messages[0].content + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if (message.role == \"user\") or (message.role == \"system\" and not loop.first) %}\n {{- '<|im_start|>' + message.role + '\\n' + message.content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set content = message.content %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is defined and message.reasoning_content is not none %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- else %}\n {%- if '</think>' in message.content %}\n {%- set content = message.content.split('</think>')[-1].lstrip('\\n') %}\n {%- set reasoning_content = message.content.split('</think>')[0].rstrip('\\n').split('<think>')[-1].lstrip('\\n') %}\n {%- endif %}\n {%- endif %}\n {%- if loop.index0 > ns.last_query_index %}\n {%- if loop.last or (not loop.last and reasoning_content) %}\n {{- '<|im_start|>' + message.role + '\\n<think>\\n' + reasoning_content.strip('\\n') + '\\n</think>\\n\\n' + content.lstrip('\\n') }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls %}\n {%- for tool_call in message.tool_calls %}\n {%- if (loop.first and content) or (not loop.first) %}\n {{- '\\n' }}\n {%- endif %}\n {%- if tool_call.function %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {{- '<tool_call>\\n{\"name\": \"' }}\n {{- tool_call.name }}\n {{- '\", \"arguments\": ' }}\n {%- if tool_call.arguments is string %}\n {{- tool_call.arguments }}\n {%- else %}\n {{- tool_call.arguments | tojson }}\n {%- endif %}\n {{- '}\\n</tool_call>' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.first or (messages[loop.index0 - 1].role != \"tool\") %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n<tool_response>\\n' }}\n {{- message.content }}\n {{- '\\n</tool_response>' }}\n {%- if loop.last or (messages[loop.index0 + 1].role != \"tool\") %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '<think>\\n\\n</think>\\n\\n' }}\n {%- endif %}\n{%- endif %}",
|
| 231 |
+
"clean_up_tokenization_spaces": false,
|
| 232 |
+
"eos_token": "<|endoftext|>",
|
| 233 |
+
"errors": "replace",
|
| 234 |
+
"model_max_length": 131072,
|
| 235 |
+
"pad_token": "<|endoftext|>",
|
| 236 |
+
"split_special_tokens": false,
|
| 237 |
+
"tokenizer_class": "Qwen2Tokenizer",
|
| 238 |
+
"unk_token": null
|
| 239 |
+
}
|
vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|