Instructions to use Arc53/arc-decide-0.6b-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Arc53/arc-decide-0.6b-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="Arc53/arc-decide-0.6b-v0.1")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Arc53/arc-decide-0.6b-v0.1") model = AutoModelForCausalLM.from_pretrained("Arc53/arc-decide-0.6b-v0.1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
arc-decide-0.6b-v0.1
One 0.6B model for the small decisions a RAG or agent app makes on every request. Give it a state (text or JSON) and a few typed questions, and it returns calibrated probabilities in one forward pass, with no generated text. It covers relevance, prompt injection, grounding, routing, intent, conversation analytics, custom tags and tool use, in English, Spanish, Chinese and Russian.
| arc-decide-0.6b-v0.1 (0.6B) | Laya (421M) | Qwen3-4B-Instruct (4B) | |
|---|---|---|---|
| Conversation analytics: answer status / reaction (acc) | 83.3 / 85.4 | 26.4 / 34.0 | 60.5 / 56.0 |
| Frustrated / goal met (AUROC) | 0.987 / 0.981 | 0.715 / 0.646 | 0.910 / 0.885 |
| Custom tags, unseen definitions (AUROC) | 0.997 | 0.823 | 0.987 |
| Tools: right tool / result used (AUROC) | 0.811 / 0.954 | 0.456 / 0.516 | 0.712 / 0.791 |
| Intent: Banking77 / CLINC unseen / MASSIVE unseen (acc) | 87.5 / 97.2 / 91.9 | 86.5 / 98.3 / 50.9 | 88.9 / 97.8 / 91.5 |
| Injection: deepset / safeguard (AUROC) | 0.992 / 0.821 | 0.901 / 0.912 | 0.817 / 0.973 |
| Grounding: AggreFact / RAGTruth (bacc) | 0.712 / 0.706 | 0.595 / 0.482 | 0.656 / 0.532 |
| Routing: needs retrieval / clarification (AUROC) | 0.915 / 0.930 | 0.528 / 0.383 | – / – |
| Relevance nDCG@10: DocsGPT docs / SciFact / NQ | 0.905 / 0.725 / 0.553 | 0.397 / 0.378 / 0.365 | 0.916 / 0.592 / 0.482 |
| Relevance nDCG@10: MIRACL es / zh / ru | 0.564 / 0.579 / 0.679 | 0.305 / 0.222 / 0.270 | – / – / – |
All three models get the same held-out sets and the same questions. arc-decide was trained on these kinds of question but never on these sets. Laya (a general-purpose System One checkpoint) and Qwen3-4B-Instruct (read through its label-token probabilities) answer zero-shot. Relevance is on each set's sampled "lite" view. The best of the three is in bold; – means not run.
It powers the decisions in DocsGPT.
Quickstart
from arc_decide import Decider # arc_decide.py ships in this repo (torch + transformers)
d = Decider("arc53/arc-decide-0.6b-v0.1")
chunk = """## Rotating API keys
To rotate a key, open **Settings → API keys**, select the key and click **Rotate**. A new key is issued
immediately; the old key keeps working for 24 hours so you can update your clients, then it is revoked."""
d.decide(
{"query": "How do I rotate an API key?", "passage": chunk},
{"relevant": {"type": "noul", "instructions": "Is the passage relevant to the query: does it contain "
"information that answers or directly addresses it?"}},
)
# {'relevant': {'type': 'noul', 'noul': 0.98}}
Put every question about a state in one call: the state is encoded once and each question reuses its cache. Questions follow TypeSafe's System One shapes:
noul: yes/no, answered with P(yes);choice: one of 2–16 options you pass in;score: a 3- or 4-level rubric.
Use the canonical questions. questions.json has every question the model was trained on, verbatim, under
the name that selects its fitted temperature. They are:
| Family | Question names |
|---|---|
| retrieval | relevant, grade |
| guardrails | injection, severity, supported |
| routing | needs_retrieval, needs_clarification, complexity |
| conversation analytics | answer_status, unanswered_reason, reaction, frustrated, asks_human, sentiment, request_type, goal_met |
| tool analytics | tool_needed, right_tool, tool_succeeded, result_used, induced_by_untrusted |
| options you supply | topic, tag, intent |
States use the key names shown there:
{"query", "passage"}for relevance;{"text"}for injection;{"source", "claim"}for grounding, one claim per call;{"message"}for routing.
For a check of your own, the practised form is
{"type": "noul", "instructions": "Based on the passage, answer this yes/no question:\n<question>?", "criteria": {"true": "Yes", "false": "No"}}.
On CPU, onnx/ holds an fp32 ONNX export. arc_decide_onnx.py runs it with only onnxruntime (1.23 or later),
tokenizers and numpy. Its probabilities match PyTorch to 2e-6. On 4 vCPU of an AMD EPYC:
- an injection check takes 0.3 s;
- an analytics turn (8 questions) takes 6.7 s;
- a 40-chunk relevance prescreen takes 97 s, so send that call to a GPU.
On a GPU server, stock vLLM serves the model, because the heads are folded into the LM head. Start it with
--logprobs-mode processed_logprobs --max-logprobs 16. For each question:
- Send the prompt as token ids, with
max_tokens: 1andallowed_token_idsset to the head's labels fromarcd_serving.json. - Apply the bias and temperature from that file, as
arc_decide.pydoes.
Good to know
- Calibration. Temperatures are fitted per question on held-out data, so 0.5 means even odds. Every analytics
label is within 5.3 points of calibrated on held-out conversations, except topic (11).
- This version has not been re-measured on production traffic.
- Its predecessor was overconfident there on why-unanswered, reaction and request type. Report those three as rates, not thresholded per conversation, until you have checked them on your traffic.
- Not a sole security control. Injection detection is strong on known attack styles (0.99 AUROC on deepset) and weaker on novel ones (0.82–0.85). It flags 12% of benign-but-alarming text at 0.5. Use a threshold of at least 0.9 and other defences too.
- Passage length. Relevance was trained on documentation chunks, from a paragraph up to about 1,400 tokens. One-line snippets score markedly lower, so rank them against each other rather than thresholding them.
- Languages. Most labels are within 1.5 points of English in every language. The exceptions:
- Russian topic is 9 points behind;
- answer status and why-unanswered are 3–5 behind;
- Spanish MIRACL relevance (0.56) is the weakest multilingual score.
- Limits.
choicetakes up to 16 options andscore3 or 4 levels. Questions far from the canonical ones are much less reliable, so validate them on your own data.
Every benchmark, against same-size specialists
Baselines were run on the same harness, following their model cards. The DocsGPT-native and analytics sets come from ten documentation sites held out of all training data. Analytics gold labels are Kimi K2.5 and DeepSeek V4 Pro where the two agree decisively, so they measure agreement with strong models, not human judgement.
Retrieval relevance (nDCG@10)
| set | arc-decide-0.6b-v0.1 (0.6B) | bge-reranker-v2-m3 (568M) | Qwen3-Reranker-0.6B (0.6B) | Qwen3-Reranker-4B (4B) |
|---|---|---|---|---|
| docsgpt-native | 0.916 | 0.881 | 0.874 | 0.912 |
| beir-scifact | 0.756 | 0.731 | 0.762 | 0.783 |
| beir-fiqa | 0.338 | 0.398 | 0.390 | 0.457 |
| beir-nq | 0.539 | 0.582 | 0.530 | 0.586 |
| beir-hotpotqa | 0.759 | 0.787 | 0.756 | 0.780 |
| beir-trec-covid | 0.765 | 0.788 | 0.848 | 0.854 |
Prompt injection (AUROC)
| set | arc-decide-0.6b-v0.1 (0.6B) | Llama Prompt Guard 2 86M (86M) | ProtectAI DeBERTa-v3 injection v2 (184M) | Laya (System One encoder) (421M) |
|---|---|---|---|---|
| inj-deepset | 0.992 | 0.910 | 0.905 | 0.901 |
| inj-jailbreakhub | 0.851 | 0.811 | 0.890 | 0.664 |
| inj-safeguard | 0.821 | 0.974 | 0.988 | 0.912 |
| inj-llmail | 0.997 | 0.986 | 0.986 | 0.990 |
Over-refusal: NotInject false-positive rate at 0.5 (lower is better)
| set | arc-decide-0.6b-v0.1 (0.6B) | Llama Prompt Guard 2 86M (86M) | ProtectAI DeBERTa-v3 injection v2 (184M) | Laya (System One encoder) (421M) |
|---|---|---|---|---|
| inj-notinject | 12.1 | 5.0 | 43.1 | 19.8 |
Grounding: is the claim supported by the source (balanced accuracy)
| set | arc-decide-0.6b-v0.1 (0.6B) | HHEM-2.1-Open (~0.25B) | MiniCheck-RoBERTa-L (355M) | Bespoke-MiniCheck-7B (7B) | Laya (System One encoder) (421M) |
|---|---|---|---|---|---|
| gr-aggrefact | 0.712 | 0.715 | 0.748 | 0.768 | 0.595 |
| gr-ragtruth | 0.706 | 0.660 | 0.593 | 0.721 | 0.482 |
Intent over options given at request time (accuracy)
| set | arc-decide-0.6b-v0.1 (0.6B) | GLiClass modern-base v3 (~150M) | Qwen3-4B-Instruct (label logits) (4B) | Laya (System One encoder) (421M) |
|---|---|---|---|---|
| ch-banking77 | 87.5 | 76.9 | 88.9 | 86.5 |
| ch-clinc-unseen | 97.2 | 85.6 | 97.8 | 98.3 |
| ch-massive-unseen | 91.9 | 49.2 | 91.5 | 50.9 |
Conversation analytics (held-out doc sites, four languages; macro-F1 for choices, balanced accuracy for yes/no; ECE in points)
| label | arc-decide-0.6b-v0.1 | embedding + LR (TnT-LLM), same labels | Qwen3-4B-Instruct (label logits) | Laya (System One encoder) |
|---|---|---|---|---|
| answer_status | 81.8 / 3.3 | 52.6 / 1.9 | 50.8 / 38.5 | 16.9 / 5.1 |
| unanswered_reason | 69.0 / 1.4 | 41.2 / 2.8 | 51.4 / 28.9 | 8.6 / 25.7 |
| reaction | 64.5 / 1.5 | 27.9 / 1.9 | 40.9 / 40.9 | 22.9 / 10.8 |
| request_type | 79.5 / 3.5 | – | – | 30.5 / 20.4 |
| sentiment | 91.1 / 1.7 | 56.5 / 3.0 | 71.2 / 27.7 | 53.3 / 8.3 |
| topic | 52.8 / 11.1 | 22.6 / 6.6 | 37.5 / 56.5 | 24.3 / 21.5 |
| frustrated | 91.1 / 1.5 | 62.3 / 9.4 | 81.2 / 6.8 | 58.6 / 45.2 |
| asks_human | 96.3 / 0.6 | 77.6 / 3.7 | 89.9 / 3.5 | 70.5 / 15.4 |
| goal_met | 88.9 / 5.3 | 72.8 / 2.7 | 79.6 / 14.2 | 54.9 / 4.1 |
| tag | 98.3 / 2.5 | 69.3 / 4.9 | 96.7 / 3.1 | 74.8 / 9.3 |
| right_tool | 68.6 / 3.2 | 57.4 / 5.2 | 64.6 / 23.5 | 50.0 / 23.6 |
| tool_succeeded | 98.1 / 3.7 | 83.9 / 3.3 | 93.0 / 6.2 | 58.7 / 3.4 |
| result_used | 75.5 / 6.4 | 50.4 / 4.3 | 70.0 / 21.8 | 51.3 / 13.1 |
Training
- Model: Qwen3-0.6B-Base with typed linear heads (yes/no, 3- and 4-level scores, a 16-way choice) read at the last token. The questions about one state share its encoding through a prefix fork.
- Objective: trained on soft teacher labels, then temperature-scaled per question on held-out rows.
- Relevance: 640k public retrieval pairs plus questions over 23 documentation sites, labelled by Qwen3.6-27B.
- Guardrails, grounding, routing and QA: public sets, translated to Spanish, Chinese and Russian.
- Analytics:
- about 50k simulated conversations and 10k agent traces over documentation sites, generated with open models and labelled by Gemma 4 and gpt-oss-120b;
- plus WildChat reactions.
- No real user conversations were used.
DATA.md lists every training source, its licence, and every model whose outputs are in the data.
Licence
MIT. The base model, Qwen3-0.6B-Base, is Apache-2.0; its licence ships as LICENSE-Qwen3.
@misc{arc53-arc-decide,
title = {arc-decide: a small calibrated decision model for RAG and agents},
author = {Arc53},
year = {2026},
url = {https://huggingface.co/arc53/arc-decide-0.6b-v0.1}
}
- Downloads last month
- 34
Model tree for Arc53/arc-decide-0.6b-v0.1
Base model
Qwen/Qwen3-0.6B-Base
