arc-decide-0.6b-v0.1

One 0.6B model for the small decisions a RAG or agent app makes on every request. Give it a state (text or JSON) and a few typed questions, and it returns calibrated probabilities in one forward pass, with no generated text. It covers relevance, prompt injection, grounding, routing, intent, conversation analytics, custom tags and tool use, in English, Spanish, Chinese and Russian.

Same questions, same held-out sets

arc-decide-0.6b-v0.1 (0.6B) Laya (421M) Qwen3-4B-Instruct (4B)
Conversation analytics: answer status / reaction (acc) 83.3 / 85.4 26.4 / 34.0 60.5 / 56.0
Frustrated / goal met (AUROC) 0.987 / 0.981 0.715 / 0.646 0.910 / 0.885
Custom tags, unseen definitions (AUROC) 0.997 0.823 0.987
Tools: right tool / result used (AUROC) 0.811 / 0.954 0.456 / 0.516 0.712 / 0.791
Intent: Banking77 / CLINC unseen / MASSIVE unseen (acc) 87.5 / 97.2 / 91.9 86.5 / 98.3 / 50.9 88.9 / 97.8 / 91.5
Injection: deepset / safeguard (AUROC) 0.992 / 0.821 0.901 / 0.912 0.817 / 0.973
Grounding: AggreFact / RAGTruth (bacc) 0.712 / 0.706 0.595 / 0.482 0.656 / 0.532
Routing: needs retrieval / clarification (AUROC) 0.915 / 0.930 0.528 / 0.383 – / –
Relevance nDCG@10: DocsGPT docs / SciFact / NQ 0.905 / 0.725 / 0.553 0.397 / 0.378 / 0.365 0.916 / 0.592 / 0.482
Relevance nDCG@10: MIRACL es / zh / ru 0.564 / 0.579 / 0.679 0.305 / 0.222 / 0.270 – / – / –

All three models get the same held-out sets and the same questions. arc-decide was trained on these kinds of question but never on these sets. Laya (a general-purpose System One checkpoint) and Qwen3-4B-Instruct (read through its label-token probabilities) answer zero-shot. Relevance is on each set's sampled "lite" view. The best of the three is in bold; – means not run.

One model vs five specialists

It powers the decisions in DocsGPT.

Quickstart

from arc_decide import Decider   # arc_decide.py ships in this repo (torch + transformers)

d = Decider("arc53/arc-decide-0.6b-v0.1")
chunk = """## Rotating API keys

To rotate a key, open **Settings → API keys**, select the key and click **Rotate**. A new key is issued
immediately; the old key keeps working for 24 hours so you can update your clients, then it is revoked."""
d.decide(
    {"query": "How do I rotate an API key?", "passage": chunk},
    {"relevant": {"type": "noul", "instructions": "Is the passage relevant to the query: does it contain "
                  "information that answers or directly addresses it?"}},
)
# {'relevant': {'type': 'noul', 'noul': 0.98}}

Put every question about a state in one call: the state is encoded once and each question reuses its cache. Questions follow TypeSafe's System One shapes:

  • noul: yes/no, answered with P(yes);
  • choice: one of 2–16 options you pass in;
  • score: a 3- or 4-level rubric.

Use the canonical questions. questions.json has every question the model was trained on, verbatim, under the name that selects its fitted temperature. They are:

Family Question names
retrieval relevant, grade
guardrails injection, severity, supported
routing needs_retrieval, needs_clarification, complexity
conversation analytics answer_status, unanswered_reason, reaction, frustrated, asks_human, sentiment, request_type, goal_met
tool analytics tool_needed, right_tool, tool_succeeded, result_used, induced_by_untrusted
options you supply topic, tag, intent

States use the key names shown there:

  • {"query", "passage"} for relevance;
  • {"text"} for injection;
  • {"source", "claim"} for grounding, one claim per call;
  • {"message"} for routing.

For a check of your own, the practised form is {"type": "noul", "instructions": "Based on the passage, answer this yes/no question:\n<question>?", "criteria": {"true": "Yes", "false": "No"}}.

On CPU, onnx/ holds an fp32 ONNX export. arc_decide_onnx.py runs it with only onnxruntime (1.23 or later), tokenizers and numpy. Its probabilities match PyTorch to 2e-6. On 4 vCPU of an AMD EPYC:

  • an injection check takes 0.3 s;
  • an analytics turn (8 questions) takes 6.7 s;
  • a 40-chunk relevance prescreen takes 97 s, so send that call to a GPU.

On a GPU server, stock vLLM serves the model, because the heads are folded into the LM head. Start it with --logprobs-mode processed_logprobs --max-logprobs 16. For each question:

  1. Send the prompt as token ids, with max_tokens: 1 and allowed_token_ids set to the head's labels from arcd_serving.json.
  2. Apply the bias and temperature from that file, as arc_decide.py does.

Good to know

  • Calibration. Temperatures are fitted per question on held-out data, so 0.5 means even odds. Every analytics label is within 5.3 points of calibrated on held-out conversations, except topic (11).
    • This version has not been re-measured on production traffic.
    • Its predecessor was overconfident there on why-unanswered, reaction and request type. Report those three as rates, not thresholded per conversation, until you have checked them on your traffic.
  • Not a sole security control. Injection detection is strong on known attack styles (0.99 AUROC on deepset) and weaker on novel ones (0.82–0.85). It flags 12% of benign-but-alarming text at 0.5. Use a threshold of at least 0.9 and other defences too.
  • Passage length. Relevance was trained on documentation chunks, from a paragraph up to about 1,400 tokens. One-line snippets score markedly lower, so rank them against each other rather than thresholding them.
  • Languages. Most labels are within 1.5 points of English in every language. The exceptions:
    • Russian topic is 9 points behind;
    • answer status and why-unanswered are 3–5 behind;
    • Spanish MIRACL relevance (0.56) is the weakest multilingual score.
  • Limits. choice takes up to 16 options and score 3 or 4 levels. Questions far from the canonical ones are much less reliable, so validate them on your own data.
Every benchmark, against same-size specialists

Baselines were run on the same harness, following their model cards. The DocsGPT-native and analytics sets come from ten documentation sites held out of all training data. Analytics gold labels are Kimi K2.5 and DeepSeek V4 Pro where the two agree decisively, so they measure agreement with strong models, not human judgement.

Retrieval relevance (nDCG@10)

set arc-decide-0.6b-v0.1 (0.6B) bge-reranker-v2-m3 (568M) Qwen3-Reranker-0.6B (0.6B) Qwen3-Reranker-4B (4B)
docsgpt-native 0.916 0.881 0.874 0.912
beir-scifact 0.756 0.731 0.762 0.783
beir-fiqa 0.338 0.398 0.390 0.457
beir-nq 0.539 0.582 0.530 0.586
beir-hotpotqa 0.759 0.787 0.756 0.780
beir-trec-covid 0.765 0.788 0.848 0.854

Prompt injection (AUROC)

set arc-decide-0.6b-v0.1 (0.6B) Llama Prompt Guard 2 86M (86M) ProtectAI DeBERTa-v3 injection v2 (184M) Laya (System One encoder) (421M)
inj-deepset 0.992 0.910 0.905 0.901
inj-jailbreakhub 0.851 0.811 0.890 0.664
inj-safeguard 0.821 0.974 0.988 0.912
inj-llmail 0.997 0.986 0.986 0.990

Over-refusal: NotInject false-positive rate at 0.5 (lower is better)

set arc-decide-0.6b-v0.1 (0.6B) Llama Prompt Guard 2 86M (86M) ProtectAI DeBERTa-v3 injection v2 (184M) Laya (System One encoder) (421M)
inj-notinject 12.1 5.0 43.1 19.8

Grounding: is the claim supported by the source (balanced accuracy)

set arc-decide-0.6b-v0.1 (0.6B) HHEM-2.1-Open (~0.25B) MiniCheck-RoBERTa-L (355M) Bespoke-MiniCheck-7B (7B) Laya (System One encoder) (421M)
gr-aggrefact 0.712 0.715 0.748 0.768 0.595
gr-ragtruth 0.706 0.660 0.593 0.721 0.482

Intent over options given at request time (accuracy)

set arc-decide-0.6b-v0.1 (0.6B) GLiClass modern-base v3 (~150M) Qwen3-4B-Instruct (label logits) (4B) Laya (System One encoder) (421M)
ch-banking77 87.5 76.9 88.9 86.5
ch-clinc-unseen 97.2 85.6 97.8 98.3
ch-massive-unseen 91.9 49.2 91.5 50.9

Conversation analytics (held-out doc sites, four languages; macro-F1 for choices, balanced accuracy for yes/no; ECE in points)

label arc-decide-0.6b-v0.1 embedding + LR (TnT-LLM), same labels Qwen3-4B-Instruct (label logits) Laya (System One encoder)
answer_status 81.8 / 3.3 52.6 / 1.9 50.8 / 38.5 16.9 / 5.1
unanswered_reason 69.0 / 1.4 41.2 / 2.8 51.4 / 28.9 8.6 / 25.7
reaction 64.5 / 1.5 27.9 / 1.9 40.9 / 40.9 22.9 / 10.8
request_type 79.5 / 3.5 – – 30.5 / 20.4
sentiment 91.1 / 1.7 56.5 / 3.0 71.2 / 27.7 53.3 / 8.3
topic 52.8 / 11.1 22.6 / 6.6 37.5 / 56.5 24.3 / 21.5
frustrated 91.1 / 1.5 62.3 / 9.4 81.2 / 6.8 58.6 / 45.2
asks_human 96.3 / 0.6 77.6 / 3.7 89.9 / 3.5 70.5 / 15.4
goal_met 88.9 / 5.3 72.8 / 2.7 79.6 / 14.2 54.9 / 4.1
tag 98.3 / 2.5 69.3 / 4.9 96.7 / 3.1 74.8 / 9.3
right_tool 68.6 / 3.2 57.4 / 5.2 64.6 / 23.5 50.0 / 23.6
tool_succeeded 98.1 / 3.7 83.9 / 3.3 93.0 / 6.2 58.7 / 3.4
result_used 75.5 / 6.4 50.4 / 4.3 70.0 / 21.8 51.3 / 13.1

Training

  • Model: Qwen3-0.6B-Base with typed linear heads (yes/no, 3- and 4-level scores, a 16-way choice) read at the last token. The questions about one state share its encoding through a prefix fork.
  • Objective: trained on soft teacher labels, then temperature-scaled per question on held-out rows.
  • Relevance: 640k public retrieval pairs plus questions over 23 documentation sites, labelled by Qwen3.6-27B.
  • Guardrails, grounding, routing and QA: public sets, translated to Spanish, Chinese and Russian.
  • Analytics:
    • about 50k simulated conversations and 10k agent traces over documentation sites, generated with open models and labelled by Gemma 4 and gpt-oss-120b;
    • plus WildChat reactions.
    • No real user conversations were used.

DATA.md lists every training source, its licence, and every model whose outputs are in the data.

Licence

MIT. The base model, Qwen3-0.6B-Base, is Apache-2.0; its licence ships as LICENSE-Qwen3.

@misc{arc53-arc-decide,
  title  = {arc-decide: a small calibrated decision model for RAG and agents},
  author = {Arc53},
  year   = {2026},
  url    = {https://huggingface.co/arc53/arc-decide-0.6b-v0.1}
}
Downloads last month
34
Safetensors
Model size
0.8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Arc53/arc-decide-0.6b-v0.1

Quantized
(94)
this model