Instructions to use Qwen/Qwen3.8-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen3.8-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Qwen/Qwen3.8-27B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Qwen/Qwen3.8-27B") model = AutoModelForMultimodalLM.from_pretrained("Qwen/Qwen3.8-27B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- HuggingChat
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Qwen/Qwen3.8-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen3.8-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Qwen/Qwen3.8-27B
- SGLang
How to use Qwen/Qwen3.8-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.8-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.8-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Qwen/Qwen3.8-27B with Docker Model Runner:
docker model run hf.co/Qwen/Qwen3.8-27B
NVFP4 Shootout (Quality and Speed)
I picked some of the most popular NVFP4 quantizations of Qwen3.8-27B and compared them in quality and speed.
Contestants
These are the models tested in these benchmarks:
| Tag | Checkpoint | Size | Notes |
|---|---|---|---|
BF16 |
Qwen/Qwen3.8-27B | 51.77 GB | reference (unquantized) |
FP8 |
Qwen/Qwen3.8-27B-FP8 | 28.77 GB | official FP8 (near-lossless anchor) |
quasar |
QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 | 19.17 GB | QAT, all-W4A4 |
unsloth |
unsloth/Qwen3.8-27B-NVFP4 | 21.83 GB | PTQ, mixed-precision (FP8 attn + NVFP4 MLP) |
radixark_bf16lm |
RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead | 22.14 GB | ModelOpt PTQ, BF16 LM head |
radixark |
RadixArk/Qwen3.8-27B-NVFP4 | 20.44 GB | ModelOpt PTQ, NVFP4 LM head |
nvidia |
nvidia/Qwen3.8-27B-NVFP4 | 21.92 GB | ModelOpt v0.48.0, mixed NVFP4 (MLP + lm_head) / FP8 (attention) |
Note: I also had RedhatAI's NVFP4 quantization on the list, but then found out that it's actually just a byte-identical copy of the unsloth quantization and therefore dropped it.
Setup
Hardware: single NVIDIA DGX Spark (GB10, 128 GB unified memory)
Engine: vLLM (nightly snapshot)
Every variant was served with an identical config: --max-model-len 262144 (native, no YaRN), --kv-cache-dtype fp8, --gpu-memory-utilization 0.86, MTP speculative decoding {"method":"mtp","num_speculative_tokens":5}, --max-logprobs 256. All generation tiers ran at temperature 0 (deterministic); the agentic tier ran at the deployment regime (temp 1.0, thinking-on).
--gpu-memory-utilizationwas set to 0.86 for the benchmark (vs. 0.94 in production) to leave headroom for the extra transient memory the logprob-based tiers (PPL/KLD) and the stability probes place on the unified 128 GB.
Benchmarks executed
I benchmarked the five NVFP4 quantizations of Qwen3.8-27B against the official BF16 (reference) and FP8 checkpoints, on a coding/agent-oriented panel (perplexity, distribution fidelity, code generation, math, agentic tool-use, throughput):
| Tier | Measures | Setting |
|---|---|---|
| PPL (wikitext-2, 323 windows) | next-token perplexity | greedy, prompt_logprobs |
| KLD (95 prompts: 49 text + 46 code) | per-position KL vs BF16 reference (p50, robust) | greedy, top-128 prompt_logprobs |
| Codegen (MBPP 257 + HumanEval 164) | unit-test pass rate | thinking OFF |
| GSM8K (500) | exact-match math accuracy | thinking OFF |
| Tool-use (20 agentic scenarios) | task success + tool-call validity | thinking ON (xhigh), multi-turn |
| Perf (10 prompts) | small-context decode throughput | thinking OFF |
Results
Deltas are vs. the BF16 reference (unquantized) model.
| Variant | PPL wiki ↓ | KLD text ↓ | KLD code ↓ | MBPP % ↑ | HumanEval % ↑ | GSM8K % ↑ | Tool succ % ↑ | tok/s ↑ | MTP acceptance ↑ |
|---|---|---|---|---|---|---|---|---|---|
| Qwen BF16 (ref) | 7.993 | — | — | 70.8 | 93.3 | 97.0 | 100.0 | 14.6 | 0.914 |
| Qwen FP8 | 8.029 (+0.037) | 0.0117 | N/A | 69.3 (−1.6) | 95.1 (+1.8) | 97.4 (+0.4) | 100.0 | 20.6 | 0.959 |
| QUASAR-QAT NVFP4 | 8.247 (+0.254) | 0.0682 | 0.0272 | 68.9 (−2.0) | 93.9 (+0.6) | 96.8 (−0.2) | 100.0 | 25.4 | 0.906 |
| unsloth NVFP4 | 8.131 (+0.138) | 0.0475 | 0.0197 | 68.9 (−2.0) | 95.1 (+1.8) | 97.2 (+0.2) | 100.0 | 31.2 | 0.884 |
| RadixArk (BF16 LMHead) | 8.286 (+0.293) | 0.0592 | 0.0263 | 70.0 (−0.8) | 91.5 (−1.8) | 97.0 (0.0) | 100.0 | 22.7 | 0.917 |
| RadixArk (NVFP4 LMHead) | 8.385 (+0.392) | 0.0952 | 0.0335 | 70.0 (−0.8) | 93.9 (+0.6) | 97.4 (+0.4) | 100.0 | 34.6 | 0.912 |
| NVIDIA NVFP4 | 8.139 (+0.146) | 0.0758 | 0.0297 | 70.4 (−0.4) | 94.5 (+1.2) | 97.2 (+0.2) | 100.0 | 38.6 | 0.897 |
Notes:
- ↓ lower is better.
- ↑ higher is better.
- KLD is reported as p50 (median per-position KL), robust to a few high-divergence positions.
- KLD text (49 wikitext prompts) is the apples-to-apples column (clean for every variant).
- KLD code (46 code prompts) is clean for the NVFP4 candidates.
- FP8 code-prompt NaN defect (upstream finding, not a harness issue): the official FP8 checkpoint emits NaN
prompt_logprobson code-like text (14 NaN + 7 extreme + 18 elevated, all code; text clean). Generation is unaffected. Consequently FP8 KLD code is N/A; its KLD text (0.0117) confirms near-lossless. - MTP acceptance is a short-sample health-gate estimate; the operator's real-time dashboard is authoritative. All variants show healthy acceptance (0.88–0.96).
- Codegen scoring was validated against ground truth (correct solutions pass, broken ones fail). MBPP prompts are given the function signature (the dataset's natural-language prompts omit it).
- PPL is measured per-variant with each variant's own tokenizer (standard); the wikitext-2 test split is the corpus.
- unsloth's KLD is valid despite a different tokenizer file. Its tokenizer's core vocabulary is byte-identical to the reference (verified 7121/7121 positions align), so its KLD is comparable.
Fidelity Rankings
Closest to the BF16 reference first (ranked by KLD text, the most sensitive metric; PPL shown for reference):
| Rank | Variant | KLD text ↓ | PPL wiki ↓ |
|---|---|---|---|
| 1 | Qwen FP8 | 0.0117 | 8.029 |
| 2 | unsloth NVFP4 | 0.0475 | 8.131 |
| 3 | RadixArk (BF16 LMHead) | 0.0592 | 8.286 |
| 4 | QUASAR-QAT NVFP4 | 0.0682 | 8.247 |
| 5 | NVIDIA NVFP4 | 0.0758 | 8.139 |
| 6 | RadixArk (NVFP4) | 0.0952 | 8.385 |
Speed Rankings
Fastest first (single-stream, small-context decode; min / mean / peak):
| Rank | Variant | min tok/s | mean tok/s | peak tok/s |
|---|---|---|---|---|
| 1 | NVIDIA NVFP4 | 32.85 | 38.55 | 45.85 |
| 2 | RadixArk (NVFP4) | 30.32 | 34.62 | 42.31 |
| 3 | unsloth NVFP4 | 26.91 | 31.21 | 36.74 |
| 4 | QUASAR-QAT NVFP4 | 23.04 | 25.35 | 28.94 |
| 5 | RadixArk (BF16 LMHead) | 20.24 | 22.74 | 27.28 |
| 6 | Qwen FP8 | 18.47 | 20.58 | 24.73 |
| — | Qwen BF16 (ref) | 13.17 | 14.61 | 16.61 |
The ordering is stable across min / mean / peak — NVIDIA leads on all three.
Bottom line
- All five NVFP4 quants are near-lossless on task accuracy — within ~1–2 points of BF16 on MBPP, HumanEval, GSM8K, and agentic tool-use.
- unsloth is closest in quality to reference BF16 (best PPL and KLD) and 3rd fastest.
- NVIDIA is now fastest (38.6 tok/s).
- RadixArk (NVFP4 LMHead) is 2nd fastest but farthest in quality from BF16.
- The two RadixArk models (LM-head ablation): Quantizing the LM head to NVFP4 (vs keeping it BF16) costs ~0.1 PPL and ~0.04 KLD-text but buys ~12 tok/s (34.6 vs 22.7) and ~1.8 GB. A clear accuracy↔speed tradeoff localized to the output head.
- Speed is not purely size-driven. The NVFP4 quants (19–22 GB) run 1.7–2.6× faster than BF16 (51.77 GB). Within the NVFP4 set, NVIDIA (21.92 GB) is fastest (38.6 tok/s) and QUASAR (19.17 GB, the smallest) is just 4th (25.4 tok/s) - so speed scales with checkpoint size and kernel path (i.e. the kernel path and mixed-precision recipe also matter).
Recommendation
The numbers narrow it to two clear leaders; the other NVFP4 builds are mid-pack and are not the best on any axis.
- unsloth NVFP4 — the overall leader. Best fidelity (closest to BF16 on PPL +0.138 and KLD 0.0475 text / 0.0197 code), competitive on all task metrics, and 3rd-fastest (31.2 tok/s). The default pick when fidelity matters most.
- NVIDIA — the speed leader. Fastest (38.6 tok/s), while yet retaining excellent fidelity and decent overall quality results.
So it is that simple, right? Not quite. Because during my tests I encountered some stability issues. The unsloth model appears to suffer from numerical instability, causing issues like garbled / gibberish output under rare situations. Furthermore all NVFP4 quantizations appear to be prone to reasoning loops, some maybe more often than others, but I could not clearly tell which one is better or worse on this regard yet. It would make sense to run a separate benchmark that just covers these stability issues. In general I have the impression though, that this Qwen3.8-27 base model is much more sensitive to NVFP4 quantizations than the previous generations.
Note: these were taken under the assumption of using MTP (speculative decoding). If you don't want to or cannot use MTP, then the picture might be different. I have not tested without MTP, because in production MTP is typically enabled and is a huge performance lever, without sacrificing quality (equal sample distribution).
History:
- 2026-09-09: Initial test results (four NVFP4 quants: unsloth, RadixArk (BF16 LMHead), RadixArk (NVFP4 LMHead), QUASAR-QAT).
- 2026-09-12: Added results of Nvidia NVFP4 quantization, and fidelity/speed ranking tables.
Thanks for putting this together, really useful comparison! I’m one of the people behind the QUASAR checkpoint.
Two things interesting in your results: QUASAR quantizes all 496/496 transformer linear layers to NVFP4 W4A4 yet the model still keeps strong GSM8K/HumanEval performance and full tool-use success.
The current checkpoint also includes the MTP head and at 19.2 GB it’s the most compressed model in the comparison.
Really appreciate you testing it independently alongside the other NVFP4 releases.
NVIDIA just released an official Qwen3.8-27B NVFP4 checkpoint too, would be really interesting to add it to this comparison:
https://huggingface.co/nvidia/Qwen3.8-27B-NVFP4
It’s a somewhat different point on the tradeoff: 21.9 GB and mixed NVFP4/FP8 (~5.5 bits), with attention / linear-attention kept in FP8.
@vincentcounathe Thanks!
I just added the results of the Nvidia NVFP4 quantization - which is the new speed leader.
The big question: whether the stability issues can still be addressed. I mean I have deployed these NVFP4 against real world coding tasks and they all solved the tasks with similar success rate. So a pick probably rather boils down to speed and reliability, and not sub-digit differences in KLD et al. I first thought only some checkpoints like unsloth would suffer from reliability issues, but I reproduced at least the reasoning loops with all NVFP4 checkpoints now. Is that something you've been aware of already?
Yes we’ve seen some related reports, including occasional early stopping on the QUASAR model. But similar looping issues also seem to happen with other Qwen3.8 quants so it’s not clear yet that NVFP4 is the cause.
If you have a few prompts where BF16/FP8 works but NVFP4 loops, that would be very useful to test.
https://huggingface.co/RedHatAI/Qwen3.8-27B-INT4 is it possible to compare with int4 quant, this one for example?
@vincentcounathe Yes, I can imagine it might happen with other quants as well. I'm still trying to make sense of it how these reasoning loops are triggered. So far it seems content/task/prompt agnostic and rather more of a setup dependent issue. For instance with OpenCode as harness, reasoning effort "low" and temperature 0.6 I can reproduce it in average in 30 - 90 turns (no matter if easy or hard goals).
With OpenCode, but reasoning effort "xhigh" and/or temperature 1.0 I could not reproduce it at all OTOH.
@igor255 I could run one INT4 checkpoint in comparison, but I have other priorities right now. What is it that you are trying to learn from an INT4 comparison?
I try to learn what is best quant for single user and how much accuracy is damaged by using 4 bit activators as it is in NVFP4.
For single user nvfp brings benefit only on prefill speed, not token generation AFAIK.
From my research best quality int4 quant may be https://huggingface.co/cyankiwi/Qwen3.8-27B-AWQ-INT4/tree/main - as it has group size 32 asymmetric quantization. Other are usually group size 128 symmetric.
If you run just one - i recommend to use this one.
Thanks in advance.
Thanks a lot for doing this!
One suggestion: DeepSWE 1.1 or Terminal Bench 4.0 could be better benchmarks, since the ones you used are very saturated and easy for this model at almost any quant.
Thank you so much for doing this. I came here from the QUASAR page and this is very insightful, thanks for sharing this.
I have a 5090 and 64GB of system RAM. Up until now, I was just using the Q8_0 GGUF from unsloth on llama.cpp with q8_0, q8_0 and it was working awesome. With the issue that I can't go over 95K context which makes it really challenging on some larger projects. pi still works, because it keeps compacting and so far, it seems to have worked okay.
I've found compared to Q8_0, even Q6 feels like a lobotomized version of the model. Making silly mistakes here and there. Granted, this is my experience carried over from Qwen3.6 and I haven't tried it in Qwen3.8. But NVFP4 allows me the full 256K context which is just mind boggling.
I've only just switched to NVFP4 so it's too soon to say, but does anyone here have anecdotal evidence about how NVFP4 is in actual practice compared to Q8_0? Mainly looking at it from a coding point of view, but other viewpoints are also welcome.
If it helps, here are the commands I use for both:
Q8 on llama.cpp:
build/bin/llama-server \
-m ~/myp/models/unsloth/qwen3.8/Qwen3.8-27B-Q8_0.gguf \
--temp 1.0 \
--top_p 0.95 \
--top_k 20 \
--min_p 0.0 \
--repeat-penalty 1.0 \
--presence-penalty 0.0 \
-c 95000 \
-t 16 \
-ngl 99 \
--flash-attn on \
--host 0.0.0.0 --port 8080 \
--mmproj ~/myp/models/unsloth/qwen3.8/mmproj-F16.gguf \
--no-mmproj-offload \
--spec-type draft-mtp --spec-draft-n-max 4 --parallel 1 -kvo \
-ctk q8_0 -ctv q8_0 -b 1024 -ub 256
NVFP4 on vllm:
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
MAX_JOBS=4 \
vllm serve /home/lenny/myp/models/quasar-qat \
--max-model-len 262144 \
--max-num-seqs=4 \
--kv-cache-memory=10244183819 \
--max-num-batched-tokens 4096 \
--kv-cache-dtype fp8 \
--served-model-name Qwen3.8-27B-NVFP4 \
--host=0.0.0.0 \
--port 8080 \
--language-model-only \
--enable-prefix-caching \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser=qwen3
Quality wise, the best results will be exl3 quants imo, u can run 200k context with 6 bit quant.
https://huggingface.co/turboderp/Qwen3.8-27B-exl3
But i currently run 4bit quant in vllm from cyanwiki (tweaked) - i dont see very big difference quality wise, but its almost 50% faster for me.
https://huggingface.co/igor255/Qwen3.8-27B-AWQ-INT4-H8-MTP4