Instructions to use JANGQ-AI/GLM-5.3-Flash-JANGH2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use JANGQ-AI/GLM-5.3-Flash-JANGH2 with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("JANGQ-AI/GLM-5.3-Flash-JANGH2") config = load_config("JANGQ-AI/GLM-5.3-Flash-JANGH2") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use JANGQ-AI/GLM-5.3-Flash-JANGH2 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/GLM-5.3-Flash-JANGH2"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "JANGQ-AI/GLM-5.3-Flash-JANGH2" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use JANGQ-AI/GLM-5.3-Flash-JANGH2 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/GLM-5.3-Flash-JANGH2"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default JANGQ-AI/GLM-5.3-Flash-JANGH2
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use JANGQ-AI/GLM-5.3-Flash-JANGH2 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/GLM-5.3-Flash-JANGH2"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "JANGQ-AI/GLM-5.3-Flash-JANGH2" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- JANGQ-AI/GLM-5.3-Flash-JANGH2
- Revision 2: long reasoning now ends on its own
- JANGH vs JANGTQ
- Quality vs the official FP8 release
- Fidelity vs the bf16 model on real transcripts
- Agentic fidelity vs the bf16 model itself
- Live behavior (served, temperature 0)
- Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3)
- What's in the bundle
- Serving contract
- Build details
- Revision 2: long reasoning now ends on its own
Run JANG models in MLX Studio / vMLX
✅ Runtime: supported in vMLX (Python). Use vMLX 1.6.68 or newer (native JANGH loading; the vMLX 1.6.71 release was validated with GLM JANGH on native reasoning efforts and chat tool/media continuations). Older vMLX builds refuse the bundle at load time rather than producing wrong output. Known limitation (vMLX 1.6.71 notes): video perception is imperfect — the model can describe overlay or transition text that is not in the clip.
JANGQ-AI/GLM-5.3-Flash-JANGH2
GLM-5.3-Flash for 128 GB Macs. Same size as our previous affine release, with 2.9x lower median KL, +5.3 points top-1, and faster decode.
A JANGH bundle of zai-org/GLM-5.3-Flash: a 300B-class MoE (288 routed experts, top-8 + shared expert) with KDA linear attention, sparse attention, and vision + video towers.
- Routed experts: JANGH at 2-3 bits. Codebook-quantized with per-row scales and a blockwise Hadamard rotation, calibrated with GPTQ + a per-expert importance matrix. Bits are placed per layer by measurement.
- Everything else: 8-bit (MXFP8 where the source weights sit on an FP8-MX grid, affine 8-bit elsewhere) or full precision.
- Vision tower: kept in bf16.
Revision 2: long reasoning now ends on its own
The first upload could not end long reasoning. After a few thousand reasoning tokens the probability of </think> was
hundreds to thousands of times too low, so the model kept thinking, and the text eventually degraded into repetition.
With the default effort (max) that range is reached on ordinary tasks.
Cause: error-minimizing row scales shrink every low-bit matrix a little (12% at 2 bits). Three matrices per expert and 42 MoE layers deep, that weakens exactly the strong, late decisions, and "stop thinking" is one of them. Revision 2 changes only the row scales of the routed experts (126 small tensors). The quantized codes, the size, the format and the speed are unchanged, and no runtime change is needed.
| where the reference model ends its reasoning | bf16 | revision 2 | revision 1 | affine JANG (previous) |
|---|---|---|---|---|
P(</think>), held-out design task, after 8,025 reasoning tokens |
0.82 | 0.75 | 0.03 | 0.02 |
P(</think>), coding task, after 3,288 reasoning tokens |
0.99 | 0.94 | 0.19 | - |
| long design prompt, served, vendor sampling: reasoning ended by the model | revision 2 | revision 1 |
|---|---|---|
effort low |
3 / 3 | 0 / 5 |
effort high |
2 / 2 | - |
effort max (40k-token budget; the reference model needs 17-23k reasoning tokens here) |
1 / 1 | 0 / 2 |
Your copy is revision 2 if jang_config.json has a scale_correction entry.
This replaces JANGQ-AI/GLM-5.3-Flash-JANG and -JANG-MTP (and was briefly published as GLM-5.3-Flash-JANGTQ2).
JANGH vs JANGTQ
JANGH replaces our earlier JANGTQ (TurboQuant) expert format. The TurboQuant rotation is no longer used: no seeded random-sign rotation and no per-dimension Beta-distribution codebook, so runtimes do not need to reproduce any RNG or dimension-specific tables. JANGH uses a fixed blockwise Hadamard-32 transform (no random signs) and a two-constant codebook per bit width, stored bit-for-bit in MLX's own packing layout, so the runtime reuses MLX's kernel structure for decode and prefill.
For compatibility with runtimes already being built against it, the on-disk identifiers keep their original names:
config.json → jangtq block (version: 2), per-module "mode": "jangtq2", and tensors *.tq2_packed /
*.tq2_scales. These are the JANGH format; they are not loadable by, and must not be routed to, JANGTQ v1 loaders.
Quality vs the official FP8 release
15,830 teacher-forced positions on held-out prompts, top-128 renormalized KL.
| Bundle | Size | median KL ↓ | mean KL ↓ | p90 / p95 / p99 ↓ | top-1 ↑ | top-5 ↑ | top-10 ↑ |
|---|---|---|---|---|---|---|---|
| GLM-5.3-Flash-JANGH2 | 95.89 GiB | 0.0301 | 0.376 | 0.98 / 1.97 / 5.24 | 83.9% | 96.8% | 98.3% |
| GLM-5.3-Flash-JANG (affine, previous release) | 95.35 GiB | 0.0882 | 0.528 | 1.50 / 2.55 / 5.67 | 78.6% | 94.6% | 96.8% |
orcarouter GLM-5.3-Flash-MLX 2bit-lite ¹ |
95.4 GiB | 0.2122 | 0.83 | — | 71.4% | 90.8% | 94.3% |
¹ Measured earlier on the same 20 prompts (15,850 positions), loading its shipped quantized weights natively.
Fidelity vs the bf16 model on real transcripts
Scored against the full bf16 model, run layer by layer from disk. 16 held-out transcripts written by the model itself in its native format: long design and coding reasoning, and multi-step tool conversations with tool results (60,243 assistant positions).
| JANGH2 | affine JANG (previous) | |
|---|---|---|
| top-1 agreement with bf16 | 82.6% | 77.1% |
| median KL vs bf16 | 0.069 | 0.158 |
| mean KL vs bf16 | 0.218 | 0.356 |
tool-call points: <tool_call> is the top choice (bf16 agrees on all 41) |
40 / 41 | 35 / 41 |
median P(<tool_call>) at those points |
0.9998 | 0.968 |
lowest P(<tool_call>) at those points |
0.165 | 0.245 |
Agentic fidelity vs the bf16 model itself
Scored against the full bf16 model, run layer by layer from disk. 72 held-out tool-use conversations (25,867 positions; tools and phrasings disjoint from calibration), including 114 "call a tool or answer?" decision points.
| JANGH2 | affine JANG (previous) | |
|---|---|---|
| median KL vs bf16 | 0.579 | 0.595 |
| top-1 agreement with bf16 | 57.3% | 54.9% |
| tool-call decisions: same choice as bf16 | 50 / 50 | 50 / 50 |
median P(<tool_call>) at call points (bf16: 0.998) |
0.999 | 0.981 |
lowest P(<tool_call>) at a call point |
0.914 | 0.784 |
| answer decisions: same next token as bf16 | 87.5% | 50.0% |
| decisions flipped call ↔ answer vs bf16 | 0 | 3 |
These are synthetic agent transcripts rendered with empty think blocks, so absolute KL is high for every bundle; read the columns against each other. Revision 2 gives up fidelity on this set (revision 1: median KL 0.145, top-1 68.7%) in exchange for ending long reasoning; on assistant output alone top-1 is 92.3% (revision 1: 94.4%). Tool decisions are unchanged. The calibration set includes agent-style conversations from the same generator (different tools and wording), so part of this gain is in-distribution. The FP8 table above is independent of that.
Live behavior (served, temperature 0)
| suite | JANGH2 | affine JANG (previous) |
|---|---|---|
| 48-case tool-use eval (required / auto × thinking on / off) | 48 / 48 | 48 / 48 |
| 16 behavior probes: math, follow-up, tool round trip, efforts low/high/max, image, video | 16 / 16 | 16 / 16 |
| same 16 probes after a server restart (SSD prefix-cache restore) | 16 / 16 | — |
Both bundles saturate these suites, so they check behavior rather than rank the two.
Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3)
| JANGH2 | affine JANG (previous) | |
|---|---|---|
| decode (tok/s) | 27.71 / 27.71 | 26.81 / 26.85 |
| prefill, ~5.2k-token prompt (tok/s) | 339 / 361 | 421 / 416 |
| peak memory while serving | 96.7 GiB | 96.1 GiB |
Two interleaved runs per bundle, same runtime build, each run the median of 3 probes on a never-seen prompt. Decode is ~3% faster than the affine bundle; long-prompt prefill is currently ~15% slower.
What's in the bundle
- Vision + video: full bf16 vision tower + the consolidated image/video processor config.
- No MTP: layer 45 is omitted; its bytes went into expert precision.
- Thinking + agentic:
- Thinking is ON by default (the template opens
<think>). - Reasoning efforts are
low/high/max(defaultmax). There is nomedium: the template renders any other value, and a missing value, as Max. - Budget for thinking: the model reasons for 8-10k tokens at
lowand 17-23k atmaxon a large design task, and for 4-6k atmaxeven on a small coding task. Useloworhighfor interactive work and a generousmax_tokens. clear_thinking=falsepreserves thinking in history.
- Thinking is ON by default (the template opens
- Tool calls: GLM's XML dialect (
<tool_call>name<arg_key>…</arg_key><arg_value>…</arg_value></tool_call>), declared astool_parser: glm_xml_args; tool results render as<|observation|>. Hermes-style JSON parsers will not work. - Self-describing:
config.jsoncarries the JANGH format block (codebook, packing, rotation, method) and a per-modulequantizationmap (126 expert projections, 147 MXFP8, 214 affine 8-bit).jang_config.jsonrecords calibration and per-layer expert bits.- Raw evaluation results are in
evaluation/.
- Memory: 95.89 GiB weights, 96.7 GiB peak while serving. Fixed-size linear-attention state + compressed-latent KV (~6 KB/token), so long contexts do not balloon memory.
- Every shard is alignment-safe (zero-copy memory mapping).
Serving contract
- Sampling:
temperature=1.0, top_p=0.95(vendor defaults), no repetition penalty - EOS:
[154820, 154827, 154829]· context: 1M native - Reasoning:
reasoning_effortchat-template kwarg (low/high/max), defaultmax - Thinking off: the template always opens
<think>. Runtimes must close it in GLM's native form,<think></think>with no whitespace; an R1-style\n</think>\n\nmeasurably degrades thinking-off tool decisions.
Build details
- Source:
zai-org/GLM-5.3-Flash-BF16@a5b45eb - Calibration:
- 600k tokens (web / code / multi-turn chat incl. tool transcripts / math), referenced to the official FP8 release
- plus a bf16 agentic capture (448 GLM-template tool conversations), with evaluation prompts held out
- Experts: JANGH (odd-cubic codebook, fp16 per-row scale, Hadamard-32 rotation), GPTQ on all 42 MoE layers with a per-expert importance matrix, bit allocation measured per layer (gate/up 2-3 bit, down 2-3 bit)
- Revision 2: row scales corrected to unit gain along the source row (
jang_config.json→scale_correction)
Quantized and validated by Jinho Jang — eric@jangq.ai
- Downloads last month
- 399
8-bit
Model tree for JANGQ-AI/GLM-5.3-Flash-JANGH2
Base model
zai-org/GLM-5.3-Flash
