Instructions to use ZhangPY/kev-snake-text-grpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ZhangPY/kev-snake-text-grpo with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-4B-Base") model = PeftModel.from_pretrained(base_model, "ZhangPY/kev-snake-text-grpo") - Notebooks
- Google Colab
- Kaggle
kev-snake-text-grpo — a text System-1 decision model (SFT → GRPO)
English | 简体中文
A Kev-style System-1 decision model that plays Snake from a structured text state (grid size, head-to-tail coordinates, food position, current direction). One forward pass, one typed decision (up / down / left / right) with calibrated probabilities — TypeSafe /v1/systemone semantics, served with the unmodified kev.serve stack at ~39 ms/decision.
Trained as Kev-4B → behavior-cloning SFT → GRPO: 22.31 ± 6.01 points over 200 held-out seeds vs 22.11 for the SFT parent (+0.9%).
Left: this model. Right: the hand-written BFS teacher. Same seeds, closed-loop.
Why publish a +0.9% improvement?
Because it is the negative control of this project's flagship finding: RL gain correlates monotonically with policy entropy (8 experiments: entropy 0.14→+0.9%, ~0.2→+2.9%, 0.29→+17%, 1.15→+24% on the vision twin). This model's SFT parent reached 93.7% step-level accuracy, collapsing its policy entropy to 0.14 — the distribution was already sharp, so GRPO had nothing to reshape. Training confirms it: entropy stayed at 0.14–0.15 for all 24 iterations while group return rose only marginally.
This reproduces, on a live game environment, the "sharpening an already-cloned policy" failure mode reported in Kev's own RL pilots — and pins the root cause to SFT-induced entropy collapse. The paired vision model (kev-snake-vision-grpo, entropy 1.15→+24% RL gain at V1.0) completes the 2×2 modality×algorithm matrix.
Results (200 held-out seeds, greedy decoding)
| Policy | Score (mean±std) | Median | Steps | Step acc | Latency (median) |
|---|---|---|---|---|---|
| Random-safe | 1.49 | — | — | — | — |
SFT parent (p1-init) |
22.11 ± 6.24 | 22.0 | 207 | 93.7% | 27.3 ms |
| GRPO (this model) | 22.31 ± 6.01 | 23.0 | 213 | 93.6% | 38.6 ms |
| BFS teacher | 29.03 | — | — | 100% | — |
- Death causes (this model): self-trap 83%, survived-to-timeout 17% — wall and reverse deaths eliminated entirely (SFT parent: 0.5% wall).
- Calibration: single-temperature fit T=1.95 → ECE 0.040 → 0.006 on held-out data.
- Catastrophic forgetting check: 8/8 general-knowledge answers unchanged after RL.
Architecture
Same recipe as Kev-4B, which this model continues from:
structured text state ──► Qwen3.5-4B backbone (frozen)
+ LoRA r=16 α=32 dropout=0.05, 33.8M trainable
(targets incl. DeltaNet in_proj_qkv/z/a/b, out_proj)
▼
pointer head (1.31M, `head.pt`) → P(up/down/left/right)
- Question format:
choicewith 4 criteria-scored options; options can't read each other (System-1 semantics). - Special tokens reuse Qwen's vocab (
<|fim_prefix|>,<|box_start|>, …). - Non-autoregressive single forward per decision.
Usage
Serve with kev (recommended) — this checkpoint is a drop-in --run target:
git clone https://github.com/jaredpalmer/kev.git && cd kev
uv sync --extra serve
uv run --extra serve python -m kev.serve --run ZhangPY/kev-snake-text-grpo --port 8009
Play 200 closed-loop games against the env bundled in code/:
python code/rollout.py --policy model --url http://127.0.0.1:8009 \
--seeds 200 --out closed.json
One decision over HTTP (TypeSafe SDK-compatible):
import json, urllib.request
from snake_env import SnakeEnv # code/
from encode import api_request # builds the /v1/systemone request
env = SnakeEnv(seed=0)
req = urllib.request.Request("http://127.0.0.1:8009/v1/systemone",
data=json.dumps(api_request(env.state())).encode(),
headers={"content-type": "application/json"})
answer = json.loads(urllib.request.urlopen(req).read())["answers"]["move"]
action = max(answer["probabilities"], key=answer["probabilities"].get)
env.step(action)
Training recipe
Stage 0 — init: jaredpalmer/kev-4b (the general Kev-4B decision model). ✅ verified zero-forgetting path: general QA unchanged through all later stages.
Stage 1 — SFT (behavior cloning), 43 min on 1×A100
| Item | Value |
|---|---|
| Data | 8,750 records from 400 BFS-teacher games (length-stratified initial lengths 3/8/15) |
| Command | kev.train --data snake_train.jsonl --base Qwen/Qwen3.5-4B-Base --init_from kev-4b --epochs 2 --lr 2e-5 --batch 4 --accum 2 --bf16 |
| Result | step acc 93.7%, closed-loop 22.11 = 76% of teacher |
Stage 2 — GRPO, 24 iterations × 96 rollouts, ~10 h on 1×A100 (kev.rl, branch devin/1790377716-agentic-rl)
| Hyperparameter | Value |
|---|---|
| group size × episodes | 8 × 12 = 96 rollouts/iter (same-world groups) |
| lr / head lr | 1e-5 / 0 |
| KL anchor to SFT (w) | 0.05 |
| reward | score − tau·death, tau = 2.0 (exploration-first death penalty) |
| replay guard | 8,750 SFT records, w = 0.5 |
| precision | bf16 + gradient checkpointing |
Training log (rl_log.json, 24 iters): group return 10.11 → 12.46; entropy 0.141 → 0.149 (flat — the collapse signal); KL 0.0003 → 0.017; wall-death rate 20.8% → 2.1%.
Files
adapter_config.json / adapter_model.safetensors # PEFT LoRA (load with PeftModel)
head.pt # pointer head (1.31M)
tokenizer.json / tokenizer_config.json / chat_template.jinja
code/ # env + teacher + encoding + evaluator
snake_env.py teacher.py encode.py rollout.py gen_data.py replay_video.py
metrics/
closed_init.json (SFT parent) closed_rl.json (this model, 200 seeds)
rl_config.json / rl_log.json / evals.json / report.json
demo_rl_vs_teacher.gif # same-seed comparison vs BFS teacher
Limitations
- Snake-specific; the state encoding (
code/encode.py) is part of the contract. - +0.9% RL gain is the point (entropy-collapsed policy), but it means this checkpoint is not evidence that GRPO can't help text policies — only that it can't help this one. See the vision twin for the high-entropy counter-example.
- 22.31 vs teacher 29.03: residual gap is long-horizon self-trapping, not perception.
License & attribution
- Apache-2.0 (matching kev and the Qwen base).
- Part of the vla-sft-rl-exam project (P1/P2 reports contain the full SFT double-run and GRPO studies).
- Built on Kev (jaredpalmer) — an open reproduction of TypeSafe's Jev System-1 decision-model API.
- Downloads last month
- -
Model tree for ZhangPY/kev-snake-text-grpo
Base model
Qwen/Qwen3.5-4B-Base