kev-snake-text-grpo — a text System-1 decision model (SFT → GRPO)

English | 简体中文

A Kev-style System-1 decision model that plays Snake from a structured text state (grid size, head-to-tail coordinates, food position, current direction). One forward pass, one typed decision (up / down / left / right) with calibrated probabilities — TypeSafe /v1/systemone semantics, served with the unmodified kev.serve stack at ~39 ms/decision.

Trained as Kev-4B → behavior-cloning SFT → GRPO: 22.31 ± 6.01 points over 200 held-out seeds vs 22.11 for the SFT parent (+0.9%).

RL vs BFS teacher, same seeds Left: this model. Right: the hand-written BFS teacher. Same seeds, closed-loop.

Why publish a +0.9% improvement?

Because it is the negative control of this project's flagship finding: RL gain correlates monotonically with policy entropy (8 experiments: entropy 0.14→+0.9%, ~0.2→+2.9%, 0.29→+17%, 1.15→+24% on the vision twin). This model's SFT parent reached 93.7% step-level accuracy, collapsing its policy entropy to 0.14 — the distribution was already sharp, so GRPO had nothing to reshape. Training confirms it: entropy stayed at 0.14–0.15 for all 24 iterations while group return rose only marginally.

This reproduces, on a live game environment, the "sharpening an already-cloned policy" failure mode reported in Kev's own RL pilots — and pins the root cause to SFT-induced entropy collapse. The paired vision model (kev-snake-vision-grpo, entropy 1.15→+24% RL gain at V1.0) completes the 2×2 modality×algorithm matrix.

Results (200 held-out seeds, greedy decoding)

Policy Score (mean±std) Median Steps Step acc Latency (median)
Random-safe 1.49 — — — —
SFT parent (p1-init) 22.11 ± 6.24 22.0 207 93.7% 27.3 ms
GRPO (this model) 22.31 ± 6.01 23.0 213 93.6% 38.6 ms
BFS teacher 29.03 — — 100% —
  • Death causes (this model): self-trap 83%, survived-to-timeout 17% — wall and reverse deaths eliminated entirely (SFT parent: 0.5% wall).
  • Calibration: single-temperature fit T=1.95 → ECE 0.040 → 0.006 on held-out data.
  • Catastrophic forgetting check: 8/8 general-knowledge answers unchanged after RL.

Architecture

Same recipe as Kev-4B, which this model continues from:

structured text state ──► Qwen3.5-4B backbone (frozen)
                            + LoRA r=16 α=32 dropout=0.05, 33.8M trainable
                            (targets incl. DeltaNet in_proj_qkv/z/a/b, out_proj)
                            ▼
                          pointer head (1.31M, `head.pt`) → P(up/down/left/right)
  • Question format: choice with 4 criteria-scored options; options can't read each other (System-1 semantics).
  • Special tokens reuse Qwen's vocab (<|fim_prefix|>, <|box_start|>, …).
  • Non-autoregressive single forward per decision.

Usage

Serve with kev (recommended) — this checkpoint is a drop-in --run target:

git clone https://github.com/jaredpalmer/kev.git && cd kev
uv sync --extra serve
uv run --extra serve python -m kev.serve --run ZhangPY/kev-snake-text-grpo --port 8009

Play 200 closed-loop games against the env bundled in code/:

python code/rollout.py --policy model --url http://127.0.0.1:8009 \
    --seeds 200 --out closed.json

One decision over HTTP (TypeSafe SDK-compatible):

import json, urllib.request
from snake_env import SnakeEnv           # code/
from encode import api_request           # builds the /v1/systemone request

env = SnakeEnv(seed=0)
req = urllib.request.Request("http://127.0.0.1:8009/v1/systemone",
                             data=json.dumps(api_request(env.state())).encode(),
                             headers={"content-type": "application/json"})
answer = json.loads(urllib.request.urlopen(req).read())["answers"]["move"]
action = max(answer["probabilities"], key=answer["probabilities"].get)
env.step(action)

Training recipe

Stage 0 — init: jaredpalmer/kev-4b (the general Kev-4B decision model). ✅ verified zero-forgetting path: general QA unchanged through all later stages.

Stage 1 — SFT (behavior cloning), 43 min on 1×A100

Item Value
Data 8,750 records from 400 BFS-teacher games (length-stratified initial lengths 3/8/15)
Command kev.train --data snake_train.jsonl --base Qwen/Qwen3.5-4B-Base --init_from kev-4b --epochs 2 --lr 2e-5 --batch 4 --accum 2 --bf16
Result step acc 93.7%, closed-loop 22.11 = 76% of teacher

Stage 2 — GRPO, 24 iterations × 96 rollouts, ~10 h on 1×A100 (kev.rl, branch devin/1790377716-agentic-rl)

Hyperparameter Value
group size × episodes 8 × 12 = 96 rollouts/iter (same-world groups)
lr / head lr 1e-5 / 0
KL anchor to SFT (w) 0.05
reward score − tau·death, tau = 2.0 (exploration-first death penalty)
replay guard 8,750 SFT records, w = 0.5
precision bf16 + gradient checkpointing

Training log (rl_log.json, 24 iters): group return 10.11 → 12.46; entropy 0.141 → 0.149 (flat — the collapse signal); KL 0.0003 → 0.017; wall-death rate 20.8% → 2.1%.

Files

adapter_config.json / adapter_model.safetensors   # PEFT LoRA (load with PeftModel)
head.pt                                           # pointer head (1.31M)
tokenizer.json / tokenizer_config.json / chat_template.jinja
code/                                             # env + teacher + encoding + evaluator
  snake_env.py  teacher.py  encode.py  rollout.py  gen_data.py  replay_video.py
metrics/
  closed_init.json (SFT parent)   closed_rl.json (this model, 200 seeds)
rl_config.json / rl_log.json / evals.json / report.json
demo_rl_vs_teacher.gif                            # same-seed comparison vs BFS teacher

Limitations

  • Snake-specific; the state encoding (code/encode.py) is part of the contract.
  • +0.9% RL gain is the point (entropy-collapsed policy), but it means this checkpoint is not evidence that GRPO can't help text policies — only that it can't help this one. See the vision twin for the high-entropy counter-example.
  • 22.31 vs teacher 29.03: residual gap is long-horizon self-trapping, not perception.

License & attribution

  • Apache-2.0 (matching kev and the Qwen base).
  • Part of the vla-sft-rl-exam project (P1/P2 reports contain the full SFT double-run and GRPO studies).
  • Built on Kev (jaredpalmer) — an open reproduction of TypeSafe's Jev System-1 decision-model API.
Downloads last month
-
Video Preview
loading

Model tree for ZhangPY/kev-snake-text-grpo

Adapter
(78)
this model