Qwen2.5-Coder-14B TSP RL Policy, Seed 101

This is the frozen step-2500 policy for seed 101 from the bounded TSP experiment in "Beyond Inference-Time Search: Reinforcement Learning Synthesizes Reusable Solvers." Training starts directly from the Qwen base model and uses the same GRPO solver-synthesis method as the SDS and JSSP experiments, without an SFT stage.

Public code commit: https://github.com/IDEALLab/neural-solver-synthesis/commit/07a798e7d7eca736cd1ef13a15209d402d401ef6

Final aggregate evidence and correction history: https://huggingface.co/datasets/IDEALLab/Neural-Solver-Synthesis-Final-Evidence-v1

Limitations

The predeclared TSP quality/stability gate failed, and native 2-opt and OR-Tools remained stronger. A parser correction was applied after outcomes were observed, so this experiment is boundary evidence rather than blind confirmation. It does not establish tuning-free transfer or native-solver dominance.

files.sha256.json records SHA-256 hashes of every inference file uploaded from the frozen checkpoint. Optimizer, scheduler, RNG, and trainer state are excluded.

Downloads last month
26
Safetensors
Model size
15B params
Tensor type
BF16
·
Video Preview
loading

Model tree for IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed101

Base model

Qwen/Qwen2.5-14B
Finetuned
(126)
this model

Collection including IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed101