The ultimate guide to multi-harness RL
Train open models with RL inside real agent harnesses
Try the three environments, follow the tutorial, and explore the datasets, models and historical RL/SFT results.
Train open models with RL inside real agent harnesses
Note The article: The ultimate guide to multi-harness RL
Answer dataset questions and view scoring results
Note Start here: solve a data task in the SETA whitebox playground. These same Bash tools are used by GRPO.
Run OpenCode tasks and view grading results
Note Deprecated native OpenCode interface, kept for the historical comparison. Use Harbor with harness="opencode" for new runs.
Visualize RL agent comparison metrics in a web dashboard
Note RL training dashboard: LFM and Qwen reward curves, per-harness checkpoint pass@1, tool/token efficiency and evaluation coverage. Opens the LFM comparison by default.
Explore and manage tasks in the Harbor agentic environment
Note Harbor multi-harness blackbox. Browse tasks, connect a model, run a harness and inspect live traces.
Visualize SFT experiment metrics from a JSON file
Note SFT training dashboard: training loss, token accuracy and checkpoint evaluations. Use the LFM OpenCode and multi-harness runs for the article’s SFT comparison.
Note Base model for the main RL and SFT runs
Note Base model for the earlier Qwen runs
Note SmolDataEnvs training tasks in Harbor format
Note SmolDataEnvs held-out test tasks (250) in Harbor format
Note SmolDataEnvs source tasks
Note Multi-harness SFT data from Qwen3.8-27B rollouts, per harness and pre-tokenized
Note Raw successful teacher rollouts across four harnesses
Browse Harbor task specs from HF Hub, GitHub, or local
Note Browse the Harbor tasks file by file
Note TRL Async GRPO; main: step 1,000, 54.2% pass@1 across four harnesses. Best: [step-700](https://huggingface.co/FineEnvs/LFM2.5-2.6B-multiharness-RL/tree/step-700), 54.6%.
Note TRL Async GRPO; main: step 1,000, 52.3% pass@1 across four harnesses. Best: [step-900](https://huggingface.co/FineEnvs/LFM2.5-2.6B-opencode-RL/tree/step-900), 52.4%.
Note TRL SFT; main: step 4,484, 43.1% pass@1 across four harnesses.
Note TRL SFT; main: step 1,208, 47.5% pass@1 across four harnesses.
Note TRL Async GRPO; main: step 500, 37.0% pass@1 across four harnesses.
Note TRL Async GRPO; main: step 700, 39.5% pass@1 across four harnesses.
Note TRL Async GRPO; main: step 1,000, 29.8% pass@1 across four harnesses.