Add LHTB (Long-Horizon Terminal-Bench) eval result: 37.8

#174

Adds the Long-Horizon Terminal-Bench (LHTB) result for Kimi K3 so it appears on the LHTB benchmark leaderboard.

Score: 37.8 (mean reward x100 over 46 tasks, partial credit)

Metric Value
Mean reward (46 tasks) 0.3778
Solved (reward >= 0.95) 6 / 46
Perfect (reward = 1.0) 5 / 46
Agent terminus-2 (official LHTB Harbor harness)
Budget 90 min per task, 1 trial per task

LHTB measures how well agents sustain useful work in a containerized terminal
over hundreds of steps, graded by hidden rebuild-from-artifact verifiers. This
places Kimi K3 2nd among open-weight models currently on the board.

Complete run artifacts (per-trial configs, results, verifier outputs and terminal
recordings) are published so the score can be audited without rerunning the suite:
https://huggingface.co/datasets/IntelligenceLab/LHTB-leaderboard/tree/main/submissions/long-horizon-terminal-bench/1.0/terminus-2__api_moonshot_kimi-k3

Submitted by the LHTB maintainers; happy to adjust the formatting or withdraw if you prefer.

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment