Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR Paper • 2609.08650 • Published 8 days ago • 12