β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Paper • 2607.28582 • Published 20 days ago • 24
Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs Paper • 2605.12460 • Published May 12 • 17
Open-Reasoner-Zero/Open-Reasoner-Zero-7B Reinforcement Learning • 8B • Updated Apr 7, 2025 • 116 • 34