Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains Paper • 2608.09873 • Published 2 days ago • 26
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published 13 days ago • 68
UEmbed: Unified Sparse and Dense Multimodal Embeddings Paper • 2608.02583 • Published 9 days ago • 50
Metacognition in LLMs: Foundations, Progress, and Opportunities Paper • 2607.11881 • Published about 1 month ago • 30
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 116
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs Paper • 2606.32032 • Published Jun 30 • 29
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents Paper • 2606.24551 • Published Jun 22 • 28
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 157
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models Paper • 2606.16140 • Published Jun 15 • 125
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Paper • 2606.13662 • Published Jun 11 • 31
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales Paper • 2606.12736 • Published Jun 10 • 5
VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding Paper • 2606.05259 • Published Jun 3 • 39
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research Paper • 2605.26114 • Published May 25 • 67
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents Paper • 2605.25624 • Published May 25 • 35
OpenComputer: Verifiable Software Worlds for Computer-Use Agents Paper • 2605.19769 • Published May 19 • 89
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems Paper • 2605.04018 • Published May 5 • 41