AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
datasets 25
benchflow/frontierphysics-pr889-evidence
Updated • 48
benchflow/frontierphysics-pr887-evidence
Updated • 60
benchflow/frontierphysics-pr888-evidence
Updated • 65
benchflow/frontierphysics-pr885-evidence
Updated • 64
benchflow/frontierphysics-pr884-evidence
Updated • 59
benchflow/frontierphysics-pr883-evidence
Updated • 63
benchflow/frontierphysics-pr881-evidence
Updated • 62
benchflow/frontierphysics-pr882-evidence
Updated • 94
benchflow/frontierphysics-pr879-evidence
Updated • 62
benchflow/frontierphysics-pr880-evidence
Updated • 50