OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software Paper • 2609.39903 • Published 3 days ago • 37
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work Paper • 2609.11977 • Published 29 days ago • 117
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales Paper • 2606.12736 • Published Jun 10 • 6
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models Paper • 2606.01961 • Published Jun 3 • 27