World Editing: Intervening on Executable Worlds at Increasing Depth Paper • 2610.02331 • Published 10 days ago • 30
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It Paper • 2609.36585 • Published 12 days ago • 70
Persona Dosing: Calibrated Activation Steering for Graded Trait Control Paper • 2609.36388 • Published 13 days ago • 56
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation Paper • 2609.38597 • Published 12 days ago • 33
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 24 days ago • 139
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 27 days ago • 215
WebWorld: The Browser as a World Model for Self-Improving Web Code Paper • 2608.30530 • Published Aug 31 • 10
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published Aug 31 • 148
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 569
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published Aug 26 • 333
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published Aug 24 • 34