view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 26 days ago • 139
Agora: Git as Shared Memory for Collective AutoResearch Paper • 2609.18094 • Published 13 days ago • 58
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 12 days ago • 57
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 14 days ago • 76
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 21 days ago • 81
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 13 days ago • 82
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Paper • 2609.08572 • Published 21 days ago • 107
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments Paper • 2609.15364 • Published 15 days ago • 84
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 12 days ago • 91
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 25 days ago • 119
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 26 days ago • 103
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 28 days ago • 121
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 12 days ago • 137
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 15 days ago • 156