view article Article One sandbox per rollout, or how labs run RL for agents in 2026 sergiopaniego • 6 days ago • 10
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization Paper • 2609.11682 • Published 7 days ago • 39
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models Paper • 2609.08418 • Published 9 days ago • 144
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training Paper • 2609.15051 • Published 3 days ago • 9
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work Paper • 2609.11977 • Published 13 days ago • 98
NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference Paper • 2609.01657 • Published 17 days ago • 33
NeoMME Collection Meet NeoMME: a family of 260M and 800M Multimodal-Native Multilingual Encoders • 12 items • Updated 14 days ago • 33
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 14 days ago • 103
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Paper • 2609.09153 • Published 9 days ago • 41
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 14 days ago • 118
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 14 days ago • 68
Kraken PP-OCRv6 text recognition models Collection Hub mirrors of Benjamin Kiessling's multilingual PP-OCRv6 line-recognition family for Kraken: tiny, small, and medium. • 3 items • Updated 14 days ago • 7
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published 21 days ago • 35
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Paper • 2608.14905 • Published Aug 14 • 31