Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards Paper • 2609.03181 • Published 18 days ago • 9
The Router Within: Eliciting Native Skill Routing from a Frozen LLM Paper • 2609.15982 • Published 6 days ago • 6
sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 Sentence Similarity • 0.1B • Updated Jan 28 • 45.7M • • 1.41k
view article Article One sandbox per rollout, or how labs run RL for agents in 2026 sergiopaniego • 9 days ago • 10
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF Image-Text-to-Text • 177B • Updated 2 days ago • 43k • 178
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization Paper • 2609.11682 • Published 10 days ago • 43
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models Paper • 2609.08418 • Published 12 days ago • 134
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training Paper • 2609.15051 • Published 6 days ago • 13