Announcement_13
We introduce ContextRL, a context-aware RL objective that rewards models for selecting the context that supports an answer, improving long-horizon agentic and multimodal reasoning by +2.2% and +1.8% over standard GRPO.
We introduce ContextRL, a context-aware RL objective that rewards models for selecting the context that supports an answer, improving long-horizon agentic and multimodal reasoning by +2.2% and +1.8% over standard GRPO.