One line. Many voicesSeek and you shall find

github.com faviconGitHub - THUDM/CaRR: This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards".

kept by

Citation-aware rubric rewards (CaRR) and C-GRPO training improve deep search agents' reasoning quality and factual grounding over standard binary outcome rewards.

read later

For all the tabs you promised to read.
Save to read. Read to clear.

Close tabs. Keep links.