Apple Machine Learning ResearchT1·68 pts
RISED: Rubrics for Agentic Multi-Env Self-Distillation
Original: RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation
- Addresses limitations of scalar rewards in multi-env RL by introducing rubric-based textual feedback to handle batches with uniform outcomes.
- Repurposes rubrics not just as rewards but to guide online data selection and policy supervision via an LLM judge.
- Experiments show RISED achieves the highest mean pass rate across environments and ranks first or second individually.