Today for AI
HOT RADAR
llmHEAT 5.7°

AI CLUSTERED EVENT · 10/6/2026

RISED: Rubrics for Agentic Multi-Env Self-Distillation

1 reports archived1 independent sourcesupdated 10/6/2026, 12:00:00 AM
Synthesis & Latest Updates
1 Sources Cross-Validated

The new RISED framework replaces scalar rewards with LLM-generated structured rubrics to address data selection and lack of intra-group contrast in multi-environment RL. By using textual feedback to guide online data filtering and policy supervision, it achieves the highest mean pass rates across diverse interactive environments.

LATEST/The new RISED framework replaces scalar rewards with LLM-generated structured rubrics to address data selection and lack of intra-group contrast in multi-environment RL. By using textual feedback to guide online data filtering and policy supervision, it achieves the highest mean pass rates across diverse interactive environments.

TIMELINECoverage timeline

Total 1 reports · Latest first
  1. Apple Machine Learning ResearchT1·68 pts

    RISED: Rubrics for Agentic Multi-Env Self-Distillation

    Original: RISED: Rubrics for Agentic Multi-Environment Selection and Self-Distillation

    • Addresses limitations of scalar rewards in multi-env RL by introducing rubric-based textual feedback to handle batches with uniform outcomes.
    • Repurposes rubrics not just as rewards but to guide online data selection and policy supervision via an LLM judge.
    • Experiments show RISED achieves the highest mean pass rate across environments and ranks first or second individually.
RISED: Rubrics for Agentic Multi-Env Self-Distillation | AI Clustered Intelligence | Today for AI