1. Research PapersScore78

    RiVER: Reinforcement Learning for LLMs Without Ground-Truth Solutions

    Researchers have introduced RiVER, a novel framework for training Large Language Models (LLMs) using reinforcement learning without requiring ground-truth solutions. This approach leverages score-based optimization tasks and deterministic execution feedback, addressing challenges like scale and frequency dominance in reward calibration.

    Source: arXiv preprint 'Reinforcement Learning without Ground-Truth Solutions can Improve LLMs' by Lin et al. (2026). Full analysis