TL;DR
A reinforcement learning feedback mechanism that evaluates and assigns scalar scores to an AI model's entire generation trajectory based strictly on the final, verifiable result.
Outcome reward models assess whether a complete reasoning trajectory or response is correct, using signals such as binary success, semantic agreement, or program execution checks. Unlike process-supervised reward models which evaluate each intermediate step, outcome reward models provide a coarser, sequence-level evaluation. They are highly effective for domains with easily verifiable final targets, such as mathematics, logic, and code generation, where they help rank candidate outputs during decoding or best-of-N selection.
Why this matters for your business
They offer a computationally cheaper and more easily automated alternative to step-by-step human labeling, driving scalable preference alignment for reasoning-intensive tasks.