TL;DR
A super-efficient fine-tuning method where a model generates multiple candidate responses, which are then filtered using a reward model to keep only the highest-quality outputs for training.
Rejection-sampling fine-tuning circumvents the computational complexity of traditional reinforcement learning algorithms. The model acts as its own data generator, producing a set of parallel completions for each training prompt. These completions are evaluated by a critic or verifier, and only the successful or high-scoring samples are selected for standard supervised fine-tuning.
Why this matters for your business
It simplifies the alignment pipeline for specialized reasoning and coding tasks, allowing teams to dramatically improve model performance with lower training overhead.