@@ -214,8 +214,7 @@ Although incorporating CoT rationales into generative reward models improves pre
...
@@ -214,8 +214,7 @@ Although incorporating CoT rationales into generative reward models improves pre
\subsubsection{Rubric-based Reward Models}
\subsubsection{Rubric-based Reward Models}
The generative reward models introduced above improve preference modeling by using the reasoning capability of LLMs. However, they still typically produce an overall preference judgment for each response or response pair. For complex open-ended tasks, such holistic evaluation may be insufficient because response quality often depends on multiple instruction-specific aspects. To address this issue, we can employ rubric-based reward modeling. The idea is to introduce a set of task-specific criteria to guide the reward model on \emph{what aspects to evaluate} when predicting preferences.
While generative reward models have shown strong performance in preference prediction, they typically produce an overall judgment for each response or response pair. For complex open-ended tasks, such holistic evaluation may be insufficient because response quality often depends on multiple instruction-specific aspects. To address this issue, we can employ rubric-based reward modeling. The basic idea is to introduce a set of task-specific criteria that explicitly guide the reward model on \emph{what aspects to evaluate} when predicting preferences. In practice, these criteria can be organized into different rubric formats, as follows: