Commit fde942f1 by wangchenglong

update.

parent 3d77765b
...@@ -214,6 +214,8 @@ Although incorporating CoT rationales into generative reward models improves pre ...@@ -214,6 +214,8 @@ Although incorporating CoT rationales into generative reward models improves pre
\subsubsection{Rubric-based Reward Models} \subsubsection{Rubric-based Reward Models}
For open-ended tasks, output quality is often multi-dimensional and highly dependent on the specific instruction. As a result, directly making a holistic judgment requires the reward model to implicitly determine both \emph{what aspects should be evaluated} and \emph{how well the output satisfies them}. This is challenging. In practice, such holistic evaluation can lead to biased, inconsistent, and less transparent judgments. For example, when evaluating an output to a writing task, the model may need to simultaneously consider factuality, relevance, clarity, and style. A single holistic reward score (or preference label) may overemphasize one aspect, such as fluency, while overlooking others. To provide comprehensive preference judgments, we can employ rubric-based reward modeling. This method is ....
......
Markdown 格式
0%
您添加了 0 到此讨论。请谨慎行事。
请先完成此评论的编辑!
注册 或者 后发表评论