Commit 7e3220a0 by wangchenglong

Initial commit

parents
# Template
Template and style files for CoLM 2025
% fancyhdr.sty version 3.2
% Fancy headers and footers for LaTeX.
% Piet van Oostrum,
% Dept of Computer and Information Sciences, University of Utrecht,
% Padualaan 14, P.O. Box 80.089, 3508 TB Utrecht, The Netherlands
% Telephone: +31 30 2532180. Email: piet@cs.uu.nl
% ========================================================================
% LICENCE:
% This file may be distributed under the terms of the LaTeX Project Public
% License, as described in lppl.txt in the base LaTeX distribution.
% Either version 1 or, at your option, any later version.
% ========================================================================
% MODIFICATION HISTORY:
% Sep 16, 1994
% version 1.4: Correction for use with \reversemargin
% Sep 29, 1994:
% version 1.5: Added the \iftopfloat, \ifbotfloat and \iffloatpage commands
% Oct 4, 1994:
% version 1.6: Reset single spacing in headers/footers for use with
% setspace.sty or doublespace.sty
% Oct 4, 1994:
% version 1.7: changed \let\@mkboth\markboth to
% \def\@mkboth{\protect\markboth} to make it more robust
% Dec 5, 1994:
% version 1.8: corrections for amsbook/amsart: define \@chapapp and (more
% importantly) use the \chapter/sectionmark definitions from ps@headings if
% they exist (which should be true for all standard classes).
% May 31, 1995:
% version 1.9: The proposed \renewcommand{\headrulewidth}{\iffloatpage...
% construction in the doc did not work properly with the fancyplain style.
% June 1, 1995:
% version 1.91: The definition of \@mkboth wasn't restored on subsequent
% \pagestyle{fancy}'s.
% June 1, 1995:
% version 1.92: The sequence \pagestyle{fancyplain} \pagestyle{plain}
% \pagestyle{fancy} would erroneously select the plain version.
% June 1, 1995:
% version 1.93: \fancypagestyle command added.
% Dec 11, 1995:
% version 1.94: suggested by Conrad Hughes <chughes@maths.tcd.ie>
% CJCH, Dec 11, 1995: added \footruleskip to allow control over footrule
% position (old hardcoded value of .3\normalbaselineskip is far too high
% when used with very small footer fonts).
% Jan 31, 1996:
% version 1.95: call \@normalsize in the reset code if that is defined,
% otherwise \normalsize.
% this is to solve a problem with ucthesis.cls, as this doesn't
% define \@currsize. Unfortunately for latex209 calling \normalsize doesn't
% work as this is optimized to do very little, so there \@normalsize should
% be called. Hopefully this code works for all versions of LaTeX known to
% mankind.
% April 25, 1996:
% version 1.96: initialize \headwidth to a magic (negative) value to catch
% most common cases that people change it before calling \pagestyle{fancy}.
% Note it can't be initialized when reading in this file, because
% \textwidth could be changed afterwards. This is quite probable.
% We also switch to \MakeUppercase rather than \uppercase and introduce a
% \nouppercase command for use in headers. and footers.
% May 3, 1996:
% version 1.97: Two changes:
% 1. Undo the change in version 1.8 (using the pagestyle{headings} defaults
% for the chapter and section marks. The current version of amsbook and
% amsart classes don't seem to need them anymore. Moreover the standard
% latex classes don't use \markboth if twoside isn't selected, and this is
% confusing as \leftmark doesn't work as expected.
% 2. include a call to \ps@empty in ps@@fancy. This is to solve a problem
% in the amsbook and amsart classes, that make global changes to \topskip,
% which are reset in \ps@empty. Hopefully this doesn't break other things.
% May 7, 1996:
% version 1.98:
% Added % after the line \def\nouppercase
% May 7, 1996:
% version 1.99: This is the alpha version of fancyhdr 2.0
% Introduced the new commands \fancyhead, \fancyfoot, and \fancyhf.
% Changed \headrulewidth, \footrulewidth, \footruleskip to
% macros rather than length parameters, In this way they can be
% conditionalized and they don't consume length registers. There is no need
% to have them as length registers unless you want to do calculations with
% them, which is unlikely. Note that this may make some uses of them
% incompatible (i.e. if you have a file that uses \setlength or \xxxx=)
% May 10, 1996:
% version 1.99a:
% Added a few more % signs
% May 10, 1996:
% version 1.99b:
% Changed the syntax of \f@nfor to be resistent to catcode changes of :=
% Removed the [1] from the defs of \lhead etc. because the parameter is
% consumed by the \@[xy]lhead etc. macros.
% June 24, 1997:
% version 1.99c:
% corrected \nouppercase to also include the protected form of \MakeUppercase
% \global added to manipulation of \headwidth.
% \iffootnote command added.
% Some comments added about \@fancyhead and \@fancyfoot.
% Aug 24, 1998
% version 1.99d
% Changed the default \ps@empty to \ps@@empty in order to allow
% \fancypagestyle{empty} redefinition.
% Oct 11, 2000
% version 2.0
% Added LPPL license clause.
%
% A check for \headheight is added. An errormessage is given (once) if the
% header is too large. Empty headers don't generate the error even if
% \headheight is very small or even 0pt.
% Warning added for the use of 'E' option when twoside option is not used.
% In this case the 'E' fields will never be used.
%
% Mar 10, 2002
% version 2.1beta
% New command: \fancyhfoffset[place]{length}
% defines offsets to be applied to the header/footer to let it stick into
% the margins (if length > 0).
% place is like in fancyhead, except that only E,O,L,R can be used.
% This replaces the old calculation based on \headwidth and the marginpar
% area.
% \headwidth will be dynamically calculated in the headers/footers when
% this is used.
%
% Mar 26, 2002
% version 2.1beta2
% \fancyhfoffset now also takes h,f as possible letters in the argument to
% allow the header and footer widths to be different.
% New commands \fancyheadoffset and \fancyfootoffset added comparable to
% \fancyhead and \fancyfoot.
% Errormessages and warnings have been made more informative.
%
% Dec 9, 2002
% version 2.1
% The defaults for \footrulewidth, \plainheadrulewidth and
% \plainfootrulewidth are changed from \z@skip to 0pt. In this way when
% someone inadvertantly uses \setlength to change any of these, the value
% of \z@skip will not be changed, rather an errormessage will be given.
% March 3, 2004
% Release of version 3.0
% Oct 7, 2004
% version 3.1
% Added '\endlinechar=13' to \fancy@reset to prevent problems with
% includegraphics in header when verbatiminput is active.
% March 22, 2005
% version 3.2
% reset \everypar (the real one) in \fancy@reset because spanish.ldf does
% strange things with \everypar between << and >>.
\def\ifancy@mpty#1{\def\temp@a{#1}\ifx\temp@a\@empty}
\def\fancy@def#1#2{\ifancy@mpty{#2}\fancy@gbl\def#1{\leavevmode}\else
\fancy@gbl\def#1{#2\strut}\fi}
\let\fancy@gbl\global
\def\@fancyerrmsg#1{%
\ifx\PackageError\undefined
\errmessage{#1}\else
\PackageError{Fancyhdr}{#1}{}\fi}
\def\@fancywarning#1{%
\ifx\PackageWarning\undefined
\errmessage{#1}\else
\PackageWarning{Fancyhdr}{#1}{}\fi}
% Usage: \@forc \var{charstring}{command to be executed for each char}
% This is similar to LaTeX's \@tfor, but expands the charstring.
\def\@forc#1#2#3{\expandafter\f@rc\expandafter#1\expandafter{#2}{#3}}
\def\f@rc#1#2#3{\def\temp@ty{#2}\ifx\@empty\temp@ty\else
\f@@rc#1#2\f@@rc{#3}\fi}
\def\f@@rc#1#2#3\f@@rc#4{\def#1{#2}#4\f@rc#1{#3}{#4}}
% Usage: \f@nfor\name:=list\do{body}
% Like LaTeX's \@for but an empty list is treated as a list with an empty
% element
\newcommand{\f@nfor}[3]{\edef\@fortmp{#2}%
\expandafter\@forloop#2,\@nil,\@nil\@@#1{#3}}
% Usage: \def@ult \cs{defaults}{argument}
% sets \cs to the characters from defaults appearing in argument
% or defaults if it would be empty. All characters are lowercased.
\newcommand\def@ult[3]{%
\edef\temp@a{\lowercase{\edef\noexpand\temp@a{#3}}}\temp@a
\def#1{}%
\@forc\tmpf@ra{#2}%
{\expandafter\if@in\tmpf@ra\temp@a{\edef#1{#1\tmpf@ra}}{}}%
\ifx\@empty#1\def#1{#2}\fi}
%
% \if@in <char><set><truecase><falsecase>
%
\newcommand{\if@in}[4]{%
\edef\temp@a{#2}\def\temp@b##1#1##2\temp@b{\def\temp@b{##1}}%
\expandafter\temp@b#2#1\temp@b\ifx\temp@a\temp@b #4\else #3\fi}
\newcommand{\fancyhead}{\@ifnextchar[{\f@ncyhf\fancyhead h}%
{\f@ncyhf\fancyhead h[]}}
\newcommand{\fancyfoot}{\@ifnextchar[{\f@ncyhf\fancyfoot f}%
{\f@ncyhf\fancyfoot f[]}}
\newcommand{\fancyhf}{\@ifnextchar[{\f@ncyhf\fancyhf{}}%
{\f@ncyhf\fancyhf{}[]}}
% New commands for offsets added
\newcommand{\fancyheadoffset}{\@ifnextchar[{\f@ncyhfoffs\fancyheadoffset h}%
{\f@ncyhfoffs\fancyheadoffset h[]}}
\newcommand{\fancyfootoffset}{\@ifnextchar[{\f@ncyhfoffs\fancyfootoffset f}%
{\f@ncyhfoffs\fancyfootoffset f[]}}
\newcommand{\fancyhfoffset}{\@ifnextchar[{\f@ncyhfoffs\fancyhfoffset{}}%
{\f@ncyhfoffs\fancyhfoffset{}[]}}
% The header and footer fields are stored in command sequences with
% names of the form: \f@ncy<x><y><z> with <x> for [eo], <y> from [lcr]
% and <z> from [hf].
\def\f@ncyhf#1#2[#3]#4{%
\def\temp@c{}%
\@forc\tmpf@ra{#3}%
{\expandafter\if@in\tmpf@ra{eolcrhf,EOLCRHF}%
{}{\edef\temp@c{\temp@c\tmpf@ra}}}%
\ifx\@empty\temp@c\else
\@fancyerrmsg{Illegal char `\temp@c' in \string#1 argument:
[#3]}%
\fi
\f@nfor\temp@c{#3}%
{\def@ult\f@@@eo{eo}\temp@c
\if@twoside\else
\if\f@@@eo e\@fancywarning
{\string#1's `E' option without twoside option is useless}\fi\fi
\def@ult\f@@@lcr{lcr}\temp@c
\def@ult\f@@@hf{hf}{#2\temp@c}%
\@forc\f@@eo\f@@@eo
{\@forc\f@@lcr\f@@@lcr
{\@forc\f@@hf\f@@@hf
{\expandafter\fancy@def\csname
f@ncy\f@@eo\f@@lcr\f@@hf\endcsname
{#4}}}}}}
\def\f@ncyhfoffs#1#2[#3]#4{%
\def\temp@c{}%
\@forc\tmpf@ra{#3}%
{\expandafter\if@in\tmpf@ra{eolrhf,EOLRHF}%
{}{\edef\temp@c{\temp@c\tmpf@ra}}}%
\ifx\@empty\temp@c\else
\@fancyerrmsg{Illegal char `\temp@c' in \string#1 argument:
[#3]}%
\fi
\f@nfor\temp@c{#3}%
{\def@ult\f@@@eo{eo}\temp@c
\if@twoside\else
\if\f@@@eo e\@fancywarning
{\string#1's `E' option without twoside option is useless}\fi\fi
\def@ult\f@@@lcr{lr}\temp@c
\def@ult\f@@@hf{hf}{#2\temp@c}%
\@forc\f@@eo\f@@@eo
{\@forc\f@@lcr\f@@@lcr
{\@forc\f@@hf\f@@@hf
{\expandafter\setlength\csname
f@ncyO@\f@@eo\f@@lcr\f@@hf\endcsname
{#4}}}}}%
\fancy@setoffs}
% Fancyheadings version 1 commands. These are more or less deprecated,
% but they continue to work.
\newcommand{\lhead}{\@ifnextchar[{\@xlhead}{\@ylhead}}
\def\@xlhead[#1]#2{\fancy@def\f@ncyelh{#1}\fancy@def\f@ncyolh{#2}}
\def\@ylhead#1{\fancy@def\f@ncyelh{#1}\fancy@def\f@ncyolh{#1}}
\newcommand{\chead}{\@ifnextchar[{\@xchead}{\@ychead}}
\def\@xchead[#1]#2{\fancy@def\f@ncyech{#1}\fancy@def\f@ncyoch{#2}}
\def\@ychead#1{\fancy@def\f@ncyech{#1}\fancy@def\f@ncyoch{#1}}
\newcommand{\rhead}{\@ifnextchar[{\@xrhead}{\@yrhead}}
\def\@xrhead[#1]#2{\fancy@def\f@ncyerh{#1}\fancy@def\f@ncyorh{#2}}
\def\@yrhead#1{\fancy@def\f@ncyerh{#1}\fancy@def\f@ncyorh{#1}}
\newcommand{\lfoot}{\@ifnextchar[{\@xlfoot}{\@ylfoot}}
\def\@xlfoot[#1]#2{\fancy@def\f@ncyelf{#1}\fancy@def\f@ncyolf{#2}}
\def\@ylfoot#1{\fancy@def\f@ncyelf{#1}\fancy@def\f@ncyolf{#1}}
\newcommand{\cfoot}{\@ifnextchar[{\@xcfoot}{\@ycfoot}}
\def\@xcfoot[#1]#2{\fancy@def\f@ncyecf{#1}\fancy@def\f@ncyocf{#2}}
\def\@ycfoot#1{\fancy@def\f@ncyecf{#1}\fancy@def\f@ncyocf{#1}}
\newcommand{\rfoot}{\@ifnextchar[{\@xrfoot}{\@yrfoot}}
\def\@xrfoot[#1]#2{\fancy@def\f@ncyerf{#1}\fancy@def\f@ncyorf{#2}}
\def\@yrfoot#1{\fancy@def\f@ncyerf{#1}\fancy@def\f@ncyorf{#1}}
\newlength{\fancy@headwidth}
\let\headwidth\fancy@headwidth
\newlength{\f@ncyO@elh}
\newlength{\f@ncyO@erh}
\newlength{\f@ncyO@olh}
\newlength{\f@ncyO@orh}
\newlength{\f@ncyO@elf}
\newlength{\f@ncyO@erf}
\newlength{\f@ncyO@olf}
\newlength{\f@ncyO@orf}
\newcommand{\headrulewidth}{0.4pt}
\newcommand{\footrulewidth}{0pt}
\newcommand{\footruleskip}{.3\normalbaselineskip}
% Fancyplain stuff shouldn't be used anymore (rather
% \fancypagestyle{plain} should be used), but it must be present for
% compatibility reasons.
\newcommand{\plainheadrulewidth}{0pt}
\newcommand{\plainfootrulewidth}{0pt}
\newif\if@fancyplain \@fancyplainfalse
\def\fancyplain#1#2{\if@fancyplain#1\else#2\fi}
\headwidth=-123456789sp %magic constant
% Command to reset various things in the headers:
% a.o. single spacing (taken from setspace.sty)
% and the catcode of ^^M (so that epsf files in the header work if a
% verbatim crosses a page boundary)
% It also defines a \nouppercase command that disables \uppercase and
% \Makeuppercase. It can only be used in the headers and footers.
\let\fnch@everypar\everypar% save real \everypar because of spanish.ldf
\def\fancy@reset{\fnch@everypar{}\restorecr\endlinechar=13
\def\baselinestretch{1}%
\def\nouppercase##1{{\let\uppercase\relax\let\MakeUppercase\relax
\expandafter\let\csname MakeUppercase \endcsname\relax##1}}%
\ifx\undefined\@newbaseline% NFSS not present; 2.09 or 2e
\ifx\@normalsize\undefined \normalsize % for ucthesis.cls
\else \@normalsize \fi
\else% NFSS (2.09) present
\@newbaseline%
\fi}
% Initialization of the head and foot text.
% The default values still contain \fancyplain for compatibility.
\fancyhf{} % clear all
% lefthead empty on ``plain'' pages, \rightmark on even, \leftmark on odd pages
% evenhead empty on ``plain'' pages, \leftmark on even, \rightmark on odd pages
\if@twoside
\fancyhead[el,or]{\fancyplain{}{\sl\rightmark}}
\fancyhead[er,ol]{\fancyplain{}{\sl\leftmark}}
\else
\fancyhead[l]{\fancyplain{}{\sl\rightmark}}
\fancyhead[r]{\fancyplain{}{\sl\leftmark}}
\fi
\fancyfoot[c]{\rm\thepage} % page number
% Use box 0 as a temp box and dimen 0 as temp dimen.
% This can be done, because this code will always
% be used inside another box, and therefore the changes are local.
\def\@fancyvbox#1#2{\setbox0\vbox{#2}\ifdim\ht0>#1\@fancywarning
{\string#1 is too small (\the#1): ^^J Make it at least \the\ht0.^^J
We now make it that large for the rest of the document.^^J
This may cause the page layout to be inconsistent, however\@gobble}%
\dimen0=#1\global\setlength{#1}{\ht0}\ht0=\dimen0\fi
\box0}
% Put together a header or footer given the left, center and
% right text, fillers at left and right and a rule.
% The \lap commands put the text into an hbox of zero size,
% so overlapping text does not generate an errormessage.
% These macros have 5 parameters:
% 1. LEFTSIDE BEARING % This determines at which side the header will stick
% out. When \fancyhfoffset is used this calculates \headwidth, otherwise
% it is \hss or \relax (after expansion).
% 2. \f@ncyolh, \f@ncyelh, \f@ncyolf or \f@ncyelf. This is the left component.
% 3. \f@ncyoch, \f@ncyech, \f@ncyocf or \f@ncyecf. This is the middle comp.
% 4. \f@ncyorh, \f@ncyerh, \f@ncyorf or \f@ncyerf. This is the right component.
% 5. RIGHTSIDE BEARING. This is always \relax or \hss (after expansion).
\def\@fancyhead#1#2#3#4#5{#1\hbox to\headwidth{\fancy@reset
\@fancyvbox\headheight{\hbox
{\rlap{\parbox[b]{\headwidth}{\raggedright#2}}\hfill
\parbox[b]{\headwidth}{\centering#3}\hfill
\llap{\parbox[b]{\headwidth}{\raggedleft#4}}}\headrule}}#5}
\def\@fancyfoot#1#2#3#4#5{#1\hbox to\headwidth{\fancy@reset
\@fancyvbox\footskip{\footrule
\hbox{\rlap{\parbox[t]{\headwidth}{\raggedright#2}}\hfill
\parbox[t]{\headwidth}{\centering#3}\hfill
\llap{\parbox[t]{\headwidth}{\raggedleft#4}}}}}#5}
\def\headrule{{\if@fancyplain\let\headrulewidth\plainheadrulewidth\fi
\hrule\@height\headrulewidth\@width\headwidth \vskip-\headrulewidth}}
\def\footrule{{\if@fancyplain\let\footrulewidth\plainfootrulewidth\fi
\vskip-\footruleskip\vskip-\footrulewidth
\hrule\@width\headwidth\@height\footrulewidth\vskip\footruleskip}}
\def\ps@fancy{%
\@ifundefined{@chapapp}{\let\@chapapp\chaptername}{}%for amsbook
%
% Define \MakeUppercase for old LaTeXen.
% Note: we used \def rather than \let, so that \let\uppercase\relax (from
% the version 1 documentation) will still work.
%
\@ifundefined{MakeUppercase}{\def\MakeUppercase{\uppercase}}{}%
\@ifundefined{chapter}{\def\sectionmark##1{\markboth
{\MakeUppercase{\ifnum \c@secnumdepth>\z@
\thesection\hskip 1em\relax \fi ##1}}{}}%
\def\subsectionmark##1{\markright {\ifnum \c@secnumdepth >\@ne
\thesubsection\hskip 1em\relax \fi ##1}}}%
{\def\chaptermark##1{\markboth {\MakeUppercase{\ifnum \c@secnumdepth>\m@ne
\@chapapp\ \thechapter. \ \fi ##1}}{}}%
\def\sectionmark##1{\markright{\MakeUppercase{\ifnum \c@secnumdepth >\z@
\thesection. \ \fi ##1}}}}%
%\csname ps@headings\endcsname % use \ps@headings defaults if they exist
\ps@@fancy
\gdef\ps@fancy{\@fancyplainfalse\ps@@fancy}%
% Initialize \headwidth if the user didn't
%
\ifdim\headwidth<0sp
%
% This catches the case that \headwidth hasn't been initialized and the
% case that the user added something to \headwidth in the expectation that
% it was initialized to \textwidth. We compensate this now. This loses if
% the user intended to multiply it by a factor. But that case is more
% likely done by saying something like \headwidth=1.2\textwidth.
% The doc says you have to change \headwidth after the first call to
% \pagestyle{fancy}. This code is just to catch the most common cases were
% that requirement is violated.
%
\global\advance\headwidth123456789sp\global\advance\headwidth\textwidth
\fi}
\def\ps@fancyplain{\ps@fancy \let\ps@plain\ps@plain@fancy}
\def\ps@plain@fancy{\@fancyplaintrue\ps@@fancy}
\let\ps@@empty\ps@empty
\def\ps@@fancy{%
\ps@@empty % This is for amsbook/amsart, which do strange things with \topskip
\def\@mkboth{\protect\markboth}%
\def\@oddhead{\@fancyhead\fancy@Oolh\f@ncyolh\f@ncyoch\f@ncyorh\fancy@Oorh}%
\def\@oddfoot{\@fancyfoot\fancy@Oolf\f@ncyolf\f@ncyocf\f@ncyorf\fancy@Oorf}%
\def\@evenhead{\@fancyhead\fancy@Oelh\f@ncyelh\f@ncyech\f@ncyerh\fancy@Oerh}%
\def\@evenfoot{\@fancyfoot\fancy@Oelf\f@ncyelf\f@ncyecf\f@ncyerf\fancy@Oerf}%
}
% Default definitions for compatibility mode:
% These cause the header/footer to take the defined \headwidth as width
% And to shift in the direction of the marginpar area
\def\fancy@Oolh{\if@reversemargin\hss\else\relax\fi}
\def\fancy@Oorh{\if@reversemargin\relax\else\hss\fi}
\let\fancy@Oelh\fancy@Oorh
\let\fancy@Oerh\fancy@Oolh
\let\fancy@Oolf\fancy@Oolh
\let\fancy@Oorf\fancy@Oorh
\let\fancy@Oelf\fancy@Oelh
\let\fancy@Oerf\fancy@Oerh
% New definitions for the use of \fancyhfoffset
% These calculate the \headwidth from \textwidth and the specified offsets.
\def\fancy@offsolh{\headwidth=\textwidth\advance\headwidth\f@ncyO@olh
\advance\headwidth\f@ncyO@orh\hskip-\f@ncyO@olh}
\def\fancy@offselh{\headwidth=\textwidth\advance\headwidth\f@ncyO@elh
\advance\headwidth\f@ncyO@erh\hskip-\f@ncyO@elh}
\def\fancy@offsolf{\headwidth=\textwidth\advance\headwidth\f@ncyO@olf
\advance\headwidth\f@ncyO@orf\hskip-\f@ncyO@olf}
\def\fancy@offself{\headwidth=\textwidth\advance\headwidth\f@ncyO@elf
\advance\headwidth\f@ncyO@erf\hskip-\f@ncyO@elf}
\def\fancy@setoffs{%
% Just in case \let\headwidth\textwidth was used
\fancy@gbl\let\headwidth\fancy@headwidth
\fancy@gbl\let\fancy@Oolh\fancy@offsolh
\fancy@gbl\let\fancy@Oelh\fancy@offselh
\fancy@gbl\let\fancy@Oorh\hss
\fancy@gbl\let\fancy@Oerh\hss
\fancy@gbl\let\fancy@Oolf\fancy@offsolf
\fancy@gbl\let\fancy@Oelf\fancy@offself
\fancy@gbl\let\fancy@Oorf\hss
\fancy@gbl\let\fancy@Oerf\hss}
\newif\iffootnote
\let\latex@makecol\@makecol
\def\@makecol{\ifvoid\footins\footnotetrue\else\footnotefalse\fi
\let\topfloat\@toplist\let\botfloat\@botlist\latex@makecol}
\def\iftopfloat#1#2{\ifx\topfloat\empty #2\else #1\fi}
\def\ifbotfloat#1#2{\ifx\botfloat\empty #2\else #1\fi}
\def\iffloatpage#1#2{\if@fcolmade #1\else #2\fi}
\newcommand{\fancypagestyle}[2]{%
\@namedef{ps@#1}{\let\fancy@gbl\relax#2\relax\ps@fancy}}
%%%%% NEW MATH DEFINITIONS %%%%%
\usepackage{amsmath,amsfonts,bm}
% Mark sections of captions for referring to divisions of figures
\newcommand{\figleft}{{\em (Left)}}
\newcommand{\figcenter}{{\em (Center)}}
\newcommand{\figright}{{\em (Right)}}
\newcommand{\figtop}{{\em (Top)}}
\newcommand{\figbottom}{{\em (Bottom)}}
\newcommand{\captiona}{{\em (a)}}
\newcommand{\captionb}{{\em (b)}}
\newcommand{\captionc}{{\em (c)}}
\newcommand{\captiond}{{\em (d)}}
% Highlight a newly defined term
\newcommand{\newterm}[1]{{\bf #1}}
% Figure reference, lower-case.
\def\figref#1{figure~\ref{#1}}
% Figure reference, capital. For start of sentence
\def\Figref#1{Figure~\ref{#1}}
\def\twofigref#1#2{figures \ref{#1} and \ref{#2}}
\def\quadfigref#1#2#3#4{figures \ref{#1}, \ref{#2}, \ref{#3} and \ref{#4}}
% Section reference, lower-case.
\def\secref#1{section~\ref{#1}}
% Section reference, capital.
\def\Secref#1{Section~\ref{#1}}
% Reference to two sections.
\def\twosecrefs#1#2{sections \ref{#1} and \ref{#2}}
% Reference to three sections.
\def\secrefs#1#2#3{sections \ref{#1}, \ref{#2} and \ref{#3}}
% Reference to an equation, lower-case.
\def\eqref#1{equation~\ref{#1}}
% Reference to an equation, upper case
\def\Eqref#1{Equation~\ref{#1}}
% A raw reference to an equation---avoid using if possible
\def\plaineqref#1{\ref{#1}}
% Reference to a chapter, lower-case.
\def\chapref#1{chapter~\ref{#1}}
% Reference to an equation, upper case.
\def\Chapref#1{Chapter~\ref{#1}}
% Reference to a range of chapters
\def\rangechapref#1#2{chapters\ref{#1}--\ref{#2}}
% Reference to an algorithm, lower-case.
\def\algref#1{algorithm~\ref{#1}}
% Reference to an algorithm, upper case.
\def\Algref#1{Algorithm~\ref{#1}}
\def\twoalgref#1#2{algorithms \ref{#1} and \ref{#2}}
\def\Twoalgref#1#2{Algorithms \ref{#1} and \ref{#2}}
% Reference to a part, lower case
\def\partref#1{part~\ref{#1}}
% Reference to a part, upper case
\def\Partref#1{Part~\ref{#1}}
\def\twopartref#1#2{parts \ref{#1} and \ref{#2}}
\def\ceil#1{\lceil #1 \rceil}
\def\floor#1{\lfloor #1 \rfloor}
\def\1{\bm{1}}
\newcommand{\train}{\mathcal{D}}
\newcommand{\valid}{\mathcal{D_{\mathrm{valid}}}}
\newcommand{\test}{\mathcal{D_{\mathrm{test}}}}
\def\eps{{\epsilon}}
% Random variables
\def\reta{{\textnormal{$\eta$}}}
\def\ra{{\textnormal{a}}}
\def\rb{{\textnormal{b}}}
\def\rc{{\textnormal{c}}}
\def\rd{{\textnormal{d}}}
\def\re{{\textnormal{e}}}
\def\rf{{\textnormal{f}}}
\def\rg{{\textnormal{g}}}
\def\rh{{\textnormal{h}}}
\def\ri{{\textnormal{i}}}
\def\rj{{\textnormal{j}}}
\def\rk{{\textnormal{k}}}
\def\rl{{\textnormal{l}}}
% rm is already a command, just don't name any random variables m
\def\rn{{\textnormal{n}}}
\def\ro{{\textnormal{o}}}
\def\rp{{\textnormal{p}}}
\def\rq{{\textnormal{q}}}
\def\rr{{\textnormal{r}}}
\def\rs{{\textnormal{s}}}
\def\rt{{\textnormal{t}}}
\def\ru{{\textnormal{u}}}
\def\rv{{\textnormal{v}}}
\def\rw{{\textnormal{w}}}
\def\rx{{\textnormal{x}}}
\def\ry{{\textnormal{y}}}
\def\rz{{\textnormal{z}}}
% Random vectors
\def\rvepsilon{{\mathbf{\epsilon}}}
\def\rvtheta{{\mathbf{\theta}}}
\def\rva{{\mathbf{a}}}
\def\rvb{{\mathbf{b}}}
\def\rvc{{\mathbf{c}}}
\def\rvd{{\mathbf{d}}}
\def\rve{{\mathbf{e}}}
\def\rvf{{\mathbf{f}}}
\def\rvg{{\mathbf{g}}}
\def\rvh{{\mathbf{h}}}
\def\rvu{{\mathbf{i}}}
\def\rvj{{\mathbf{j}}}
\def\rvk{{\mathbf{k}}}
\def\rvl{{\mathbf{l}}}
\def\rvm{{\mathbf{m}}}
\def\rvn{{\mathbf{n}}}
\def\rvo{{\mathbf{o}}}
\def\rvp{{\mathbf{p}}}
\def\rvq{{\mathbf{q}}}
\def\rvr{{\mathbf{r}}}
\def\rvs{{\mathbf{s}}}
\def\rvt{{\mathbf{t}}}
\def\rvu{{\mathbf{u}}}
\def\rvv{{\mathbf{v}}}
\def\rvw{{\mathbf{w}}}
\def\rvx{{\mathbf{x}}}
\def\rvy{{\mathbf{y}}}
\def\rvz{{\mathbf{z}}}
% Elements of random vectors
\def\erva{{\textnormal{a}}}
\def\ervb{{\textnormal{b}}}
\def\ervc{{\textnormal{c}}}
\def\ervd{{\textnormal{d}}}
\def\erve{{\textnormal{e}}}
\def\ervf{{\textnormal{f}}}
\def\ervg{{\textnormal{g}}}
\def\ervh{{\textnormal{h}}}
\def\ervi{{\textnormal{i}}}
\def\ervj{{\textnormal{j}}}
\def\ervk{{\textnormal{k}}}
\def\ervl{{\textnormal{l}}}
\def\ervm{{\textnormal{m}}}
\def\ervn{{\textnormal{n}}}
\def\ervo{{\textnormal{o}}}
\def\ervp{{\textnormal{p}}}
\def\ervq{{\textnormal{q}}}
\def\ervr{{\textnormal{r}}}
\def\ervs{{\textnormal{s}}}
\def\ervt{{\textnormal{t}}}
\def\ervu{{\textnormal{u}}}
\def\ervv{{\textnormal{v}}}
\def\ervw{{\textnormal{w}}}
\def\ervx{{\textnormal{x}}}
\def\ervy{{\textnormal{y}}}
\def\ervz{{\textnormal{z}}}
% Random matrices
\def\rmA{{\mathbf{A}}}
\def\rmB{{\mathbf{B}}}
\def\rmC{{\mathbf{C}}}
\def\rmD{{\mathbf{D}}}
\def\rmE{{\mathbf{E}}}
\def\rmF{{\mathbf{F}}}
\def\rmG{{\mathbf{G}}}
\def\rmH{{\mathbf{H}}}
\def\rmI{{\mathbf{I}}}
\def\rmJ{{\mathbf{J}}}
\def\rmK{{\mathbf{K}}}
\def\rmL{{\mathbf{L}}}
\def\rmM{{\mathbf{M}}}
\def\rmN{{\mathbf{N}}}
\def\rmO{{\mathbf{O}}}
\def\rmP{{\mathbf{P}}}
\def\rmQ{{\mathbf{Q}}}
\def\rmR{{\mathbf{R}}}
\def\rmS{{\mathbf{S}}}
\def\rmT{{\mathbf{T}}}
\def\rmU{{\mathbf{U}}}
\def\rmV{{\mathbf{V}}}
\def\rmW{{\mathbf{W}}}
\def\rmX{{\mathbf{X}}}
\def\rmY{{\mathbf{Y}}}
\def\rmZ{{\mathbf{Z}}}
% Elements of random matrices
\def\ermA{{\textnormal{A}}}
\def\ermB{{\textnormal{B}}}
\def\ermC{{\textnormal{C}}}
\def\ermD{{\textnormal{D}}}
\def\ermE{{\textnormal{E}}}
\def\ermF{{\textnormal{F}}}
\def\ermG{{\textnormal{G}}}
\def\ermH{{\textnormal{H}}}
\def\ermI{{\textnormal{I}}}
\def\ermJ{{\textnormal{J}}}
\def\ermK{{\textnormal{K}}}
\def\ermL{{\textnormal{L}}}
\def\ermM{{\textnormal{M}}}
\def\ermN{{\textnormal{N}}}
\def\ermO{{\textnormal{O}}}
\def\ermP{{\textnormal{P}}}
\def\ermQ{{\textnormal{Q}}}
\def\ermR{{\textnormal{R}}}
\def\ermS{{\textnormal{S}}}
\def\ermT{{\textnormal{T}}}
\def\ermU{{\textnormal{U}}}
\def\ermV{{\textnormal{V}}}
\def\ermW{{\textnormal{W}}}
\def\ermX{{\textnormal{X}}}
\def\ermY{{\textnormal{Y}}}
\def\ermZ{{\textnormal{Z}}}
% Vectors
\def\vzero{{\bm{0}}}
\def\vone{{\bm{1}}}
\def\vmu{{\bm{\mu}}}
\def\vtheta{{\bm{\theta}}}
\def\va{{\bm{a}}}
\def\vb{{\bm{b}}}
\def\vc{{\bm{c}}}
\def\vd{{\bm{d}}}
\def\ve{{\bm{e}}}
\def\vf{{\bm{f}}}
\def\vg{{\bm{g}}}
\def\vh{{\bm{h}}}
\def\vi{{\bm{i}}}
\def\vj{{\bm{j}}}
\def\vk{{\bm{k}}}
\def\vl{{\bm{l}}}
\def\vm{{\bm{m}}}
\def\vn{{\bm{n}}}
\def\vo{{\bm{o}}}
\def\vp{{\bm{p}}}
\def\vq{{\bm{q}}}
\def\vr{{\bm{r}}}
\def\vs{{\bm{s}}}
\def\vt{{\bm{t}}}
\def\vu{{\bm{u}}}
\def\vv{{\bm{v}}}
\def\vw{{\bm{w}}}
\def\vx{{\bm{x}}}
\def\vy{{\bm{y}}}
\def\vz{{\bm{z}}}
% Elements of vectors
\def\evalpha{{\alpha}}
\def\evbeta{{\beta}}
\def\evepsilon{{\epsilon}}
\def\evlambda{{\lambda}}
\def\evomega{{\omega}}
\def\evmu{{\mu}}
\def\evpsi{{\psi}}
\def\evsigma{{\sigma}}
\def\evtheta{{\theta}}
\def\eva{{a}}
\def\evb{{b}}
\def\evc{{c}}
\def\evd{{d}}
\def\eve{{e}}
\def\evf{{f}}
\def\evg{{g}}
\def\evh{{h}}
\def\evi{{i}}
\def\evj{{j}}
\def\evk{{k}}
\def\evl{{l}}
\def\evm{{m}}
\def\evn{{n}}
\def\evo{{o}}
\def\evp{{p}}
\def\evq{{q}}
\def\evr{{r}}
\def\evs{{s}}
\def\evt{{t}}
\def\evu{{u}}
\def\evv{{v}}
\def\evw{{w}}
\def\evx{{x}}
\def\evy{{y}}
\def\evz{{z}}
% Matrix
\def\mA{{\bm{A}}}
\def\mB{{\bm{B}}}
\def\mC{{\bm{C}}}
\def\mD{{\bm{D}}}
\def\mE{{\bm{E}}}
\def\mF{{\bm{F}}}
\def\mG{{\bm{G}}}
\def\mH{{\bm{H}}}
\def\mI{{\bm{I}}}
\def\mJ{{\bm{J}}}
\def\mK{{\bm{K}}}
\def\mL{{\bm{L}}}
\def\mM{{\bm{M}}}
\def\mN{{\bm{N}}}
\def\mO{{\bm{O}}}
\def\mP{{\bm{P}}}
\def\mQ{{\bm{Q}}}
\def\mR{{\bm{R}}}
\def\mS{{\bm{S}}}
\def\mT{{\bm{T}}}
\def\mU{{\bm{U}}}
\def\mV{{\bm{V}}}
\def\mW{{\bm{W}}}
\def\mX{{\bm{X}}}
\def\mY{{\bm{Y}}}
\def\mZ{{\bm{Z}}}
\def\mBeta{{\bm{\beta}}}
\def\mPhi{{\bm{\Phi}}}
\def\mLambda{{\bm{\Lambda}}}
\def\mSigma{{\bm{\Sigma}}}
% Tensor
\DeclareMathAlphabet{\mathsfit}{\encodingdefault}{\sfdefault}{m}{sl}
\SetMathAlphabet{\mathsfit}{bold}{\encodingdefault}{\sfdefault}{bx}{n}
\newcommand{\tens}[1]{\bm{\mathsfit{#1}}}
\def\tA{{\tens{A}}}
\def\tB{{\tens{B}}}
\def\tC{{\tens{C}}}
\def\tD{{\tens{D}}}
\def\tE{{\tens{E}}}
\def\tF{{\tens{F}}}
\def\tG{{\tens{G}}}
\def\tH{{\tens{H}}}
\def\tI{{\tens{I}}}
\def\tJ{{\tens{J}}}
\def\tK{{\tens{K}}}
\def\tL{{\tens{L}}}
\def\tM{{\tens{M}}}
\def\tN{{\tens{N}}}
\def\tO{{\tens{O}}}
\def\tP{{\tens{P}}}
\def\tQ{{\tens{Q}}}
\def\tR{{\tens{R}}}
\def\tS{{\tens{S}}}
\def\tT{{\tens{T}}}
\def\tU{{\tens{U}}}
\def\tV{{\tens{V}}}
\def\tW{{\tens{W}}}
\def\tX{{\tens{X}}}
\def\tY{{\tens{Y}}}
\def\tZ{{\tens{Z}}}
% Graph
\def\gA{{\mathcal{A}}}
\def\gB{{\mathcal{B}}}
\def\gC{{\mathcal{C}}}
\def\gD{{\mathcal{D}}}
\def\gE{{\mathcal{E}}}
\def\gF{{\mathcal{F}}}
\def\gG{{\mathcal{G}}}
\def\gH{{\mathcal{H}}}
\def\gI{{\mathcal{I}}}
\def\gJ{{\mathcal{J}}}
\def\gK{{\mathcal{K}}}
\def\gL{{\mathcal{L}}}
\def\gM{{\mathcal{M}}}
\def\gN{{\mathcal{N}}}
\def\gO{{\mathcal{O}}}
\def\gP{{\mathcal{P}}}
\def\gQ{{\mathcal{Q}}}
\def\gR{{\mathcal{R}}}
\def\gS{{\mathcal{S}}}
\def\gT{{\mathcal{T}}}
\def\gU{{\mathcal{U}}}
\def\gV{{\mathcal{V}}}
\def\gW{{\mathcal{W}}}
\def\gX{{\mathcal{X}}}
\def\gY{{\mathcal{Y}}}
\def\gZ{{\mathcal{Z}}}
% Sets
\def\sA{{\mathbb{A}}}
\def\sB{{\mathbb{B}}}
\def\sC{{\mathbb{C}}}
\def\sD{{\mathbb{D}}}
% Don't use a set called E, because this would be the same as our symbol
% for expectation.
\def\sF{{\mathbb{F}}}
\def\sG{{\mathbb{G}}}
\def\sH{{\mathbb{H}}}
\def\sI{{\mathbb{I}}}
\def\sJ{{\mathbb{J}}}
\def\sK{{\mathbb{K}}}
\def\sL{{\mathbb{L}}}
\def\sM{{\mathbb{M}}}
\def\sN{{\mathbb{N}}}
\def\sO{{\mathbb{O}}}
\def\sP{{\mathbb{P}}}
\def\sQ{{\mathbb{Q}}}
\def\sR{{\mathbb{R}}}
\def\sS{{\mathbb{S}}}
\def\sT{{\mathbb{T}}}
\def\sU{{\mathbb{U}}}
\def\sV{{\mathbb{V}}}
\def\sW{{\mathbb{W}}}
\def\sX{{\mathbb{X}}}
\def\sY{{\mathbb{Y}}}
\def\sZ{{\mathbb{Z}}}
% Entries of a matrix
\def\emLambda{{\Lambda}}
\def\emA{{A}}
\def\emB{{B}}
\def\emC{{C}}
\def\emD{{D}}
\def\emE{{E}}
\def\emF{{F}}
\def\emG{{G}}
\def\emH{{H}}
\def\emI{{I}}
\def\emJ{{J}}
\def\emK{{K}}
\def\emL{{L}}
\def\emM{{M}}
\def\emN{{N}}
\def\emO{{O}}
\def\emP{{P}}
\def\emQ{{Q}}
\def\emR{{R}}
\def\emS{{S}}
\def\emT{{T}}
\def\emU{{U}}
\def\emV{{V}}
\def\emW{{W}}
\def\emX{{X}}
\def\emY{{Y}}
\def\emZ{{Z}}
\def\emSigma{{\Sigma}}
% entries of a tensor
% Same font as tensor, without \bm wrapper
\newcommand{\etens}[1]{\mathsfit{#1}}
\def\etLambda{{\etens{\Lambda}}}
\def\etA{{\etens{A}}}
\def\etB{{\etens{B}}}
\def\etC{{\etens{C}}}
\def\etD{{\etens{D}}}
\def\etE{{\etens{E}}}
\def\etF{{\etens{F}}}
\def\etG{{\etens{G}}}
\def\etH{{\etens{H}}}
\def\etI{{\etens{I}}}
\def\etJ{{\etens{J}}}
\def\etK{{\etens{K}}}
\def\etL{{\etens{L}}}
\def\etM{{\etens{M}}}
\def\etN{{\etens{N}}}
\def\etO{{\etens{O}}}
\def\etP{{\etens{P}}}
\def\etQ{{\etens{Q}}}
\def\etR{{\etens{R}}}
\def\etS{{\etens{S}}}
\def\etT{{\etens{T}}}
\def\etU{{\etens{U}}}
\def\etV{{\etens{V}}}
\def\etW{{\etens{W}}}
\def\etX{{\etens{X}}}
\def\etY{{\etens{Y}}}
\def\etZ{{\etens{Z}}}
% The true underlying data generating distribution
\newcommand{\pdata}{p_{\rm{data}}}
% The empirical distribution defined by the training set
\newcommand{\ptrain}{\hat{p}_{\rm{data}}}
\newcommand{\Ptrain}{\hat{P}_{\rm{data}}}
% The model distribution
\newcommand{\pmodel}{p_{\rm{model}}}
\newcommand{\Pmodel}{P_{\rm{model}}}
\newcommand{\ptildemodel}{\tilde{p}_{\rm{model}}}
% Stochastic autoencoder distributions
\newcommand{\pencode}{p_{\rm{encoder}}}
\newcommand{\pdecode}{p_{\rm{decoder}}}
\newcommand{\precons}{p_{\rm{reconstruct}}}
\newcommand{\laplace}{\mathrm{Laplace}} % Laplace distribution
\newcommand{\E}{\mathbb{E}}
\newcommand{\Ls}{\mathcal{L}}
\newcommand{\R}{\mathbb{R}}
\newcommand{\emp}{\tilde{p}}
\newcommand{\lr}{\alpha}
\newcommand{\reg}{\lambda}
\newcommand{\rect}{\mathrm{rectifier}}
\newcommand{\softmax}{\mathrm{softmax}}
\newcommand{\sigmoid}{\sigma}
\newcommand{\softplus}{\zeta}
\newcommand{\KL}{D_{\mathrm{KL}}}
\newcommand{\Var}{\mathrm{Var}}
\newcommand{\standarderror}{\mathrm{SE}}
\newcommand{\Cov}{\mathrm{Cov}}
% Wolfram Mathworld says $L^2$ is for function spaces and $\ell^2$ is for vectors
% But then they seem to use $L^2$ for vectors throughout the site, and so does
% wikipedia.
\newcommand{\normlzero}{L^0}
\newcommand{\normlone}{L^1}
\newcommand{\normltwo}{L^2}
\newcommand{\normlp}{L^p}
\newcommand{\normmax}{L^\infty}
\newcommand{\parents}{Pa} % See usage in notation.tex. Chosen to match Daphne's book.
\DeclareMathOperator*{\argmax}{arg\,max}
\DeclareMathOperator*{\argmin}{arg\,min}
\DeclareMathOperator{\sign}{sign}
\DeclareMathOperator{\Tr}{Tr}
\let\ab\allowbreak
%%
%% This is file `natbib.sty',
%% generated with the docstrip utility.
%%
%% The original source files were:
%%
%% natbib.dtx (with options: `package,all')
%% =============================================
%% IMPORTANT NOTICE:
%%
%% This program can be redistributed and/or modified under the terms
%% of the LaTeX Project Public License Distributed from CTAN
%% archives in directory macros/latex/base/lppl.txt; either
%% version 1 of the License, or any later version.
%%
%% This is a generated file.
%% It may not be distributed without the original source file natbib.dtx.
%%
%% Full documentation can be obtained by LaTeXing that original file.
%% Only a few abbreviated comments remain here to describe the usage.
%% =============================================
%% Copyright 1993-2009 Patrick W Daly
%% Max-Planck-Institut f\"ur Sonnensystemforschung
%% Max-Planck-Str. 2
%% D-37191 Katlenburg-Lindau
%% Germany
%% E-mail: daly@mps.mpg.de
\NeedsTeXFormat{LaTeX2e}[1995/06/01]
\ProvidesPackage{natbib}
[2009/07/16 8.31 (PWD, AO)]
% This package reimplements the LaTeX \cite command to be used for various
% citation styles, both author-year and numerical. It accepts BibTeX
% output intended for many other packages, and therefore acts as a
% general, all-purpose citation-style interface.
%
% With standard numerical .bst files, only numerical citations are
% possible. With an author-year .bst file, both numerical and
% author-year citations are possible.
%
% If author-year citations are selected, \bibitem must have one of the
% following forms:
% \bibitem[Jones et al.(1990)]{key}...
% \bibitem[Jones et al.(1990)Jones, Baker, and Williams]{key}...
% \bibitem[Jones et al., 1990]{key}...
% \bibitem[\protect\citeauthoryear{Jones, Baker, and Williams}{Jones
% et al.}{1990}]{key}...
% \bibitem[\protect\citeauthoryear{Jones et al.}{1990}]{key}...
% \bibitem[\protect\astroncite{Jones et al.}{1990}]{key}...
% \bibitem[\protect\citename{Jones et al., }1990]{key}...
% \harvarditem[Jones et al.]{Jones, Baker, and Williams}{1990}{key}...
%
% This is either to be made up manually, or to be generated by an
% appropriate .bst file with BibTeX.
% Author-year mode || Numerical mode
% Then, \citet{key} ==>> Jones et al. (1990) || Jones et al. [21]
% \citep{key} ==>> (Jones et al., 1990) || [21]
% Multiple citations as normal:
% \citep{key1,key2} ==>> (Jones et al., 1990; Smith, 1989) || [21,24]
% or (Jones et al., 1990, 1991) || [21,24]
% or (Jones et al., 1990a,b) || [21,24]
% \cite{key} is the equivalent of \citet{key} in author-year mode
% and of \citep{key} in numerical mode
% Full author lists may be forced with \citet* or \citep*, e.g.
% \citep*{key} ==>> (Jones, Baker, and Williams, 1990)
% Optional notes as:
% \citep[chap. 2]{key} ==>> (Jones et al., 1990, chap. 2)
% \citep[e.g.,][]{key} ==>> (e.g., Jones et al., 1990)
% \citep[see][pg. 34]{key}==>> (see Jones et al., 1990, pg. 34)
% (Note: in standard LaTeX, only one note is allowed, after the ref.
% Here, one note is like the standard, two make pre- and post-notes.)
% \citealt{key} ==>> Jones et al. 1990
% \citealt*{key} ==>> Jones, Baker, and Williams 1990
% \citealp{key} ==>> Jones et al., 1990
% \citealp*{key} ==>> Jones, Baker, and Williams, 1990
% Additional citation possibilities (both author-year and numerical modes)
% \citeauthor{key} ==>> Jones et al.
% \citeauthor*{key} ==>> Jones, Baker, and Williams
% \citeyear{key} ==>> 1990
% \citeyearpar{key} ==>> (1990)
% \citetext{priv. comm.} ==>> (priv. comm.)
% \citenum{key} ==>> 11 [non-superscripted]
% Note: full author lists depends on whether the bib style supports them;
% if not, the abbreviated list is printed even when full requested.
%
% For names like della Robbia at the start of a sentence, use
% \Citet{dRob98} ==>> Della Robbia (1998)
% \Citep{dRob98} ==>> (Della Robbia, 1998)
% \Citeauthor{dRob98} ==>> Della Robbia
%
%
% Citation aliasing is achieved with
% \defcitealias{key}{text}
% \citetalias{key} ==>> text
% \citepalias{key} ==>> (text)
%
% Defining the citation mode and punctual (citation style)
% \setcitestyle{<comma-separated list of keywords, same
% as the package options>}
% Example: \setcitestyle{square,semicolon}
% Alternatively:
% Use \bibpunct with 6 mandatory arguments:
% 1. opening bracket for citation
% 2. closing bracket
% 3. citation separator (for multiple citations in one \cite)
% 4. the letter n for numerical styles, s for superscripts
% else anything for author-year
% 5. punctuation between authors and date
% 6. punctuation between years (or numbers) when common authors missing
% One optional argument is the character coming before post-notes. It
% appears in square braces before all other arguments. May be left off.
% Example (and default) \bibpunct[, ]{(}{)}{;}{a}{,}{,}
%
% To make this automatic for a given bib style, named newbib, say, make
% a local configuration file, natbib.cfg, with the definition
% \newcommand{\bibstyle@newbib}{\bibpunct...}
% Then the \bibliographystyle{newbib} will cause \bibstyle@newbib to
% be called on THE NEXT LATEX RUN (via the aux file).
%
% Such preprogrammed definitions may be invoked anywhere in the text
% by calling \citestyle{newbib}. This is only useful if the style specified
% differs from that in \bibliographystyle.
%
% With \citeindextrue and \citeindexfalse, one can control whether the
% \cite commands make an automatic entry of the citation in the .idx
% indexing file. For this, \makeindex must also be given in the preamble.
%
% Package Options: (for selecting punctuation)
% round - round parentheses are used (default)
% square - square brackets are used [option]
% curly - curly braces are used {option}
% angle - angle brackets are used <option>
% semicolon - multiple citations separated by semi-colon (default)
% colon - same as semicolon, an earlier confusion
% comma - separated by comma
% authoryear - selects author-year citations (default)
% numbers- selects numerical citations
% super - numerical citations as superscripts
% sort - sorts multiple citations according to order in ref. list
% sort&compress - like sort, but also compresses numerical citations
% compress - compresses without sorting
% longnamesfirst - makes first citation full author list
% sectionbib - puts bibliography in a \section* instead of \chapter*
% merge - allows the citation key to have a * prefix,
% signifying to merge its reference with that of the previous citation.
% elide - if references are merged, repeated portions of later ones may be removed.
% mcite - recognizes and ignores the * prefix for merging.
% Punctuation so selected dominates over any predefined ones.
% Package options are called as, e.g.
% \usepackage[square,comma]{natbib}
% LaTeX the source file natbib.dtx to obtain more details
% or the file natnotes.tex for a brief reference sheet.
%-----------------------------------------------------------
\providecommand\@ifxundefined[1]{%
\ifx#1\@undefined\expandafter\@firstoftwo\else\expandafter\@secondoftwo\fi
}%
\providecommand\@ifnum[1]{%
\ifnum#1\expandafter\@firstoftwo\else\expandafter\@secondoftwo\fi
}%
\providecommand\@ifx[1]{%
\ifx#1\expandafter\@firstoftwo\else\expandafter\@secondoftwo\fi
}%
\providecommand\appdef[2]{%
\toks@\expandafter{#1}\@temptokena{#2}%
\edef#1{\the\toks@\the\@temptokena}%
}%
\@ifclassloaded{agu2001}{\PackageError{natbib}
{The agu2001 class already includes natbib coding,\MessageBreak
so you should not add it explicitly}
{Type <Return> for now, but then later remove\MessageBreak
the command \protect\usepackage{natbib} from the document}
\endinput}{}
\@ifclassloaded{agutex}{\PackageError{natbib}
{The AGUTeX class already includes natbib coding,\MessageBreak
so you should not add it explicitly}
{Type <Return> for now, but then later remove\MessageBreak
the command \protect\usepackage{natbib} from the document}
\endinput}{}
\@ifclassloaded{aguplus}{\PackageError{natbib}
{The aguplus class already includes natbib coding,\MessageBreak
so you should not add it explicitly}
{Type <Return> for now, but then later remove\MessageBreak
the command \protect\usepackage{natbib} from the document}
\endinput}{}
\@ifclassloaded{nlinproc}{\PackageError{natbib}
{The nlinproc class already includes natbib coding,\MessageBreak
so you should not add it explicitly}
{Type <Return> for now, but then later remove\MessageBreak
the command \protect\usepackage{natbib} from the document}
\endinput}{}
\@ifclassloaded{egs}{\PackageError{natbib}
{The egs class already includes natbib coding,\MessageBreak
so you should not add it explicitly}
{Type <Return> for now, but then later remove\MessageBreak
the command \protect\usepackage{natbib} from the document}
\endinput}{}
\@ifclassloaded{egu}{\PackageError{natbib}
{The egu class already includes natbib coding,\MessageBreak
so you should not add it explicitly}
{Type <Return> for now, but then later remove\MessageBreak
the command \protect\usepackage{natbib} from the document}
\endinput}{}
% Define citation punctuation for some author-year styles
% One may add and delete at this point
% Or put additions into local configuration file natbib.cfg
\newcommand\bibstyle@chicago{\bibpunct{(}{)}{;}{a}{,}{,}}
\newcommand\bibstyle@named{\bibpunct{[}{]}{;}{a}{,}{,}}
\newcommand\bibstyle@agu{\bibpunct{[}{]}{;}{a}{,}{,~}}%Amer. Geophys. Union
\newcommand\bibstyle@copernicus{\bibpunct{(}{)}{;}{a}{,}{,}}%Copernicus Publications
\let\bibstyle@egu=\bibstyle@copernicus
\let\bibstyle@egs=\bibstyle@copernicus
\newcommand\bibstyle@agsm{\bibpunct{(}{)}{,}{a}{}{,}\gdef\harvardand{\&}}
\newcommand\bibstyle@kluwer{\bibpunct{(}{)}{,}{a}{}{,}\gdef\harvardand{\&}}
\newcommand\bibstyle@dcu{\bibpunct{(}{)}{;}{a}{;}{,}\gdef\harvardand{and}}
\newcommand\bibstyle@aa{\bibpunct{(}{)}{;}{a}{}{,}} %Astronomy & Astrophysics
\newcommand\bibstyle@pass{\bibpunct{(}{)}{;}{a}{,}{,}}%Planet. & Space Sci
\newcommand\bibstyle@anngeo{\bibpunct{(}{)}{;}{a}{,}{,}}%Annales Geophysicae
\newcommand\bibstyle@nlinproc{\bibpunct{(}{)}{;}{a}{,}{,}}%Nonlin.Proc.Geophys.
% Define citation punctuation for some numerical styles
\newcommand\bibstyle@cospar{\bibpunct{/}{/}{,}{n}{}{}%
\gdef\bibnumfmt##1{##1.}}
\newcommand\bibstyle@esa{\bibpunct{(Ref.~}{)}{,}{n}{}{}%
\gdef\bibnumfmt##1{##1.\hspace{1em}}}
\newcommand\bibstyle@nature{\bibpunct{}{}{,}{s}{}{\textsuperscript{,}}%
\gdef\bibnumfmt##1{##1.}}
% The standard LaTeX styles
\newcommand\bibstyle@plain{\bibpunct{[}{]}{,}{n}{}{,}}
\let\bibstyle@alpha=\bibstyle@plain
\let\bibstyle@abbrv=\bibstyle@plain
\let\bibstyle@unsrt=\bibstyle@plain
% The author-year modifications of the standard styles
\newcommand\bibstyle@plainnat{\bibpunct{[}{]}{,}{a}{,}{,}}
\let\bibstyle@abbrvnat=\bibstyle@plainnat
\let\bibstyle@unsrtnat=\bibstyle@plainnat
\newif\ifNAT@numbers \NAT@numbersfalse
\newif\ifNAT@super \NAT@superfalse
\let\NAT@merge\z@
\DeclareOption{numbers}{\NAT@numberstrue
\ExecuteOptions{square,comma,nobibstyle}}
\DeclareOption{super}{\NAT@supertrue\NAT@numberstrue
\renewcommand\NAT@open{}\renewcommand\NAT@close{}
\ExecuteOptions{nobibstyle}}
\DeclareOption{authoryear}{\NAT@numbersfalse
\ExecuteOptions{round,semicolon,bibstyle}}
\DeclareOption{round}{%
\renewcommand\NAT@open{(} \renewcommand\NAT@close{)}
\ExecuteOptions{nobibstyle}}
\DeclareOption{square}{%
\renewcommand\NAT@open{[} \renewcommand\NAT@close{]}
\ExecuteOptions{nobibstyle}}
\DeclareOption{angle}{%
\renewcommand\NAT@open{$<$} \renewcommand\NAT@close{$>$}
\ExecuteOptions{nobibstyle}}
\DeclareOption{curly}{%
\renewcommand\NAT@open{\{} \renewcommand\NAT@close{\}}
\ExecuteOptions{nobibstyle}}
\DeclareOption{comma}{\renewcommand\NAT@sep{,}
\ExecuteOptions{nobibstyle}}
\DeclareOption{semicolon}{\renewcommand\NAT@sep{;}
\ExecuteOptions{nobibstyle}}
\DeclareOption{colon}{\ExecuteOptions{semicolon}}
\DeclareOption{nobibstyle}{\let\bibstyle=\@gobble}
\DeclareOption{bibstyle}{\let\bibstyle=\@citestyle}
\newif\ifNAT@openbib \NAT@openbibfalse
\DeclareOption{openbib}{\NAT@openbibtrue}
\DeclareOption{sectionbib}{\def\NAT@sectionbib{on}}
\def\NAT@sort{\z@}
\def\NAT@cmprs{\z@}
\DeclareOption{sort}{\def\NAT@sort{\@ne}}
\DeclareOption{compress}{\def\NAT@cmprs{\@ne}}
\DeclareOption{sort&compress}{\def\NAT@sort{\@ne}\def\NAT@cmprs{\@ne}}
\DeclareOption{mcite}{\let\NAT@merge\@ne}
\DeclareOption{merge}{\@ifnum{\NAT@merge<\tw@}{\let\NAT@merge\tw@}{}}
\DeclareOption{elide}{\@ifnum{\NAT@merge<\thr@@}{\let\NAT@merge\thr@@}{}}
\@ifpackageloaded{cite}{\PackageWarningNoLine{natbib}
{The `cite' package should not be used\MessageBreak
with natbib. Use option `sort' instead}\ExecuteOptions{sort}}{}
\@ifpackageloaded{mcite}{\PackageWarningNoLine{natbib}
{The `mcite' package should not be used\MessageBreak
with natbib. Use option `merge' instead}\ExecuteOptions{merge}}{}
\@ifpackageloaded{citeref}{\PackageError{natbib}
{The `citeref' package must be loaded after natbib}%
{Move \protect\usepackage{citeref} to after \string\usepackage{natbib}}}{}
\newif\ifNAT@longnames\NAT@longnamesfalse
\DeclareOption{longnamesfirst}{\NAT@longnamestrue}
\DeclareOption{nonamebreak}{\def\NAT@nmfmt#1{\mbox{\NAT@up#1}}}
\def\NAT@nmfmt#1{{\NAT@up#1}}
\renewcommand\bibstyle[1]{\csname bibstyle@#1\endcsname}
\AtBeginDocument{\global\let\bibstyle=\@gobble}
\let\@citestyle\bibstyle
\newcommand\citestyle[1]{\@citestyle{#1}\let\bibstyle\@gobble}
\newcommand\bibpunct[7][, ]%
{\gdef\NAT@open{#2}\gdef\NAT@close{#3}\gdef
\NAT@sep{#4}\global\NAT@numbersfalse
\ifx #5n\global\NAT@numberstrue\global\NAT@superfalse
\else
\ifx #5s\global\NAT@numberstrue\global\NAT@supertrue
\fi\fi
\gdef\NAT@aysep{#6}\gdef\NAT@yrsep{#7}%
\gdef\NAT@cmt{#1}%
\NAT@@setcites
}
\newcommand\setcitestyle[1]{
\@for\@tempa:=#1\do
{\def\@tempb{round}\ifx\@tempa\@tempb
\renewcommand\NAT@open{(}\renewcommand\NAT@close{)}\fi
\def\@tempb{square}\ifx\@tempa\@tempb
\renewcommand\NAT@open{[}\renewcommand\NAT@close{]}\fi
\def\@tempb{angle}\ifx\@tempa\@tempb
\renewcommand\NAT@open{$<$}\renewcommand\NAT@close{$>$}\fi
\def\@tempb{curly}\ifx\@tempa\@tempb
\renewcommand\NAT@open{\{}\renewcommand\NAT@close{\}}\fi
\def\@tempb{semicolon}\ifx\@tempa\@tempb
\renewcommand\NAT@sep{;}\fi
\def\@tempb{colon}\ifx\@tempa\@tempb
\renewcommand\NAT@sep{;}\fi
\def\@tempb{comma}\ifx\@tempa\@tempb
\renewcommand\NAT@sep{,}\fi
\def\@tempb{authoryear}\ifx\@tempa\@tempb
\NAT@numbersfalse\fi
\def\@tempb{numbers}\ifx\@tempa\@tempb
\NAT@numberstrue\NAT@superfalse\fi
\def\@tempb{super}\ifx\@tempa\@tempb
\NAT@numberstrue\NAT@supertrue\fi
\expandafter\NAT@find@eq\@tempa=\relax\@nil
\if\@tempc\relax\else
\expandafter\NAT@rem@eq\@tempc
\def\@tempb{open}\ifx\@tempa\@tempb
\xdef\NAT@open{\@tempc}\fi
\def\@tempb{close}\ifx\@tempa\@tempb
\xdef\NAT@close{\@tempc}\fi
\def\@tempb{aysep}\ifx\@tempa\@tempb
\xdef\NAT@aysep{\@tempc}\fi
\def\@tempb{yysep}\ifx\@tempa\@tempb
\xdef\NAT@yrsep{\@tempc}\fi
\def\@tempb{notesep}\ifx\@tempa\@tempb
\xdef\NAT@cmt{\@tempc}\fi
\def\@tempb{citesep}\ifx\@tempa\@tempb
\xdef\NAT@sep{\@tempc}\fi
\fi
}%
\NAT@@setcites
}
\def\NAT@find@eq#1=#2\@nil{\def\@tempa{#1}\def\@tempc{#2}}
\def\NAT@rem@eq#1={\def\@tempc{#1}}
\def\NAT@@setcites{\global\let\bibstyle\@gobble}
\AtBeginDocument{\let\NAT@@setcites\NAT@set@cites}
\newcommand\NAT@open{(} \newcommand\NAT@close{)}
\newcommand\NAT@sep{;}
\ProcessOptions
\newcommand\NAT@aysep{,} \newcommand\NAT@yrsep{,}
\newcommand\NAT@cmt{, }
\newcommand\NAT@cite%
[3]{\ifNAT@swa\NAT@@open\if*#2*\else#2\NAT@spacechar\fi
#1\if*#3*\else\NAT@cmt#3\fi\NAT@@close\else#1\fi\endgroup}
\newcommand\NAT@citenum%
[3]{\ifNAT@swa\NAT@@open\if*#2*\else#2\NAT@spacechar\fi
#1\if*#3*\else\NAT@cmt#3\fi\NAT@@close\else#1\fi\endgroup}
\newcommand\NAT@citesuper[3]{\ifNAT@swa
\if*#2*\else#2\NAT@spacechar\fi
\unskip\kern\p@\textsuperscript{\NAT@@open#1\NAT@@close}%
\if*#3*\else\NAT@spacechar#3\fi\else #1\fi\endgroup}
\providecommand\textsuperscript[1]{\mbox{$^{\mbox{\scriptsize#1}}$}}
\begingroup \catcode`\_=8
\gdef\NAT@ifcat@num#1{%
\ifcat_\ifnum\z@<0#1_\else A\fi
\expandafter\@firstoftwo
\else
\expandafter\@secondoftwo
\fi
}%
\endgroup
\providecommand\@firstofone[1]{#1}
\newcommand\NAT@citexnum{}
\def\NAT@citexnum[#1][#2]#3{%
\NAT@reset@parser
\NAT@sort@cites{#3}%
\NAT@reset@citea
\@cite{\def\NAT@num{-1}\let\NAT@last@yr\relax\let\NAT@nm\@empty
\@for\@citeb:=\NAT@cite@list\do
{\@safe@activestrue
\edef\@citeb{\expandafter\@firstofone\@citeb\@empty}%
\@safe@activesfalse
\@ifundefined{b@\@citeb\@extra@b@citeb}{%
{\reset@font\bfseries?}
\NAT@citeundefined\PackageWarning{natbib}%
{Citation `\@citeb' on page \thepage \space undefined}}%
{\let\NAT@last@num\NAT@num\let\NAT@last@nm\NAT@nm
\NAT@parse{\@citeb}%
\ifNAT@longnames\@ifundefined{bv@\@citeb\@extra@b@citeb}{%
\let\NAT@name=\NAT@all@names
\global\@namedef{bv@\@citeb\@extra@b@citeb}{}}{}%
\fi
\ifNAT@full\let\NAT@nm\NAT@all@names\else
\let\NAT@nm\NAT@name\fi
\ifNAT@swa
\@ifnum{\NAT@ctype>\@ne}{%
\@citea
\NAT@hyper@{\@ifnum{\NAT@ctype=\tw@}{\NAT@test{\NAT@ctype}}{\NAT@alias}}%
}{%
\@ifnum{\NAT@cmprs>\z@}{%
\NAT@ifcat@num\NAT@num
{\let\NAT@nm=\NAT@num}%
{\def\NAT@nm{-2}}%
\NAT@ifcat@num\NAT@last@num
{\@tempcnta=\NAT@last@num\relax}%
{\@tempcnta\m@ne}%
\@ifnum{\NAT@nm=\@tempcnta}{%
\@ifnum{\NAT@merge>\@ne}{}{\NAT@last@yr@mbox}%
}{%
\advance\@tempcnta by\@ne
\@ifnum{\NAT@nm=\@tempcnta}{%
\ifx\NAT@last@yr\relax
\def@NAT@last@yr{\@citea}%
\else
\def@NAT@last@yr{--\NAT@penalty}%
\fi
}{%
\NAT@last@yr@mbox
}%
}%
}{%
\@tempswatrue
\@ifnum{\NAT@merge>\@ne}{\@ifnum{\NAT@last@num=\NAT@num\relax}{\@tempswafalse}{}}{}%
\if@tempswa\NAT@citea@mbox\fi
}%
}%
\NAT@def@citea
\else
\ifcase\NAT@ctype
\ifx\NAT@last@nm\NAT@nm \NAT@yrsep\NAT@penalty\NAT@space\else
\@citea \NAT@test{\@ne}\NAT@spacechar\NAT@mbox{\NAT@super@kern\NAT@@open}%
\fi
\if*#1*\else#1\NAT@spacechar\fi
\NAT@mbox{\NAT@hyper@{{\citenumfont{\NAT@num}}}}%
\NAT@def@citea@box
\or
\NAT@hyper@citea@space{\NAT@test{\NAT@ctype}}%
\or
\NAT@hyper@citea@space{\NAT@test{\NAT@ctype}}%
\or
\NAT@hyper@citea@space\NAT@alias
\fi
\fi
}%
}%
\@ifnum{\NAT@cmprs>\z@}{\NAT@last@yr}{}%
\ifNAT@swa\else
\@ifnum{\NAT@ctype=\z@}{%
\if*#2*\else\NAT@cmt#2\fi
}{}%
\NAT@mbox{\NAT@@close}%
\fi
}{#1}{#2}%
}%
\def\NAT@citea@mbox{%
\@citea\mbox{\NAT@hyper@{{\citenumfont{\NAT@num}}}}%
}%
\def\NAT@hyper@#1{%
\hyper@natlinkstart{\@citeb\@extra@b@citeb}#1\hyper@natlinkend
}%
\def\NAT@hyper@citea#1{%
\@citea
\NAT@hyper@{#1}%
\NAT@def@citea
}%
\def\NAT@hyper@citea@space#1{%
\@citea
\NAT@hyper@{#1}%
\NAT@def@citea@space
}%
\def\def@NAT@last@yr#1{%
\protected@edef\NAT@last@yr{%
#1%
\noexpand\mbox{%
\noexpand\hyper@natlinkstart{\@citeb\@extra@b@citeb}%
{\noexpand\citenumfont{\NAT@num}}%
\noexpand\hyper@natlinkend
}%
}%
}%
\def\NAT@last@yr@mbox{%
\NAT@last@yr\let\NAT@last@yr\relax
\NAT@citea@mbox
}%
\newcommand\NAT@test[1]{%
\@ifnum{#1=\@ne}{%
\ifx\NAT@nm\NAT@noname
\begingroup\reset@font\bfseries(author?)\endgroup
\PackageWarning{natbib}{%
Author undefined for citation`\@citeb' \MessageBreak on page \thepage%
}%
\else \NAT@nm
\fi
}{%
\if\relax\NAT@date\relax
\begingroup\reset@font\bfseries(year?)\endgroup
\PackageWarning{natbib}{%
Year undefined for citation`\@citeb' \MessageBreak on page \thepage%
}%
\else \NAT@date
\fi
}%
}%
\let\citenumfont=\@empty
\newcommand\NAT@citex{}
\def\NAT@citex%
[#1][#2]#3{%
\NAT@reset@parser
\NAT@sort@cites{#3}%
\NAT@reset@citea
\@cite{\let\NAT@nm\@empty\let\NAT@year\@empty
\@for\@citeb:=\NAT@cite@list\do
{\@safe@activestrue
\edef\@citeb{\expandafter\@firstofone\@citeb\@empty}%
\@safe@activesfalse
\@ifundefined{b@\@citeb\@extra@b@citeb}{\@citea%
{\reset@font\bfseries ?}\NAT@citeundefined
\PackageWarning{natbib}%
{Citation `\@citeb' on page \thepage \space undefined}\def\NAT@date{}}%
{\let\NAT@last@nm=\NAT@nm\let\NAT@last@yr=\NAT@year
\NAT@parse{\@citeb}%
\ifNAT@longnames\@ifundefined{bv@\@citeb\@extra@b@citeb}{%
\let\NAT@name=\NAT@all@names
\global\@namedef{bv@\@citeb\@extra@b@citeb}{}}{}%
\fi
\ifNAT@full\let\NAT@nm\NAT@all@names\else
\let\NAT@nm\NAT@name\fi
\ifNAT@swa\ifcase\NAT@ctype
\if\relax\NAT@date\relax
\@citea\NAT@hyper@{\NAT@nmfmt{\NAT@nm}\NAT@date}%
\else
\ifx\NAT@last@nm\NAT@nm\NAT@yrsep
\ifx\NAT@last@yr\NAT@year
\def\NAT@temp{{?}}%
\ifx\NAT@temp\NAT@exlab\PackageWarningNoLine{natbib}%
{Multiple citation on page \thepage: same authors and
year\MessageBreak without distinguishing extra
letter,\MessageBreak appears as question mark}\fi
\NAT@hyper@{\NAT@exlab}%
\else\unskip\NAT@spacechar
\NAT@hyper@{\NAT@date}%
\fi
\else
\@citea\NAT@hyper@{%
\NAT@nmfmt{\NAT@nm}%
\hyper@natlinkbreak{%
\NAT@aysep\NAT@spacechar}{\@citeb\@extra@b@citeb
}%
\NAT@date
}%
\fi
\fi
\or\@citea\NAT@hyper@{\NAT@nmfmt{\NAT@nm}}%
\or\@citea\NAT@hyper@{\NAT@date}%
\or\@citea\NAT@hyper@{\NAT@alias}%
\fi \NAT@def@citea
\else
\ifcase\NAT@ctype
\if\relax\NAT@date\relax
\@citea\NAT@hyper@{\NAT@nmfmt{\NAT@nm}}%
\else
\ifx\NAT@last@nm\NAT@nm\NAT@yrsep
\ifx\NAT@last@yr\NAT@year
\def\NAT@temp{{?}}%
\ifx\NAT@temp\NAT@exlab\PackageWarningNoLine{natbib}%
{Multiple citation on page \thepage: same authors and
year\MessageBreak without distinguishing extra
letter,\MessageBreak appears as question mark}\fi
\NAT@hyper@{\NAT@exlab}%
\else
\unskip\NAT@spacechar
\NAT@hyper@{\NAT@date}%
\fi
\else
\@citea\NAT@hyper@{%
\NAT@nmfmt{\NAT@nm}%
\hyper@natlinkbreak{\NAT@spacechar\NAT@@open\if*#1*\else#1\NAT@spacechar\fi}%
{\@citeb\@extra@b@citeb}%
\NAT@date
}%
\fi
\fi
\or\@citea\NAT@hyper@{\NAT@nmfmt{\NAT@nm}}%
\or\@citea\NAT@hyper@{\NAT@date}%
\or\@citea\NAT@hyper@{\NAT@alias}%
\fi
\if\relax\NAT@date\relax
\NAT@def@citea
\else
\NAT@def@citea@close
\fi
\fi
}}\ifNAT@swa\else\if*#2*\else\NAT@cmt#2\fi
\if\relax\NAT@date\relax\else\NAT@@close\fi\fi}{#1}{#2}}
\def\NAT@spacechar{\ }%
\def\NAT@separator{\NAT@sep\NAT@penalty}%
\def\NAT@reset@citea{\c@NAT@ctr\@ne\let\@citea\@empty}%
\def\NAT@def@citea{\def\@citea{\NAT@separator\NAT@space}}%
\def\NAT@def@citea@space{\def\@citea{\NAT@separator\NAT@spacechar}}%
\def\NAT@def@citea@close{\def\@citea{\NAT@@close\NAT@separator\NAT@space}}%
\def\NAT@def@citea@box{\def\@citea{\NAT@mbox{\NAT@@close}\NAT@separator\NAT@spacechar}}%
\newif\ifNAT@par \NAT@partrue
\newcommand\NAT@@open{\ifNAT@par\NAT@open\fi}
\newcommand\NAT@@close{\ifNAT@par\NAT@close\fi}
\newcommand\NAT@alias{\@ifundefined{al@\@citeb\@extra@b@citeb}{%
{\reset@font\bfseries(alias?)}\PackageWarning{natbib}
{Alias undefined for citation `\@citeb'
\MessageBreak on page \thepage}}{\@nameuse{al@\@citeb\@extra@b@citeb}}}
\let\NAT@up\relax
\newcommand\NAT@Up[1]{{\let\protect\@unexpandable@protect\let~\relax
\expandafter\NAT@deftemp#1}\expandafter\NAT@UP\NAT@temp}
\newcommand\NAT@deftemp[1]{\xdef\NAT@temp{#1}}
\newcommand\NAT@UP[1]{\let\@tempa\NAT@UP\ifcat a#1\MakeUppercase{#1}%
\let\@tempa\relax\else#1\fi\@tempa}
\newcommand\shortcites[1]{%
\@bsphack\@for\@citeb:=#1\do
{\@safe@activestrue
\edef\@citeb{\expandafter\@firstofone\@citeb\@empty}%
\@safe@activesfalse
\global\@namedef{bv@\@citeb\@extra@b@citeb}{}}\@esphack}
\newcommand\NAT@biblabel[1]{\hfill}
\newcommand\NAT@biblabelnum[1]{\bibnumfmt{#1}}
\let\bibnumfmt\@empty
\providecommand\@biblabel[1]{[#1]}
\AtBeginDocument{\ifx\bibnumfmt\@empty\let\bibnumfmt\@biblabel\fi}
\newcommand\NAT@bibsetnum[1]{\settowidth\labelwidth{\@biblabel{#1}}%
\setlength{\leftmargin}{\labelwidth}\addtolength{\leftmargin}{\labelsep}%
\setlength{\itemsep}{\bibsep}\setlength{\parsep}{\z@}%
\ifNAT@openbib
\addtolength{\leftmargin}{\bibindent}%
\setlength{\itemindent}{-\bibindent}%
\setlength{\listparindent}{\itemindent}%
\setlength{\parsep}{0pt}%
\fi
}
\newlength{\bibhang}
\setlength{\bibhang}{1em}
\newlength{\bibsep}
{\@listi \global\bibsep\itemsep \global\advance\bibsep by\parsep}
\newcommand\NAT@bibsetup%
[1]{\setlength{\leftmargin}{\bibhang}\setlength{\itemindent}{-\leftmargin}%
\setlength{\itemsep}{\bibsep}\setlength{\parsep}{\z@}}
\newcommand\NAT@set@cites{%
\ifNAT@numbers
\ifNAT@super \let\@cite\NAT@citesuper
\def\NAT@mbox##1{\unskip\nobreak\textsuperscript{##1}}%
\let\citeyearpar=\citeyear
\let\NAT@space\relax
\def\NAT@super@kern{\kern\p@}%
\else
\let\NAT@mbox=\mbox
\let\@cite\NAT@citenum
\let\NAT@space\NAT@spacechar
\let\NAT@super@kern\relax
\fi
\let\@citex\NAT@citexnum
\let\@biblabel\NAT@biblabelnum
\let\@bibsetup\NAT@bibsetnum
\renewcommand\NAT@idxtxt{\NAT@name\NAT@spacechar\NAT@open\NAT@num\NAT@close}%
\def\natexlab##1{}%
\def\NAT@penalty{\penalty\@m}%
\else
\let\@cite\NAT@cite
\let\@citex\NAT@citex
\let\@biblabel\NAT@biblabel
\let\@bibsetup\NAT@bibsetup
\let\NAT@space\NAT@spacechar
\let\NAT@penalty\@empty
\renewcommand\NAT@idxtxt{\NAT@name\NAT@spacechar\NAT@open\NAT@date\NAT@close}%
\def\natexlab##1{##1}%
\fi}
\AtBeginDocument{\NAT@set@cites}
\AtBeginDocument{\ifx\SK@def\@undefined\else
\ifx\SK@cite\@empty\else
\SK@def\@citex[#1][#2]#3{\SK@\SK@@ref{#3}\SK@@citex[#1][#2]{#3}}\fi
\ifx\SK@citeauthor\@undefined\def\HAR@checkdef{}\else
\let\citeauthor\SK@citeauthor
\let\citefullauthor\SK@citefullauthor
\let\citeyear\SK@citeyear\fi
\fi}
\newif\ifNAT@full\NAT@fullfalse
\newif\ifNAT@swa
\DeclareRobustCommand\citet
{\begingroup\NAT@swafalse\let\NAT@ctype\z@\NAT@partrue
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\newcommand\NAT@citetp{\@ifnextchar[{\NAT@@citetp}{\NAT@@citetp[]}}
\newcommand\NAT@@citetp{}
\def\NAT@@citetp[#1]{\@ifnextchar[{\@citex[#1]}{\@citex[][#1]}}
\DeclareRobustCommand\citep
{\begingroup\NAT@swatrue\let\NAT@ctype\z@\NAT@partrue
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\DeclareRobustCommand\cite
{\begingroup\let\NAT@ctype\z@\NAT@partrue\NAT@swatrue
\@ifstar{\NAT@fulltrue\NAT@cites}{\NAT@fullfalse\NAT@cites}}
\newcommand\NAT@cites{\@ifnextchar [{\NAT@@citetp}{%
\ifNAT@numbers\else
\NAT@swafalse
\fi
\NAT@@citetp[]}}
\DeclareRobustCommand\citealt
{\begingroup\NAT@swafalse\let\NAT@ctype\z@\NAT@parfalse
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\DeclareRobustCommand\citealp
{\begingroup\NAT@swatrue\let\NAT@ctype\z@\NAT@parfalse
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\DeclareRobustCommand\citenum
{\begingroup
\NAT@swatrue\let\NAT@ctype\z@\NAT@parfalse\let\textsuperscript\NAT@spacechar
\NAT@citexnum[][]}
\DeclareRobustCommand\citeauthor
{\begingroup\NAT@swafalse\let\NAT@ctype\@ne\NAT@parfalse
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\DeclareRobustCommand\Citet
{\begingroup\NAT@swafalse\let\NAT@ctype\z@\NAT@partrue
\let\NAT@up\NAT@Up
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\DeclareRobustCommand\Citep
{\begingroup\NAT@swatrue\let\NAT@ctype\z@\NAT@partrue
\let\NAT@up\NAT@Up
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\DeclareRobustCommand\Citealt
{\begingroup\NAT@swafalse\let\NAT@ctype\z@\NAT@parfalse
\let\NAT@up\NAT@Up
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\DeclareRobustCommand\Citealp
{\begingroup\NAT@swatrue\let\NAT@ctype\z@\NAT@parfalse
\let\NAT@up\NAT@Up
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\DeclareRobustCommand\Citeauthor
{\begingroup\NAT@swafalse\let\NAT@ctype\@ne\NAT@parfalse
\let\NAT@up\NAT@Up
\@ifstar{\NAT@fulltrue\NAT@citetp}{\NAT@fullfalse\NAT@citetp}}
\DeclareRobustCommand\citeyear
{\begingroup\NAT@swafalse\let\NAT@ctype\tw@\NAT@parfalse\NAT@citetp}
\DeclareRobustCommand\citeyearpar
{\begingroup\NAT@swatrue\let\NAT@ctype\tw@\NAT@partrue\NAT@citetp}
\newcommand\citetext[1]{\NAT@open#1\NAT@close}
\DeclareRobustCommand\citefullauthor
{\citeauthor*}
\newcommand\defcitealias[2]{%
\@ifundefined{al@#1\@extra@b@citeb}{}
{\PackageWarning{natbib}{Overwriting existing alias for citation #1}}
\@namedef{al@#1\@extra@b@citeb}{#2}}
\DeclareRobustCommand\citetalias{\begingroup
\NAT@swafalse\let\NAT@ctype\thr@@\NAT@parfalse\NAT@citetp}
\DeclareRobustCommand\citepalias{\begingroup
\NAT@swatrue\let\NAT@ctype\thr@@\NAT@partrue\NAT@citetp}
\renewcommand\nocite[1]{\@bsphack
\@for\@citeb:=#1\do{%
\@safe@activestrue
\edef\@citeb{\expandafter\@firstofone\@citeb\@empty}%
\@safe@activesfalse
\if@filesw\immediate\write\@auxout{\string\citation{\@citeb}}\fi
\if*\@citeb\else
\@ifundefined{b@\@citeb\@extra@b@citeb}{%
\NAT@citeundefined \PackageWarning{natbib}%
{Citation `\@citeb' undefined}}{}\fi}%
\@esphack}
\newcommand\NAT@parse[1]{%
\begingroup
\let\protect=\@unexpandable@protect
\let~\relax
\let\active@prefix=\@gobble
\edef\NAT@temp{\csname b@#1\@extra@b@citeb\endcsname}%
\aftergroup\NAT@split
\expandafter
\endgroup
\NAT@temp{}{}{}{}{}@@%
\expandafter\NAT@parse@date\NAT@date??????@@%
\ifciteindex\NAT@index\fi
}%
\def\NAT@split#1#2#3#4#5@@{%
\gdef\NAT@num{#1}\gdef\NAT@name{#3}\gdef\NAT@date{#2}%
\gdef\NAT@all@names{#4}%
\ifx\NAT@num\@empty\gdef\NAT@num{0}\fi
\ifx\NAT@noname\NAT@all@names \gdef\NAT@all@names{#3}\fi
}%
\def\NAT@reset@parser{%
\global\let\NAT@num\@empty
\global\let\NAT@name\@empty
\global\let\NAT@date\@empty
\global\let\NAT@all@names\@empty
}%
\newcommand\NAT@parse@date{}
\def\NAT@parse@date#1#2#3#4#5#6@@{%
\ifnum\the\catcode`#1=11\def\NAT@year{}\def\NAT@exlab{#1}\else
\ifnum\the\catcode`#2=11\def\NAT@year{#1}\def\NAT@exlab{#2}\else
\ifnum\the\catcode`#3=11\def\NAT@year{#1#2}\def\NAT@exlab{#3}\else
\ifnum\the\catcode`#4=11\def\NAT@year{#1#2#3}\def\NAT@exlab{#4}\else
\def\NAT@year{#1#2#3#4}\def\NAT@exlab{{#5}}\fi\fi\fi\fi}
\newcommand\NAT@index{}
\let\NAT@makeindex=\makeindex
\renewcommand\makeindex{\NAT@makeindex
\renewcommand\NAT@index{\@bsphack\begingroup
\def~{\string~}\@wrindex{\NAT@idxtxt}}}
\newcommand\NAT@idxtxt{\NAT@name\NAT@spacechar\NAT@open\NAT@date\NAT@close}
\@ifxundefined\@indexfile{}{\let\NAT@makeindex\relax\makeindex}
\newif\ifciteindex \citeindexfalse
\newcommand\citeindextype{default}
\newcommand\NAT@index@alt{{\let\protect=\noexpand\let~\relax
\xdef\NAT@temp{\NAT@idxtxt}}\expandafter\NAT@exp\NAT@temp\@nil}
\newcommand\NAT@exp{}
\def\NAT@exp#1\@nil{\index[\citeindextype]{#1}}
\AtBeginDocument{%
\@ifpackageloaded{index}{\let\NAT@index=\NAT@index@alt}{}}
\newcommand\NAT@ifcmd{\futurelet\NAT@temp\NAT@ifxcmd}
\newcommand\NAT@ifxcmd{\ifx\NAT@temp\relax\else\expandafter\NAT@bare\fi}
\def\NAT@bare#1(#2)#3(@)#4\@nil#5{%
\if @#2
\expandafter\NAT@apalk#1, , \@nil{#5}%
\else
\NAT@wrout{\the\c@NAT@ctr}{#2}{#1}{#3}{#5}%
\fi
}
\newcommand\NAT@wrout[5]{%
\if@filesw
{\let\protect\noexpand\let~\relax
\immediate
\write\@auxout{\string\bibcite{#5}{{#1}{#2}{{#3}}{{#4}}}}}\fi
\ignorespaces}
\def\NAT@noname{{}}
\renewcommand\bibitem{\@ifnextchar[{\@lbibitem}{\@lbibitem[]}}%
\let\NAT@bibitem@first@sw\@secondoftwo
\def\@lbibitem[#1]#2{%
\if\relax\@extra@b@citeb\relax\else
\@ifundefined{br@#2\@extra@b@citeb}{}{%
\@namedef{br@#2}{\@nameuse{br@#2\@extra@b@citeb}}%
}%
\fi
\@ifundefined{b@#2\@extra@b@citeb}{%
\def\NAT@num{}%
}{%
\NAT@parse{#2}%
}%
\def\NAT@tmp{#1}%
\expandafter\let\expandafter\bibitemOpen\csname NAT@b@open@#2\endcsname
\expandafter\let\expandafter\bibitemShut\csname NAT@b@shut@#2\endcsname
\@ifnum{\NAT@merge>\@ne}{%
\NAT@bibitem@first@sw{%
\@firstoftwo
}{%
\@ifundefined{NAT@b*@#2}{%
\@firstoftwo
}{%
\expandafter\def\expandafter\NAT@num\expandafter{\the\c@NAT@ctr}%
\@secondoftwo
}%
}%
}{%
\@firstoftwo
}%
{%
\global\advance\c@NAT@ctr\@ne
\@ifx{\NAT@tmp\@empty}{\@firstoftwo}{%
\@secondoftwo
}%
{%
\expandafter\def\expandafter\NAT@num\expandafter{\the\c@NAT@ctr}%
\global\NAT@stdbsttrue
}{}%
\bibitem@fin
\item[\hfil\NAT@anchor{#2}{\NAT@num}]%
\global\let\NAT@bibitem@first@sw\@secondoftwo
\NAT@bibitem@init
}%
{%
\NAT@anchor{#2}{}%
\NAT@bibitem@cont
\bibitem@fin
}%
\@ifx{\NAT@tmp\@empty}{%
\NAT@wrout{\the\c@NAT@ctr}{}{}{}{#2}%
}{%
\expandafter\NAT@ifcmd\NAT@tmp(@)(@)\@nil{#2}%
}%
}%
\def\bibitem@fin{%
\@ifxundefined\@bibstop{}{\csname bibitem@\@bibstop\endcsname}%
}%
\def\NAT@bibitem@init{%
\let\@bibstop\@undefined
}%
\def\NAT@bibitem@cont{%
\let\bibitem@Stop\bibitemStop
\let\bibitem@NoStop\bibitemContinue
}%
\def\BibitemOpen{%
\bibitemOpen
}%
\def\BibitemShut#1{%
\bibitemShut
\def\@bibstop{#1}%
\let\bibitem@Stop\bibitemStop
\let\bibitem@NoStop\bibitemNoStop
}%
\def\bibitemStop{}%
\def\bibitemNoStop{.\spacefactor\@mmm\space}%
\def\bibitemContinue{\spacefactor\@mmm\space}%
\mathchardef\@mmm=3000 %
\providecommand{\bibAnnote}[3]{%
\BibitemShut{#1}%
\def\@tempa{#3}\@ifx{\@tempa\@empty}{}{%
\begin{quotation}\noindent
\textsc{Key:}\ #2\\\textsc{Annotation:}\ \@tempa
\end{quotation}%
}%
}%
\providecommand{\bibAnnoteFile}[2]{%
\IfFileExists{#2}{%
\bibAnnote{#1}{#2}{\input{#2}}%
}{%
\bibAnnote{#1}{#2}{}%
}%
}%
\let\bibitemOpen\relax
\let\bibitemShut\relax
\def\bibfield{\@ifnum{\NAT@merge>\tw@}{\@bibfield}{\@secondoftwo}}%
\def\@bibfield#1#2{%
\begingroup
\let\Doi\@gobble
\let\bibinfo\relax
\let\restore@protect\@empty
\protected@edef\@tempa{#2}%
\aftergroup\def\aftergroup\@tempa
\expandafter\endgroup\expandafter{\@tempa}%
\expandafter\@ifx\expandafter{\csname @bib#1\endcsname\@tempa}{%
\expandafter\let\expandafter\@tempa\csname @bib@X#1\endcsname
}{%
\expandafter\let\csname @bib#1\endcsname\@tempa
\expandafter\let\expandafter\@tempa\csname @bib@Y#1\endcsname
}%
\@ifx{\@tempa\relax}{\let\@tempa\@firstofone}{}%
\@tempa{#2}%
}%
\def\bibinfo#1{%
\expandafter\let\expandafter\@tempa\csname bibinfo@X@#1\endcsname
\@ifx{\@tempa\relax}{\@firstofone}{\@tempa}%
}%
\def\@bib@Xauthor#1{\let\@bib@Xjournal\@gobble}%
\def\@bib@Xjournal#1{\begingroup\let\bibinfo@X@journal\@bib@Z@journal#1\endgroup}%
\def\@bibibid@#1{\textit{ibid}.}%
\appdef\NAT@bibitem@init{%
\let\@bibauthor \@empty
\let\@bibjournal \@empty
\let\@bib@Z@journal\@bibibid@
}%
\ifx\SK@lbibitem\@undefined\else
\let\SK@lbibitem\@lbibitem
\def\@lbibitem[#1]#2{%
\SK@lbibitem[#1]{#2}\SK@\SK@@label{#2}\ignorespaces}\fi
\newif\ifNAT@stdbst \NAT@stdbstfalse
\AtEndDocument{%
\ifNAT@stdbst\if@filesw
\immediate\write\@auxout{%
\string\providecommand\string\NAT@force@numbers{}%
\string\NAT@force@numbers
}%
\fi\fi
}
\newcommand\NAT@force@numbers{%
\ifNAT@numbers\else
\PackageError{natbib}{Bibliography not compatible with author-year
citations.\MessageBreak
Press <return> to continue in numerical citation style}
{Check the bibliography entries for non-compliant syntax,\MessageBreak
or select author-year BibTeX style, e.g. plainnat}%
\global\NAT@numberstrue\fi}
\providecommand\bibcite{}
\renewcommand\bibcite[2]{%
\@ifundefined{b@#1\@extra@binfo}{\relax}{%
\NAT@citemultiple
\PackageWarningNoLine{natbib}{Citation `#1' multiply defined}%
}%
\global\@namedef{b@#1\@extra@binfo}{#2}%
}%
\AtEndDocument{\NAT@swatrue\let\bibcite\NAT@testdef}
\newcommand\NAT@testdef[2]{%
\def\NAT@temp{#2}%
\expandafter \ifx \csname b@#1\@extra@binfo\endcsname\NAT@temp
\else
\ifNAT@swa \NAT@swafalse
\PackageWarningNoLine{natbib}{%
Citation(s) may have changed.\MessageBreak
Rerun to get citations correct%
}%
\fi
\fi
}%
\newcommand\NAT@apalk{}
\def\NAT@apalk#1, #2, #3\@nil#4{%
\if\relax#2\relax
\global\NAT@stdbsttrue
\NAT@wrout{#1}{}{}{}{#4}%
\else
\NAT@wrout{\the\c@NAT@ctr}{#2}{#1}{}{#4}%
\fi
}%
\newcommand\citeauthoryear{}
\def\citeauthoryear#1#2#3(@)(@)\@nil#4{%
\if\relax#3\relax
\NAT@wrout{\the\c@NAT@ctr}{#2}{#1}{}{#4}%
\else
\NAT@wrout{\the\c@NAT@ctr}{#3}{#2}{#1}{#4}%
\fi
}%
\newcommand\citestarts{\NAT@open}%
\newcommand\citeends{\NAT@close}%
\newcommand\betweenauthors{and}%
\newcommand\astroncite{}
\def\astroncite#1#2(@)(@)\@nil#3{%
\NAT@wrout{\the\c@NAT@ctr}{#2}{#1}{}{#3}%
}%
\newcommand\citename{}
\def\citename#1#2(@)(@)\@nil#3{\expandafter\NAT@apalk#1#2, \@nil{#3}}
\newcommand\harvarditem[4][]{%
\if\relax#1\relax
\bibitem[#2(#3)]{#4}%
\else
\bibitem[#1(#3)#2]{#4}%
\fi
}%
\newcommand\harvardleft{\NAT@open}
\newcommand\harvardright{\NAT@close}
\newcommand\harvardyearleft{\NAT@open}
\newcommand\harvardyearright{\NAT@close}
\AtBeginDocument{\providecommand{\harvardand}{and}}
\newcommand\harvardurl[1]{\textbf{URL:} \textit{#1}}
\providecommand\bibsection{}
\@ifundefined{chapter}{%
\renewcommand\bibsection{%
\section*{\refname\@mkboth{\MakeUppercase{\refname}}{\MakeUppercase{\refname}}}%
}%
}{%
\@ifxundefined\NAT@sectionbib{%
\renewcommand\bibsection{%
\chapter*{\bibname\@mkboth{\MakeUppercase{\bibname}}{\MakeUppercase{\bibname}}}%
}%
}{%
\renewcommand\bibsection{%
\section*{\bibname\ifx\@mkboth\@gobbletwo\else\markright{\MakeUppercase{\bibname}}\fi}%
}%
}%
}%
\@ifclassloaded{amsart}{\renewcommand\bibsection{\section*{\refname}}}{}%
\@ifclassloaded{amsbook}{\renewcommand\bibsection{\chapter*{\bibname}}}{}%
\@ifxundefined\bib@heading{}{\let\bibsection\bib@heading}%
\newcounter{NAT@ctr}
\renewenvironment{thebibliography}[1]{%
\bibsection
\parindent\z@
\bibpreamble
\bibfont
\list{\@biblabel{\the\c@NAT@ctr}}{\@bibsetup{#1}\global\c@NAT@ctr\z@}%
\ifNAT@openbib
\renewcommand\newblock{\par}%
\else
\renewcommand\newblock{\hskip .11em \@plus.33em \@minus.07em}%
\fi
\sloppy\clubpenalty4000\widowpenalty4000
\sfcode`\.\@m
\let\NAT@bibitem@first@sw\@firstoftwo
\let\citeN\cite \let\shortcite\cite
\let\citeasnoun\cite
}{%
\bibitem@fin
\bibpostamble
\def\@noitemerr{%
\PackageWarning{natbib}{Empty `thebibliography' environment}%
}%
\endlist
\bibcleanup
}%
\let\bibfont\@empty
\let\bibpreamble\@empty
\let\bibpostamble\@empty
\def\bibcleanup{\vskip-\lastskip}%
\providecommand\reset@font{\relax}
\providecommand\bibname{Bibliography}
\providecommand\refname{References}
\newcommand\NAT@citeundefined{\gdef \NAT@undefined {%
\PackageWarningNoLine{natbib}{There were undefined citations}}}
\let \NAT@undefined \relax
\newcommand\NAT@citemultiple{\gdef \NAT@multiple {%
\PackageWarningNoLine{natbib}{There were multiply defined citations}}}
\let \NAT@multiple \relax
\AtEndDocument{\NAT@undefined\NAT@multiple}
\providecommand\@mkboth[2]{}
\providecommand\MakeUppercase{\uppercase}
\providecommand{\@extra@b@citeb}{}
\gdef\@extra@binfo{}
\def\NAT@anchor#1#2{%
\hyper@natanchorstart{#1\@extra@b@citeb}%
\def\@tempa{#2}\@ifx{\@tempa\@empty}{}{\@biblabel{#2}}%
\hyper@natanchorend
}%
\providecommand\hyper@natanchorstart[1]{}%
\providecommand\hyper@natanchorend{}%
\providecommand\hyper@natlinkstart[1]{}%
\providecommand\hyper@natlinkend{}%
\providecommand\hyper@natlinkbreak[2]{#1}%
\AtBeginDocument{%
\@ifpackageloaded{babel}{%
\let\org@@citex\@citex}{}}
\providecommand\@safe@activestrue{}%
\providecommand\@safe@activesfalse{}%
\newcommand\NAT@sort@cites[1]{%
\let\NAT@cite@list\@empty
\@for\@citeb:=#1\do{\expandafter\NAT@star@cite\@citeb\@@}%
\if@filesw
\expandafter\immediate\expandafter\write\expandafter\@auxout
\expandafter{\expandafter\string\expandafter\citation\expandafter{\NAT@cite@list}}%
\fi
\@ifnum{\NAT@sort>\z@}{%
\expandafter\NAT@sort@cites@\expandafter{\NAT@cite@list}%
}{}%
}%
\def\NAT@star@cite{%
\let\NAT@star@sw\@secondoftwo
\@ifnum{\NAT@merge>\z@}{%
\@ifnextchar*{%
\let\NAT@star@sw\@firstoftwo
\NAT@star@cite@star
}{%
\NAT@star@cite@nostar
}%
}{%
\NAT@star@cite@noextension
}%
}%
\def\NAT@star@cite@star*{%
\NAT@star@cite@nostar
}%
\def\NAT@star@cite@nostar{%
\let\nat@keyopt@open\@empty
\let\nat@keyopt@shut\@empty
\@ifnextchar[{\NAT@star@cite@pre}{\NAT@star@cite@pre[]}%
}%
\def\NAT@star@cite@pre[#1]{%
\def\nat@keyopt@open{#1}%
\@ifnextchar[{\NAT@star@cite@post}{\NAT@star@cite@post[]}%
}%
\def\NAT@star@cite@post[#1]#2\@@{%
\def\nat@keyopt@shut{#1}%
\NAT@star@sw{\expandafter\global\expandafter\let\csname NAT@b*@#2\endcsname\@empty}{}%
\NAT@cite@list@append{#2}%
}%
\def\NAT@star@cite@noextension#1\@@{%
\let\nat@keyopt@open\@empty
\let\nat@keyopt@shut\@empty
\NAT@cite@list@append{#1}%
}%
\def\NAT@cite@list@append#1{%
\edef\@citeb{\@firstofone#1\@empty}%
\if@filesw\@ifxundefined\@cprwrite{}{\expandafter\@cprwrite\@citeb=}\fi
\if\relax\nat@keyopt@open\relax\else
\global\expandafter\let\csname NAT@b@open@\@citeb\endcsname\nat@keyopt@open
\fi
\if\relax\nat@keyopt@shut\relax\else
\global\expandafter\let\csname NAT@b@shut@\@citeb\endcsname\nat@keyopt@shut
\fi
\toks@\expandafter{\NAT@cite@list}%
\ifx\NAT@cite@list\@empty
\@temptokena\expandafter{\@citeb}%
\else
\@temptokena\expandafter{\expandafter,\@citeb}%
\fi
\edef\NAT@cite@list{\the\toks@\the\@temptokena}%
}%
\newcommand\NAT@sort@cites@[1]{%
\count@\z@
\@tempcntb\m@ne
\let\@celt\delimiter
\def\NAT@num@list{}%
\let\NAT@cite@list\@empty
\let\NAT@nonsort@list\@empty
\@for \@citeb:=#1\do{\NAT@make@cite@list}%
\ifx\NAT@nonsort@list\@empty\else
\protected@edef\NAT@cite@list{\NAT@cite@list\NAT@nonsort@list}%
\fi
\ifx\NAT@cite@list\@empty\else
\protected@edef\NAT@cite@list{\expandafter\NAT@xcom\NAT@cite@list @@}%
\fi
}%
\def\NAT@make@cite@list{%
\advance\count@\@ne
\@safe@activestrue
\edef\@citeb{\expandafter\@firstofone\@citeb\@empty}%
\@safe@activesfalse
\@ifundefined{b@\@citeb\@extra@b@citeb}%
{\def\NAT@num{A}}%
{\NAT@parse{\@citeb}}%
\NAT@ifcat@num\NAT@num
{\@tempcnta\NAT@num \relax
\@ifnum{\@tempcnta<\@tempcntb}{%
\let\NAT@@cite@list=\NAT@cite@list
\let\NAT@cite@list\@empty
\begingroup\let\@celt=\NAT@celt\NAT@num@list\endgroup
\protected@edef\NAT@num@list{%
\expandafter\NAT@num@celt \NAT@num@list \@gobble @%
}%
}{%
\protected@edef\NAT@num@list{\NAT@num@list \@celt{\NAT@num}}%
\protected@edef\NAT@cite@list{\NAT@cite@list\@citeb,}%
\@tempcntb\@tempcnta
}%
}%
{\protected@edef\NAT@nonsort@list{\NAT@nonsort@list\@citeb,}}%
}%
\def\NAT@celt#1{%
\@ifnum{#1>\@tempcnta}{%
\xdef\NAT@cite@list{\NAT@cite@list\@citeb,\NAT@@cite@list}%
\let\@celt\@gobble
}{%
\expandafter\def@NAT@cite@lists\NAT@@cite@list\@@
}%
}%
\def\NAT@num@celt#1#2{%
\ifx#1\@celt
\@ifnum{#2>\@tempcnta}{%
\@celt{\number\@tempcnta}%
\@celt{#2}%
}{%
\@celt{#2}%
\expandafter\NAT@num@celt
}%
\fi
}%
\def\def@NAT@cite@lists#1,#2\@@{%
\xdef\NAT@cite@list{\NAT@cite@list#1,}%
\xdef\NAT@@cite@list{#2}%
}%
\def\NAT@nextc#1,#2@@{#1,}
\def\NAT@restc#1,#2{#2}
\def\NAT@xcom#1,@@{#1}
\InputIfFileExists{natbib.cfg}
{\typeout{Local config file natbib.cfg used}}{}
%%
%% <<<<< End of generated file <<<<<<
%%
%% End of file `natbib.sty'.
\documentclass{article} % For LaTeX2e
\usepackage[final]{rlhf-introduction}
\usepackage{microtype}
\usepackage{hyperref}
\usepackage{url}
\usepackage{booktabs}
\usepackage{lineno}
\usepackage{CJK}
\definecolor{darkblue}{rgb}{0, 0, 0.5}
\hypersetup{colorlinks=true, citecolor=darkblue, linkcolor=darkblue, urlcolor=darkblue}
\title{Reinforcement Learning without Tears: An Introduction for Large Language Model Researchers}
% Authors must not appear in the submitted version. They should be hidden
% as long as the \colmfinalcopy macro remains commented out below.
% Non-anonymous submissions will be rejected without review.
\author{Antiquus S.~Hippocampus, Natalia Cerebro \& Amelie P. Amygdale \thanks{ Use footnote for providing further information
about author (webpage, alternative address)---\emph{not} for acknowledging
funding agencies. Funding acknowledgements go at the end of the paper.} \\
Department of Computer Science\\
Cranberry-Lemon University\\
Pittsburgh, PA 15213, USA \\
\texttt{\{hippo,brain,jen\}@cs.cranberry-lemon.edu} \\
\And
Ji Q. Ren \& Yevgeny LeNet \\
Department of Computational Neuroscience \\
University of the Witwatersrand \\
Joburg, South Africa \\
\texttt{\{robot,net\}@wits.ac.za} \\
\AND
Coauthor \\
Affiliation \\
Address \\
\texttt{email}
}
% The \author macro works with any number of authors. There are two commands
% used to separate the names and addresses of multiple authors: \And and \AND.
%
% Using \And between authors leaves it to \LaTeX{} to determine where to break
% the lines. Using \AND forces a linebreak at that point. So, if \LaTeX{}
% puts 3 of 4 authors names on the first line, and the last on the second
% line, try using \AND instead of \And before the third author name.
\newcommand{\fix}{\marginpar{FIX}}
\newcommand{\new}{\marginpar{NEW}}
\begin{document}
\ifcolmsubmission
\linenumbers
\fi
\maketitle
\begin{abstract}
abstract.
\end{abstract}
\section{Introduction}
Introduction.
\section{Basics of Reinforcement Learning}
\begin{CJK}{UTF8}{gbsn}
这部分主要讲述强化学习技术和一些符号的定义。\\
(1)介绍RL总体的架构和一些重要的元素(奖励,状态,环境) \\
(2)算法包括REINFORCE,A2C.另外提一下它们的应用场景和优缺点,这样可以为后面“训练LLM”部分做更好的衔接。
\end{CJK}
\section{Training LLMs with Reinforcement Learning}
\begin{CJK}{UTF8}{gbsn}
这部分主要讲述强化学习在大语言模型训练过程中是如何进行建模的。\\
(1)这里首先需要简单说一下大语言模型部分,pre-training,instruction tuning. \\
(2)如何进行优化LLM? 对应到大模型的时候,策略是什么?状态对应的什么?如何进行采样?如何进行参数更新的? \\
(3)奖励如何设计的?(训练奖励模型 or 一些奖励规则--训练math推理的时候)\\
(4)另外结束的时候需要加个notation的表格。讲述各种符号分别对应的关系(在强化学习-》在大模型训练的过程)。
\end{CJK}
\subsection{Challenges}
\begin{CJK}{UTF8}{gbsn}
这里主要进行对训练LLM的强化学习训练的关键组件进行拆解,进行介绍当下的研究点:数据,效率,稳定性,以及可解释性(这部分考虑融合到上面一个章节,作为一个子章节)。
\end{CJK}
\section{Improved Reinforcement Learning for LLMs}
\begin{CJK}{UTF8}{gbsn}
对应上面的挑战,来去说现在的一些主流的方法提升强化学习在LLM中的工作。\\
(1)偏好数据、偏好泛化 \\
(2)更好的奖励建模 \\
(3)简化优化框架(dynamic RL,GRPO,remax等工作)\\
(3)直接偏好优化
% \include{section4-DPO}
\end{CJK}
\section{Reinforcement Learning for LLM Inference}
\begin{CJK}{UTF8}{gbsn}
除了直接训练LLMs之后,强化学习还从什么方面来去优化了LLMs? \\
(1)inference-time alignment \\
(2)test-time ranking, o1, R1
\end{CJK}
\section{Reinforcement Learning for Multi-Modal Language Models}
\begin{CJK}{UTF8}{gbsn}
单独的使用一个章节来去说多模态中的RL。不过这个章节不用像之前文本那样描述的比较细致,更多是一些扩展就好。\\
(1)视觉(图生文)\\
(2)使用RL训练diffusion model \\
(3)音频和视频。 \\
\end{CJK}
\section{Conclusion}
\begin{CJK}{UTF8}{gbsn}
总结和一些方向,,是否还要放一些分析(加一些实验)?
\end{CJK}
\section{Some Tools, Datasets, or Systems}
\begin{CJK}{UTF8}{gbsn}
(1)一些RL训练系统 \\
(2)一些教程 \\
(3)可视化分析工具(我记得有一个) \\
\end{CJK}
\bibliography{rlhf-introduction}
\bibliographystyle{rlhf-introduction}
\appendix
\section{Appendix}
This is an appendix.
\end{document}
@article{sun-etal:2023aligning,
title={Aligning large multimodal models with factually augmented rlhf},
author={Sun, Zhiqing and Shen, Sheng and Cao, Shengcao and Liu, Haotian and Li, Chunyuan and Shen, Yikang and Gan, Chuang and Gui, Liang-Yan and Wang, Yu-Xiong and Yang, Yiming and others},
journal={arXiv preprint arXiv:2309.14525},
year={2023}
}
@article{ji2025safe,
title={Safe RLHF-V: Safe Reinforcement Learning from Human Feedback in Multimodal Large Language Models},
author={Ji, Jiaming and Chen, Xinyu and Pan, Rui and Zhu, Han and Zhang, Conghui and Li, Jiahao and Hong, Donghai and Chen, Boyuan and Zhou, Jiayi and Wang, Kaile and others},
journal={arXiv preprint arXiv:2503.17682},
year={2025}
}
@article{zang2025internlm,
title={InternLM-XComposer2. 5-Reward: A Simple Yet Effective Multi-Modal Reward Model},
author={Zang, Yuhang and Dong, Xiaoyi and Zhang, Pan and Cao, Yuhang and Liu, Ziyu and Ding, Shengyuan and Wu, Shenxi and Ma, Yubo and Duan, Haodong and Zhang, Wenwei and others},
journal={arXiv preprint arXiv:2501.12368},
year={2025}
}
@article{cheng-etal:2023everyone,
title={Everyone deserves a reward: Learning customized human preferences},
author={Cheng, Pengyu and Xie, Jiawen and Bai, Ke and Dai, Yong and Du, Nan},
journal={arXiv preprint arXiv:2309.03126},
year={2023}
}
@article{wu-etal:2024reuse,
title={Reuse your rewards: Reward model transfer for zero-shot cross-lingual alignment},
author={Wu, Zhaofeng and Balashankar, Ananth and Kim, Yoon and Eisenstein, Jacob and Beirami, Ahmad},
journal={arXiv preprint arXiv:2404.12318},
year={2024}
}
@article{xia-etal:2024less,
title={Less: Selecting influential data for targeted instruction tuning},
author={Xia, Mengzhou and Malladi, Sadhika and Gururangan, Suchin and Arora, Sanjeev and Chen, Danqi},
journal={arXiv preprint arXiv:2402.04333},
year={2024}
}
@article{wang-etal:2024rovrm,
title={RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data},
author={Wang, Chenglong and Gan, Yang and Huo, Yifu and Mu, Yongyu and Yang, Murun and He, Qiaozhi and Xiao, Tong and Zhang, Chunliang and Liu, Tongran and Du, Quan and others},
journal={arXiv preprint arXiv:2408.12109},
year={2024}
}
@article{yu-etal:2024rlaif,
title={Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness},
author={Yu, Tianyu and Zhang, Haoye and Yao, Yuan and Dang, Yunkai and Chen, Da and Lu, Xiaoman and Cui, Ganqu and He, Taiwen and Liu, Zhiyuan and Chua, Tat-Seng and others},
journal={arXiv preprint arXiv:2405.17220},
year={2024}
}
@article{zhang-etal:2024mm,
title={Mm-llms: Recent advances in multimodal large language models},
author={Zhang, Duzhen and Yu, Yahan and Dong, Jiahua and Li, Chenxing and Su, Dan and Chu, Chenhui and Yu, Dong},
journal={arXiv preprint arXiv:2401.13601},
year={2024}
}
@inproceedings{yu-etal:2024rlhf,
title={Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback},
author={Yu, Tianyu and Yao, Yuan and Zhang, Haoye and He, Taiwen and Han, Yifeng and Cui, Ganqu and Hu, Jinyi and Liu, Zhiyuan and Zheng, Hai-Tao and Sun, Maosong and others},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={13807--13816},
year={2024}
}
@article{grattafiori-etal:2024llama,
title={The llama 3 herd of models},
author={Grattafiori, Aaron and Dubey, Abhimanyu and Jauhri, Abhinav and Pandey, Abhinav and Kadian, Abhishek and Al-Dahle, Ahmad and Letman, Aiesha and Mathur, Akhil and Schelten, Alan and Vaughan, Alex and others},
journal={arXiv preprint arXiv:2407.21783},
year={2024}
}
@misc{qwenTeam:2025qwen2.5-VL,
title = {Qwen2.5-VL},
url = {https://qwenlm.github.io/blog/qwen2.5-vl/},
author = {Qwen Team},
month = {January},
year = {2025}
}
@article{xu-etal:2025qwen2,
title={Qwen2. 5-Omni Technical Report},
author={Xu, Jin and Guo, Zhifang and He, Jinzheng and Hu, Hangrui and He, Ting and Bai, Shuai and Chen, Keqin and Wang, Jialin and Fan, Yang and Dang, Kai and others},
journal={arXiv preprint arXiv:2503.20215},
year={2025}
}
@misc{radford-etal:2021learning,
title={Learning Transferable Visual Models From Natural Language Supervision},
author={Alec Radford and Jong Wook Kim and Chris Hallacy and Aditya Ramesh and Gabriel Goh and Sandhini Agarwal and Girish Sastry and Amanda Askell and Pamela Mishkin and Jack Clark and Gretchen Krueger and Ilya Sutskever},
year={2021},
eprint={2103.00020},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2103.00020},
}
@article{baltruvsaitis-etal:2018multimodal,
title={Multimodal machine learning: A survey and taxonomy},
author={Baltru{\v{s}}aitis, Tadas and Ahuja, Chaitanya and Morency, Louis-Philippe},
journal={IEEE transactions on pattern analysis and machine intelligence},
volume={41},
number={2},
pages={423--443},
year={2018},
publisher={IEEE}
}
@incollection{deb-etal:2016multi,
title={Multi-objective optimization},
author={Deb, Kalyanmoy and Sindhya, Karthik and Hakanen, Jussi},
booktitle={Decision sciences},
pages={161--200},
year={2016},
publisher={CRC Press}
}
@incollection{miettinen-etal:2008introduction,
title={Introduction to multiobjective optimization: interactive approaches},
author={Miettinen, Kaisa and Ruiz, Francisco and Wierzbicki, Andrzej P},
booktitle={Multiobjective optimization: interactive and evolutionary approaches},
pages={27--57},
year={2008},
publisher={Springer}
}
@article{gorbatovski-etal:2024learn,
author = {Gorbatovski, Alexey and Shaposhnikov, Boris and Malakhov, Alexey and Surnachev, Nikita and Aksenov, Yaroslav and Maksimov, Ian and Balagansky, Nikita and Gavrilov, Daniil},
journal = {ArXiv preprint},
title = {Learn your reference model for real good alignment},
url = {https://arxiv.org/abs/2404.09656},
volume = {abs/2404.09656},
year = {2024}
}
@inproceedings{yuan-etal:2024selfrewarding,
author = {Weizhe Yuan and
Richard Yuanzhe Pang and
Kyunghyun Cho and
Xian Li and
Sainbayar Sukhbaatar and
Jing Xu and
Jason Weston},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/icml/YuanPCLSXW24.bib},
booktitle = {Forty-first International Conference on Machine Learning, {ICML} 2024,
Vienna, Austria, July 21-27, 2024},
publisher = {OpenReview.net},
timestamp = {Mon, 02 Sep 2024 01:00:00 +0200},
title = {Self-Rewarding Language Models},
url = {https://openreview.net/forum?id=0NphYCmgua},
year = {2024}
}
@article{zeng-etal:2025revisiting,
author = {Zeng, Zhiyuan and Cheng, Qinyuan and Yin, Zhangyue and Zhou, Yunhua and Qiu, Xipeng},
journal = {ArXiv preprint},
title = {Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?},
url = {https://arxiv.org/abs/2502.12215},
volume = {abs/2502.12215},
year = {2025}
}
@article{yuan-etal:2023scaling,
author = {Yuan, Zheng and Yuan, Hongyi and Li, Chengpeng and Dong, Guanting and Lu, Keming and Tan, Chuanqi and Zhou, Chang and Zhou, Jingren},
journal = {ArXiv preprint},
title = {Scaling relationship on learning mathematical reasoning with large language models},
url = {https://arxiv.org/abs/2308.01825},
volume = {abs/2308.01825},
year = {2023}
}
@inproceedings{yao-etal:2023tree,
author = {Shunyu Yao and
Dian Yu and
Jeffrey Zhao and
Izhak Shafran and
Tom Griffiths and
Yuan Cao and
Karthik Narasimhan},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/YaoYZS00N23.bib},
booktitle = {Advances in Neural Information Processing Systems 36: Annual Conference
on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans,
LA, USA, December 10 - 16, 2023},
editor = {Alice Oh and
Tristan Naumann and
Amir Globerson and
Kate Saenko and
Moritz Hardt and
Sergey Levine},
timestamp = {Fri, 01 Mar 2024 00:00:00 +0100},
title = {Tree of Thoughts: Deliberate Problem Solving with Large Language Models},
url = {http://papers.nips.cc/paper\_files/paper/2023/hash/271db9922b8d1f4dd7aaef84ed5ac703-Abstract-Conference.html},
year = {2023}
}
@article{xie-etal:2025logic,
author = {Xie, Tian and Gao, Zitian and Ren, Qingnan and Luo, Haoming and Hong, Yuqian and Dai, Bryan and Zhou, Joey and Qiu, Kai and Wu, Zhirong and Luo, Chong},
journal = {ArXiv preprint},
title = {Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning},
url = {https://arxiv.org/abs/2502.14768},
volume = {abs/2502.14768},
year = {2025}
}
@article{wang-etal:2023math,
author = {Wang, Peiyi and Li, Lei and Shao, Zhihong and Xu, RX and Dai, Damai and Li, Yifei and Chen, Deli and Wu, Yu and Sui, Zhifang},
journal = {ArXiv preprint},
title = {Math-shepherd: Verify and reinforce llms step-by-step without human annotations},
url = {https://arxiv.org/abs/2312.08935},
volume = {abs/2312.08935},
year = {2023}
}
@inproceedings{lightman-etal:2024lets,
author = {Hunter Lightman and
Vineet Kosaraju and
Yuri Burda and
Harrison Edwards and
Bowen Baker and
Teddy Lee and
Jan Leike and
John Schulman and
Ilya Sutskever and
Karl Cobbe},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/iclr/LightmanKBEBLLS24.bib},
booktitle = {The Twelfth International Conference on Learning Representations,
{ICLR} 2024, Vienna, Austria, May 7-11, 2024},
publisher = {OpenReview.net},
timestamp = {Wed, 07 Aug 2024 01:00:00 +0200},
title = {Let's Verify Step by Step},
url = {https://openreview.net/forum?id=v8L0pN6EOi},
year = {2024}
}
@article{setlur-etal:2024rewarding,
author = {Setlur, Amrith and Nagpal, Chirag and Fisch, Adam and Geng, Xinyang and Eisenstein, Jacob and Agarwal, Rishabh and Agarwal, Alekh and Berant, Jonathan and Kumar, Aviral},
journal = {ArXiv preprint},
title = {Rewarding progress: Scaling automated process verifiers for llm reasoning},
url = {https://arxiv.org/abs/2410.08146},
volume = {abs/2410.08146},
year = {2024}
}
@article{touvron-etal:2023llama2,
author = {Hugo Touvron and Louis Martin and Kevin Stone and Peter Albert and Amjad Almahairi and Yasmine Babaei and Nikolay Bashlykov and Soumya Batra and Prajjwal Bhargava and Shruti Bhosale and Dan Bikel and Lukas Blecher and Cristian Canton Ferrer and Moya Chen and Guillem Cucurull and David Esiobu and Jude Fernandes and Jeremy Fu and Wenyin Fu and Brian Fuller and Cynthia Gao and Vedanuj Goswami and Naman Goyal and Anthony Hartshorn and Saghar Hosseini and Rui Hou and Hakan Inan and Marcin Kardas and Viktor Kerkez and Madian Khabsa and Isabel Kloumann and Artem Korenev and Punit Singh Koura and Marie-Anne Lachaux and Thibaut Lavril and Jenya Lee and Diana Liskovich and Yinghai Lu and Yuning Mao and Xavier Martinet and Todor Mihaylov and Pushkar Mishra and Igor Molybog and Yixin Nie and Andrew Poulton and Jeremy Reizenstein and Rashi Rungta and Kalyan Saladi and Alan Schelten and Ruan Silva and Eric Michael Smith and Ranjan Subramanian and Xiaoqing Ellen Tan and Binh Tang and Ross Taylor and Adina Williams and Jian Xiang Kuan and Puxin Xu and Zheng Yan and Iliyan Zarov and Yuchen Zhang and Angela Fan and Melanie Kambadur and Sharan Narang and Aurelien Rodriguez and Robert Stojnic and Sergey Edunov and Thomas Scialom},
journal = {ArXiv preprint},
title = {Llama 2: Open foundation and fine-tuned chat models},
url = {https://arxiv.org/abs/2307.09288},
volume = {abs/2307.09288},
year = {2023}
}
@article{nakano-etal:2021webgpt,
author = {Reiichiro Nakano and Jacob Hilton and Suchir Balaji and Jeff Wu and Long Ouyang and Christina Kim and Christopher Hesse and Shantanu Jain and Vineet Kosaraju and William Saunders and Xu Jiang and Karl Cobbe and Tyna Eloundou and Gretchen Krueger and Kevin Button and Matthew Knight and Benjamin Chess and John Schulman},
journal = {ArXiv preprint},
title = {Webgpt: Browser-assisted question-answering with human feedback},
url = {https://arxiv.org/abs/2112.09332},
volume = {abs/2112.09332},
year = {2021}
}
@article{xiao-and-zhu:2025foundations,
author = {Xiao, Tong and Zhu, Jingbo},
journal = {ArXiv preprint},
title = {Foundations of Large Language Models},
url = {https://arxiv.org/abs/2501.09223},
volume = {abs/2501.09223},
year = {2025}
}
@article{vinyals-rtal:2019grandmaster,
author = {Vinyals, Oriol and Babuschkin, Igor and Czarnecki, Wojciech M and Mathieu, Micha{\"e}l and Dudzik, Andrew and Chung, Junyoung and Choi, David H and Powell, Richard and Ewalds, Timo and Georgiev, Petko and others},
journal = {nature},
number = {7782},
pages = {350--354},
publisher = {Nature Publishing Group},
title = {Grandmaster level in StarCraft II using multi-agent reinforcement learning},
volume = {575},
year = {2019}
}
@article{silver-etal:2016mastering,
author = {Silver, David and Huang, Aja and Maddison, Chris J and Guez, Arthur and Sifre, Laurent and Van Den Driessche, George and Schrittwieser, Julian and Antonoglou, Ioannis and Panneershelvam, Veda and Lanctot, Marc and others},
journal = {nature},
number = {7587},
pages = {484--489},
publisher = {Nature Publishing Group},
title = {Mastering the game of Go with deep neural networks and tree search},
volume = {529},
year = {2016}
}
@article{kimi-team:2025kimi,
author = {Team, Kimi and Du, Angang and Gao, Bofei and Xing, Bowei and Jiang, Changjiu and Chen, Cheng and Li, Cheng and Xiao, Chenjun and Du, Chenzhuang and Liao, Chonghua and others},
journal = {ArXiv preprint},
title = {Kimi k1. 5: Scaling reinforcement learning with llms},
url = {https://arxiv.org/abs/2501.12599},
volume = {abs/2501.12599},
year = {2025}
}
@article{xiao-etal:2023introduction,
author = {Xiao, Tong and Zhu, Jingbo},
journal = {ArXiv preprint},
title = {Introduction to transformers: an nlp perspective},
url = {https://arxiv.org/abs/2311.17633},
volume = {abs/2311.17633},
year = {2023}
}
@inproceedings{wang-etal:2023far,
author = {Yizhong Wang and
Hamish Ivison and
Pradeep Dasigi and
Jack Hessel and
Tushar Khot and
Khyathi Chandu and
David Wadden and
Kelsey MacMillan and
Noah A. Smith and
Iz Beltagy and
Hannaneh Hajishirzi},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/WangIDHKCWMSBH23.bib},
booktitle = {Advances in Neural Information Processing Systems 36: Annual Conference
on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans,
LA, USA, December 10 - 16, 2023},
editor = {Alice Oh and
Tristan Naumann and
Amir Globerson and
Kate Saenko and
Moritz Hardt and
Sergey Levine},
timestamp = {Fri, 01 Mar 2024 00:00:00 +0100},
title = {How Far Can Camels Go? Exploring the State of Instruction Tuning on
Open Resources},
url = {http://papers.nips.cc/paper\_files/paper/2023/hash/ec6413875e4ab08d7bc4d8e225263398-Abstract-Datasets\_and\_Benchmarks.html},
year = {2023}
}
@article{morimura-etal:2024filtered,
author = {Morimura, Tetsuro and Sakamoto, Mitsuki and Jinnai, Yuu and Abe, Kenshi and Ariu, Kaito},
journal = {ArXiv preprint},
title = {Filtered direct preference optimization},
url = {https://arxiv.org/abs/2404.13846},
volume = {abs/2404.13846},
year = {2024}
}
@inproceedings{cui-etal:2023ultrafeedback,
author = {Ganqu Cui and
Lifan Yuan and
Ning Ding and
Guanming Yao and
Bingxiang He and
Wei Zhu and
Yuan Ni and
Guotong Xie and
Ruobing Xie and
Yankai Lin and
Zhiyuan Liu and
Maosong Sun},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/icml/CuiY0YH0NXXL0024.bib},
booktitle = {Forty-first International Conference on Machine Learning, {ICML} 2024,
Vienna, Austria, July 21-27, 2024},
publisher = {OpenReview.net},
timestamp = {Mon, 02 Sep 2024 01:00:00 +0200},
title = {{ULTRAFEEDBACK:} Boosting Language Models with Scaled {AI} Feedback},
url = {https://openreview.net/forum?id=BOorDpKHiJ},
year = {2024}
}
@article{wang-etal:2023learning,
author = {Wang, Chenglong and Zhou, Hang and Chang, Kaiyan and Liu, Tongran and Zhang, Chunliang and Du, Quan and Xiao, Tong and Zhu, Jingbo},
journal = {ArXiv preprint},
title = {Learning evaluation models from large language models for sequence generation},
url = {https://arxiv.org/abs/2308.04386},
volume = {abs/2308.04386},
year = {2023}
}
@inproceedings{sun-etal:2023simple,
author = {Mingjie Sun and
Zhuang Liu and
Anna Bair and
J. Zico Kolter},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/iclr/Sun0BK24.bib},
booktitle = {The Twelfth International Conference on Learning Representations,
{ICLR} 2024, Vienna, Austria, May 7-11, 2024},
publisher = {OpenReview.net},
timestamp = {Wed, 07 Aug 2024 01:00:00 +0200},
title = {A Simple and Effective Pruning Approach for Large Language Models},
url = {https://openreview.net/forum?id=PxoFut3dWW},
year = {2024}
}
@article{masoudnia-etal:2014mixture,
author = {Masoudnia, Saeed and Ebrahimpour, Reza},
journal = {Artificial Intelligence Review},
pages = {275--293},
publisher = {Springer},
title = {Mixture of experts: a literature survey},
volume = {42},
year = {2014}
}
@inproceedings{li-etal:2023remax,
author = {Ziniu Li and
Tian Xu and
Yushun Zhang and
Zhihang Lin and
Yang Yu and
Ruoyu Sun and
Zhi{-}Quan Luo},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/icml/LiXZL00L24.bib},
booktitle = {Forty-first International Conference on Machine Learning, {ICML} 2024,
Vienna, Austria, July 21-27, 2024},
publisher = {OpenReview.net},
timestamp = {Sat, 14 Dec 2024 00:00:00 +0100},
title = {ReMax: {A} Simple, Effective, and Efficient Reinforcement Learning
Method for Aligning Large Language Models},
url = {https://openreview.net/forum?id=Stn8hXkpe6},
year = {2024}
}
@article{hu:2025reinforce++,
author = {Hu, Jian},
journal = {ArXiv preprint},
title = {REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models},
url = {https://arxiv.org/abs/2501.03262},
volume = {abs/2501.03262},
year = {2025}
}
@article{wang-etal:2024hybrid,
author = {Wang, Chenglong and Zhou, Hang and Chang, Kaiyan and Li, Bei and Mu, Yongyu and Xiao, Tong and Liu, Tongran and Zhu, Jingbo},
journal = {ArXiv preprint},
title = {Hybrid Alignment Training for Large Language Models},
url = {https://arxiv.org/abs/2406.15178},
volume = {abs/2406.15178},
year = {2024}
}
@article{sutton-and-richard:1988learning,
author = {Sutton, Richard S},
journal = {Machine learning},
pages = {9--44},
publisher = {Springer},
title = {Learning to predict by the methods of temporal differences},
volume = {3},
year = {1988}
}
@inproceedings{bahdanau-etal:2016actor,
author = {Dzmitry Bahdanau and
Philemon Brakel and
Kelvin Xu and
Anirudh Goyal and
Ryan Lowe and
Joelle Pineau and
Aaron C. Courville and
Yoshua Bengio},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/iclr/BahdanauBXGLPCB17.bib},
booktitle = {5th International Conference on Learning Representations, {ICLR} 2017,
Toulon, France, April 24-26, 2017, Conference Track Proceedings},
publisher = {OpenReview.net},
timestamp = {Thu, 25 Jul 2019 01:00:00 +0200},
title = {An Actor-Critic Algorithm for Sequence Prediction},
url = {https://openreview.net/forum?id=SJDaqqveg},
year = {2017}
}
@article{kumar-etal:2024training,
author = {Kumar, Aviral and Zhuang, Vincent and Agarwal, Rishabh and Su, Yi and Co-Reyes, John D and Singh, Avi and Baumli, Kate and Iqbal, Shariq and Bishop, Colton and Roelofs, Rebecca and others},
journal = {ArXiv preprint},
title = {Training language models to self-correct via reinforcement learning},
url = {https://arxiv.org/abs/2409.12917},
volume = {abs/2409.12917},
year = {2024}
}
@inproceedings{wu-etal:2023fine,
author = {Zeqiu Wu and
Yushi Hu and
Weijia Shi and
Nouha Dziri and
Alane Suhr and
Prithviraj Ammanabrolu and
Noah A. Smith and
Mari Ostendorf and
Hannaneh Hajishirzi},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/WuHSDSASOH23.bib},
booktitle = {Advances in Neural Information Processing Systems 36: Annual Conference
on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans,
LA, USA, December 10 - 16, 2023},
editor = {Alice Oh and
Tristan Naumann and
Amir Globerson and
Kate Saenko and
Moritz Hardt and
Sergey Levine},
timestamp = {Fri, 01 Mar 2024 00:00:00 +0100},
title = {Fine-Grained Human Feedback Gives Better Rewards for Language Model
Training},
url = {http://papers.nips.cc/paper\_files/paper/2023/hash/b8c90b65739ae8417e61eadb521f63d5-Abstract-Conference.html},
year = {2023}
}
@inproceedings{lightman-etal:2023let,
author = {Hunter Lightman and
Vineet Kosaraju and
Yuri Burda and
Harrison Edwards and
Bowen Baker and
Teddy Lee and
Jan Leike and
John Schulman and
Ilya Sutskever and
Karl Cobbe},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/iclr/LightmanKBEBLLS24.bib},
booktitle = {The Twelfth International Conference on Learning Representations,
{ICLR} 2024, Vienna, Austria, May 7-11, 2024},
publisher = {OpenReview.net},
timestamp = {Wed, 07 Aug 2024 01:00:00 +0200},
title = {Let's Verify Step by Step},
url = {https://openreview.net/forum?id=v8L0pN6EOi},
year = {2024}
}
@article{bromley-etal:1993signature,
author = {Bromley, Jane and Guyon, Isabelle and LeCun, Yann and S{\"a}ckinger, Eduard and Shah, Roopak},
journal = {Advances in neural information processing systems},
title = {Signature verification using a" siamese" time delay neural network},
volume = {6},
year = {1993}
}
@article{singhal-etal:2023long,
author = {Singhal, Prasann and Goyal, Tanya and Xu, Jiacheng and Durrett, Greg},
journal = {ArXiv preprint},
title = {A long way to go: Investigating length correlations in rlhf},
url = {https://arxiv.org/abs/2310.03716},
volume = {abs/2310.03716},
year = {2023}
}
@article{havrilla-etal:2024teaching,
author = {Havrilla, Alex and Du, Yuqing and Raparthy, Sharath Chandra and Nalmpantis, Christoforos and Dwivedi-Yu, Jane and Zhuravinskyi, Maksym and Hambro, Eric and Sukhbaatar, Sainbayar and Raileanu, Roberta},
journal = {ArXiv preprint},
title = {Teaching large language models to reason with reinforcement learning},
url = {https://arxiv.org/abs/2403.04642},
volume = {abs/2403.04642},
year = {2024}
}
@inproceedings{wang-etal:2021selective,
address = {Online},
author = {Wang, Fusheng and
Yan, Jianhao and
Meng, Fandong and
Zhou, Jie},
booktitle = {Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)},
doi = {10.18653/v1/2021.acl-long.504},
editor = {Zong, Chengqing and
Xia, Fei and
Li, Wenjie and
Navigli, Roberto},
pages = {6456--6466},
publisher = {Association for Computational Linguistics},
title = {Selective Knowledge Distillation for Neural Machine Translation},
url = {https://aclanthology.org/2021.acl-long.504},
year = {2021}
}
@inproceedings{chen-etal:2023alpagasus,
author = {Lichang Chen and
Shiyang Li and
Jun Yan and
Hai Wang and
Kalpa Gunaratna and
Vikas Yadav and
Zheng Tang and
Vijay Srinivasan and
Tianyi Zhou and
Heng Huang and
Hongxia Jin},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/iclr/ChenLYWGYTS0HJ24.bib},
booktitle = {The Twelfth International Conference on Learning Representations,
{ICLR} 2024, Vienna, Austria, May 7-11, 2024},
publisher = {OpenReview.net},
timestamp = {Wed, 07 Aug 2024 01:00:00 +0200},
title = {AlpaGasus: Training a Better Alpaca with Fewer Data},
url = {https://openreview.net/forum?id=FdVXgSJhvz},
year = {2024}
}
@article{li-etal:2025limr,
author = {Li, Xuefeng and Zou, Haoyang and Liu, Pengfei},
journal = {ArXiv preprint},
title = {LIMR: Less is More for RL Scaling},
url = {https://arxiv.org/abs/2502.11886},
volume = {abs/2502.11886},
year = {2025}
}
@inproceedings{wang-and-zhou:2024chain,
author = {Xuezhi Wang and
Denny Zhou},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/0002Z24.bib},
booktitle = {Advances in Neural Information Processing Systems 38: Annual Conference
on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver,
BC, Canada, December 10 - 15, 2024},
editor = {Amir Globersons and
Lester Mackey and
Danielle Belgrave and
Angela Fan and
Ulrich Paquet and
Jakub M. Tomczak and
Cheng Zhang},
timestamp = {Thu, 13 Feb 2025 00:00:00 +0100},
title = {Chain-of-Thought Reasoning Without Prompting},
url = {http://papers.nips.cc/paper\_files/paper/2024/hash/7a8e7fd295aa04eac4b470ae27f8785c-Abstract-Conference.html},
year = {2024}
}
@article{chen-etal:2023accelerating,
author = {Chen, Charlie and Borgeaud, Sebastian and Irving, Geoffrey and Lespiau, Jean-Baptiste and Sifre, Laurent and Jumper, John},
journal = {ArXiv preprint},
title = {Accelerating large language model decoding with speculative sampling},
url = {https://arxiv.org/abs/2302.01318},
volume = {abs/2302.01318},
year = {2023}
}
@article{zhao-etal:2024atom,
author = {Zhao, Yilong and Lin, Chien-Yu and Zhu, Kan and Ye, Zihao and Chen, Lequn and Zheng, Size and Ceze, Luis and Krishnamurthy, Arvind and Chen, Tianqi and Kasikci, Baris},
journal = {Proceedings of Machine Learning and Systems},
pages = {196--209},
title = {Atom: Low-bit quantization for efficient and accurate llm serving},
volume = {6},
year = {2024}
}
@article{pope-etal:2023efficiently,
author = {Pope, Reiner and Douglas, Sholto and Chowdhery, Aakanksha and Devlin, Jacob and Bradbury, James and Heek, Jonathan and Xiao, Kefan and Agrawal, Shivani and Dean, Jeff},
journal = {Proceedings of Machine Learning and Systems},
pages = {606--624},
title = {Efficiently scaling transformer inference},
volume = {5},
year = {2023}
}
@inproceedings{vaswani-etal:2017attention,
author = {Ashish Vaswani and
Noam Shazeer and
Niki Parmar and
Jakob Uszkoreit and
Llion Jones and
Aidan N. Gomez and
Lukasz Kaiser and
Illia Polosukhin},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/VaswaniSPUJGKP17.bib},
booktitle = {Advances in Neural Information Processing Systems 30: Annual Conference
on Neural Information Processing Systems 2017, December 4-9, 2017,
Long Beach, CA, {USA}},
editor = {Isabelle Guyon and
Ulrike von Luxburg and
Samy Bengio and
Hanna M. Wallach and
Rob Fergus and
S. V. N. Vishwanathan and
Roman Garnett},
pages = {5998--6008},
timestamp = {Thu, 21 Jan 2021 00:00:00 +0100},
title = {Attention is All you Need},
url = {https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html},
year = {2017}
}
@inproceedings{schulman-etal:2015high,
author = {John Schulman and
Philipp Moritz and
Sergey Levine and
Michael I. Jordan and
Pieter Abbeel},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/journals/corr/SchulmanMLJA15.bib},
booktitle = {4th International Conference on Learning Representations, {ICLR} 2016,
San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings},
editor = {Yoshua Bengio and
Yann LeCun},
timestamp = {Thu, 25 Jul 2019 01:00:00 +0200},
title = {High-Dimensional Continuous Control Using Generalized Advantage Estimation},
url = {http://arxiv.org/abs/1506.02438},
year = {2016}
}
@article{wang-etal:2023large,
author = {Wang, Peiyi and Li, Lei and Chen, Liang and Cai, Zefan and Zhu, Dawei and Lin, Binghuai and Cao, Yunbo and Liu, Qi and Liu, Tianyu and Sui, Zhifang},
journal = {ArXiv preprint},
title = {Large language models are not fair evaluators},
year = {2023}
}
@article{mahan-etal:2024generative,
author = {Mahan, Dakota and Van Phung, Duy and Rafailov, Rafael and Blagden, Chase and Lile, Nathan and Castricato, Louis and Fr{\"a}nken, Jan-Philipp and Finn, Chelsea and Albalak, Alon},
journal = {ArXiv preprint},
title = {Generative reward models},
url = {https://arxiv.org/abs/2410.12832},
volume = {abs/2410.12832},
year = {2024}
}
@article{zhang-etal:2024generative,
author = {Zhang, Lunjun and Hosseini, Arian and Bansal, Hritik and Kazemi, Mehran and Kumar, Aviral and Agarwal, Rishabh},
journal = {ArXiv preprint},
title = {Generative verifiers: Reward modeling as next-token prediction},
url = {https://arxiv.org/abs/2408.15240},
volume = {abs/2408.15240},
year = {2024}
}
@inproceedings{yang-etal:2024regularizing,
author = {Rui Yang and
Ruomeng Ding and
Yong Lin and
Huan Zhang and
Tong Zhang},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/YangDLZZ24.bib},
booktitle = {Advances in Neural Information Processing Systems 38: Annual Conference
on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver,
BC, Canada, December 10 - 15, 2024},
editor = {Amir Globersons and
Lester Mackey and
Danielle Belgrave and
Angela Fan and
Ulrich Paquet and
Jakub M. Tomczak and
Cheng Zhang},
timestamp = {Thu, 13 Feb 2025 00:00:00 +0100},
title = {Regularizing Hidden States Enables Learning Generalizable Reward Model
for LLMs},
url = {http://papers.nips.cc/paper\_files/paper/2024/hash/71f7154547c748c8041505521ca433ab-Abstract-Conference.html},
year = {2024}
}
@book{Bishop:2006,
author = {Christopher M. Bishop},
publisher = {Springer},
title = {Pattern Recognition and Machine Learning},
year = {2006}
}
@book{miettinen:1999nonlinear,
author = {Miettinen, Kaisa},
publisher = {Springer Science \& Business Media},
title = {Nonlinear multiobjective optimization},
volume = {12},
year = {1999}
}
@article{lin-etal:2024dogerm,
author = {Lin, Tzu-Han and Li, Chen-An and Lee, Hung-yi and Chen, Yun-Nung},
journal = {ArXiv preprint},
title = {Dogerm: Equipping reward models with domain knowledge through model merging},
url = {https://arxiv.org/abs/2407.01470},
volume = {abs/2407.01470},
year = {2024}
}
@inproceedings{coste-etal:2024reward,
author = {Thomas Coste and
Usman Anwar and
Robert Kirk and
David Krueger},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/iclr/CosteAK024.bib},
booktitle = {The Twelfth International Conference on Learning Representations,
{ICLR} 2024, Vienna, Austria, May 7-11, 2024},
publisher = {OpenReview.net},
timestamp = {Wed, 07 Aug 2024 01:00:00 +0200},
title = {Reward Model Ensembles Help Mitigate Overoptimization},
url = {https://openreview.net/forum?id=dcjtMYkpXx},
year = {2024}
}
@article{eisenstein-etal:2023helping,
author = {Eisenstein, Jacob and Nagpal, Chirag and Agarwal, Alekh and Beirami, Ahmad and D'Amour, Alex and Dvijotham, DJ and Fisch, Adam and Heller, Katherine and Pfohl, Stephen and Ramachandran, Deepak and others},
journal = {ArXiv preprint},
title = {Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking},
url = {https://arxiv.org/abs/2312.09244},
volume = {abs/2312.09244},
year = {2023}
}
@inproceedings{gao-etal:2023scaling,
author = {Leo Gao and
John Schulman and
Jacob Hilton},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/icml/GaoSH23.bib},
booktitle = {International Conference on Machine Learning, {ICML} 2023, 23-29 July
2023, Honolulu, Hawaii, {USA}},
editor = {Andreas Krause and
Emma Brunskill and
Kyunghyun Cho and
Barbara Engelhardt and
Sivan Sabato and
Jonathan Scarlett},
pages = {10835--10866},
publisher = {{PMLR}},
series = {Proceedings of Machine Learning Research},
timestamp = {Mon, 28 Aug 2023 01:00:00 +0200},
title = {Scaling Laws for Reward Model Overoptimization},
url = {https://proceedings.mlr.press/v202/gao23h.html},
volume = {202},
year = {2023}
}
@inproceedings{liu-etal:2023GEval,
address = {Singapore},
author = {Liu, Yang and
Iter, Dan and
Xu, Yichong and
Wang, Shuohang and
Xu, Ruochen and
Zhu, Chenguang},
booktitle = {Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing},
doi = {10.18653/v1/2023.emnlp-main.153},
editor = {Bouamor, Houda and
Pino, Juan and
Bali, Kalika},
pages = {2511--2522},
publisher = {Association for Computational Linguistics},
title = {{G}-Eval: {NLG} Evaluation using Gpt-4 with Better Human Alignment},
url = {https://aclanthology.org/2023.emnlp-main.153},
year = {2023}
}
@inproceedings{zheng-etal:2023judging,
author = {Lianmin Zheng and
Wei{-}Lin Chiang and
Ying Sheng and
Siyuan Zhuang and
Zhanghao Wu and
Yonghao Zhuang and
Zi Lin and
Zhuohan Li and
Dacheng Li and
Eric P. Xing and
Hao Zhang and
Joseph E. Gonzalez and
Ion Stoica},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/ZhengC00WZL0LXZ23.bib},
booktitle = {Advances in Neural Information Processing Systems 36: Annual Conference
on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans,
LA, USA, December 10 - 16, 2023},
editor = {Alice Oh and
Tristan Naumann and
Amir Globerson and
Kate Saenko and
Moritz Hardt and
Sergey Levine},
timestamp = {Thu, 04 Jul 2024 01:00:00 +0200},
title = {Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena},
url = {http://papers.nips.cc/paper\_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets\_and\_Benchmarks.html},
year = {2023}
}
@inproceedings{ng1999-etal:policy,
author = {Ng, Andrew Y and Harada, Daishi and Russell, Stuart J},
booktitle = {Proceedings of the Sixteenth International Conference on Machine Learning},
pages = {278--287},
title = {Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping},
year = {1999}
}
@inproceedings{dubois-etal:2024alpacafarm,
author = {Yann Dubois and
Chen Xuechen Li and
Rohan Taori and
Tianyi Zhang and
Ishaan Gulrajani and
Jimmy Ba and
Carlos Guestrin and
Percy Liang and
Tatsunori B. Hashimoto},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/DuboisLTZGBGLH23.bib},
booktitle = {Advances in Neural Information Processing Systems 36: Annual Conference
on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans,
LA, USA, December 10 - 16, 2023},
editor = {Alice Oh and
Tristan Naumann and
Amir Globerson and
Kate Saenko and
Moritz Hardt and
Sergey Levine},
timestamp = {Fri, 01 Mar 2024 00:00:00 +0100},
title = {AlpacaFarm: {A} Simulation Framework for Methods that Learn from Human
Feedback},
url = {http://papers.nips.cc/paper\_files/paper/2023/hash/5fc47800ee5b30b8777fdd30abcaaf3b-Abstract-Conference.html},
year = {2023}
}
@inproceedings{cui-etal:2024ultra,
author = {Ganqu Cui and
Lifan Yuan and
Ning Ding and
Guanming Yao and
Bingxiang He and
Wei Zhu and
Yuan Ni and
Guotong Xie and
Ruobing Xie and
Yankai Lin and
Zhiyuan Liu and
Maosong Sun},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/icml/CuiY0YH0NXXL0024.bib},
booktitle = {Forty-first International Conference on Machine Learning, {ICML} 2024,
Vienna, Austria, July 21-27, 2024},
publisher = {OpenReview.net},
timestamp = {Mon, 02 Sep 2024 01:00:00 +0200},
title = {{ULTRAFEEDBACK:} Boosting Language Models with Scaled {AI} Feedback},
url = {https://openreview.net/forum?id=BOorDpKHiJ},
year = {2024}
}
@article{lee-etal:2023rlaif,
author = {Lee, Harrison and Phatale, Samrat and Mansoor, Hassan and Lu, Kellie Ren and Mesnard, Thomas and Ferret, Johan and Bishop, Colton and Hall, Ethan and Carbune, Victor and Rastogi, Abhinav},
journal = {ArXiv preprint},
title = {Rlaif: Scaling reinforcement learning from human feedback with AI feedback},
url = {https://arxiv.org/abs/2309.00267},
volume = {abs/2309.00267},
year = {2023}
}
@article{shao-etal:2024deepseekmath,
author = {Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Bi, Xiao and Zhang, Haowei and Zhang, Mingchuan and Li, YK and Wu, Y and others},
journal = {ArXiv preprint},
title = {Deepseekmath: Pushing the limits of mathematical reasoning in open language models},
url = {https://arxiv.org/abs/2402.03300},
volume = {abs/2402.03300},
year = {2024}
}
@inproceedings{wang-etal:2024esrl,
author = {Chenglong Wang and
Hang Zhou and
Yimin Hu and
Yifu Huo and
Bei Li and
Tongran Liu and
Tong Xiao and
Jingbo Zhu},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/aaai/WangZHHLLXZ24.bib},
booktitle = {Thirty-Eighth {AAAI} Conference on Artificial Intelligence, {AAAI}
2024, Thirty-Sixth Conference on Innovative Applications of Artificial
Intelligence, {IAAI} 2024, Fourteenth Symposium on Educational Advances
in Artificial Intelligence, {EAAI} 2014, February 20-27, 2024, Vancouver,
Canada},
doi = {10.1609/AAAI.V38I17.29878},
editor = {Michael J. Wooldridge and
Jennifer G. Dy and
Sriraam Natarajan},
pages = {19107--19115},
publisher = {{AAAI} Press},
timestamp = {Tue, 02 Apr 2024 01:00:00 +0200},
title = {{ESRL:} Efficient Sampling-Based Reinforcement Learning for Sequence
Generation},
url = {https://doi.org/10.1609/aaai.v38i17.29878},
year = {2024}
}
@inproceedings{schulman-etal:2015trust,
author = {John Schulman and
Sergey Levine and
Pieter Abbeel and
Michael I. Jordan and
Philipp Moritz},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/icml/SchulmanLAJM15.bib},
booktitle = {Proceedings of the 32nd International Conference on Machine Learning,
{ICML} 2015, Lille, France, 6-11 July 2015},
editor = {Francis R. Bach and
David M. Blei},
pages = {1889--1897},
publisher = {JMLR.org},
series = {{JMLR} Workshop and Conference Proceedings},
timestamp = {Wed, 29 May 2019 01:00:00 +0200},
title = {Trust Region Policy Optimization},
url = {http://proceedings.mlr.press/v37/schulman15.html},
volume = {37},
year = {2015}
}
@article{schulman-etal:2017proximal,
author = {Schulman, John and Wolski, Filip and Dhariwal, Prafulla and Radford, Alec and Klimov, Oleg},
journal = {ArXiv preprint},
title = {Proximal policy optimization algorithms},
url = {https://arxiv.org/abs/1707.06347},
volume = {abs/1707.06347},
year = {2017}
}
@article{bradley-and-terry:rank,
author = {Ralph Allan Bradley and Milton E. Terry},
journal = {Biometrika},
number = {3/4},
pages = {324--345},
title = {Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons},
volume = {39},
year = {1952}
}
@inproceedings{stiennon-etal:2020learning,
author = {Nisan Stiennon and
Long Ouyang and
Jeffrey Wu and
Daniel M. Ziegler and
Ryan Lowe and
Chelsea Voss and
Alec Radford and
Dario Amodei and
Paul F. Christiano},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/StiennonO0ZLVRA20.bib},
booktitle = {Advances in Neural Information Processing Systems 33: Annual Conference
on Neural Information Processing Systems 2020, NeurIPS 2020, December
6-12, 2020, virtual},
editor = {Hugo Larochelle and
Marc'Aurelio Ranzato and
Raia Hadsell and
Maria{-}Florina Balcan and
Hsuan{-}Tien Lin},
timestamp = {Tue, 19 Jan 2021 00:00:00 +0100},
title = {Learning to summarize with human feedback},
url = {https://proceedings.neurips.cc/paper/2020/hash/1f89885d556929e98d3ef9b86448f951-Abstract.html},
year = {2020}
}
@inproceedings{christiano-etal:2017deep,
author = {Paul F. Christiano and
Jan Leike and
Tom B. Brown and
Miljan Martic and
Shane Legg and
Dario Amodei},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/ChristianoLBMLA17.bib},
booktitle = {Advances in Neural Information Processing Systems 30: Annual Conference
on Neural Information Processing Systems 2017, December 4-9, 2017,
Long Beach, CA, {USA}},
editor = {Isabelle Guyon and
Ulrike von Luxburg and
Samy Bengio and
Hanna M. Wallach and
Rob Fergus and
S. V. N. Vishwanathan and
Roman Garnett},
pages = {4299--4307},
timestamp = {Thu, 21 Jan 2021 00:00:00 +0100},
title = {Deep Reinforcement Learning from Human Preferences},
url = {https://proceedings.neurips.cc/paper/2017/hash/d5e2c0adad503c91f91df240d0cd4e49-Abstract.html},
year = {2017}
}
@inproceedings{wei-etal:2022finetuned,
author = {Jason Wei and
Maarten Bosma and
Vincent Y. Zhao and
Kelvin Guu and
Adams Wei Yu and
Brian Lester and
Nan Du and
Andrew M. Dai and
Quoc V. Le},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/iclr/WeiBZGYLDDL22.bib},
booktitle = {The Tenth International Conference on Learning Representations, {ICLR}
2022, Virtual Event, April 25-29, 2022},
publisher = {OpenReview.net},
timestamp = {Sat, 20 Aug 2022 01:00:00 +0200},
title = {Finetuned Language Models are Zero-Shot Learners},
url = {https://openreview.net/forum?id=gEZrGCozdqR},
year = {2022}
}
@article{havrilla:2024teaching,
author = {Havrilla, Alex and Du, Yuqing and Raparthy, Sharath Chandra and Nalmpantis, Christoforos and Dwivedi-Yu, Jane and Zhuravinskyi, Maksym and Hambro, Eric and Sukhbaatar, Sainbayar and Raileanu, Roberta},
journal = {ArXiv preprint},
title = {Teaching large language models to reason with reinforcement learning},
url = {https://arxiv.org/abs/2403.04642},
volume = {abs/2403.04642},
year = {2024}
}
@inproceedings{shen:2015minimum,
address = {Berlin, Germany},
author = {Shen, Shiqi and
Cheng, Yong and
He, Zhongjun and
He, Wei and
Wu, Hua and
Sun, Maosong and
Liu, Yang},
booktitle = {Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
doi = {10.18653/v1/P16-1159},
editor = {Erk, Katrin and
Smith, Noah A.},
pages = {1683--1692},
publisher = {Association for Computational Linguistics},
title = {Minimum Risk Training for Neural Machine Translation},
url = {https://aclanthology.org/P16-1159},
year = {2016}
}
@article{guo:2025deepseek,
author = {Guo, Daya and Yang, Dejian and Zhang, Haowei and Song, Junxiao and Zhang, Ruoyu and Xu, Runxin and Zhu, Qihao and Ma, Shirong and Wang, Peiyi and Bi, Xiao and others},
journal = {ArXiv preprint},
title = {Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning},
url = {https://arxiv.org/abs/2501.12948},
volume = {abs/2501.12948},
year = {2025}
}
@article{touvron:2023llama,
author = {Touvron, Hugo and Martin, Louis and Stone, Kevin and Albert, Peter and Almahairi, Amjad and Babaei, Yasmine and Bashlykov, Nikolay and Batra, Soumya and Bhargava, Prajjwal and Bhosale, Shruti and others},
journal = {ArXiv preprint},
title = {Llama 2: Open foundation and fine-tuned chat models},
url = {https://arxiv.org/abs/2307.09288},
volume = {abs/2307.09288},
year = {2023}
}
@inproceedings{ouyang:2022training,
author = {Long Ouyang and
Jeffrey Wu and
Xu Jiang and
Diogo Almeida and
Carroll L. Wainwright and
Pamela Mishkin and
Chong Zhang and
Sandhini Agarwal and
Katarina Slama and
Alex Ray and
John Schulman and
Jacob Hilton and
Fraser Kelton and
Luke Miller and
Maddie Simens and
Amanda Askell and
Peter Welinder and
Paul F. Christiano and
Jan Leike and
Ryan Lowe},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/Ouyang0JAWMZASR22.bib},
booktitle = {Advances in Neural Information Processing Systems 35: Annual Conference
on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans,
LA, USA, November 28 - December 9, 2022},
editor = {Sanmi Koyejo and
S. Mohamed and
A. Agarwal and
Danielle Belgrave and
K. Cho and
A. Oh},
timestamp = {Mon, 08 Jan 2024 00:00:00 +0100},
title = {Training language models to follow instructions with human feedback},
url = {http://papers.nips.cc/paper\_files/paper/2022/hash/b1efde53be364a73914f58805a001731-Abstract-Conference.html},
year = {2022}
}
@inproceedings{bahdanau:2016actor,
author = {Dzmitry Bahdanau and
Philemon Brakel and
Kelvin Xu and
Anirudh Goyal and
Ryan Lowe and
Joelle Pineau and
Aaron C. Courville and
Yoshua Bengio},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/iclr/BahdanauBXGLPCB17.bib},
booktitle = {5th International Conference on Learning Representations, {ICLR} 2017,
Toulon, France, April 24-26, 2017, Conference Track Proceedings},
publisher = {OpenReview.net},
timestamp = {Thu, 25 Jul 2019 01:00:00 +0200},
title = {An Actor-Critic Algorithm for Sequence Prediction},
url = {https://openreview.net/forum?id=SJDaqqveg},
year = {2017}
}
@inproceedings{mnih-etal:2016asynchronous,
author = {Volodymyr Mnih and
Adri{\`{a}} Puigdom{\`{e}}nech Badia and
Mehdi Mirza and
Alex Graves and
Timothy P. Lillicrap and
Tim Harley and
David Silver and
Koray Kavukcuoglu},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/icml/MnihBMGLHSK16.bib},
booktitle = {Proceedings of the 33nd International Conference on Machine Learning,
{ICML} 2016, New York City, NY, USA, June 19-24, 2016},
editor = {Maria{-}Florina Balcan and
Kilian Q. Weinberger},
pages = {1928--1937},
publisher = {JMLR.org},
series = {{JMLR} Workshop and Conference Proceedings},
timestamp = {Wed, 29 May 2019 01:00:00 +0200},
title = {Asynchronous Methods for Deep Reinforcement Learning},
url = {http://proceedings.mlr.press/v48/mniha16.html},
volume = {48},
year = {2016}
}
@article{szepesvari:2010algorithms,
author = {Szepesv{\'a}ri, Csaba},
journal = {Synthesis Lectures on Artificial Intelligence and Machine Learning},
number = {1},
pages = {1--103},
publisher = {Springer Science and Business Media LLC},
title = {Algorithms for Reinforcement Learning},
volume = {4},
year = {2010}
}
@article{williams:1992simple,
author = {Williams, Ronald J},
journal = {Machine learning},
pages = {229--256},
publisher = {Springer},
title = {Simple statistical gradient-following algorithms for connectionist reinforcement learning},
volume = {8},
year = {1992}
}
@book{Sutton-and-Barto:2018RL,
author = {Richard S. Sutton and Andrew G. Barto},
publisher = {The MIT Press},
title = {Reinforcement Learning: An Introduction (2nd ed.)},
year = {2018}
}
@article{bradley:1952rank,
author = {Bradley, Ralph Allan and Terry, Milton E},
journal = {Biometrika},
number = {3/4},
pages = {324--345},
publisher = {JSTOR},
title = {Rank analysis of incomplete block designs: I. The method of paired comparisons},
volume = {39},
year = {1952}
}
@inproceedings{rafailov:2023direct,
author = {Rafael Rafailov and
Archit Sharma and
Eric Mitchell and
Christopher D. Manning and
Stefano Ermon and
Chelsea Finn},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/RafailovSMMEF23.bib},
booktitle = {Advances in Neural Information Processing Systems 36: Annual Conference
on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans,
LA, USA, December 10 - 16, 2023},
editor = {Alice Oh and
Tristan Naumann and
Amir Globerson and
Kate Saenko and
Moritz Hardt and
Sergey Levine},
timestamp = {Fri, 01 Mar 2024 00:00:00 +0100},
title = {Direct Preference Optimization: Your Language Model is Secretly a
Reward Model},
url = {http://papers.nips.cc/paper\_files/paper/2023/hash/a85b405ed65c6477a4fe8302b5e06ce7-Abstract-Conference.html},
year = {2023}
}
@inproceedings{azar:2024general,
author = {Mohammad Gheshlaghi Azar and
Zhaohan Daniel Guo and
Bilal Piot and
R{\'{e}}mi Munos and
Mark Rowland and
Michal Valko and
Daniele Calandriello},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/aistats/AzarGPMRVC24.bib},
booktitle = {International Conference on Artificial Intelligence and Statistics,
2-4 May 2024, Palau de Congressos, Valencia, Spain},
editor = {Sanjoy Dasgupta and
Stephan Mandt and
Yingzhen Li},
pages = {4447--4455},
publisher = {{PMLR}},
series = {Proceedings of Machine Learning Research},
timestamp = {Mon, 13 May 2024 01:00:00 +0200},
title = {A General Theoretical Paradigm to Understand Learning from Human Preferences},
url = {https://proceedings.mlr.press/v238/gheshlaghi-azar24a.html},
volume = {238},
year = {2024}
}
@inproceedings{xu:2024contrastive,
author = {Haoran Xu and
Amr Sharaf and
Yunmo Chen and
Weiting Tan and
Lingfeng Shen and
Benjamin Van Durme and
Kenton Murray and
Young Jin Kim},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/icml/XuSCTSDM024.bib},
booktitle = {Forty-first International Conference on Machine Learning, {ICML} 2024,
Vienna, Austria, July 21-27, 2024},
publisher = {OpenReview.net},
timestamp = {Mon, 02 Sep 2024 01:00:00 +0200},
title = {Contrastive Preference Optimization: Pushing the Boundaries of {LLM}
Performance in Machine Translation},
url = {https://openreview.net/forum?id=51iwkioZpn},
year = {2024}
}
@article{ethayarajh:2024kto,
author = {Ethayarajh, Kawin and Xu, Winnie and Muennighoff, Niklas and Jurafsky, Dan and Kiela, Douwe},
journal = {ArXiv preprint},
title = {Kto: Model alignment as prospect theoretic optimization},
url = {https://arxiv.org/abs/2402.01306},
volume = {abs/2402.01306},
year = {2024}
}
@article{hong:2024orpo,
author = {Hong, Jiwoo and Lee, Noah and Thorne, James},
journal = {ArXiv preprint},
title = {Orpo: Monolithic preference optimization without reference model},
url = {https://arxiv.org/abs/2403.07691},
volume = {abs/2403.07691},
year = {2024}
}
@article{gallego:2024refined,
author = {Gallego, V{\'\i}ctor},
journal = {ArXiv preprint},
title = {Refined direct preference optimization with synthetic data for behavioral alignment of llms},
url = {https://arxiv.org/abs/2402.08005},
volume = {abs/2402.08005},
year = {2024}
}
@inproceedings{meng:2025simpo,
author = {Yu Meng and
Mengzhou Xia and
Danqi Chen},
bibsource = {dblp computer science bibliography, https://dblp.org},
biburl = {https://dblp.org/rec/conf/nips/0001X024.bib},
booktitle = {Advances in Neural Information Processing Systems 38: Annual Conference
on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver,
BC, Canada, December 10 - 15, 2024},
editor = {Amir Globersons and
Lester Mackey and
Danielle Belgrave and
Angela Fan and
Ulrich Paquet and
Jakub M. Tomczak and
Cheng Zhang},
timestamp = {Thu, 13 Feb 2025 00:00:00 +0100},
title = {SimPO: Simple Preference Optimization with a Reference-Free Reward},
url = {http://papers.nips.cc/paper\_files/paper/2024/hash/e099c1c9699814af0be873a175361713-Abstract-Conference.html},
year = {2024}
}
@inproceedings{zhou:2024prior,
author = {Zhou, Hang and Wang, Chenglong and Hu, Yimin and Xiao, Tong and Zhang, Chunliang and Zhu, Jingbo},
booktitle = {China National Conference on Chinese Computational Linguistics},
organization = {Springer},
pages = {555--570},
title = {Prior constraints-based reward model training for aligning large language models},
year = {2024}
}
@article{shao:2025earlier,
author = {Shao, Ruichen and Li, Bei and Liu, Gangao and Chen, Yang and Zhou, Xiang and Wang, Jingang and Cai, Xunliang and Li, Peng},
journal = {ArXiv preprint},
title = {Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective},
url = {https://arxiv.org/abs/2502.14340},
volume = {abs/2502.14340},
year = {2025}
}
@misc{openai:2024learning,
title = {Learning to Reason with LLMs},
url = {https://openai.com/index/learning-to-reason-with-llms/},
author = {OpenAI},
month = {September},
year = {2024}
}
@article{deepseek:2025r1,
title={Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning},
author={Deepseek},
journal={arXiv preprint arXiv:2501.12948},
year={2025}
}
\ No newline at end of file
%% File: `iclr2024.bst'
%% A copy of iclm2010.bst, which is a modification of `plainnl.bst' for use with natbib package
%%
%% Copyright 2010 Hal Daum\'e III
%% Modified by J. Fürnkranz
%% - Changed labels from (X and Y, 2000) to (X & Y, 2000)
%%
%% Copyright 1993-2007 Patrick W Daly
%% Max-Planck-Institut f\"ur Sonnensystemforschung
%% Max-Planck-Str. 2
%% D-37191 Katlenburg-Lindau
%% Germany
%% E-mail: daly@mps.mpg.de
%%
%% This program can be redistributed and/or modified under the terms
%% of the LaTeX Project Public License Distributed from CTAN
%% archives in directory macros/latex/base/lppl.txt; either
%% version 1 of the License, or any later version.
%%
% Version and source file information:
% \ProvidesFile{icml2010.mbs}[2007/11/26 1.93 (PWD)]
%
% BibTeX `plainnat' family
% version 0.99b for BibTeX versions 0.99a or later,
% for LaTeX versions 2.09 and 2e.
%
% For use with the `natbib.sty' package; emulates the corresponding
% member of the `plain' family, but with author-year citations.
%
% With version 6.0 of `natbib.sty', it may also be used for numerical
% citations, while retaining the commands \citeauthor, \citefullauthor,
% and \citeyear to print the corresponding information.
%
% For version 7.0 of `natbib.sty', the KEY field replaces missing
% authors/editors, and the date is left blank in \bibitem.
%
% Includes field EID for the sequence/citation number of electronic journals
% which is used instead of page numbers.
%
% Includes fields ISBN and ISSN.
%
% Includes field URL for Internet addresses.
%
% Includes field DOI for Digital Object Idenfifiers.
%
% Works best with the url.sty package of Donald Arseneau.
%
% Works with identical authors and year are further sorted by
% citation key, to preserve any natural sequence.
%
ENTRY
{ address
author
booktitle
chapter
doi
eid
edition
editor
howpublished
institution
isbn
issn
journal
key
month
note
number
organization
pages
publisher
school
series
title
type
url
volume
year
}
{}
{ label extra.label sort.label short.list }
INTEGERS { output.state before.all mid.sentence after.sentence after.block }
FUNCTION {init.state.consts}
{ #0 'before.all :=
#1 'mid.sentence :=
#2 'after.sentence :=
#3 'after.block :=
}
STRINGS { s t }
FUNCTION {output.nonnull}
{ 's :=
output.state mid.sentence =
{ ", " * write$ }
{ output.state after.block =
{ add.period$ write$
newline$
"\newblock " write$
}
{ output.state before.all =
'write$
{ add.period$ " " * write$ }
if$
}
if$
mid.sentence 'output.state :=
}
if$
s
}
FUNCTION {output}
{ duplicate$ empty$
'pop$
'output.nonnull
if$
}
FUNCTION {output.check}
{ 't :=
duplicate$ empty$
{ pop$ "empty " t * " in " * cite$ * warning$ }
'output.nonnull
if$
}
FUNCTION {fin.entry}
{ add.period$
write$
newline$
}
FUNCTION {new.block}
{ output.state before.all =
'skip$
{ after.block 'output.state := }
if$
}
FUNCTION {new.sentence}
{ output.state after.block =
'skip$
{ output.state before.all =
'skip$
{ after.sentence 'output.state := }
if$
}
if$
}
FUNCTION {not}
{ { #0 }
{ #1 }
if$
}
FUNCTION {and}
{ 'skip$
{ pop$ #0 }
if$
}
FUNCTION {or}
{ { pop$ #1 }
'skip$
if$
}
FUNCTION {new.block.checka}
{ empty$
'skip$
'new.block
if$
}
FUNCTION {new.block.checkb}
{ empty$
swap$ empty$
and
'skip$
'new.block
if$
}
FUNCTION {new.sentence.checka}
{ empty$
'skip$
'new.sentence
if$
}
FUNCTION {new.sentence.checkb}
{ empty$
swap$ empty$
and
'skip$
'new.sentence
if$
}
FUNCTION {field.or.null}
{ duplicate$ empty$
{ pop$ "" }
'skip$
if$
}
FUNCTION {emphasize}
{ duplicate$ empty$
{ pop$ "" }
{ "\emph{" swap$ * "}" * }
if$
}
INTEGERS { nameptr namesleft numnames }
FUNCTION {format.names}
{ 's :=
#1 'nameptr :=
s num.names$ 'numnames :=
numnames 'namesleft :=
{ namesleft #0 > }
{ s nameptr "{ff~}{vv~}{ll}{, jj}" format.name$ 't :=
nameptr #1 >
{ namesleft #1 >
{ ", " * t * }
{ numnames #2 >
{ "," * }
'skip$
if$
t "others" =
{ " et~al." * }
{ " and " * t * }
if$
}
if$
}
't
if$
nameptr #1 + 'nameptr :=
namesleft #1 - 'namesleft :=
}
while$
}
FUNCTION {format.key}
{ empty$
{ key field.or.null }
{ "" }
if$
}
FUNCTION {format.authors}
{ author empty$
{ "" }
{ author format.names }
if$
}
FUNCTION {format.editors}
{ editor empty$
{ "" }
{ editor format.names
editor num.names$ #1 >
{ " (eds.)" * }
{ " (ed.)" * }
if$
}
if$
}
FUNCTION {format.isbn}
{ isbn empty$
{ "" }
{ new.block "ISBN " isbn * }
if$
}
FUNCTION {format.issn}
{ issn empty$
{ "" }
{ new.block "ISSN " issn * }
if$
}
FUNCTION {format.url}
{ url empty$
{ "" }
{ new.block "URL \url{" url * "}" * }
if$
}
FUNCTION {format.doi}
{ doi empty$
{ "" }
{ new.block "\doi{" doi * "}" * }
if$
}
FUNCTION {format.title}
{ title empty$
{ "" }
{ title "t" change.case$ }
if$
}
FUNCTION {format.full.names}
{'s :=
#1 'nameptr :=
s num.names$ 'numnames :=
numnames 'namesleft :=
{ namesleft #0 > }
{ s nameptr
"{vv~}{ll}" format.name$ 't :=
nameptr #1 >
{
namesleft #1 >
{ ", " * t * }
{
numnames #2 >
{ "," * }
'skip$
if$
t "others" =
{ " et~al." * }
{ " and " * t * }
if$
}
if$
}
't
if$
nameptr #1 + 'nameptr :=
namesleft #1 - 'namesleft :=
}
while$
}
FUNCTION {author.editor.full}
{ author empty$
{ editor empty$
{ "" }
{ editor format.full.names }
if$
}
{ author format.full.names }
if$
}
FUNCTION {author.full}
{ author empty$
{ "" }
{ author format.full.names }
if$
}
FUNCTION {editor.full}
{ editor empty$
{ "" }
{ editor format.full.names }
if$
}
FUNCTION {make.full.names}
{ type$ "book" =
type$ "inbook" =
or
'author.editor.full
{ type$ "proceedings" =
'editor.full
'author.full
if$
}
if$
}
FUNCTION {output.bibitem}
{ newline$
"\bibitem[" write$
label write$
")" make.full.names duplicate$ short.list =
{ pop$ }
{ * }
if$
"]{" * write$
cite$ write$
"}" write$
newline$
""
before.all 'output.state :=
}
FUNCTION {n.dashify}
{ 't :=
""
{ t empty$ not }
{ t #1 #1 substring$ "-" =
{ t #1 #2 substring$ "--" = not
{ "--" *
t #2 global.max$ substring$ 't :=
}
{ { t #1 #1 substring$ "-" = }
{ "-" *
t #2 global.max$ substring$ 't :=
}
while$
}
if$
}
{ t #1 #1 substring$ *
t #2 global.max$ substring$ 't :=
}
if$
}
while$
}
FUNCTION {format.date}
{ year duplicate$ empty$
{ "empty year in " cite$ * warning$
pop$ "" }
'skip$
if$
month empty$
'skip$
{ month
" " * swap$ *
}
if$
extra.label *
}
FUNCTION {format.btitle}
{ title emphasize
}
FUNCTION {tie.or.space.connect}
{ duplicate$ text.length$ #3 <
{ "~" }
{ " " }
if$
swap$ * *
}
FUNCTION {either.or.check}
{ empty$
'pop$
{ "can't use both " swap$ * " fields in " * cite$ * warning$ }
if$
}
FUNCTION {format.bvolume}
{ volume empty$
{ "" }
{ "volume" volume tie.or.space.connect
series empty$
'skip$
{ " of " * series emphasize * }
if$
"volume and number" number either.or.check
}
if$
}
FUNCTION {format.number.series}
{ volume empty$
{ number empty$
{ series field.or.null }
{ output.state mid.sentence =
{ "number" }
{ "Number" }
if$
number tie.or.space.connect
series empty$
{ "there's a number but no series in " cite$ * warning$ }
{ " in " * series * }
if$
}
if$
}
{ "" }
if$
}
FUNCTION {format.edition}
{ edition empty$
{ "" }
{ output.state mid.sentence =
{ edition "l" change.case$ " edition" * }
{ edition "t" change.case$ " edition" * }
if$
}
if$
}
INTEGERS { multiresult }
FUNCTION {multi.page.check}
{ 't :=
#0 'multiresult :=
{ multiresult not
t empty$ not
and
}
{ t #1 #1 substring$
duplicate$ "-" =
swap$ duplicate$ "," =
swap$ "+" =
or or
{ #1 'multiresult := }
{ t #2 global.max$ substring$ 't := }
if$
}
while$
multiresult
}
FUNCTION {format.pages}
{ pages empty$
{ "" }
{ pages multi.page.check
{ "pp.\ " pages n.dashify tie.or.space.connect }
{ "pp.\ " pages tie.or.space.connect }
if$
}
if$
}
FUNCTION {format.eid}
{ eid empty$
{ "" }
{ "art." eid tie.or.space.connect }
if$
}
FUNCTION {format.vol.num.pages}
{ volume field.or.null
number empty$
'skip$
{ "\penalty0 (" number * ")" * *
volume empty$
{ "there's a number but no volume in " cite$ * warning$ }
'skip$
if$
}
if$
pages empty$
'skip$
{ duplicate$ empty$
{ pop$ format.pages }
{ ":\penalty0 " * pages n.dashify * }
if$
}
if$
}
FUNCTION {format.vol.num.eid}
{ volume field.or.null
number empty$
'skip$
{ "\penalty0 (" number * ")" * *
volume empty$
{ "there's a number but no volume in " cite$ * warning$ }
'skip$
if$
}
if$
eid empty$
'skip$
{ duplicate$ empty$
{ pop$ format.eid }
{ ":\penalty0 " * eid * }
if$
}
if$
}
FUNCTION {format.chapter.pages}
{ chapter empty$
'format.pages
{ type empty$
{ "chapter" }
{ type "l" change.case$ }
if$
chapter tie.or.space.connect
pages empty$
'skip$
{ ", " * format.pages * }
if$
}
if$
}
FUNCTION {format.in.ed.booktitle}
{ booktitle empty$
{ "" }
{ editor empty$
{ "In " booktitle emphasize * }
{ "In " format.editors * ", " * booktitle emphasize * }
if$
}
if$
}
FUNCTION {empty.misc.check}
{ author empty$ title empty$ howpublished empty$
month empty$ year empty$ note empty$
and and and and and
key empty$ not and
{ "all relevant fields are empty in " cite$ * warning$ }
'skip$
if$
}
FUNCTION {format.thesis.type}
{ type empty$
'skip$
{ pop$
type "t" change.case$
}
if$
}
FUNCTION {format.tr.number}
{ type empty$
{ "Technical Report" }
'type
if$
number empty$
{ "t" change.case$ }
{ number tie.or.space.connect }
if$
}
FUNCTION {format.article.crossref}
{ key empty$
{ journal empty$
{ "need key or journal for " cite$ * " to crossref " * crossref *
warning$
""
}
{ "In \emph{" journal * "}" * }
if$
}
{ "In " }
if$
" \citet{" * crossref * "}" *
}
FUNCTION {format.book.crossref}
{ volume empty$
{ "empty volume in " cite$ * "'s crossref of " * crossref * warning$
"In "
}
{ "Volume" volume tie.or.space.connect
" of " *
}
if$
editor empty$
editor field.or.null author field.or.null =
or
{ key empty$
{ series empty$
{ "need editor, key, or series for " cite$ * " to crossref " *
crossref * warning$
"" *
}
{ "\emph{" * series * "}" * }
if$
}
'skip$
if$
}
'skip$
if$
" \citet{" * crossref * "}" *
}
FUNCTION {format.incoll.inproc.crossref}
{ editor empty$
editor field.or.null author field.or.null =
or
{ key empty$
{ booktitle empty$
{ "need editor, key, or booktitle for " cite$ * " to crossref " *
crossref * warning$
""
}
{ "In \emph{" booktitle * "}" * }
if$
}
{ "In " }
if$
}
{ "In " }
if$
" \citet{" * crossref * "}" *
}
FUNCTION {article}
{ output.bibitem
format.authors "author" output.check
author format.key output
new.block
format.title "title" output.check
new.block
crossref missing$
{ journal emphasize "journal" output.check
eid empty$
{ format.vol.num.pages output }
{ format.vol.num.eid output }
if$
format.date "year" output.check
}
{ format.article.crossref output.nonnull
eid empty$
{ format.pages output }
{ format.eid output }
if$
}
if$
format.issn output
format.doi output
format.url output
new.block
note output
fin.entry
}
FUNCTION {book}
{ output.bibitem
author empty$
{ format.editors "author and editor" output.check
editor format.key output
}
{ format.authors output.nonnull
crossref missing$
{ "author and editor" editor either.or.check }
'skip$
if$
}
if$
new.block
format.btitle "title" output.check
crossref missing$
{ format.bvolume output
new.block
format.number.series output
new.sentence
publisher "publisher" output.check
address output
}
{ new.block
format.book.crossref output.nonnull
}
if$
format.edition output
format.date "year" output.check
format.isbn output
format.doi output
format.url output
new.block
note output
fin.entry
}
FUNCTION {booklet}
{ output.bibitem
format.authors output
author format.key output
new.block
format.title "title" output.check
howpublished address new.block.checkb
howpublished output
address output
format.date output
format.isbn output
format.doi output
format.url output
new.block
note output
fin.entry
}
FUNCTION {inbook}
{ output.bibitem
author empty$
{ format.editors "author and editor" output.check
editor format.key output
}
{ format.authors output.nonnull
crossref missing$
{ "author and editor" editor either.or.check }
'skip$
if$
}
if$
new.block
format.btitle "title" output.check
crossref missing$
{ format.bvolume output
format.chapter.pages "chapter and pages" output.check
new.block
format.number.series output
new.sentence
publisher "publisher" output.check
address output
}
{ format.chapter.pages "chapter and pages" output.check
new.block
format.book.crossref output.nonnull
}
if$
format.edition output
format.date "year" output.check
format.isbn output
format.doi output
format.url output
new.block
note output
fin.entry
}
FUNCTION {incollection}
{ output.bibitem
format.authors "author" output.check
author format.key output
new.block
format.title "title" output.check
new.block
crossref missing$
{ format.in.ed.booktitle "booktitle" output.check
format.bvolume output
format.number.series output
format.chapter.pages output
new.sentence
publisher "publisher" output.check
address output
format.edition output
format.date "year" output.check
}
{ format.incoll.inproc.crossref output.nonnull
format.chapter.pages output
}
if$
format.isbn output
format.doi output
format.url output
new.block
note output
fin.entry
}
FUNCTION {inproceedings}
{ output.bibitem
format.authors "author" output.check
author format.key output
new.block
format.title "title" output.check
new.block
crossref missing$
{ format.in.ed.booktitle "booktitle" output.check
format.bvolume output
format.number.series output
format.pages output
address empty$
{ organization publisher new.sentence.checkb
organization output
publisher output
format.date "year" output.check
}
{ address output.nonnull
format.date "year" output.check
new.sentence
organization output
publisher output
}
if$
}
{ format.incoll.inproc.crossref output.nonnull
format.pages output
}
if$
format.isbn output
format.doi output
format.url output
new.block
note output
fin.entry
}
FUNCTION {conference} { inproceedings }
FUNCTION {manual}
{ output.bibitem
format.authors output
author format.key output
new.block
format.btitle "title" output.check
organization address new.block.checkb
organization output
address output
format.edition output
format.date output
format.url output
new.block
note output
fin.entry
}
FUNCTION {mastersthesis}
{ output.bibitem
format.authors "author" output.check
author format.key output
new.block
format.title "title" output.check
new.block
"Master's thesis" format.thesis.type output.nonnull
school "school" output.check
address output
format.date "year" output.check
format.url output
new.block
note output
fin.entry
}
FUNCTION {misc}
{ output.bibitem
format.authors output
author format.key output
title howpublished new.block.checkb
format.title output
howpublished new.block.checka
howpublished output
format.date output
format.issn output
format.url output
new.block
note output
fin.entry
empty.misc.check
}
FUNCTION {phdthesis}
{ output.bibitem
format.authors "author" output.check
author format.key output
new.block
format.btitle "title" output.check
new.block
"PhD thesis" format.thesis.type output.nonnull
school "school" output.check
address output
format.date "year" output.check
format.url output
new.block
note output
fin.entry
}
FUNCTION {proceedings}
{ output.bibitem
format.editors output
editor format.key output
new.block
format.btitle "title" output.check
format.bvolume output
format.number.series output
address output
format.date "year" output.check
new.sentence
organization output
publisher output
format.isbn output
format.doi output
format.url output
new.block
note output
fin.entry
}
FUNCTION {techreport}
{ output.bibitem
format.authors "author" output.check
author format.key output
new.block
format.title "title" output.check
new.block
format.tr.number output.nonnull
institution "institution" output.check
address output
format.date "year" output.check
format.url output
new.block
note output
fin.entry
}
FUNCTION {unpublished}
{ output.bibitem
format.authors "author" output.check
author format.key output
new.block
format.title "title" output.check
new.block
note "note" output.check
format.date output
format.url output
fin.entry
}
FUNCTION {default.type} { misc }
MACRO {jan} {"January"}
MACRO {feb} {"February"}
MACRO {mar} {"March"}
MACRO {apr} {"April"}
MACRO {may} {"May"}
MACRO {jun} {"June"}
MACRO {jul} {"July"}
MACRO {aug} {"August"}
MACRO {sep} {"September"}
MACRO {oct} {"October"}
MACRO {nov} {"November"}
MACRO {dec} {"December"}
MACRO {acmcs} {"ACM Computing Surveys"}
MACRO {acta} {"Acta Informatica"}
MACRO {cacm} {"Communications of the ACM"}
MACRO {ibmjrd} {"IBM Journal of Research and Development"}
MACRO {ibmsj} {"IBM Systems Journal"}
MACRO {ieeese} {"IEEE Transactions on Software Engineering"}
MACRO {ieeetc} {"IEEE Transactions on Computers"}
MACRO {ieeetcad}
{"IEEE Transactions on Computer-Aided Design of Integrated Circuits"}
MACRO {ipl} {"Information Processing Letters"}
MACRO {jacm} {"Journal of the ACM"}
MACRO {jcss} {"Journal of Computer and System Sciences"}
MACRO {scp} {"Science of Computer Programming"}
MACRO {sicomp} {"SIAM Journal on Computing"}
MACRO {tocs} {"ACM Transactions on Computer Systems"}
MACRO {tods} {"ACM Transactions on Database Systems"}
MACRO {tog} {"ACM Transactions on Graphics"}
MACRO {toms} {"ACM Transactions on Mathematical Software"}
MACRO {toois} {"ACM Transactions on Office Information Systems"}
MACRO {toplas} {"ACM Transactions on Programming Languages and Systems"}
MACRO {tcs} {"Theoretical Computer Science"}
READ
FUNCTION {sortify}
{ purify$
"l" change.case$
}
INTEGERS { len }
FUNCTION {chop.word}
{ 's :=
'len :=
s #1 len substring$ =
{ s len #1 + global.max$ substring$ }
's
if$
}
FUNCTION {format.lab.names}
{ 's :=
s #1 "{vv~}{ll}" format.name$
s num.names$ duplicate$
#2 >
{ pop$ " et~al." * }
{ #2 <
'skip$
{ s #2 "{ff }{vv }{ll}{ jj}" format.name$ "others" =
{ " et~al." * }
{ " \& " * s #2 "{vv~}{ll}" format.name$ * }
if$
}
if$
}
if$
}
FUNCTION {author.key.label}
{ author empty$
{ key empty$
{ cite$ #1 #3 substring$ }
'key
if$
}
{ author format.lab.names }
if$
}
FUNCTION {author.editor.key.label}
{ author empty$
{ editor empty$
{ key empty$
{ cite$ #1 #3 substring$ }
'key
if$
}
{ editor format.lab.names }
if$
}
{ author format.lab.names }
if$
}
FUNCTION {author.key.organization.label}
{ author empty$
{ key empty$
{ organization empty$
{ cite$ #1 #3 substring$ }
{ "The " #4 organization chop.word #3 text.prefix$ }
if$
}
'key
if$
}
{ author format.lab.names }
if$
}
FUNCTION {editor.key.organization.label}
{ editor empty$
{ key empty$
{ organization empty$
{ cite$ #1 #3 substring$ }
{ "The " #4 organization chop.word #3 text.prefix$ }
if$
}
'key
if$
}
{ editor format.lab.names }
if$
}
FUNCTION {calc.short.authors}
{ type$ "book" =
type$ "inbook" =
or
'author.editor.key.label
{ type$ "proceedings" =
'editor.key.organization.label
{ type$ "manual" =
'author.key.organization.label
'author.key.label
if$
}
if$
}
if$
'short.list :=
}
FUNCTION {calc.label}
{ calc.short.authors
short.list
"("
*
year duplicate$ empty$
short.list key field.or.null = or
{ pop$ "" }
'skip$
if$
*
'label :=
}
FUNCTION {sort.format.names}
{ 's :=
#1 'nameptr :=
""
s num.names$ 'numnames :=
numnames 'namesleft :=
{ namesleft #0 > }
{
s nameptr "{vv{ } }{ll{ }}{ ff{ }}{ jj{ }}" format.name$ 't :=
nameptr #1 >
{
" " *
namesleft #1 = t "others" = and
{ "zzzzz" * }
{ numnames #2 > nameptr #2 = and
{ "zz" * year field.or.null * " " * }
'skip$
if$
t sortify *
}
if$
}
{ t sortify * }
if$
nameptr #1 + 'nameptr :=
namesleft #1 - 'namesleft :=
}
while$
}
FUNCTION {sort.format.title}
{ 't :=
"A " #2
"An " #3
"The " #4 t chop.word
chop.word
chop.word
sortify
#1 global.max$ substring$
}
FUNCTION {author.sort}
{ author empty$
{ key empty$
{ "to sort, need author or key in " cite$ * warning$
""
}
{ key sortify }
if$
}
{ author sort.format.names }
if$
}
FUNCTION {author.editor.sort}
{ author empty$
{ editor empty$
{ key empty$
{ "to sort, need author, editor, or key in " cite$ * warning$
""
}
{ key sortify }
if$
}
{ editor sort.format.names }
if$
}
{ author sort.format.names }
if$
}
FUNCTION {author.organization.sort}
{ author empty$
{ organization empty$
{ key empty$
{ "to sort, need author, organization, or key in " cite$ * warning$
""
}
{ key sortify }
if$
}
{ "The " #4 organization chop.word sortify }
if$
}
{ author sort.format.names }
if$
}
FUNCTION {editor.organization.sort}
{ editor empty$
{ organization empty$
{ key empty$
{ "to sort, need editor, organization, or key in " cite$ * warning$
""
}
{ key sortify }
if$
}
{ "The " #4 organization chop.word sortify }
if$
}
{ editor sort.format.names }
if$
}
FUNCTION {presort}
{ calc.label
label sortify
" "
*
type$ "book" =
type$ "inbook" =
or
'author.editor.sort
{ type$ "proceedings" =
'editor.organization.sort
{ type$ "manual" =
'author.organization.sort
'author.sort
if$
}
if$
}
if$
" "
*
year field.or.null sortify
*
" "
*
cite$
*
#1 entry.max$ substring$
'sort.label :=
sort.label *
#1 entry.max$ substring$
'sort.key$ :=
}
ITERATE {presort}
SORT
STRINGS { longest.label last.label next.extra }
INTEGERS { longest.label.width last.extra.num number.label }
FUNCTION {initialize.longest.label}
{ "" 'longest.label :=
#0 int.to.chr$ 'last.label :=
"" 'next.extra :=
#0 'longest.label.width :=
#0 'last.extra.num :=
#0 'number.label :=
}
FUNCTION {forward.pass}
{ last.label label =
{ last.extra.num #1 + 'last.extra.num :=
last.extra.num int.to.chr$ 'extra.label :=
}
{ "a" chr.to.int$ 'last.extra.num :=
"" 'extra.label :=
label 'last.label :=
}
if$
number.label #1 + 'number.label :=
}
FUNCTION {reverse.pass}
{ next.extra "b" =
{ "a" 'extra.label := }
'skip$
if$
extra.label 'next.extra :=
extra.label
duplicate$ empty$
'skip$
{ "{\natexlab{" swap$ * "}}" * }
if$
'extra.label :=
label extra.label * 'label :=
}
EXECUTE {initialize.longest.label}
ITERATE {forward.pass}
REVERSE {reverse.pass}
FUNCTION {bib.sort.order}
{ sort.label 'sort.key$ :=
}
ITERATE {bib.sort.order}
SORT
FUNCTION {begin.bib}
{ preamble$ empty$
'skip$
{ preamble$ write$ newline$ }
if$
"\begin{thebibliography}{" number.label int.to.str$ * "}" *
write$ newline$
"\providecommand{\natexlab}[1]{#1}"
write$ newline$
"\providecommand{\url}[1]{\texttt{#1}}"
write$ newline$
"\expandafter\ifx\csname urlstyle\endcsname\relax"
write$ newline$
" \providecommand{\doi}[1]{doi: #1}\else"
write$ newline$
" \providecommand{\doi}{doi: \begingroup \urlstyle{rm}\Url}\fi"
write$ newline$
}
EXECUTE {begin.bib}
EXECUTE {init.state.consts}
ITERATE {call.type$}
FUNCTION {end.bib}
{ newline$
"\end{thebibliography}" write$ newline$
}
EXECUTE {end.bib}
%%%% COLM Macros (LaTex)
%%%% Adapted by Yoav Artzi and Sasha Rush from Hugo Larochelle's adaptation for ICLR, which has been adaptated from the NIPS stylefile Macros
%%%% Style File
%%%% Dec 12, 1990 Rev Aug 14, 1991; Sept, 1995; April, 1997; April, 1999; October 2014
% This file can be used with Latex2e whether running in main mode, or
% 2.09 compatibility mode.
%
% If using main mode, you need to include the commands
% \documentclass{article}
% \usepackage{colm14submit_e}
%
% Define options
\newif\ifcolmsubmission
\newif\ifcolmpreprint
\newif\ifcolmfinal
% Set submission as default
\colmsubmissiontrue
\colmpreprintfalse
\colmfinalfalse
% Define option handling
\DeclareOption{submission}{\colmsubmissiontrue\colmpreprintfalse\colmfinalfalse}
\DeclareOption{preprint}{\colmsubmissionfalse\colmpreprinttrue\colmfinalfalse}
\DeclareOption{final}{\colmsubmissionfalse\colmpreprintfalse\colmfinaltrue}
\ProcessOptions\relax
% Palatino font
\RequirePackage{tgpagella} % text only
\RequirePackage{mathpazo} % math & text
\RequirePackage{inconsolata} % for tt font
% Change the overall width of the page. If these parameters are
% changed, they will require corresponding changes in the
% maketitle section.
%
\usepackage{eso-pic} % used by \AddToShipoutPicture
\RequirePackage{fancyhdr}
\RequirePackage{natbib}
% modification to natbib citations
\setcitestyle{authoryear,round,citesep={;},aysep={,},yysep={;}}
\renewcommand{\topfraction}{0.95} % let figure take up nearly whole page
\renewcommand{\textfraction}{0.05} % let figure take up nearly whole page
% Specify the dimensions of each page
\setlength{\paperheight}{11in}
\setlength{\paperwidth}{8.5in}
\oddsidemargin .09in % Note \oddsidemargin = \evensidemargin
\evensidemargin .09in
\marginparwidth 0.07 true in
%\marginparwidth 0.75 true in
%\topmargin 0 true pt % Nominal distance from top of page to top of
%\topmargin 0.125in
\topmargin -0.625in
\addtolength{\headsep}{0.25in}
\textheight 9.0 true in % Height of text (including footnotes & figures)
\textwidth 6.4 true in % Width of text line.
\widowpenalty=10000
\clubpenalty=10000
% \thispagestyle{empty} \pagestyle{empty}
\flushbottom \sloppy
% We're never going to need a table of contents, so just flush it to
% save space --- suggested by drstrip@sandia-2
\def\addcontentsline#1#2#3{}
% Title stuff, taken from deproc.
\def\maketitle{\par
\begingroup
\def\thefootnote{\fnsymbol{footnote}}
\def\@makefnmark{\hbox to 0pt{$^{\@thefnmark}$\hss}} % for perfect author
% name centering
% The footnote-mark was overlapping the footnote-text,
% added the following to fix this problem (MK)
\long\def\@makefntext##1{\parindent 1em\noindent
\hbox to1.8em{\hss $\m@th ^{\@thefnmark}$}##1}
\@maketitle \@thanks
\endgroup
\setcounter{footnote}{0}
\let\maketitle\relax \let\@maketitle\relax
\gdef\@thanks{}\gdef\@author{}\gdef\@title{}\let\thanks\relax}
% The toptitlebar has been raised to top-justify the first page
\usepackage{fancyhdr}
\pagestyle{fancy}
% \renewcommand{\headrulewidth}{1.5pt}
\renewcommand{\headrulewidth}{0pt}
\fancyhead{}
% Title (includes both anonymized and non-anonymized versions)
\def\@maketitle{\vbox{\hsize\textwidth
%\linewidth\hsize \vskip 0.1in \toptitlebar \centering
{\Large\bf \@title\par}
%\bottomtitlebar % \vskip 0.1in % minus
\ifcolmfinal
% \lhead{Published as a conference paper at COLM 2025}
\def\And{\end{tabular}\hfil\linebreak[0]\hfil
\begin{tabular}[t]{l}\bf\rule{\z@}{24pt}\ignorespaces}%
\def\AND{\end{tabular}\hfil\linebreak[4]\hfil
\begin{tabular}[t]{l}\bf\rule{\z@}{24pt}\ignorespaces}%
\begin{tabular}[t]{l}\bf\rule{\z@}{24pt}\@author\end{tabular}%
\else\ifcolmpreprint
\lhead{Preprint. Under review.}
\def\And{\end{tabular}\hfil\linebreak[0]\hfil
\begin{tabular}[t]{l}\bf\rule{\z@}{24pt}\ignorespaces}%
\def\AND{\end{tabular}\hfil\linebreak[4]\hfil
\begin{tabular}[t]{l}\bf\rule{\z@}{24pt}\ignorespaces}%
\begin{tabular}[t]{l}\bf\rule{\z@}{24pt}\@author\end{tabular}%
\else
\lhead{Under review as a conference paper at COLM 2025}
\def\And{\end{tabular}\hfil\linebreak[0]\hfil
\begin{tabular}[t]{l}\bf\rule{\z@}{24pt}\ignorespaces}%
\def\AND{\end{tabular}\hfil\linebreak[4]\hfil
\begin{tabular}[t]{l}\bf\rule{\z@}{24pt}\ignorespaces}%
\begin{tabular}[t]{l}\bf\rule{\z@}{24pt}Anonymous authors\\Paper under double-blind review\end{tabular}%
\fi\fi
\vskip 0.3in minus 0.1in}}
\renewenvironment{abstract}{\vskip.075in\centerline{\large\bf
Abstract}\vspace{0.5ex}\begin{quote}}{\par\end{quote}\vskip 1ex}
% Less leading in most fonts (due to the narrow columns)
% The choices were between 1-pt and 1.5-pt leading
%\def\@normalsize{\@setsize\normalsize{11pt}\xpt\@xpt} % got rid of @ (MK)
\def\normalsize{\@setsize\normalsize{12.5pt}\xpt\@xpt}
\def\small{\@setsize\small{10pt}\ixpt\@ixpt}
\def\footnotesize{\@setsize\footnotesize{10pt}\ixpt\@ixpt}
\def\scriptsize{\@setsize\scriptsize{8pt}\viipt\@viipt}
\def\tiny{\@setsize\tiny{7pt}\vipt\@vipt}
\def\large{\@setsize\large{14pt}\xiipt\@xiipt}
\def\Large{\@setsize\Large{16pt}\xivpt\@xivpt}
\def\LARGE{\@setsize\LARGE{20pt}\xviipt\@xviipt}
\def\huge{\@setsize\huge{23pt}\xxpt\@xxpt}
\def\Huge{\@setsize\Huge{28pt}\xxvpt\@xxvpt}
% sections with less space
\def\section{\@startsection {section}{1}{\z@}{-2.0ex plus
-0.5ex minus -.2ex}{1.5ex plus 0.3ex
minus0.2ex}{\large\bf\raggedright}}
\def\subsection{\@startsection{subsection}{2}{\z@}{-1.8ex plus
-0.5ex minus -.2ex}{0.8ex plus .2ex}{\normalsize\bf\raggedright}}
\def\subsubsection{\@startsection{subsubsection}{3}{\z@}{-1.5ex
plus -0.5ex minus -.2ex}{0.5ex plus
.2ex}{\normalsize\bf\itshape\raggedright}}
\def\paragraph{\@startsection{paragraph}{4}{\z@}{1.5ex plus
0.5ex minus .2ex}{-1em}{\normalsize\bf}}
\def\subparagraph{\@startsection{subparagraph}{5}{\z@}{1.5ex plus
0.5ex minus .2ex}{-1em}{\normalsize\it}}
\def\subsubsubsection{\vskip
5pt{\noindent\normalsize\raggedright}}
% Footnotes
\footnotesep 6.65pt %
\skip\footins 9pt plus 4pt minus 2pt
\def\footnoterule{\kern-3pt \hrule width 12pc \kern 2.6pt }
\setcounter{footnote}{0}
% Lists and paragraphs
\parindent 0pt
\topsep 4pt plus 1pt minus 2pt
\partopsep 1pt plus 0.5pt minus 0.5pt
\itemsep 2pt plus 1pt minus 0.5pt
\parsep 2pt plus 1pt minus 0.5pt
\parskip .6pc
%\leftmargin2em
\leftmargin3pc
\leftmargini\leftmargin \leftmarginii 2em
\leftmarginiii 1.5em \leftmarginiv 1.0em \leftmarginv .5em
%\labelsep \labelsep 5pt
\def\@listi{\leftmargin\leftmargini}
\def\@listii{\leftmargin\leftmarginii
\labelwidth\leftmarginii\advance\labelwidth-\labelsep
\topsep 2pt plus 1pt minus 0.5pt
\parsep 1pt plus 0.5pt minus 0.5pt
\itemsep \parsep}
\def\@listiii{\leftmargin\leftmarginiii
\labelwidth\leftmarginiii\advance\labelwidth-\labelsep
\topsep 1pt plus 0.5pt minus 0.5pt
\parsep \z@ \partopsep 0.5pt plus 0pt minus 0.5pt
\itemsep \topsep}
\def\@listiv{\leftmargin\leftmarginiv
\labelwidth\leftmarginiv\advance\labelwidth-\labelsep}
\def\@listv{\leftmargin\leftmarginv
\labelwidth\leftmarginv\advance\labelwidth-\labelsep}
\def\@listvi{\leftmargin\leftmarginvi
\labelwidth\leftmarginvi\advance\labelwidth-\labelsep}
\abovedisplayskip 7pt plus2pt minus5pt%
\belowdisplayskip \abovedisplayskip
\abovedisplayshortskip 0pt plus3pt%
\belowdisplayshortskip 4pt plus3pt minus3pt%
\def\toptitlebar{\hrule height4pt\vskip .25in\vskip-\parskip}
\def\bottomtitlebar{\vskip .29in\vskip-\parskip\hrule height1pt\vskip
.09in} %
%Reduced second vskip to compensate for adding the strut in \@author
\documentclass{article} % For LaTeX2e
\usepackage[final]{rl-introduction}
\usepackage{microtype}
\usepackage{hyperref}
\usepackage{url}
\usepackage{booktabs}
\usepackage{fontawesome5}
\usepackage{lineno}
\usepackage{bbding}
% Write outlines on Chinese
\usepackage{CJK}
%%%%%%%%%%%%%%%%%%%%%% algorithm %%%%%%%%%%%%%%%%%%%%%%%
\usepackage{algorithm} % http://ctan.org/pkg/algorithms
\usepackage{algpseudocode}
% \makeatletter
% \renewcommand{\ALG@beginalgorithmic}{\normal}
% \makeatother
\algnewcommand\algorithmicinput{\textbf{Input:}}
\algnewcommand\algorithmicoutput{\textbf{Output:}}
\algnewcommand\INPUT{\item[\algorithmicinput]}
\algnewcommand\OUTPUT{\item[\algorithmicoutput]}
%%%%%%%%%%%%%%%%%%%%%% figures %%%%%%%%%%%%%%%%%%%%%%%
\usepackage{tikz-qtree}
\usepackage{tikz}
\usepackage{tcolorbox}
\usetikzlibrary {arrows.meta,bending,positioning}
\usetikzlibrary{angles,quotes}
\usetikzlibrary{fit}
\usetikzlibrary{backgrounds}
\usetikzlibrary{shadows}
\usetikzlibrary{calc}
\usetikzlibrary{matrix}
\usepackage{pgfplots}
\usepackage{soul}
\usepackage{xcolor}
\usepackage{makecell}
\usepackage[normalem]{ulem}
\usepackage{ulem}
\usepackage{multirow}
\usepackage{setspace}
\usepackage{fontawesome5}
\usepackage{varwidth}
\usetikzlibrary{shadows}
\definecolor{ocrebase}{RGB}{243,102,25}
\definecolor{ocre}{RGB}{0,0,0}
\definecolor{amber}{rgb}{1.0, 0.75, 0.0}
\definecolor{ublue}{rgb}{0.152,0.250,0.545}
\definecolor{ugreen}{rgb}{0,0.5,0}
\definecolor{lgreen}{rgb}{0.9,1,0.8}
\definecolor{lightgreen}{rgb}{0.56, 0.93, 0.56}
\definecolor{kellygreen}{rgb}{0.3, 0.73, 0.09}
\definecolor{xtgreen}{rgb}{0.914,0.945,0.902}
\definecolor{lightgray}{gray}{0.85}
\definecolor{darkblue}{rgb}{0, 0, 0.5}
\definecolor{shadecolor}{rgb}{0.96,0.96,0.93}
\definecolor{lightcyan}{RGB}{194,232,247}
\definecolor{lightorange}{RGB}{255,226,187}
\definecolor{lightpink}{RGB}{252,224,225}
\definecolor{lightgreen}{RGB}{204,231,207}
% some colors
% from https://www.webdesignrankings.com/resources/lolcolors/
\definecolor{lolgreen}{RGB}{103,213,181}
\definecolor{lolred}{RGB}{238,119,133}
\definecolor{lolpurple}{RGB}{200,158,196}
\definecolor{lolblue}{RGB}{132,177,237}
\definecolor{lolorange}{RGB}{246,179,82}
\definecolor{lightsalmon}{RGB}{255,160,122}
\definecolor{lightskyblue}{RGB}{135,206,250}
%%%%%%%%%%%%%%%%%%%%%% new commands %%%%%%%%%%%%%%%%%%%%%%%
\newcommand{\ctext}[3][RGB]{%
\begingroup
\definecolor{hlcolor}{#1}{#2}\sethlcolor{hlcolor}%
\hl{#3}%
\endgroup
}
\newcommand{\mindex}[1]{\textbf{#1}\index{#1}}
\newcommand*\circled[1]{\tikz[baseline=(char.base)]{
\node[shape=circle,draw,inner sep=1pt] (char) {#1};}}
%%%%%%%%%%%%%%%%%%%%%% math %%%%%%%%%%%%%%%%%%%%%%%
\usepackage{amsmath}
\usepackage{amssymb}
\DeclareMathOperator*{\argmax}{arg\,max}
\DeclareMathOperator*{\argmin}{arg\,min}
\DeclareSymbolFont{EulerExtension}{U}{euex}{m}{n}%将积分号修改为正体
\DeclareMathSymbol{\euintop}{\mathop} {EulerExtension}{"52}
%\DeclareMathSymbol{\euointop}{\mathop} {EulerExtension}{"48}
\let\intop\euintop
%\let\ointop\euointop
\newcommand{\intoo}[2]{\mathopen{]}#1\,;#2\mathclose{[}}
\newcommand{\ud}{\mathop{\mathrm{{}d}}\mathopen{}}
\newcommand{\intff}[2]{\mathopen{[}#1\,;#2\mathclose{]}}
\definecolor{darkblue}{rgb}{0, 0, 0.5}
\hypersetup{colorlinks=true, citecolor=darkblue, linkcolor=darkblue, urlcolor=darkblue}
\title{Reinforcement Learning without Tears:\\ An Introduction for Large Language Model Researchers}
% Authors must not appear in the submitted version. They should be hidden
% as long as the \colmfinalcopy macro remains commented out below.
% Non-anonymous submissions will be rejected without review.
\author{Chenglong Wang, Hang Zhou, Tong Xiao \& Jingbo Zhu \\
% \thanks{ Use footnote for providing further information about author (webpage, alternative address)---\emph{not} for acknowledging funding agencies. Funding acknowledgements go at the end of the paper.} \\
School of Computer Science and Engineering\\
Northeastern University\\
Shenyang, China \\
\texttt{\{clwang1119,stceum\}@gmail.com}\\
\texttt{\{xiaotong,zhujingbo\}@mail.neu.edu.cn} \\
\AND
Tongran Liu \\
CAS Key Laboratory of Behavioral Science \\
Institute of Psychology \\
Chinese Academy of Sciences\\
Beijing, China \\
\texttt{liutr@psych.ac.cn} \\
}
% The \author macro works with any number of authors. There are two commands
% used to separate the names and addresses of multiple authors: \And and \AND.
%
% Using \And between authors leaves it to \LaTeX{} to determine where to break
% the lines. Using \AND forces a linebreak at that point. So, if \LaTeX{}
% puts 3 of 4 authors names on the first line, and the last on the second
% line, try using \AND instead of \And before the third author name.
\newcommand{\fix}{\marginpar{FIX}}
\newcommand{\new}{\marginpar{NEW}}
\begin{document}
\ifcolmsubmission
\linenumbers
\fi
\maketitle
\begin{abstract}
abstract.
\end{abstract}
\tableofcontents
\clearpage
\section{Introduction}
Reinforcement learning (RL) is an interdisciplinary field, and has its origins in both psychology and optimal control. Over the years, the concept has been further developed and formalized within the realms of artificial intelligence and machine learning. The basic idea is that an agent learns to make decisions by receiving feedback on its actions from the environment. This approach is very general and applies to a wide range of problems, from playing complex games like chess and Go to controlling autonomous vehicles.
A recent breakthrough is the large-scale application of RL in the field of natural language processing (NLP), in particular in the development of large language models (LLMs). For example, RL has been widely considered an effective way of aligning pre-trained LLMs with human preferences and values \citep{christiano-etal:2017deep}. One such method, called reinforcement learning from human feedback (RLHF), forms the basis for the development of advanced LLMs like OpenAI's GPT-4. More recently, RL has shown great promise in enabling LLMs to perform complex reasoning \citep{openai:2024learning,deepseek:2025r1}. As a result, interest in NLP has exploded. There have been many, many conference papers discussing RL in LLMs, and even more in the past few months. The research community is highly enthusiastic, and many in the field are anticipating a new turning point where large-scale RL algorithms could further advance AI.
However, RL was not so popular in the long history of NLP, and the field has just begun to explore its potential. Although applying RL to LLMs seems a natural choice for machine learning researchers, it introduces many new concepts to the NLP community, which has traditionally been dominated by supervised learning methodologies. As NLP researchers, when we attended presentations on LLMs, we often heard other researchers discuss advanced RL techniques, such as “use a value function to estimate the expected cumulative reward” or “learn a policy via PPO”. It seems that these techniques are so simple that presenters do not even need to give any explanation. But we were lost when seeing the math of RL, such as all those expectation signs, as well as the names of various algorithms that we are unfamiliar with.
Of course, we'd like to learn about RL, which appears straightforward. However, common RL textbooks, such as \citet{Sutton-and-Barto:2018RL}'s book, are mostly based on classic robotics or control problems. This makes it difficult to align the concepts of RL with those of LLMs. On the other hand, most LLM literature lacks an in-depth discussion of RL techniques, often omitting related details. As a result, we find ourselves in an awkward middle ground between RL and LLMs, struggling to bridge the two fields.
The aim of this paper is to provide a comprehensive introduction to RL from the perspective of LLMs. We begin by introducing basic concepts and algorithms of RL using the LLM language. In particular, we illustrate their application through an example of training LLMs, making the concepts more accessible. We then discuss a series of refinements to the basic RL framework for LLM alignment, including advanced reward modeling, improved advantage estimation, efficient sampling, and direct LLM optimization without RL. Furthermore, we discuss how to apply RL techniques to LLM reasoning. These methods are closely related to recent LLMs, such as OpenAI o1/o3 and Deepseek R1, which adopt large-scale RL and test-time scaling to significantly advance their reasoning abilities. In addition, we discuss applications of RL to multimodal LLMs to demonstrate how RL can be adapted to various problems. Finally, we conclude by outlining promising future research directions and discussing relevant systems and datasets.
% \clearpage
% New section: training LLMs with reinforcement learning
% \input{section2/section2}
% \clearpage
% New section: training LLMs with reinforcement learning
% \input{section3/section3}
% \clearpage
% New section: improved reinforcement learning for LLMs
\input{section4/section4}
% \clearpage
% New section: reinforcement learning for LLM inference
% \input{section5/section5}
% \clearpage
% New section: reinforcement learning for MLLMs
% \input{section6/section6}
% \clearpage
% conclusion
% \input{section7/section7}
% \clearpage
% systems & datasets todo: ganyang
% \input{section8/section8}
\clearpage
\bibliography{rl-introduction}
\bibliographystyle{rl-introduction}
% \appendix
\end{document}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\def\ssep{0.5cm}
\def\nsize{0.3cm}
\tikzstyle{lnode} = [minimum width=7cm,minimum height=1.1cm,inner sep=2pt,draw,thick,fill=white];
%%% step t
\begin{scope}
\node [anchor=west] (x0) at (0,0) {\scriptsize{$x_0$}};
\node [anchor=west] (x1) at ([xshift=\ssep]x0) {\scriptsize{$x_1$}};
\node [anchor=west] (x2) at ([xshift=\ssep]x1) {\scriptsize{$...$}};
\node [anchor=west] (x3) at ([xshift=\ssep]x2) {\scriptsize{$x_m$}};
\node [anchor=west] (y1) at ([xshift=\ssep]x3) {\scriptsize{$y_1$}};
\node [anchor=west] (y2) at ([xshift=\ssep]y1) {\scriptsize{$...$}};
\node [anchor=west] (y3) at ([xshift=\ssep]y2) {\scriptsize{$y_{t-1}$}};
%\node [anchor=west] (y4) at ([xshift=\ssep]y3) {\scriptsize{$y_{t}$}};
%\node [anchor=west] (y5) at ([xshift=\ssep]y4) {\scriptsize{$...$}};
\node [lnode,anchor=south] (llm) at ([xshift=0,yshift=0.3cm]y1.north) {\large{Policy (LLM)}};
\draw [->] ([yshift=0.18cm]x0.center) -- ([yshift=0.47cm]x0.center);
\draw [->] ([yshift=0.18cm]x1.center) -- ([yshift=0.47cm]x1.center);
\draw [->] ([yshift=0.18cm]x3.center) -- ([yshift=0.47cm]x3.center);
\draw [->] ([yshift=0.18cm]y1.center) -- ([yshift=0.47cm]y1.center);
\draw [->] ([yshift=0.18cm]y3.center) -- ([yshift=0.47cm]y3.center);
%\draw [->] ([yshift=0.18cm]y4.center) -- ([yshift=0.47cm]y4.center);
%\node [anchor=center] (ox0) at ([yshift=2.0cm]x0.north) {\scriptsize{\color{gray!50} $x_1$}};
%\node [anchor=center] (ox1) at ([yshift=2.0cm]x1.north) {\scriptsize{\color{gray!50} $x_2$}};
%\node [anchor=center] (ox2) at ([yshift=2.0cm]x2.north) {\scriptsize{\color{gray!50} $...$}};
\node [anchor=center] (ox3) at ([yshift=2.0cm]x3.north) {\scriptsize{\color{gray!50} $y_1$}};
\node [anchor=center] (oy1) at ([yshift=2.0cm]y1.north) {\scriptsize{\color{gray!50} $y_2$}};
\node [anchor=center] (oy2) at ([yshift=2.0cm]y2.north) {\scriptsize{\color{gray!50} $...$}};
\node [anchor=center] (oy3) at ([yshift=2.0cm]y3.north) {\scriptsize{$y_t$}};
%\node [anchor=center] (oy4) at ([yshift=2.0cm]y4.north) {\scriptsize{$y_{t+1}$}};
%\node [anchor=center] (oy5) at ([yshift=2.0cm]y5.north) {\scriptsize{$...$}};
%\draw [<-,gray!50] ([yshift=-0.18cm]ox0.center) -- ([yshift=-0.47cm]ox0.center);
%\draw [<-,gray!50] ([yshift=-0.18cm]ox1.center) -- ([yshift=-0.47cm]ox1.center);
\draw [<-,gray!50] ([yshift=-0.18cm]ox3.center) -- ([yshift=-0.47cm]ox3.center);
\draw [<-,gray!50] ([yshift=-0.18cm]oy1.center) -- ([yshift=-0.47cm]oy1.center);
\draw [<-] ([yshift=-0.18cm]oy3.center) -- ([yshift=-0.47cm]oy3.center);
%\draw [<-] ([yshift=-0.18cm]oy4.center) -- ([yshift=-0.47cm]oy4.center);
\draw [decoration={brace,amplitude=4pt},decorate] ([xshift=0,yshift=-0.2cm]y3.east) -- ([xshift=0,yshift=-0.2cm]x0.west) node [pos=0.5,below,yshift=-0.2cm,fill=lightsalmon!80] (statebox) {\scriptsize{State $s_t$ ($\mathbf{x}$ and $\mathbf{y}_{<t}$)}};
\node [anchor=south,fill=lightskyblue!80] (actionbox) at ([yshift=-0.0cm]oy3.north) {\scriptsize{Action $a_t$}};
\node [anchor=south west,minimum width=3.5cm,minimum height=1.3cm,draw,thick,align=center] (rewardmodel) at ([yshift=-1.4cm,xshift=2cm]llm.south east) {\footnotesize{Reward Model}\\[-0.0cm] \small{$R(\ $\ctext[RGB]{255,160,122}{$s_t$},\ \ctext[RGB]{135,206,250}{$a_t$}\ $)$}};
\node [anchor=south west,minimum width=3.5cm,minimum height=1.3cm,draw,thick,align=center] (valuefunction) at ([yshift=1.4cm]rewardmodel.north west) {\footnotesize{Value Functions}\\[-0.0cm] \small{$V($\ \ctext[RGB]{255,160,122}{$s_t$}\ $)$} \small{and} \small{$Q($\ \ctext[RGB]{255,160,122}{$s_t$},\ \ctext[RGB]{135,206,250}{$a_t$}\ $)$}};
\draw [->,thin] ([xshift=2pt]actionbox.east) -- ([xshift=2cm]actionbox.east) -- ([xshift=2cm,yshift=-3cm]actionbox.east) -- ([xshift=3.05cm,yshift=-3cm]actionbox.east);
\draw [->,thin] ([xshift=2pt]statebox.east) -- ([xshift=4.74cm]statebox.east);
\draw [->,thick] ([yshift=2pt]rewardmodel.north) -- ([yshift=-2pt]valuefunction.south);
\draw [->,thick,dotted] ([yshift=2pt]valuefunction.north) -- ([yshift=1.2cm]valuefunction.north) -- ([yshift=1.2cm,xshift=-5.6cm]valuefunction.north) node [pos=0.5,above] {\small{Feedback}} -- ([yshift=-0.17cm,xshift=-5.6cm]valuefunction.north);
\end{scope}
\end{tikzpicture}
\end{center}
\section{Preliminary}
In this section, we introduce the fundamentals of supervised fine-tuning for LLMs and focus on presenting the key concepts and nations necessary for the discussions in future sections.
Although pre-trained LLMs possess a vast amount of general knowledge, they are typically limited to tasks related to language modeling. Our ultimate goal is to enable them to perform various specific tasks, such as summarization, question-answer, and machine translation.
One straightforward method to achieve this goal is supervised fine-tuning (SFT), in which the pre-trained LLM is further trained on a dataset comprising task-specific input (i.e., instruction+user input)\footnote{In this paper, the \textit{instruction} represents the description of a specific task, e.g., ``Summarize the following article''; the \textit{user input} represents the extra content to the instruction, e.g., ``Article: In recent years, solar energy has seen a significant increase in the amount of energy available to the public. recent years, solar energy has seen unprecedented growth, becoming the fastest-growing ...''} paired with their expected outputs \citep{ouyang:2022training,wei-etal:2022finetuned}.
Specifically, let $\mathbf{x}=x_0...x_m$ be an input (e.g., instruction + user input) and $\mathbf{y}=y_1...y_n$ be the corresponding output. In SFT, we aim to maximize the probability of the output $\mathbf{y}$ given the input $\mathbf{x}$. Consider an LLM with pre-trained parameters $\hat{\theta}$. The fine-tuning objective can then be formulated as:
\begin{eqnarray}
\tilde{\theta} & = & \argmax_{\hat{\theta}^+} \sum_{(\mathbf{x},\mathbf{y}) \in \mathcal{S}} \log \mathrm{Pr}_{\hat{\theta}^+}(\mathbf{y}|\mathbf{x}) \label{eq:instruction-fine-tuning}
\end{eqnarray}
where $\Pr(\cdot)$ denotes the probability distribution, $\mathcal{S}$ denotes the labeled data, $\tilde{\theta}$ denotes the parameters optimized via SFT, and $\hat{\theta}^+$ represents an adjustment to $\hat{\theta}$. Here, we will omit the superscript $+$ and use $\theta$ to represent $\hat{\theta}^+$ to keep the notation uncluttered. However, the reader should remember that the fine-tuning starts from the pre-trained parameters rather than randomly initialized ones.
The objective function $\log \mathrm{Pr}_{\theta}(y_i|\mathbf{x},\mathbf{y}_{<i})$ is computed by summing the log-probabilities of the tokens in $\mathbf{y}$, conditional on the input $\mathbf{x}$ and all the previous tokens $\mathbf{y}_{<i}$:
\begin{eqnarray}
\log \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) & = & \sum_{i=1}^{n} \log \mathrm{Pr}_{\theta}(y_i|\mathbf{x},\mathbf{y}_{<i})
\end{eqnarray}
This formulation is equivalent to minimizing the cross-entropy loss. In the following section, we refer to the LLM after SFT as the ``SFT LLM'' for short. Beyond SFT, the model architecture design and the pre-training process are also fundamentals of building an LLM. However, we will not discuss these topics. Interested readers can find further details in \citet{xiao-and-zhu:2025foundations}'s book.
%%%%%%%%%% Preliminary修改建议 %%%%%%%%%%
% 1. 再加一些强化学习的基础知识
% 包含对环境、动作、policy这些术语的定义,方面后面对应起来的好定义。
% 如果修改这一章节 那么3.1这部分要对应起来,有些需要删除掉。
\begin{tabular}{ll}
\toprule[1.1pt]
\multicolumn{1}{c}{\textbf{Element in Reinforcement Learning}} & \multicolumn{1}{c}{\textbf{Counterpart in LLMs}} \\ \midrule
\parbox{4cm}{
Action ($a$)
} & \parbox{7.2cm}{
The possible predicted tokens in the vocabulary of the LLM.
} \\ \midrule
Reward ($R$) & \\
& \\
& \\
& \\
\bottomrule[1.1pt]
\end{tabular}
\ No newline at end of file
\begin{center}
\def\tgtwidth{0.8\textwidth}
\begin{tikzpicture}
\tikzstyle{ynode} = [draw,rounded corners=4pt, minimum width=\tgtwidth,inner sep=4pt,align=left,text width=\tgtwidth]
\begin{scope}
\node (input) at (0,0) {$\mathbf{x}$};
\node [ynode, anchor=west] (note) at ([xshift=.2cm]input.east) {
\begin{varwidth}{\textwidth}
\scriptsize{Give me three tips to improve my accuracy in solving math problems.}
\end{varwidth}
};
\node [ynode,anchor=north west] (output1) at ([yshift=-.2cm]note.south west) {
\scriptsize{Improving accuracy in solving math problems is crucial for success in mathematics. Here are three tips that can help: \\
1.Practice regularly... \\
2.Understand the concepts... \\
\ctext[RGB]{255,204,204}{3.Double-check your work...any-one can significantly enhance their math problem-solving abilities.} \\
\textbf{\{Total Tokens: 160, Reward: 16\}}}
};
\node [anchor=east](y1) at ([xshift=-.2cm]output1.west) {$\mathbf{y}_1$};
\node [ynode,anchor=north] (output2) at ([yshift=-.2cm]output1.south) {
\scriptsize{Improving math accuracy requires careful attention to detail, strategic problem-solving techniques, and consistent practice. Here are three tips: \\
1.Avoid rushing and read carefully... \\
2.Use estimation to validate answers... \\
\ctext[RGB]{255,204,204}{3.Double-check your work...any-one can significantly enhance their math problem-solving abilities.} \\
\textbf{\{Total Tokens: 750, Reward: 75\}}}
};
\node [anchor=east](y2) at ([xshift=-.2cm]output2.west) {$\mathbf{y}_2$};
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\def\ssep{1cm}
\def\nsize{0.3cm}
\tikzstyle{lnode} = [minimum width=7cm,minimum height=1.1cm,inner sep=2pt,draw,thick,fill=white];
\begin{scope}
\node [anchor=west] (x0) at (0,0) {\footnotesize{$x_0$}};
\node [anchor=center] (x1) at ([xshift=\ssep]x0.center) {\footnotesize{$x_1$}};
\node [anchor=center] (x2) at ([xshift=\ssep]x1.center) {\footnotesize{$x_2$}};
\node [anchor=center] (x3) at ([xshift=\ssep]x2.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (x4) at ([xshift=\ssep]x3.center) {\footnotesize{$x_m$}};
\node [anchor=center] (y1) at ([xshift=\ssep]x4.center) {\footnotesize{$y_1$}};
\node [anchor=center] (y2) at ([xshift=\ssep]y1.center) {\footnotesize{$y_2$}};
\node [anchor=center] (y3) at ([xshift=\ssep]y2.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (y4) at ([xshift=\ssep]y3.center) {\footnotesize{$y_n$}};
\node [anchor=north] (y4label) at ([yshift=0.1cm]y4.south) {\scriptsize{(Last Token $\langle \mathrm{EOS} \rangle$)}};
\draw [->] ([yshift=0.2cm]x0.center) -- ([yshift=0.55cm]x0.center);
\draw [->] ([yshift=0.2cm]x1.center) -- ([yshift=0.55cm]x1.center);
\draw [->] ([yshift=0.2cm]x2.center) -- ([yshift=0.55cm]x2.center);
\draw [->] ([yshift=0.2cm]x4.center) -- ([yshift=0.55cm]x4.center);
\draw [->] ([yshift=0.2cm]y1.center) -- ([yshift=0.55cm]y1.center);
\draw [->] ([yshift=0.2cm]y2.center) -- ([yshift=0.55cm]y2.center);
\draw [->] ([yshift=0.2cm]y4.center) -- ([yshift=0.55cm]y4.center);
\node [anchor=center] (ox0) at ([yshift=2.6cm]x0.center) {\footnotesize{$\mathbf{h}_{x_0}$}};
\node [anchor=center] (ox1) at ([yshift=2.6cm]x1.center) {\footnotesize{$\mathbf{h}_{x_1}$}};
\node [anchor=center] (ox2) at ([yshift=2.6cm]x2.center) {\footnotesize{$\mathbf{h}_{x_2}$}};
\node [anchor=center] (ox3) at ([yshift=2.6cm]x3.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (ox4) at ([yshift=2.6cm]x4.center) {\footnotesize{$\mathbf{h}_{x_m}$}};
\node [anchor=center] (oy1) at ([yshift=2.6cm]y1.center) {\footnotesize{$\mathbf{h}_{y_1}$}};
\node [anchor=center] (oy2) at ([yshift=2.6cm]y2.center) {\footnotesize{$\mathbf{h}_{y_2}$}};
\node [anchor=center] (oy3) at ([yshift=2.6cm]y3.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (oy4) at ([yshift=2.6cm]y4.center) {\footnotesize{$\mathbf{h}_{\mathrm{last}}$}};
\node [anchor=south,draw,thick,minimum width=9cm,minimum height=1.3cm] (llm) at ([yshift=0.4cm]x4.north) {\large{Transformer Decoder (LLM)}};
\node [anchor=east] (representation) at ([xshift=-0.3cm]ox0.west) {\scriptsize{Representation}};
\node [anchor=north west] (representation2) at ([yshift=0.1cm]representation.south west) {\scriptsize{at Each Position}};
\draw [<-] ([yshift=-0.25cm]ox0.center) -- ([yshift=-0.6cm]ox0.center);
\draw [<-] ([yshift=-0.25cm]ox1.center) -- ([yshift=-0.6cm]ox1.center);
\draw [<-] ([yshift=-0.25cm]ox2.center) -- ([yshift=-0.6cm]ox2.center);
\draw [<-] ([yshift=-0.25cm]ox4.center) -- ([yshift=-0.6cm]ox4.center);
\draw [<-] ([yshift=-0.25cm]oy1.center) -- ([yshift=-0.6cm]oy1.center);
\draw [<-] ([yshift=-0.25cm]oy2.center) -- ([yshift=-0.6cm]oy2.center);
\draw [<-] ([yshift=-0.25cm]oy4.center) -- ([yshift=-0.6cm]oy4.center);
\filldraw [fill=red!20,draw=white] ([xshift=-0.5cm,yshift=0.2cm]oy4.north west) -- ([xshift=0.5cm,yshift=0.2cm]oy4.north east) -- ([yshift=1cm,xshift=0]oy4.north east) -- ([yshift=1cm,xshift=0]oy4.north west) -- ([xshift=-0.5cm,yshift=0.2cm]oy4.north west);
\node [anchor=south] (reward) at ([yshift=1.2cm]oy4.north) {\footnotesize{Reward (Scalar)}};
\node [anchor=south] (Wr) at ([yshift=0.3cm]oy4.north) {\footnotesize{$\mathbf{W}_r$}};
\node [anchor=west] (linear) at ([xshift=0.5cm]Wr.east) {\scriptsize{Linear Map}};
\draw [->] ([yshift=-0.1cm]oy4.north) -- ([yshift=0.18cm]oy4.north);
\draw [<-] ([yshift=0.1cm]reward.south) -- ([yshift=-0.18cm]reward.south);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\def\ssep{1.5cm}
\def\nsize{0.3cm}
\tikzstyle{lnode} = [minimum width=7cm,minimum height=1.1cm,inner sep=2pt,draw,thick,fill=white];
\begin{scope}
\node [anchor=center] (x1) at (0,0) {\footnotesize{$x_1$}};
\node [anchor=center] (xcdots) at ([xshift=\ssep]x1.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (xm) at ([xshift=\ssep]xcdots.center) {\footnotesize{$x_m$}};
\node [anchor=center] (y1) at ([xshift=\ssep]xm.center) {\footnotesize{$y_1$}};
\node [anchor=center] (ycdots) at ([xshift=\ssep]y1.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (yt) at ([xshift=\ssep]ycdots.center) {\footnotesize{$y_{T}$}};
\foreach \x in {x1, xm, y1, yt}
\draw [->] ([yshift=0.2cm]\x.center) -- ([yshift=0.55cm]\x.center);
\node [anchor=center] (oy1) at ([yshift=2.6cm]y1.center) {\footnotesize{$\mathbf{h}_{y_1}$}};
\node [anchor=center] (oycdots) at ([yshift=2.6cm]ycdots.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (oyt) at ([yshift=2.6cm]yt.center) {\footnotesize{$\mathbf{h}_{y_T}$}};
\node [anchor=south,draw,thick,minimum width=9cm,minimum height=1.3cm] (llm) at ([yshift=0.4cm]$(xm.north)!.5!(y1.north)$) {\large{Value Model (LLM)}};
\foreach \x in {oy1, oyt}
\draw [<-] ([yshift=-0.25cm]\x.center) -- ([yshift=-0.6cm]\x.center);
\filldraw [fill=red!20,draw=white] ([xshift=-0.5cm,yshift=0.2cm]oyt.north west) -- ([xshift=0.5cm,yshift=0.2cm]oyt.north east) -- ([yshift=1cm,xshift=0]oyt.north east) -- ([yshift=1cm,xshift=0]oyt.north west) -- ([xshift=-0.5cm,yshift=0.2cm]oyt.north west);
\filldraw [fill=red!20,draw=white] ([xshift=-0.5cm,yshift=0.2cm]oy1.north west) -- ([xshift=0.5cm,yshift=0.2cm]oy1.north east) -- ([yshift=1cm,xshift=0]oy1.north east) -- ([yshift=1cm,xshift=0]oy1.north west) -- ([xshift=-0.5cm,yshift=0.2cm]oy1.north west);
\node at ([yshift=0.6cm]$(oy1.north)!.5!(oyt.north)$) {\footnotesize{$\cdots$}};
\node [anchor=south] (Wr) at ([yshift=0.3cm]oyt.north) {\footnotesize{$\mathbf{W}_r$}};
\node [anchor=south] (output) at ([yshift=0.5cm]Wr.north) {\footnotesize{Value (Scalar)}};
\draw [->] ([yshift=0.2cm]Wr.north) -- (output.south);
\node [anchor=west] (linear) at ([xshift=0.5cm]Wr.east) {\scriptsize{Linear Map}};
\node [anchor=south] (Wr) at ([yshift=0.3cm]oy1.north) {\footnotesize{$\mathbf{W}_r$}};
\node [anchor=south] (output) at ([yshift=0.5cm]Wr.north) {\footnotesize{Value (Scalar)}};
\draw [->] ([yshift=0.2cm]Wr.north) -- (output.south);
\node [anchor=east,align=center] at ([xshift=-0.0cm]oy1.west) {\scriptsize{Representation}};
\draw [->] ([yshift=-0.1cm]oyt.north) -- ([yshift=0.18cm]oyt.north);
\draw [->] ([yshift=-0.1cm]oy1.north) -- ([yshift=0.18cm]oy1.north);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\tikzset{
block/.style = {
minimum width=3.5cm,minimum height=1.5cm,
draw=#1,fill=white,line width=2pt,
drop shadow={fill=gray,shadow xshift=.6ex,shadow yshift=-.6ex}
},
block/.default=black
}
\def\sep{1.5cm}
\def\textwid{7.0cm}
\begin{scope}
\node [anchor=east] at (0,0) {\large{$\cdots$}};
\node [anchor=west, block=lolred] (llm) at (\sep/2,0) {Reference Policy (LLM)};
\draw [->] (0,0) -- (\sep/2,0);
\node [anchor=north] (x) at ([yshift=-.5cm]llm.south) {\large{$\mathbf{x}$}};
\node [draw, dashed, rounded corners=2pt,anchor=south, align=left, text width=\textwid, minimum height=2.6cm] (output) at ([yshift=.5cm]llm.north) {\footnotesize{There are three tips }\footnotesize{\sethlcolor{red!20}\hl{for improving for improving for improving} \footnotesize{accuracy in solving math problems:}} \\
\footnotesize{1.Practice regularly. \\
2.Understand the concepts. \\
3.Double-check your work.}};
\node [draw, dashed, rounded corners=2pt,anchor=north, align=left, text width=\textwid] (input) at ([yshift=0cm]x.south) {\footnotesize{Give me three tips to improve my accuracy in solving math problems.}};
\node [anchor=south] (y) at ([yshift=0cm]output.north) {\large{$\mathbf{y}$}};
\node [align=left] (step) at ([yshift=-.5cm]input.south) {\footnotesize{sampling in a terrible step}};
\draw [->] (x.north) -- ([yshift=-.1cm]llm.south);
\draw [<-] ([yshift=-.1cm]output.south) -- ([yshift=.1cm]llm.north);
%%
\node [block=lolred,anchor=west] (llm1) at ([xshift=\sep*3]llm.east) {Reference Policy (LLM)};
\node [anchor=north] (x) at ([yshift=-.5cm]llm1.south) {\large{$\mathbf{x}$}};
\node [draw, dashed, rounded corners=2pt,anchor=south, align=left, text width=\textwid, minimum height=2.6cm] (output) at ([yshift=.5cm]llm1.north) {\footnotesize{There are three tips }\footnotesize{\sethlcolor{red!20}\hl{for improving improving improving improving improving improving improving improving } \footnotesize{accuracy in solving math problems:}} \\
\footnotesize{1.Practice }\sethlcolor{red!20}\hl{regularly regularly regularly regularly regularly regularly regularly }. \footnotesize{...}\\};
\node [draw, dashed, rounded corners=2pt,anchor=north, align=left, text width=\textwid] (input) at ([yshift=0cm]x.south) {\footnotesize{Give me three tips to improve my accuracy in solving math problems.}};
\node [anchor=south] (y) at ([yshift=0cm]output.north) {\large{$\mathbf{y}$}};
\node [align=left] (step) at ([yshift=-.5cm]input.south) {\footnotesize{sampling after a terrible step}};
\draw [->] (x.north) -- ([yshift=-.1cm]llm1.south);
\draw [<-] ([yshift=-.1cm]output.south) -- ([yshift=.1cm]llm1.north);
% \node [anchor=north, align=left] (step) at (llm1.south |- step.north) {\footnotesize{xxxxxx}};
\node [anchor=west] (llm2) at ([xshift=\sep/2]llm1.east) {\large{$\cdots$}};
\draw [->] (llm.east) -- node [midway, align=center] {\footnotesize{resync the reference policy} \\\footnotesize{after a terrible step}} (llm1.west);
\draw [->] (llm1.east) -- (llm2.west);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\tikzset{
block/.style = {
minimum width=3.5cm,minimum height=1.5cm,
draw=#1,fill=white,line width=2pt,
drop shadow={fill=gray,shadow xshift=.6ex,shadow yshift=-.6ex}
},
block/.default=black
}
\def\sep{1.3cm}
\begin{scope}
\node [block=lolblue] (llm) at (0,0) {Policy (LLM)};
\node [anchor=south, block=lolred] (old llm) at ([yshift=1cm]llm.north) {Reference Policy (LLM)};
\draw [->] ([yshift=.1cm]llm.north) -- node [midway,right] {copy} ([yshift=-.1cm]old llm.south);
\node [align=left] (step) at ([yshift=-.6cm]llm.south) {\scriptsize{Step (1): resync the reference policy}};
%%
\node [block=lolred,anchor=west] (old llm1) at ([xshift=\sep]$(old llm.east)!.5!(llm.east)$) {Reference Policy (LLM)};
\node [anchor=north] (x) at ([yshift=-.5cm]old llm1.south) {\large{$\mathbf{x}$}};
\node [anchor=south] (y) at ([yshift=.5cm]old llm1.north) {\large{$\mathbf{y}$}};
\draw [->] (x.north) -- ([yshift=-.1cm]old llm1.south);
\draw [<-] (y.south) -- ([yshift=.1cm]old llm1.north);
\node [anchor=north, align=left] (step) at (old llm1.south |- step.north) {\scriptsize{Step (2): sample an output from}\\ \scriptsize{the reference policy}};
%%
\node [block=lolblue,anchor=west] (llm2) at ([xshift=\sep]old llm1.east) {Policy (LLM)};
\node [anchor=north] (x) at ([xshift=-1.5cm, yshift=-.5cm]llm2.south) {\large{$\mathbf{x}$}};
\node [anchor=north] (y) at ([xshift=-1.1cm, yshift=-.5cm]llm2.south) {\large{$\mathbf{y}$}};
\node [align=center] at ([yshift=-.3cm]$(x.south)!.5!(y.south)$) {\scriptsize{1st time}\\[-2pt]\scriptsize{update}};
\draw [->] (x.north) -- ([yshift=-.1cm]x.north|-llm2.south);
\draw [->] (y.north) -- ([yshift=-.1cm]y.north|-llm2.south);
\node [anchor=north] (x) at ([xshift=-.2cm, yshift=-.5cm]llm2.south) {\large\textcolor{black!40!white}{$\mathbf{x}$}};
\node [anchor=north] (y) at ([xshift=.2cm, yshift=-.5cm]llm2.south) {\large\textcolor{black!40!white}{$\mathbf{y}$}};
\node [align=center] at ([yshift=-.3cm]$(x.south)!.5!(y.south)$) {\scriptsize{2nd time}\\[-2pt]\scriptsize{update}};
\draw [->,black!40!white] (x.north) -- ([yshift=-.1cm]x.north|-llm2.south);
\draw [->,black!40!white] (y.north) -- ([yshift=-.1cm]y.north|-llm2.south);
\node [anchor=north] (cdots) at ([xshift=1.3cm, yshift=-.5cm]llm2.south) {\large$\cdots$};
\node [align=center] at ([yshift=-.3cm]cdots.south) {\large$\cdots$};
\node [anchor=north,align=left] (step) at (llm2.south |- step.north) {\scriptsize{Step (3): update the policy multiple}\\ \scriptsize{times and go to Step (1)}};
\end{scope}
\end{tikzpicture}
\end{center}
\begin{center}
\def\tgtwidth{0.5\textwidth}
\begin{tikzpicture}
\tikzstyle{ynode} = [draw,rounded corners=4pt, minimum width=\tgtwidth,inner sep=4pt,align=left,text width=\tgtwidth]
\begin{scope}
\node (input) at (0,0) {$\mathbf{x}$};
\node [ynode, anchor=west] (note) at ([xshift=.2cm]input.east) {
\begin{varwidth}{\textwidth}
\scriptsize{Give me three tips to improve my accuracy in solving math problems.}
\end{varwidth}
};
\node [ynode,anchor=north west] (output1) at ([yshift=-.2cm]note.south west) {
\scriptsize{There are three tips for improving accuracy in solving math problems: \\
1.Practice regularly. \\
2.Understand the concepts. \\
3.Double-check your work. \\
\textbf{\{Total Tokens: 20\}}}
};
\node [anchor=east](y1) at ([xshift=-.2cm]output1.west) {$\mathbf{y}_1$};
\node [ynode,anchor=north] (output2) at ([yshift=-.2cm]output1.south) {
\scriptsize{Improving accuracy in solving math problems is crucial for success in mathematics. Here are three tips that can help: \\
1.Practice regularly: \\
~~~- Consistent practice builds fluency and familiarity with different problem types. \\
... \\
By combining your advice with these additional tips, anyone can significantly enhance their math problem-solving abilities. \\
\textbf{\{Total Tokens: 320\}}}
};
\node [anchor=east](y2) at ([xshift=-.2cm]output2.west) {$\mathbf{y}_2$};
\node [ynode,anchor=north] (output3) at ([yshift=-.2cm]output2.south) {
\scriptsize{Improving math accuracy requires careful attention to detail, strategic problem-solving techniques, and consistent practice. Here are three tips: \\
1.Practice regularly: \\
~~~- Mathematics is a skill, and like any skill, it improves with consistent practice.... \\
... \\
This is a great approach for anyone looking to improve their math skills, from students to professionals. \\
\textbf{\{Total Tokens: 800\}}}
};
\node [anchor=east](y3) at ([xshift=-.2cm]output3.west) {$\mathbf{y}_3$};
\begin{axis}[
compat=1.3,
anchor=north west,
at={(output1.north east)},
axis x line=bottom,axis y line=left,
xshift=1.4cm,
yshift=-1cm,
xtick={0,1,2,3},
xticklabels={$\mathbf{y}_1$,$\mathbf{y}_2$,$\mathbf{y}_3$,$\cdots$},
every tick label/.append style={font=\footnotesize},
every axis label/.append style={font=\footnotesize},
xtick distance=.1cm,
ymin=0,
ymax=.6,
ylabel=$\mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})$,
ylabel shift=-.1cm,
enlarge x limits=0.3,
width=5cm,
height=4cm,
ybar]
\addplot[fill=lolblue,draw=none] coordinates {
(0, .5)
(1, .3)
(2, .1)
(3, .1)
};
\end{axis}
\begin{axis}[
compat=1.3,
anchor=south west,
at={(output3.south east)},
axis x line=bottom,axis y line=left,
xshift=1.4cm,
yshift=.5cm,
xtick={0,1,2,3},
xticklabels={$\mathbf{y}_1$,$\mathbf{y}_2$,$\mathbf{y}_3$,$\cdots$},
every tick label/.append style={font=\footnotesize},
every axis label/.append style={font=\footnotesize},
xtick distance=.1cm,
ymin=0,
ymax=.6,
ylabel=$\mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})$,
ylabel shift=-.1cm,
enlarge x limits=0.3,
width=5cm,
height=4cm,
ybar]
\addplot [fill=lolblue,draw=none] coordinates {
(0, .1)
(1, .3)
(2, .5)
(3, .1)
};
\end{axis}
\matrix [
matrix anchor=north,
matrix of nodes,
nodes={inner sep=2pt},
column sep=.1cm,
inner sep=0,
] at ([xshift=3cm,yshift=0cm]note.north east) {
\hspace{-0.25cm} \footnotesize{$R(\mathbf{y}_1)=2$} \\
\footnotesize{$R(\mathbf{y}_2)=32$} \\
\footnotesize{$R(\mathbf{y}_3)=80$} \\
};
\node at ([xshift=2.03cm,yshift=-1.2cm]output1.north east) {\color{lolred}{\faTimes}};
\node at ([xshift=3.5cm,yshift=-2.85cm]output1.north east) {\color{lolgreen}{\faCheck}};
\node at ([xshift=2.03cm,yshift=1.1cm]output3.south east) {\color{lolred}{\faTimes}};
\node at ([xshift=3.5cm,yshift=+2.70cm]output3.south east) {\color{lolgreen}{\faCheck}};
\draw [->,thick] ([xshift=3cm,yshift=-4cm]output1.north east) -- node [right,align=left] {
\scriptsize{Increase the probability} \\[-1mm]
\scriptsize{of longer outputs;} \\ [-1mm]
\scriptsize{Decrease the probability} \\[-1mm]
\scriptsize{of shorter outputs}
}([xshift=3cm,yshift=3.2cm]output3.south east);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\def\ssep{.7cm}
\def\nsize{0.3cm}
\tikzstyle{lnode} = [minimum width=7cm,minimum height=1.1cm,inner sep=2pt,draw,thick,fill=white];
\begin{scope}
\node [draw,thick,minimum width=5.2cm,minimum height=1.3cm,fill=white] (llm) at (0,0) {Reward Model (LLM)};
\node [anchor=south east] (oy last) at ([yshift=.5cm,xshift=0cm]llm.north east) {\ctext[RGB]{153,204,255}{$R_{\phi}(\mathbf{x},\mathbf{y}_a)$}};
\draw [->] ([yshift=.1cm]llm.north-|oy last.south) -- (oy last.south);
\node [draw,rounded corners=2pt,inner sep=2pt,text width=5.5em, anchor=north west,font=\linespread{0.8}\selectfont, minimum height=9ex] (x) at ([yshift=-1cm,xshift=-1cm]llm.south west) {\scriptsize{Give me three tips to improve my accuracy in solving math problems.\\}};
\node [draw,rounded corners=2pt,inner sep=2pt,text width=13.5em, anchor=north east,font=\linespread{0.8}\selectfont, minimum height=9ex] (y) at ([yshift=-1cm,xshift=1cm]llm.south east) {\scriptsize{Improving math accuracy requires careful attention to detail, ... \\
1.Practice regularly: ... \\
... , anyone can significantly enhance their math problem-solving abilities.\\}};
\node [anchor=north] at ([yshift=-.2cm]x.south) {$\mathbf{x}$};
\node [anchor=north] at ([yshift=-.2cm]y.south) {$\mathbf{y}_a$};
\draw [<-] ([yshift=0cm,xshift=-.5cm]llm.south) .. controls +(0, -.95cm) and +(0,+.95cm) .. +(-2cm, -.95cm);
\draw [<-] ([yshift=0cm,xshift=.5cm]llm.south) .. controls +(0, -.95cm) and +(0,+.95cm) .. +(+2cm, -.95cm);
%%
\node [draw,thick,minimum width=5.2cm,minimum height=1.3cm,fill=white] (llm) at (+8cm,0) {Reward Model (LLM)};
\node [anchor=south east] (oy last1) at ([yshift=.5cm,xshift=0cm]llm.north east) {\ctext[RGB]{255,204,204}{$R_{\phi}(\mathbf{x},\mathbf{y}_b)$}};
\draw [->] ([yshift=.1cm]llm.north-|oy last1.south) -- (oy last1.south);
\node [draw,rounded corners=2pt,inner sep=2pt,text width=5.5em, anchor=north west,font=\linespread{0.8}\selectfont, minimum height=9ex] (x) at ([yshift=-1cm,xshift=-1cm]llm.south west) {\scriptsize{Give me three tips to improve my accuracy in solving math problems.\\}};
\node [draw,rounded corners=2pt,inner sep=2pt,text width=13.5em, anchor=north east,font=\linespread{0.8}\selectfont, minimum height=9ex] (y) at ([yshift=-1cm,xshift=1cm]llm.south east) {\scriptsize{There are three tips for improving accuracy in solving math problems: \\
1.Practice regularly. \\
2.Understand the concepts.\\
3.Double-check your work.\\}};
\node [anchor=north] at ([yshift=-.2cm]x.south) {$\mathbf{x}$};
\node [anchor=north] at ([yshift=-.2cm]y.south) {$\mathbf{y}_b$};
\draw [<-] ([yshift=0cm,xshift=-.5cm]llm.south) .. controls +(0, -.95cm) and +(0,+.95cm) .. +(-2cm, -.95cm);
\draw [<-] ([yshift=0cm,xshift=.5cm]llm.south) .. controls +(0, -.95cm) and +(0,+.95cm) .. +(+2cm, -.95cm);
\node at (4.65cm,3.3cm) {$- \log \mathrm{Sigmoid}($\ctext[RGB]{153,204,255}{$R_{\phi}(\mathbf{x},\mathbf{y}_a)$}$-$\ctext[RGB]{255,204,204}{$R_{\phi}(\mathbf{x},\mathbf{y}_b)$}$)$};
\draw [->] (oy last.north) .. controls +(0, 1cm) and +(0, -1cm) .. ([xshift=3.1cm, yshift=1.25cm]oy last.north);
\draw [->] (oy last1.north) .. controls +(0, 1cm) and +(0, -1cm) .. ([xshift=-3.1cm, yshift=1.25cm]oy last1.north);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\begin{scope}
\begin{axis}[
compat=1.3,
anchor=south east,
at={(0,0)},
axis x line=bottom,axis y line=left,
xtick={0,1,1.2,2},
xticklabels={,1,$1+\epsilon$,},
ytick=\empty,
every tick label/.append style={font=\scriptsize},
every axis label/.append style={font=\scriptsize},
xtick distance=.1cm,
ymin=0,
ymax=3,
enlarge x limits=0.0,
width=6.4cm,
height=6.5cm,
clip=false,]
\addplot [mark=none,thick] coordinates {
(0, 0)
(1, 1.5)
(1.2, 1.8)
(2, 1.8)
};
\addplot [mark=none,dashed,line width=0.5pt] coordinates {
(1, 0)
(1, 3)
};
\addplot [mark=none,dashed,lolred,line width=1.5pt] coordinates {
(1.2, 0)
(1.2, 3)
};
\node [anchor=south west] at (0,4.8cm) {$\textrm{Clip}(\cdot)$};
\node [anchor=north] at (4.8cm, 0cm) {\footnotesize{ratio}};
\node [anchor=north] at (0,0) {0};
\node (caption a) at (xticklabel cs:0.5) {};
\end{axis}
\node (tmp) at (0,0) {};
\node [anchor=north] (a) at ([yshift=-.8cm]caption a|-tmp) {\footnotesize{(a) $A_t>0$}};
\begin{axis}[
compat=1.3,
anchor=south west,
at={(+3cm,0)},
axis x line=top,axis y line=left,
xtick={0,0.8,1,2},
xticklabels={,$1-\epsilon$,1,},
ytick=\empty,
every tick label/.append style={font=\scriptsize},
every axis label/.append style={font=\scriptsize},
xtick distance=.1cm,
ymin=0,
ymax=3,
enlarge x limits=0.0,
width=6.4cm,
height=6.5cm,
y dir=reverse,
clip=false,]
\addplot [mark=none,thick] coordinates {
(0, 1.2)
(0.8, 1.2)
(1.0, 1.5)
(1.8, 2.7)
};
\addplot [mark=none,dashed,lolred,line width=1.5pt] coordinates {
(0.8, 0)
(0.8, 3)
};
\addplot [mark=none,dashed,line width=0.5pt] coordinates {
(1, 0)
(1, 3)
};
\node [anchor=south west] at (0,-5.4cm) {$\textrm{Clip}(\cdot)$};
\node [anchor=south] at (4.8cm, 0) {\footnotesize{ratio}};
\node [anchor=south] at (0,0) {0};
\node (caption b) at (xticklabel cs:0.5) {};
\end{axis}
\node [anchor=north] (b) at ([yshift=-.8cm]caption b|-tmp) {\footnotesize{(b) $A_t<0$}};
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\def\ssep{1.5cm}
\def\nsize{0.3cm}
\tikzstyle{lnode} = [minimum width=7cm,minimum height=1.1cm,inner sep=2pt,draw,thick,fill=white];
\begin{scope}
\node [anchor=center] (x1) at (0,0) {\footnotesize{$x_1$}};
\node [anchor=center] (xcdots) at ([xshift=\ssep]x1.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (xm) at ([xshift=\ssep]xcdots.center) {\footnotesize{$x_m$}};
\node [anchor=center] (y1) at ([xshift=\ssep]xm.center) {\footnotesize{$y_1$}};
\node [anchor=center] (ycdots) at ([xshift=\ssep]y1.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (yt) at ([xshift=\ssep]ycdots.center) {\footnotesize{$y_{T}$}};
\foreach \x in {x1, xm, y1, yt}
\draw [->] ([yshift=0.2cm]\x.center) -- ([yshift=0.55cm]\x.center);
\node [anchor=center] (oyt) at ([yshift=2.6cm]yt.center) {\footnotesize{$\mathbf{h}_{y_T}$}};
\node [anchor=south,draw,thick,minimum width=9cm,minimum height=1.3cm] (llm) at ([yshift=0.4cm]$(xm.north)!.5!(y1.north)$) {\large{Reward Model (LLM)}};
\foreach \x in {oyt}
\draw [<-] ([yshift=-0.25cm]\x.center) -- ([yshift=-0.6cm]\x.center);
\filldraw [fill=red!20,draw=white] ([xshift=-0.5cm,yshift=0.2cm]oyt.north west) -- ([xshift=0.5cm,yshift=0.2cm]oyt.north east) -- ([yshift=1cm,xshift=0]oyt.north east) -- ([yshift=1cm,xshift=0]oyt.north west) -- ([xshift=-0.5cm,yshift=0.2cm]oyt.north west);
\node [anchor=south] (Wr) at ([yshift=0.3cm]oyt.north) {\footnotesize{$\mathbf{W}_r$}};
\node [anchor=south] (output) at ([yshift=0.5cm]Wr.north) {\footnotesize{Reward (Scalar)}};
\node [anchor=west] (linear) at ([xshift=0.5cm]Wr.east) {\scriptsize{Linear Map}};
\node [anchor=west, align=left,xshift=-0.4cm] at (linear.west|-oyt.east) {\scriptsize{Representation}};
\draw [->] ([yshift=0.2cm]Wr.north) -- (output.south);
\draw [->] ([yshift=-0.1cm]oyt.north) -- ([yshift=0.18cm]oyt.north);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\tikzstyle{model} = [minimum width=2cm, minimum height=1.2cm,align=center,inner sep=2pt,draw,thick,drop shadow={fill=gray,shadow xshift=.6ex,shadow yshift=-.6ex},fill=white]
\tikzstyle{label} = [minimum width=2ex, minimum height=1.5ex,draw,scale=.75,fill=white]
\tikzstyle{circled} = [shape=circle,draw,inner sep=2pt]
\begin{scope}
\node [anchor=west] (x) at (-0.2cm,0) {$\mathbf{x}$};
\node [anchor=west,model] (policy model) at ([xshift=.8cm]x) {Policy\\Model};
\node [anchor=north west,label] at ([xshift=-.1cm,yshift=.15cm]policy model.north west) {To Learn};
\node [anchor=west] (y) at ([xshift=1.2cm]policy model.east) {$\mathbf{y}$};
\node [anchor=west,model] (reward model) at ([xshift=1.4cm]y.east) {Reward\\Model};
\node [anchor=north west,label] at ([xshift=-.1cm,yshift=.15cm]reward model.north west) {Fixed};
\node [anchor=south,model] (reference model) at ([yshift=.4cm]reward model.north) {Reference\\Model};
\node [anchor=north west,label] at ([xshift=-.1cm,yshift=.15cm]reference model.north west) {Fixed};
\node [anchor=north,model] (value model) at ([yshift=-.4cm]reward model.south) {Value\\Model};
\node [anchor=north west,label] at ([xshift=-.1cm,yshift=.15cm]value model.north west) {To Learn};
\node [anchor=west] (reward) at ([xshift=0.1cm]reward model.east) {$R_\phi(\mathbf{x}, \mathbf{y})$};
\node [anchor=west] (value) at ([xshift=0.1cm]value model.east) {$\{V_{\omega,1}, \cdots, V_{\omega,T}\}$};
\node [anchor=west] (reference) at ([xshift=0.1cm]reference model.east) {Penalty};
\node [anchor=west,align=center,rounded corners=5pt,inner sep=12pt,fill=lightgray] (advantage estimation) at ([xshift=1.5cm]reward) {Advantage\\Estimation};
\node [anchor=west] (advantage estimation out) at (advantage estimation.east) {$\{A_{1}, \cdots, A_{T}\}$};
\node (left bottom) at (x.west|-value model.south) {};
\node (right bottom) at (advantage estimation out.east|-value model.south) {};
\node [anchor=north,minimum width=13.6cm, minimum height=2cm,rounded corners=8pt,draw=lightgray,thick] (note) at ($(left bottom)!.5!(right bottom)+(0, -1cm)$) {};
\matrix [anchor=center, every even column/.style={column sep=1cm}] at (note.center) {
\node {\circled{1}}; & \node [anchor=west, inner sep=0pt] {sample an output $\mathbf{y}$ from $\mathrm{Pr}_{\theta_{\mathrm{ref}}}(\cdot|\mathbf{x})$} ; & \node {\circled{2}}; & \node [anchor=west, inner sep=0pt] {compute the $\mathrm{Penalty}$} ; \\
\node {\circled{3}}; & \node [anchor=west, inner sep=0pt] {compute the reward for $\mathrm{seq}_{\mathbf{x}, \mathbf{y}}$ } ; & \node {\circled{4}}; & \node [anchor=west, inner sep=0pt] {predict the value for $\mathrm{seq}_{\mathbf{x}, \mathbf{y}}$} ; \\
\node {\circled{5}}; & \node [anchor=west, inner sep=0pt] {optimize the value model $V_{\omega}(\cdot)$} ; & \node {\circled{6}}; & \node [anchor=west, inner sep=0pt] {optimize the policy model} ; \\
};
\draw [->,thick] (x.east) -- (policy model.west);
\draw [->,thick] (policy model.east) node [xshift=.5cm,yshift=.0cm,above,align=center] {\circled{1}} -- (y.west);
\draw [->,thick] (y.east) .. controls (reward model.west) and ($(y.east |- reference model.west)$) .. (reference model.west);
\draw [->,thick] (y.east) .. controls (reward model.west) and ($(y.east |- value model.west)$) .. (value model.west);
\draw [->,thick] (y.east) -- (reward model.west);
\draw [->,thick] (reference.east) node [above,xshift=.8cm] {\circled{2}} .. controls (reference.east -| advantage estimation.north) .. (advantage estimation.north);
\draw [->,thick] (reward.east) -- node [above,xshift=0.0cm] {\circled{3}} (advantage estimation.west);
\draw [->,thick] (value.east) .. controls (value.east -| advantage estimation.south) .. node [above,xshift=-.2cm] {\circled{4}}(advantage estimation.south);
% update model
\node (policy model bottom) at ($(right bottom -| policy model)+(0,-.6cm)$) {};
\node [anchor=north] (policy model bottom1) at (policy model bottom |- advantage estimation out.north) {};
\node (value model bottom) at ($(right bottom -| value model)+(0,-.6cm)$) {};
\node (advantage estimation out bottom) at ($(advantage estimation out |- right bottom)+(0,-.6cm)$) {};
\draw [->,thick,dashed] (advantage estimation out) -- ($(advantage estimation out)!.5!(advantage estimation out bottom)$) .. controls (advantage estimation out bottom) .. ($(advantage estimation out bottom.center)!.5!(policy model bottom.center)+(5em,0)$) -- ($(advantage estimation out bottom.center)!.5!(policy model bottom.center)-(5em,0)$) .. controls (policy model bottom.center) .. ($(policy model bottom1.center)!.5!(policy model bottom)$) -- node [above,xshift=0.3cm,yshift=-.2cm,solid] {\circled{6}}(policy model.south);
\draw [->,thick,dashed] ($(advantage estimation out bottom.center)!.5!(policy model bottom.center)+(2.89em,0)$) node [above,xshift=-1.5cm,solid] {\circled{5}}.. controls (value model bottom) .. (value model.south);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\def\ssep{.7cm}
\def\nsize{0.3cm}
\tikzstyle{lnode} = [minimum width=7cm,minimum height=1.1cm,inner sep=2pt,draw,thick,fill=white];
\begin{scope}
\node [anchor=west] (x0) at (0,0) {\footnotesize{$x_0$}};
\node [anchor=center] (x cdots) at ([xshift=\ssep]x0.center) {};
\node [anchor=center] (x m) at ([xshift=\ssep]x cdots.center) {};
\node [anchor=center] (y1) at ([xshift=\ssep]x m.center) {\footnotesize{$x_{m}$}};
\node (x real cdots) at ($(x cdots)!.5!(x m)$) {\footnotesize{$\cdots$}};
\node [anchor=center] (y2) at ([xshift=\ssep]y1.center) {};
\node [anchor=center] (y cdots) at ([xshift=\ssep]y2.center) {};
\node [anchor=center] (y n) at ([xshift=\ssep]y cdots.center) {};
\foreach \x in {x0,y1}
\draw [->] ([yshift=0.2cm]\x.center) -- ([yshift=0.55cm]\x.center);
\node [anchor=center] (oy1) at ([yshift=2.8cm]y1.center) {\footnotesize{$y_1$}};
\node [anchor=center] (oy2) at ([yshift=2.8cm]y2.center) {};
\node [anchor=center] (oy cdots) at ([yshift=2.8cm]y cdots.center) {};
\node [anchor=center] (oy last) at ([yshift=2.8cm]y n.center) {\footnotesize{$y_T$}};
\node (oy real cdots) at ($(oy cdots)!.5!(oy2)$) {\footnotesize{$\cdots$}};
\draw [decorate, decoration={brace,amplitude=.2cm},thick] (oy1.north west) -- (oy last.north east);
\node [anchor=center] (oy) at ([yshift=.7cm]$(oy2)!.5!(oy cdots)$) {\footnotesize{$\mathbf{y}$}};
\foreach \x in {oy1, oy last}
\draw [<-] ([yshift=-0.25cm]\x.center) -- ([yshift=-0.8cm]\x.center);
\node [anchor=south,draw,thick,minimum width=5.2cm,minimum height=1.3cm,fill=white,drop shadow={fill=gray,shadow xshift=.6ex,shadow yshift=-.6ex}] (llm) at ([yshift=.6cm]y1.center) {Policy (LLM)};
\node [anchor=center] (freeze) at ([xshift=-.25cm,yshift=.25cm]llm.south east) {\textcolor{cyan!60}{ \faSnowflake}};
\node (rside) at ([xshift=0.5cm,yshift=-1.2cm]oy last.north east|-llm.east) {};
\node (lside) at ([yshift=-1.2cm]llm.west|-y n.east) {};
\node at ($(lside)!.5!(rside)$) {\footnotesize{(a)~sample an output $\mathbf{y}$}};
\node [anchor=west] (x0) at (\textwidth/2,0) {\footnotesize{$x_0$}};
\node [anchor=center] (x cdots) at ([xshift=\ssep]x0.center) {};
\node [anchor=center] (x m) at ([xshift=\ssep]x cdots.center) {};
\node [anchor=center] (y1) at ([xshift=\ssep]x m.center) {\footnotesize{$x_{m}$}};
\node (x real cdots) at ($(x cdots)!.5!(x m)$) {\footnotesize{$\cdots$}};
\node [anchor=center] (y2) at ([xshift=\ssep]y1.center) {};
\node [anchor=center] (y cdots) at ([xshift=\ssep]y2.center) {};
\node [anchor=center] (y n) at ([xshift=\ssep]y cdots.center) {$y_{T-1}$};
\node (y real cdots) at ($(y2)!.5!(y cdots)$) {\footnotesize{$\cdots$}};
\foreach \x in {x0,y1, y n}
\draw [->] ([yshift=0.2cm]\x.center) -- ([yshift=0.55cm]\x.center);
\node [anchor=center] (oy1) at ([yshift=2.8cm]y1.center) {};
\node [anchor=center] (oy2) at ([yshift=2.8cm]y2.center) {};
\node [anchor=center] (oy cdots) at ([yshift=2.8cm]y cdots.center) {};
\node [anchor=center] (oy last) at ([yshift=2.8cm]y n.center) {};
\node [anchor=south,draw,thick,minimum width=5.2cm,minimum height=1.3cm,fill=white,drop shadow={fill=gray,shadow xshift=.6ex,shadow yshift=-.6ex}] (llm) at ([yshift=.6cm]y1.center) {Policy (LLM)};
\node [anchor=center] (freeze) at ([xshift=-.25cm,yshift=.25cm]llm.south east) {\textcolor{red!60}{ \faFire}};
\node [anchor=east,inner sep=0] (prob) at ([xshift=1em]oy last.east-|llm.east) {\ctext[RGB]{255,204,204}{$\log\left(\textrm{Pr}_\theta (y_1|\textbf{x})\times\cdots\times\textrm{Pr}_\theta (y_T|\textbf{y}_{<T},\textbf{x})\right)$}};
\node [anchor=west,inner sep=0] (reward) at ([xshift=.2cm]prob.east) {\ctext[RGB]{153,204,255}{$R(\mathbf{y})$}};
\node [anchor=center] at ($(prob.east)!.5!(reward.west)$) {$\times$};
\node (rside) at ([xshift=0.5cm,yshift=-1.2cm]oy last.north east|-llm.east) {};
\node (lside) at ([yshift=-1.2cm]llm.west|-y n.east) {};
\node at ($(lside)!.5!(rside)$) {\footnotesize{(b)~train the LLM via policy gradient}};
\draw [<-] ([yshift=-.25cm,xshift=-1.35cm]oy1.center) .. controls ([yshift=-.525cm,xshift=-1.35cm]oy1.center) .. ([yshift=-.525cm,xshift=-1cm]oy1.center) -- ([yshift=-.525cm,xshift=-.35cm]oy1.center) .. controls ([yshift=-.525cm]oy1.center) .. ([yshift=-.8cm]oy1.center);
\draw [<-] ([yshift=-.25cm, xshift=-.7cm]oy last.center) .. controls ([yshift=-.525cm,xshift=-.7cm]oy last.center) .. ([yshift=-.525cm,xshift=-.35cm]oy last.center) .. controls ([yshift=-.525cm]oy last.center) .. ([yshift=-.8cm]oy last.center);
\draw [->,thick] (oy.north) -- +(0,.6cm) -- ([yshift=.6cm]oy.north-|reward) node [midway,above] {compute the reward: \ctext[RGB]{153,204,255}{$R(\mathbf{y})$}} -- ([yshift=.2cm]reward.north);
\end{scope}
\end{tikzpicture}
\end{center}
\begin{algorithm}[t]
\caption{PPO-based Optimization Workflow}
\begin{algorithmic}[1]
\INPUT the optimized reward model $r_{\phi}(\cdot)$, the reference model $\pi_{\theta_{\mathrm{ref}}}(\cdot)$, the initialized value model $V_{\omega}(\cdot)$, the policy model $\pi_{\theta}(\cdot)$, the input-only dataset $S_{x}$;
\OUTPUT the aligned $\pi_{\theta}(\cdot)$;
\State \%\%\% suppose the batch size is set to 1
\For{$\mathbf{x} \in S_x$}
\State $\pi_{\theta_\mathrm{old}(\cdot)} = \mathrm{Clone}(\pi_{\theta}(\cdot))$
\State sample an output from $\pi_{\theta_\mathrm{old}}$: $\mathbf{y} \sim \pi_{\theta_\mathrm{old}}(\cdot|\mathbf{x})$
\State compute the reward for $[\mathbf{x},\mathbf{y}]$: $r_{\phi}(\mathbf{x},\mathbf{y})$
\State \%\%\% roll out
\For{ppo\_epoch\_idx=1 to $\mathrm{PPO\_EPOCH}$}
\State \%\%\% \textit{compute policy model loss and value model loss through advantages and returns}
\State initialize the loss values: $\mathrm{loss}_{p} = 0$, $\mathrm{loss}_{v} = 0$
\For{t=1 to T}
\State compute the penalty through Eq. (\ref{eq:penalty}): $\mathrm{Penalty}_{t}=\log \mathrm{Pr}_{\theta}(y_t|\mathbf{x},\mathbf{y}_{<t}) - \log \mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_t|\mathbf{x},\mathbf{y}_{<t})$
\State \textbf{if} t==T \textbf{then}
\State \hspace{0.5cm} $r_{t} = r_{\phi}(\mathbf{x}, \mathbf{y}) -\beta \mathrm{Penalty}_{t}$
\State \textbf{else}
\State \hspace{0.5cm} $r_{t} = -\beta \mathrm{Penalty}_{t}$
\State \textbf{end if}
\State predict the value $V_{\omega}(\mathbf{x},\mathbf{y}_{<t},y_{t})$
\State \textbf{if} t==T \textbf{then}
\State \hspace{0.5cm} $V_{\omega}(\mathbf{x},\mathbf{y}_{<t+1},y_{t+1})=0$
\State \textbf{else}
\State \hspace{0.5cm} predict the next time step value $V_{\omega}(\mathbf{x},\mathbf{y}_{<t+1},y_{t+1})$
\State \textbf{end if}
\State \%\%\% \textit{compute policy model loss}
\State compute the advantage $A(\mathbf{x},\mathbf{y}_{<t},y_{t})$ through Eq. (\ref{eq:advantage}): $A(\mathbf{x},\mathbf{y}_{<t},y_{t})=r_{t}+\gamma V_{\omega}(\mathbf{x},\mathbf{y}_{<t+1},y_{t+1})-V_{\omega}(\mathbf{x},\mathbf{y}_{<t},y_{t})$
\State compute the policy model loss through Eq. (\ref{eq:ppo-loss}) and add it to $\mathrm{loss}_{p}$: $\mathrm{loss}_{p} = \mathrm{loss}_{p} + \mathrm{Clip}\Big( \frac{\pi_{\theta}(y_t|\mathbf{x},\mathbf{y}_{<t})}{\pi_{\theta_{\mathrm{old}}}(y_t|\mathbf{x},\mathbf{y}_{<t})} \Big) A(\mathbf{x},\mathbf{y}_{<t},y_t)$
\State \%\%\% \textit{compute value model loss}
\State compute the return $\mathrm{Return}_{t}$: $\mathrm{Return}_{t}=A(\mathbf{x},\mathbf{y}_{<t},y_t)+V_{\omega}(\mathbf{x},\mathbf{y}_{<t},y_{t})$
\State compute the value model loss through Eq. (\ref{eq:value-loss}) and add it to $\mathrm{loss}_{v}$: $\mathrm{loss}_v = \mathrm{loss}_v + (\mathrm{Return}_{t}-V_{\omega}(\mathbf{x},\mathbf{y}_{<t},y_{t}))^2$
\EndFor
\State update the parameters of $\pi_{\theta}(\cdot)$ through $\mathrm{loss}_{p}$
\State update the parameters of $V_{\omega}(\cdot)$ through $\mathrm{loss}_{v}$
\EndFor
\EndFor
\State return $\pi_{\theta}$
\end{algorithmic}
\label{alg:rlhf}
\end{algorithm}
\ No newline at end of file
% An example of REINFORCE and optimization objective
\section{An Example of Using Reinforcement Learning to Train LLMs}
\label{sec:example-using-rl-training-llms}
To explain how RL can be applied to train LLMs, we begin by considering a practical scenario: using an LLM as a homework assistant. In this context, we aim to improve the capability of an SFT LLM to handle education-related inputs more effectively. Suppose we have a homework assistant powered by an SFT LLM, and a student types the input ``Give me three tips to improve my accuracy in solving math problems.'' A typical SFT LLM might generate a short output, such as
\vspace{0.1cm}
\begin{tcolorbox}[frame empty]
\begingroup
\setlength{\leftskip}{2em}
\setlength{\rightskip}{2em}
There are three tips for improving accuracy in solving math problems: \\ [1mm]
1.Practice regularly. \\ [1mm]
2.Understand the concepts. \\ [1mm]
3.Double-check your work.
\endgroup
\end{tcolorbox}
\vspace{0.5em}
Although this output adheres to the given input, it is very short and not sufficiently detailed or comprehensive for this scenario. In practice, the ``short'' feature in the generated output stems from the training approach of the SFT LLM. During SFT, the LLM learns from a large set of labeled samples that typically contain concise and factual answers to common inputs. These samples guide the LLM to prioritize brevity and clarity, so it often generates short outputs, like the one shown above. Of course, we could annotate enough additional data to fine-tune the LLM further and adjust it to generate the desired outputs. However, this approach is limited in its ability to scale. For example, in this scenario, the input involves a variety of potential answers, and the expectations of the student can be pretty diverse. Therefore, describing the ``ideal'' output would require immense annotation effort. Consequently, collecting or annotating fine-tuning data is not as straightforward as it is with SFT, particularly when attempting to cover the breadth of possible student inputs and outputs. Instead, we can use RL to enable the model to discern outputs that better align with human preferences, such as generating more detailed and comprehensive content that not only adheres to the given input but also meets the expectations of the student in this scenario.
In the following sections, we will demonstrate how to use RL to train LLMs through a specific example—enabling the LLM to generate a longer output in the context of a homework assistant. Note that the techniques discussed are not limited to this single case. Rather, they can be broadly applied to align LLM with any set of expectations.
% policy gradient
\subsection{Policy Gradient}
\label{sec:policy-gradient}
In this section, we aim to lengthen the LLM output for the ``give me three tips'' input using a classical RL algorithm called policy gradient.
To define our optimization objective, let $R(\cdot)$ be the reward function that describes the goal the algorithm aims to optimize. In this scenario, the reward is designed to align the output of the LLM with a specific behaviour, that is, generating longer outputs. Therefore, we define the reward function based on a length-related metric. More specifically, the reward can be proportional to the length of the output, that is, longer outputs can receive higher rewards:
\begin{eqnarray}
R(\mathbf{y}) = \frac{\mathrm{Length(\mathbf{y})}}{10}
\label{eq:langth-based-reward-function}
\end{eqnarray}
where $\mathrm{Length}(\cdot)$ indicates the length of the given output. We can use the number of tokens in the output as the length for simplicity. This allows us to assign an immediate reward $r_t$ for the $t$-th token generated by the LLM, which is $\frac{1}{10}$.
Next, we apply RL to encourage the LLM to generate longer outputs based on the reward function. First, let us review the optimization objective of RL: learning a policy that maximizes the agent’s cumulative reward (or return) over the long term. In this case, the agent can be defined as the LLM itself, and the policy can be defined as the probability distribution over the tokens the LLM predicts, given the preceding context tokens, i.e., $\mathrm{Pr}_{\theta}(\cdot|\mathbf{x})$\footnote{To maintain consistency with notations in RL, this policy is also often represented as $\pi_{\theta}(\cdot|\mathbf{x})$ in much of the related literature. Here, from the perspective of LLM research, we opt to use the probability notation $\mathrm{Pr}_{\theta}(\cdot|\mathbf{x})$, which is more familiar to researchers in this field.}.
The optimization objective can be given by
\begin{eqnarray}
J(\theta) & = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \sum_{\mathbf{y}\in \Omega } \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) R(\mathbf{y}) \right] \nonumber \\
& = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \sum_{\mathbf{y}\in \Omega } \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) \sum_{t=1}^{T} r_{t} \right]
\end{eqnarray}
\begin{figure*}
\centering
\input{section3/Figures/figure-update-with-policy-gradient.tex}
\caption{
Illustration of training an LLM using policy gradient. The central component is an LLM. Given an input $\mathbf{x}=x_1...x_m$, we sample an output $\mathbf{y}$ from the LLM. Note that gradients are not computed during the sampling to prevent extensive memory consumption because the sampling is performed via an autoregressive mode. Instead, by leveraging the parallel computation capability of the LLM, similar to SFT, the sampled output $\mathbf{y}$ is used for a subsequent forward pass to compute the gradients produced during the generation of each token.
}
\label{fig:update-with-policy-gradient}
\end{figure*}
where $S_{x}$ indicates the input-only dataset, $\mathbf{y}\in \Omega$ indicates that output $\mathbf{y}$ is drawn from the hypothesis space $\Omega$, $T$ indicates the length of $\mathbf{y}$. $J(\theta)$ is also called the performance function. Then the training objective is maximize $J(\theta)$:
\begin{eqnarray}
\hat{\theta} & = & \argmax_{\theta} J(\theta)
\end{eqnarray}
Note that the hypothesis space $\Omega$ is typically very large, as LLMs often have large vocabularies\footnote{The hypothesis space is large because LLMs typically have vast vocabularies and complex tokenization schemes. The number of possible combinations of tokens, even for a modest length output, grows exponentially, leading to a very large output space. This makes exhaustive enumeration of all possible outputs computationally infeasible.}. Therefore, it is not feasible to enumerate all possible outputs directly. Instead, we typically use a Monte Carlo method to estimate this expectation. This approach involves sampling outputs $\{\mathbf{y}_1,\mathbf{y}_2,\cdots,\mathbf{y}_{d}\}$ from the LLM and using those samples to approximate the hypothesis space, denoted by $\mathcal{D}$:
\begin{eqnarray}
J(\theta) & = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \sum_{\mathbf{y} \in \mathcal{D} } \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) \sum^{T}_{t=1}r_{t} \right]
\end{eqnarray}
We can further refine the process when optimizing the objective using gradient descent. First, we compute the gradient of $J(\theta)$ with respect to $\theta$:
\begin{eqnarray}
\frac{\partial J(\theta)}{\partial \theta} & = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \sum_{\mathbf{y} \in \mathcal{D}} \frac{ \partial \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) R(\mathbf{y})}{\partial \theta} \right] \nonumber \\
& = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \sum_{\mathbf{y} \in \mathcal{D}} \frac{ \partial \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})}{\partial \theta} R(\mathbf{y}) \right] \nonumber \\
& = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \sum_{\mathbf{y} \in \mathcal{D}} \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) \frac{ \partial \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})/\partial \theta}{\mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})} R(\mathbf{y}) \right] \nonumber \\
& = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \sum_{\mathbf{y} \in \mathcal{D}} \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) \frac{ \partial \log \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})}{\partial \theta} R(\mathbf{y}) \right]
\label{eq:rl-gradient-j-theta}
\end{eqnarray}
\begin{figure*}[!t]
\centering
\input{section3/Figures/figure-one-example}
\caption{
The mechanism of RL in training LLMs from the perspective of hypothesis space reranking: given an input $\mathbf{x}$, the model increases the probability of longer outputs, such as $\mathbf{y}_3$, within the hypothesis space by assigning higher rewards to them; conversely, shorter outputs like $\mathbf{y}_1$ receive lower rewards, which reduces their probabilities.
}
\label{fig:understand-policy-gradient}
\end{figure*}
We assume that every output in $\mathcal{D}$ is equally probable (i.e., $\mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) = 1/|\mathcal{D}|$). In this case we can simplify Eq. (\ref{eq:rl-gradient-j-theta}) and need only consider the terms $\frac{\partial \log \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})}{\partial \theta}$ and $R(\mathbf{y})$:
\begin{eqnarray}
\frac{\partial J(\theta)}{\partial \theta} & = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \frac{1}{d} \sum_{\mathbf{y} \in \mathcal{D}} \frac{ \partial \log \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})}{\partial \theta} R(\mathbf{y}) \right]
\label{eq:rl-j-theta-gradient-simplified}
\end{eqnarray}
Now, as illustrated in Figure \ref{fig:update-with-policy-gradient}, we have an RL approach to optimize the output of the LLM: 1) we sample some inputs from the SFT LLM; then, 2) we evaluate each input using a reward function that we define; then, 3) we update the LLM using gradient descent, following Eq. (\ref{eq:rl-j-theta-gradient-simplified}), to maximize the performance function.
We can understand this optimization process from the perspective of hypothesis space reranking, as illustrated in Figure \ref{fig:understand-policy-gradient}. The objective is to increase the probability of longer outputs in the hypothesis space by assigning a high reward to those outputs, while simultaneously decreasing the probability of shorter outputs by assigning them a lower reward. In other words, the policy gradient approach encourages the model to generate longer outputs by reinforcing those outputs that meet the desired length criteria, and it penalizes shorter outputs that do not meet the objective.
\subsection{Temporal Decomposition}
\label{sec:temporal-decomposition}
While optimizing using Eq. (\ref{eq:rl-j-theta-gradient-simplified}) is intuitive, a potential issue arises: in some cases, the sampled outputs we sample may significantly overlap, yet the rewards for the overlapping parts differ. For example, as illustrated in Figure \ref{fig:an-overlap-example}, if outputs $\mathbf{y}_1$ and $\mathbf{y}_2$ both include the same tokens in the third tip ``3.Double-check your work...any-one can significantly enhance their math problem-solving abilities.'', but due to variations in the entire outputs, they receive markedly different rewards (e.g., $R(\mathbf{y}_1)=50$ vs. $R(\mathbf{y}_2)=10$). Ideally, we would want similar generational behaviours to be rewarded consistently, without significant variation, as their contributions are equivalent to the sum of rewards.
Before discussing an approach to solve this issue, we first perform temporal decomposition for the term $\frac{\partial \log \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})}{\partial \theta} R(\mathbf{y})$ in Eq. (\ref{eq:rl-j-theta-gradient-simplified}), and obtain
\begin{eqnarray}
\frac{\partial \log \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})}{\partial \theta} R(\mathbf{y}) & = & \big(\sum_{t=1}^{T} \frac{\partial \log \mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\partial \theta}
\big)\big(\sum_{t=1}^{T} r_{t}\big) \nonumber \\
& = & \big(\sum_{t=1}^{T} \frac{ \partial \log \mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\partial \theta}\big)\big(\sum_{k=1}^{t-1} r_{k} + \sum_{k=t}^{T} r_{k}\big)
\end{eqnarray}
\begin{figure*}[!t]
\centering
\input{section3/Figures/figure-an-overlap-example}
\caption{
Example of sampled outputs demonstrating that the same tokens will receive significantly different rewards (i.e., 16 vs. 75). This introduces high gradient variance in the optimization process of the policy gradient.
}
\label{fig:an-overlap-example}
\end{figure*}
In this form, we observe that, when generating the token at $t$-th time step, it does not affect the term $\sum_{k=1}^{t-1} r_{k}$ and only affects the term $\sum_{k=t}^{T} r_{k}$. Therefore, we can safely remove the former part, $\sum_{k=1}^{t-1} r_{k}$, to prevent it from affecting the optimization of the current step. By doing so, we can achieve a more stable gradient computation for Eq. (\ref{eq:rl-j-theta-gradient-simplified}):
\begin{eqnarray}
\frac{\partial J(\theta)}{\partial \theta} & = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \frac{1}{d} \sum_{\mathbf{y} \in \mathcal{D}} \sum_{t=1}^{T} \frac{ \partial \log \mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\partial \theta}\sum_{k=t}^{T} r_{k}
\right]
\end{eqnarray}
In fact, this simplification follows the Markov process, asserting that the future is independent of the past, given the present. As a result, we focus on the reward terms from the $t$-th time step onward, which are directly influenced by the token generation at time step $t$.
\subsection{Reducing Gradient Variance}
\label{sec:reduce-gradient-variance}
Returning to the case discussed in Section \ref{sec:temporal-decomposition}, while the strategy of excluding rewards for past tokens reduces gradient variance, substantial gradient variance still occurs in practice. First, if overlapping portions of the output appear early, the modified Eq. (\ref{eq:modified-gradient-simlified}) may still experience similar problems, i.e., the same tokens receive vastly different rewards. Furthermore, the reward distribution can vary significantly between different time steps within a single output. For example, consider an output sequence of 500 tokens. According to Eq. (\ref{eq:modified-gradient-simlified}), the reward at the 10th step might be significantly higher than at the 450th step, i.e., 49 vs. 5. Such varying rewards for good and poor tokens can result in a very low total reward for the entire output, even if it includes good tokens.
One simple method for further reducing the variance of the gradient is to set a baseline $b$ and subtract it from $\sum_{k=t}^{T} r_k$, resulting in $\sum_{k=t}^{T} r_k - b$.\footnote{In fact, the use of a baseline $b$ does not change the variance of the total rewards $\sum_{t=1}^{T} r_t$. However, it is important to note that while introducing a baseline does not alter the overall variance of the rewards, it helps reduce the variance of the gradient estimates. This is because subtracting the baseline from the total rewards effectively reduces fluctuations around their mean, which makes the gradient estimates more stable. In general, the operation $\sum_{k=t}^{T} r_k - b$ centers the rewards around zero (e.g., $b$ is defined as the expected value of $\sum_{k=t}^{T} r_k$), which can lead to reduced variance in the product $\sum_{k=t}^{T} \log \mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t}) (\sum_{k=t}^{T} r_k - b)$.} Here, the baseline can be interpreted as a reference point. By centering the rewards around this baseline, we remove systematic biases in the reward.
This policy gradient model with a baseline can be given by
\begin{eqnarray}
\frac{\partial J(\theta)}{\partial \theta} & = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \frac{1}{d} \sum_{\mathbf{y} \in \mathcal{D}} \sum_{t=1}^{T} \frac{ \partial \log \mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\partial \theta} \big(\sum_{k=t}^{T} r_{k}-b\big)
\right]
\label{eq:modified-gradient-simlified}
\end{eqnarray}
There are many ways to define the baseline $b$. For example, we can set the baseline $b$ to the average length of the sampled outputs across $\mathcal{D}$, i.e., $b = (\sum_{\mathbf{y} \in \mathcal{D}} \mathrm{Length}(\mathbf{y}))/d$. While this approach can reduce the relative size of the gradient variance, it does not address the issue of varying gradient variance across different time steps within a single output. To tackle this challenge, we would want a different baseline for each time step that can specifically be adjusted to reduce variance at that step. To achieve this ideal baseline, we can introduce a value function $V(\cdot)$. In the training of the LLM, this value function calculates the expected value of the sum of future rewards (or return for short) when generating $y_{t}$, given the input $\mathbf{x}$ and the tokens previously generated $\mathbf{y}_{<t}$, i.e., $V(\mathbf{x}, \mathbf{y}_{<t}, y_{t}) = \mathbb{E}[\sum_{k=t}^T r_{k}]$. Hence we have
\begin{eqnarray}
A(\mathbf{x}, \mathbf{y}_{<t}, y_{t}) & = & \sum_{k=t}^{T} r_k - b \nonumber \\
& = & \sum_{k=t}^{T} r_k - V(\mathbf{x}, \mathbf{y}_{<t}, y_{t})
\label{eq:monte-carlo-advantage}
\end{eqnarray}
where $\sum_{k=t}^{T} r_k $ represents the actual return received, and $V(\mathbf{x}, \mathbf{y}_{<t}, y_{t})$ (or $V_{t}$ for short) represents the expected return at time step $t$. $A(\mathbf{x}, \mathbf{y}_{<t}, y_{t})$ (or $A_t$ for short) is called the advantage at time step $t$, which quantifies the relative benefit of the current generated $y_t$ compared to the expected return. Computing the advantage using Eq. (\ref{eq:monte-carlo-advantage}) is a method known as Monte Carlo-based advantage estimation.
By using the advantage function $A_{t}$, the gradient of $J(\theta)$ can be written in the form
\begin{eqnarray}
\frac{\partial J(\theta)}{\partial \theta} & = & \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \left[ \frac{1}{d} \sum_{\mathbf{y} \in \mathcal{D}} \sum_{t=1}^{T} \frac{ \partial (\log \mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})A_{t})}{\partial \theta}
\right]
\label{eq:modified-gradient-simlified-advantage}
\end{eqnarray}
Based on this training objective, the loss function of training the LLM can be written in the form
\begin{eqnarray}
\mathcal{L}(\theta) & = & - \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \mathbb{E}_{\mathbf{y} \sim \mathcal{D}} \left[ \sum_{t=1}^{T} \log \mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})A_{t}
\right]
\label{eq:loss-function-llm-update}
\end{eqnarray}
\begin{figure*}[!t]
\centering
\input{section3/Figures/figure-architecture-value-function}
\caption{
Architecture of the value model based on a pre-trained LLM. The main component of this model is still an LLM. For a given sequence $\mathrm{seq}_{\mathbf{x}, \mathbf{y}}=[\mathbf{x}, \mathbf{y}]$, the model extracts representations corresponding to the positions of the tokens in $\mathbf{y}$. These representations are then individually mapped to scalars through a linear transformation. Each scalar represents the value associated with each token in $\mathbf{y}$. Note that the linear mapping matrix used for the transformations is typically shared across all tokens.
}
\label{fig:architecture-value-function}
\end{figure*}
In practice, there are many ways to implement the value function. One simple approach is to build it based on a pre-trained LLM (e.g., an SFT LLM). In this way, the value function is also called the value model (or critic model). More specifically, we can concatenate $\mathbf{x}$ and $\mathbf{y}$ to form a single token sequence $\mathrm{seq}_{\mathbf{x},\mathbf{y}} = [\mathbf{x}, \mathbf{y}]$. We run an SFT LLM on this sequence, as usual, and at each position, we obtain a representation from the top-most Transformer layer. Then, we take the representations at the output (denoted by $\{\mathbf{h}_{y_{1}},\mathbf{h}_{y_{2}},\cdots,\mathbf{h}_{y_{T}}\}$) and map them to scalars via linear transformation, respectively. Figure \ref{fig:architecture-value-function} illustrates this architecture of the value model. Hence, at $t$-th time step, we can compute the value $V_{t}$
\begin{eqnarray}
V_{t} & = & \mathbf{h}_{y_{t}} \mathbf{W}_{v}
\end{eqnarray}
where $\mathbf{h}_{y_{t}}$ is a $d$-dimensional vector, and $\mathbf{W}{v}$ is a $d \times 1$ linear mapping matrix. This value model is typically trained using the Temporal Difference (TD) error. This training method relies on the principle that the difference between the values at adjacent time steps should correspond to the reward received at the former time step, i.e., $V_t - V_{t+1} = r_t$. Suppose the value model is parameterized by $\omega$. Given a sequence $\mathrm{seq}_{\mathbf{x}, \mathbf{y}}$, the loss function is given by
\begin{eqnarray}
\mathcal{L}_v(\omega) & = & \frac{1}{T} \sum_{t=1}^{T} \big(r_t + V_{\omega,t+1} - V_{\omega,t} \big)^2
\end{eqnarray}
or alternatively, introduce the discount factor $\gamma$ to obtain a more general form
\begin{eqnarray}
\mathcal{L}_v(\omega) & = & \frac{1}{T} \sum_{t=1}^{T} \big(r_t + \gamma V_{\omega,t+1} - V_{\omega,t} \big)^2
\end{eqnarray}
where $\gamma \in [0,1]$ is the discount factor that adjusts the importance of future rewards. When $\gamma$ is set to less than 1, it signifies that early rewards are considered more important than future rewards. This basic idea is also applied in other fields. For example, in LLMs, it has been demonstrated that early token generation plays a crucial role, as it can influence the style and accuracy of the entire output \citep{wang-and-zhou:2024chain}.
In this subsection, we have detailed the integration of a value model in training LLMs. In practice, this training approach, represented by Eq. (\ref{eq:modified-gradient-simlified-advantage}), is known as the Advantage Actor-Critic (A2C) method \citep{mnih-etal:2016asynchronous}. The A2C method facilitates an interaction where a policy model (the actor) and a value model (the critic) learn in parallel and undergo synchronous updates. However, RL is a vast field, and many technical details cannot be covered here. The interested reader can refer to RL books for more details \citep{szepesvari:2010algorithms, Sutton-and-Barto:2018RL}.
\begin{figure*}[!t]
\centering
\input{section3/Figures/figure-importance-sampling}
\caption{
Workflow of integrating important sampling in the training of LLMs with policy gradient. In Step 1, the current policy is synchronized to establish the reference policy. In Step 2, we sample an output $\mathbf{y}$ from this reference policy for a given input $\mathbf{x}$. Finally, in Step 3, the sequence $[\mathbf{x}, \mathbf{y}]$ is used to optimize the current policy multiple times. After optimization, the updated policy is synchronized to become the reference policy.
}
\label{fig:policy-gradient-with-importance-sampling}
\end{figure*}
\subsection{Importance Sampling}
\label{sec:importance-sampling}
In this subsection, we discuss methods to improve the efficiency of the policy gradient, focusing on the sampling process. Although we discussed in Section \ref{sec:policy-gradient} that Monte Carlo-based sampling can estimate the hypothesis space, it is awfully inefficient for one single update in scenarios like ours where the sampled output may consist of hundreds or thousands of tokens. A natural idea from this point is to consider whether it is possible to reuse these sampled outputs to update the LLMs multiple times. With importance sampling, this is feasible, and we can rewrite the loss function of training the LLM as
\begin{eqnarray}
\mathcal{L}(\theta) & = & - \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \mathbb{E}_{\mathbf{y}\sim \mathrm{Pr}_{\theta_{\mathrm{ref}}}(\cdot|\mathbf{x})} \left[
\sum_{t=1}^{T} \frac{\mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}A_{t}
\right]
\label{eq:loss-function-training-llm-importance-sampling}
\end{eqnarray}
where $\theta_{\mathrm{ref}}$ denotes the parameters of the previously used LLM (also called reference policy or old policy). Here, we use $\mathrm{Pr}_{\theta}(\cdot|\mathbf{x})$ or $\mathrm{Pr}_{\theta_{\mathrm{ref}}}(\cdot|\mathbf{x})$ to denote $\mathcal{D}$, distinguishing from which model the output is sampled. The ratio $\frac{\mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}$, also called the ratio function, compares the log-probability of token $y_t$ under the current and reference policies. This ratio function is used to reweight the observed rewards, reflecting how much more or less likely a token is under the current policy compared to the reference policy. When this ratio is greater than 1, it indicates that the token $y_t$ is more favored by the current policy than the reference policy. Conversely, a ratio less than 1 indicates that $y_t$ is less favored by the current policy. However, when the current used $\theta_{\mathrm{ref}}$ diverges significantly from $\theta$, the accuracy of the hypothesis space estimation decreases. Therefore, we do need to resync it regularly. As illustrated in Figure \ref{fig:policy-gradient-with-importance-sampling}, one simple way is to designate the reference policy as the LLM from which we initiate updates during a training step. Note that we can conserve computational time and memory by eliminating the need for model copying in Step 1. Specifically, rather than maintaining a separate copy of the model, we directly sample from the current policy, retain the output probabilities $\{\mathrm{Pr}_{\theta}(y_1|\mathbf{x}), \cdots, \mathrm{Pr}_{\theta}(y_T|\mathbf{x}, \mathbf{y}_{<T})\}$, and later use these stored probabilities to act as $\{\mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_1|\mathbf{x}), \cdots, \mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_T|\mathbf{x}, \mathbf{y}_{<T})\}$ when updating the policy with Eq. (\ref{eq:loss-function-training-llm-importance-sampling}).
However, an inherent issue arises when using importance sampling to update LLMs. As shown in Figure \ref{fig:terrible-step-importance-sampling}, when a particular optimization step results in poor updates (or a terrible step for short), the reference policy utilized in subsequent Step 1 of the next iteration inherits these poor characteristics. Consequently, the samples drawn in Step 2 are likely to be adversely affected, leading to a more terrible step. Unlike SFT\footnote{In supervised learning, a terrible step during a training step caused by bad samples can often be corrected in subsequent steps by good samples.}, this cycle does not correct the initial terrible step but potentially exacerbates it, creating a downward spiral in policy performance.
\begin{figure*}[!t]
\centering
\input{section3/Figures/figure-importance-sampling-bad-sampling}
\caption{
Illustration of the impact of a terrible step on sampling within an LLM training cycle using importance sampling. The left shows poor output sampled with slight repetition resulting in a terrible step. The right shows subsequent output sampled with more severe repetition after a terrible step.
}
\label{fig:terrible-step-importance-sampling}
\end{figure*}
Addressing this issue involves considering trust regions in optimization \citep{schulman-etal:2015trust}, which refers to a region around the current parameter estimate where the model is well-behaved. One approach to incorporating trust regions is to impose a constraint on the size of the policy update. This can be effectively managed by imposing a constraint that prevents significant deviations from the policy before optimization. In training LLMs, this constraint is typically achieved by adding a penalty to the objective function based on a divergence measure between current and reference policies. A simple form of such a penalty is given by the difference in the log-probability of the sampled $\mathbf{y}$ under the current policy versus the reference policy:
\begin{eqnarray}
\mathrm{Penalty} = \log \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) - \log \mathrm{Pr}_{\theta_{\mathrm{ref}}}(\mathbf{y}|\mathbf{x})
\end{eqnarray}
At the time step $t$, we can obtain the penalty as
\begin{eqnarray}
\mathrm{Penalty}_{t} = \log \mathrm{Pr}_{\theta}(y_t|\mathbf{x},\mathbf{y}_{<t}) - \log \mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_t|\mathbf{x},\mathbf{y}_{<t})
\end{eqnarray}
By including this penalty in the optimization objective, we encourage the current policy to remain close to the reference policy, limiting very large updates in a terrible step.
We can incorporate this penalty into the Eq. (\ref{eq:loss-function-training-llm-importance-sampling}), and obtain
\begin{eqnarray}
\mathcal{L}(\theta) & = & - \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \mathbb{E}_{\mathbf{y}\sim \mathrm{Pr}_{\theta_{\mathrm{ref}}}(\cdot|\mathbf{x})} \left[
\sum_{t=1}^{T} \frac{\mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}A_{t} - \beta \mathrm{Penalty}
\right]
\label{eq:importance-sampling-with-penalty}
\end{eqnarray}
where $\beta$ is the weight of the penalty. This training method is also called trust region policy optimization (TRPO) \citep{schulman-etal:2015trust}.
\subsection{Proximal Policy Optimization}
A further improvement to TRPO involves addressing the gradient variance problem outlined in Section \ref{sec:reduce-gradient-variance}.
In Eq. (\ref{eq:importance-sampling-with-penalty}), the range of the ratio function is from $[0, +\infty)$, which could lead to high gradient variance in the learning process. To mitigate this problem, clipping is commonly employed to limit the magnitude of importance weights, thereby preventing excessively large updates and promoting stability in the learning process. A clipped version can be given by
\begin{eqnarray}
\mathcal{L}(\theta) & = & - \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \mathbb{E}_{\mathbf{y}\sim \mathrm{Pr}_{\theta_{\mathrm{ref}}}(\cdot|\mathbf{x})} \left[
\sum_{t=1}^{T} \mathrm{Clip}\big(\frac{\mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}A_{t}\big) - \beta \mathrm{Penalty}
\right]
\label{eq:ppo-loss}
\end{eqnarray}
\begin{figure*}[!t]
\centering
\input{section3/Figures/figure-ppo-clip}
\caption{
PPO clipping at time step $t$. The basic idea can be represented with a hypothesis space as described in Figure \ref{fig:understand-policy-gradient}. The sub-figure (a) shows the case where the advantage $A_{t}$ is positive, suggesting that the output is preferred. Here, the probability of the sampled output is increased within a defined constraint (up to $1 + \epsilon$) in the hypothesis space. The sub-figure (b) shows the case where $A_{t}$ is negative, indicating a dispreferred output. In this case, the probability of the sample output is reduced to a minimum (down to $1 - \epsilon$) in the hypothesis space.
}
\label{fig:ppo-clip}
\end{figure*}
where the clipping function $\mathrm{Clip}(\cdot)$ can be defined by
\begin{eqnarray}
\mathrm{Clip}\big(\frac{\mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}A_{t}\big) & = & \min\Big(\frac{\mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}A_{t}, \mathrm{bound}(\frac{\mathrm{Pr}_{\theta}(y_{t}|\mathbf{x},\mathbf{y}_{<t})}{\mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_{t}|\mathbf{x},\mathbf{y}_{<t}}, 1-\epsilon, 1+\epsilon)A_{t} \Big)
\label{eq:cliping-function}
\end{eqnarray}
where the function $\mathrm{bound}(\cdot)$ constrains the ratio function to within the range $[1-\epsilon, 1+\epsilon]$. The clipping imposed by Eq. (\ref{eq:cliping-function}) is illustrated in Figure \ref{fig:ppo-clip}. This training method is also known as proximal policy optimization (PPO) \citep{schulman-etal:2017proximal}, currently the most widely used method for training LLMs with RL. Originally, PPO also introduced an adaptive $\beta$ to dynamically control this penalty. However, this adaptive approach has not been widely applied in training LLMs. This is because this process depends on the expectation of the penalty (i.e., $\mathbb{E}(\mathrm{Penalty})$), which would introduce additional computational overhead. Interested readers can refer to the original literature for more details on this adaptive approach.
% introduce reward models
\subsection{Training Reward Models}
\label{sec:training-reward-models}
While our initial reward function focused on output length, as described in Section \ref{sec:policy-gradient}, provides effective signals in the learning process, human preferences are considerably more complex and extend beyond merely preferring longer outputs. For example, in scenarios like a homework assistant, it is essential not only to generate longer outputs but also to avoid redundancy, ensure accuracy, and maintain fluency. This complexity underscores the need for a sophisticated reward modeling technique, moving beyond simple function-based methods to more effectively capture human preferences.
Given these complexities, we typically train a reward model, a neural network that maps a pair of input and output token sequences to a scalar value, to capture human preferences. Given an input $\mathbf{x}$ and an output $\mathbf{y}$, the reward is expressed as $\mathrm{Reward}(\mathbf{x},\mathbf{y})$, where $\mathrm{Reward}(\cdot)$ denotes the reward model. There are many ways to implement the reward model. One simple approach is to build the reward model based on a pre-trained LLM, similar to the value model as presented in Section \ref{sec:reduce-gradient-variance}. More specifically, we employ the sequence $\mathrm{seq}_{\mathbf{x},\mathbf{y}} = [\mathbf{x}, \mathbf{y}]$ to serve as the input. We run the LLM on this sequence and obtain a representation from the top-most Transformer layer. Then, we take the representation at the last position of the output and map it to a scalar via linear transformation\footnote{In practice, when employing an LLM for training a reward model, we utilize it primarily as an encoder to encode the sequence $\mathrm{seq}_{\mathbf{x}, \mathbf{y}}$ into a representation. Here, the representation from the last position is selected as it encapsulates the semantic information of the entire sequence.}:
\begin{eqnarray}
R_{\phi}(\mathbf{x},\mathbf{y}) & = & \mathbf{h}_{y_{T}} \mathbf{W}_{r}
\end{eqnarray}
where $\mathbf{W}{r}$ is a $d \times 1$ linear mapping matrix, and $\phi$ represents the parameters of the reward model, which includes both the parameters of the LLM and the $\mathbf{W}{r}$. This architecture of the reward model is illustrated in Figure \ref{fig:reward-model}.
\begin{figure*}[!t]
\centering
\input{section3/Figures/figure-reward-model}
\caption{
Architecture of the reward model based on an LLM. Unlike the value model depicted in Figure \ref{fig:architecture-value-function}, this model extracts the representation from the last position of the output $\mathbf{y}$ to represent the entire sequence $[\mathbf{x}, \mathbf{y}]$. This representation is then mapped to a scalar through a linear transformation, which serves as the reward for the output.
}
\label{fig:reward-model}
\end{figure*}
To train the reward model, the first step is to collect human feedback on a set of generated outputs. Given an input $\mathbf{x}$, we use the LLM to produce multiple candidate outputs $\{\mathbf{y}_1,...,\mathbf{y}_d\}$. Then, we can obtain human feedback in the following ways:
\begin{itemize}
\item \textbf{Rating}. Human experts provide a score or rating for each output. This score is often a continuous or discrete numerical value, such as a score on a scale (e.g., 1-5 stars or 1-10 points). In some cases, the rating might be binary, indicating a ``yes/no'' or ``positive/negative'' preference.
\item \textbf{Listwise Ranking}. Human experts are asked to rank or order the given set of possible outputs.
\item \textbf{Pairwise Comparison} (\textbf{Pairwise Ranking}). Given two different outputs, human experts select which one is better. This annotation approach is particularly popular because it is easier to annotate than the other two methods.
\end{itemize}
\vspace{0.5em}
Here we consider pairwise comparison feedback to train a reward model. In this setting, each time, two outputs $(\mathbf{y}_a,\mathbf{y}_b)$ are randomly drawn from the candidate output pool $\{\mathbf{y}_1,...,\mathbf{y}_d\}$. Human experts are then presented with these pairs and asked to decide which output they prefer based on specific criteria, such as information, accuracy, and fluency. The human feedback can be encoded as a binary label, $\mathbf{y}_a \succ \mathbf{y}_b$ for a preference for $\mathbf{y}_a$, and $\mathbf{y}_b \succ \mathbf{y}_a$ for a preference for $\mathbf{y}_b$.
One simple and widely used model for describing such pairwise comparisons is the Bradley-Terry model \citep{bradley-and-terry:rank}. It is a probabilistic model that estimates the probability that one item is preferred over another. Adapting this model to the notation used here, we can write the probability that $\mathbf{y}_a$ is preferred over $\mathbf{y}_b$ in the form
\begin{eqnarray}
\mathrm{Pr}_{\phi}(\mathbf{y}_a \succ \mathbf{y}_b | \mathbf{x}) & = & \frac{e^{R_{\phi}(\mathbf{x},\mathbf{y}_a)}}{e^{R_{\phi}(\mathbf{x},\mathbf{y}_a)} + e^{R_{\phi}(\mathbf{x},\mathbf{y}_b)}} \nonumber \\
& = & \frac{e^{R_{\phi}(\mathbf{x},\mathbf{y}_a)-R_{\phi}(\mathbf{x},\mathbf{y}_b)}}{e^{R_{\phi}(\mathbf{x},\mathbf{y}_a)-R_{\phi}(\mathbf{x},\mathbf{y}_b)}+1} \nonumber \\
& = & \mathrm{Sigmoid}(R_{\phi}(\mathbf{x},\mathbf{y}_a)-R_{\phi}(\mathbf{x},\mathbf{y}_b))
\end{eqnarray}
\begin{figure*}[!t]
\centering
\input{section3/Figures/figure-pairwise-reward-loss-expectation}
\caption{
An example of loss computation using the Bradley-Terry model. Given an input $\mathbf{x}$ and two outputs $(\mathbf{y}_{a}, \mathbf{y}_{b})$, where $\mathbf{y}_a \succ \mathbf{y}_b$, we initially obtain rewards $R(\mathbf{x}, \mathbf{y}_a)$ and $R(\mathbf{x}, \mathbf{y}_b)$ for two sequences $\mathrm{seq}_{\mathbf{x},\mathbf{y}_{a}}$ and $\mathrm{seq}_{\mathbf{x},\mathbf{y}_{b}}$. These rewards are subsequently used to compute the loss as described in Eq. (\ref{eq:pairwise-reward-loss-expectation}). Note that, for efficiency, both sequences are typically processed in a single batch during one forward pass to simultaneously obtain $R(\mathbf{x}, \mathbf{y}_a)$ and $R(\mathbf{x}, \mathbf{y}_b)$. In fact, this training approach is also common in Siamese networks \citep{bromley-etal:1993signature}.
}
\label{fig:training-reward-models}
\end{figure*}
When training the reward model, we want to maximize this preference probability. A loss function based on the Bradley-Terry model is given by
\begin{eqnarray}
\mathcal{L}_r(\phi) & = & -\mathbb{E}_{(\mathbf{x},\mathbf{y}_a,\mathbf{y}_b) \sim \mathcal{D}_r} \big[ \log \mathrm{Pr}_{\phi}(\mathbf{y}_a \succ \mathbf{y}_b | \mathbf{x}) \big] \nonumber \\
& = & - \mathbb{E}_{(\mathbf{x},\mathbf{y}_a,\mathbf{y}_b) \sim \mathcal{D}_r} \big[\log \mathrm{Sigmoid}(R_{\phi}(\mathbf{x},\mathbf{y}_a)-R_{\phi}(\mathbf{x},\mathbf{y}_b)) ]
\label{eq:pairwise-reward-loss-expectation}
\end{eqnarray}
where $(\mathbf{x},\mathbf{y}_a,\mathbf{y}_b)$ is drawn from a human-annotated dataset $\mathcal{D}_r$ consisting of preference pairs of outputs and their corresponding inputs. The goal of training the reward model is to find the optimal parameters $\hat{\phi}$ that minimize this loss function, given by
\begin{eqnarray}
\hat{\phi} & = & \argmin_{\phi} \mathcal{L}_r(\phi)
\end{eqnarray}
Since the reward model itself is also an LLM, we can directly reuse the Transformer training procedure to optimize the reward model. The difference from training a standard LLM is that we only need to replace the cross-entropy loss with the pairwise comparison loss as illustrated in Figure \ref{fig:training-reward-models}. After the training of the reward model, we can apply the trained reward model $r_{\hat{\phi}}(\cdot)$ to supervise the target LLM for alignment.
It is worth noting that although we train the reward model to perform pairwise ranking, we apply it to score each input-output pair independently during the alignment process. The pairwise ranking objective ensures that the reward model is sensitive to subtle differences between outputs, but we rely on the continuous scores produced by the reward model to guide the optimization of the LLM. An advantage of this approach is that we can choose from or combine various ranking loss functions and still apply the resulting reward models in the same way as we have done in Section \ref{sec:improved-reward-generalization}. However, a challenge arises with this method: the reward model can provide only sparse rewards; that is, it offers a delayed reward rather than an intermediate one. Hence, in this case, we can obtain
\begin{eqnarray}
r_{t} & = &
\begin{cases}
0 & \text{if } t < T \\
R_{\phi}(\mathbf{x}, \mathbf{y}) & \text{if } t = T
\end{cases}
\end{eqnarray}
To address this challenge, we can typically incorporate shaping rewards into the learning process \citep{wu-etal:2023fine} or train a process reward model by annotating process preference data \citep{lightman-etal:2023let}. Please refer to Section \ref{sec:step-by-step-verification} for more details.
\begin{figure}[!t]
\centering
\input{section3/Figures/figure-rlhf}
\caption{
Illustration of PPO-based optimization workflow. Given an optimized reward model and a reference model, we proceed to train both the policy and the value model. At each prediction step, we compute the sum of the PPO-based loss and update the policy parameters. This requires access to the reward model, the reference model, and the value model at hand. At the same time, we update the parameters of the value model. It is worth noting that steps 2, 3, and 4 do not follow a fixed order, and they can be executed simultaneously or in any order.
}
\label{fig:ppo-rlhf}
\end{figure}
% a complete workflow: key elements and workflow
\subsection{A Complete Workflow}
Up to this point, we have used numerous examples to discuss how to use RL to train LLMs to better align with human preferences. We have also shown how different RL algorithms are designed to address specific challenges in our scenarios. In this subsection, we will outline a complete workflow for training LLMs using RL, specifically employing the PPO. To implement PPO, we first build four models, all based on an LLM.
\begin{itemize}
\item \vspace{0.3em} \textbf{Reward Model} (denoted by $R_{\phi}(\cdot)$ where $\phi$ denotes the parameters). The reward model learns from human preference data to predict the reward for each pair of input and output token sequences. It is typically initialized with an SFT LLM or a pre-trained LLM.
\item \vspace{0.3em} \textbf{Value Model} or \textbf{Value Function} (denoted by $V_{\omega}(\cdot)$ where $\omega$ denotes the parameters). The value model receives rewards from the reward model and is trained to predict the expected sum of rewards. It is generally initialized with a reward model or an SFT LLM.
\item \vspace{0.3em} \textbf{Reference Model} or \textbf{Old Policy Model} (denoted by $\mathrm{Pr}_{\theta_{\mathrm{ref}}}(\cdot)$ where $\theta_{\mathrm{ref}}$ denotes the parameters). The reference model is the baseline LLM that serves as a starting point for policy training. Unlike the discussion in Section \ref{sec:importance-sampling}, when computing the penalty in RLHF, we typically use the previous version of the model or a model trained without human feedback to serve as the reference policy model, making a more stable learning process.
\item \vspace{0.3em} \textbf{Policy Model} (denoted by $\mathrm{Pr}_{\theta}(\cdot)$ where $\theta$ denotes the parameters). Given its context, this policy governs how the LLM decides the most appropriate next token. It is trained under the supervision of both the reward model and the value model.
\end{itemize}
In practice, these models need to be trained in a certain order. First, we need to initialize them using some other models. For example, the reward model and the value model can be initialized with a pre-trained LLM or an SFT LLM, while the reference model and the policy model can be initialized with an SFT LLM. Note that, at this point, the reference model is ready for use and will not be further updated. Second, we need to collect human preference data and train the reward model on this data. Third, both the value model and the policy are trained simultaneously using the reward model\footnote{During PPO-based optimization, a cold-start approach can be used where only the value model is updated in the initial phases, avoiding adjustments to the policy model based on potential inaccuracies in early value predictions \citep{wang-etal:2024hybrid}.}. At each position in an output sequence, we update the value model by minimizing the MSE error of value prediction, and the policy is updated by minimizing the PPO loss.
Although the RL process introduced above seems highly promising, as it can effectively align the SFT LLM with human preferences, there are significant challenges in implementing this process. For example,
\begin{itemize}
\item \vspace{0.5em} Training a reward model requires labeled preference data. Consequently, it is essential to generate multiple outputs for a given input and label them according to human preferences. For example, in our scenario, we need to annotate extensive data to accurately reflect student preferences, thereby enabling the homework assistant to produce the content students expect. Furthermore, in this process, it is crucial to ensure that the data is of high quality and accurately mirrors human preferences. This is because inaccuracies in data can lead to misguided learning, where the model might adopt undesirable biases or fail to meet user expectations, also called overoptimization or reward hacking \citep{gao-etal:2023scaling}. Therefore, such preference data must be meticulously collected and sufficiently extensive to enable the reward model to accurately capture human preferences.
\item \vspace{0.5cm} RL is computationally expensive for training LLMs. This arises from two primary factors. One is that during the RL process, we need to perform sampling in an autoregressive mode \citep{xiao-etal:2023introduction}. This is highly computationally intensive because the policy model usually has a large number of parameters. Furthermore, during the PPO-based optimization, we need to load four LLMs simultaneously, which requires significantly more GPU memory compared to the SFT. To address these challenges, there has been significant interest in developing efficient RL strategies, such as dynamic sampling \citep{wang-etal:2024esrl} and rule-based rewards \citep{shao-etal:2024deepseekmath}, which can maintain state-of-the-art performance without requiring high computational and time costs.
\item \vspace{0.5cm} RL is often an unstable produce. While RL provides a strong theoretical foundation for the design of each component, it is also highly sensitive to changes in any part of the system. Even small adjustments to the reward model, policy model, or hyperparameters can lead to significant fluctuations in performance, making it difficult to ensure consistent and stable results. This fragility highlights the importance of carefully tuning and stabilizing the various elements involved in RL, especially when applied to complex systems like LLMs. Techniques such as reward shaping, adaptive learning rate, and more sophisticated optimization strategies are often required to mitigate instability and improve the robustness of the learning process.
\end{itemize}
\begin{center}
\def\tgtwidth{0.8\textwidth}
\begin{tikzpicture}
\tikzstyle{ynode} = [fill=black!5!white,rounded corners=5pt,minimum width=\tgtwidth,inner sep=4pt,align=left,text width=\tgtwidth]
\begin{scope}
\node (input) at (0,0) {\large{$\mathbf{x}$}};
\node [ynode, anchor=west] (note) at ([xshift=.2cm]input.east) {
\begin{varwidth}{\textwidth}
\footnotesize{Give me three tips to improve my accuracy in solving math problems.}
\end{varwidth}
};
\node [ynode,anchor=north west] (output1) at ([yshift=-.2cm]note.south west) {
\setstretch{2.2}\\\footnotesize{There are three tips for improving accuracy in solving math problems: \\
1.Practice regularly. \\
2.Understand the concepts. \\
3.Double-check your work. \\}
};
\node [anchor=east](y1) at ([xshift=-.2cm]output1.west) {\large{$\mathbf{y}$}};
\node [fill=lolred!50, anchor=south west,minimum height=.45cm,inner sep=1pt] (label1) at ([xshift=6.85cm,yshift=-.65cm]output1.north west) {\scriptsize{Intermediate Reward: +1}};
\draw [-,lolred!50,line width=1.5pt] ([yshift=-.2pt]label1.north west) -- ([yshift=-.90cm]label1.north west);
\node [fill=lolred!50, anchor=south west,minimum height=.45cm,inner sep=1pt] (label1) at ([xshift=4.05cm,yshift=-2.25cm]output1.north west) {\scriptsize{Intermediate Reward: +1}};
\draw [-,line width=1.5pt,lolred!50] ([yshift=-.2pt]label1.north west) -- ([yshift=-.78cm]label1.north west);
\node [fill=lolblue!50, anchor=south west,minimum height=.45cm,inner sep=1pt] (label1) at ([xshift=4.1cm,yshift=-3.05cm]output1.north west) {\scriptsize{Delayed Reward: +10.6}};
\draw [-,line width=1.5pt,lolblue!50] ([yshift=-.2pt]label1.north west) -- ([yshift=-.75cm]label1.north west);
\end{scope}
\end{tikzpicture}
\end{center}
\begin{center}
\def\tgtwidth{0.5\textwidth}
\begin{tikzpicture}
\tikzstyle{ynode} = [draw,rounded corners=4pt, minimum width=\tgtwidth,inner sep=4pt,align=left,text width=\tgtwidth]
\begin{scope}
\node (input) at (0,0) {$\mathbf{x}$};
\node [ynode, anchor=west] (note) at ([xshift=.2cm]input.east) {
\begin{varwidth}{\textwidth}
\scriptsize{Give me three tips to improve my accuracy in solving math problems.}
\end{varwidth}
};
\node [ynode,anchor=north west] (output1) at ([yshift=-.2cm]note.south west) {
\scriptsize{Improving accuracy in math problem-solving is crucial for success. Here are three tips to help you enhance your accuracy: \\
1. Focus on Understanding Concepts: Before diving into solving problems, ensure you have a solid understanding of the underlying concepts. Break down the problem into smaller parts and relate each part to the concepts you've learned. Understanding the theory behind the problem will help you choose the right approach to solve it accurately. \\
2. Practice Regularly: Like any other skill, math proficiency improves with practice. Work on a variety of problems regularly to build your problem-solving skills and increase your accuracy. Start with easier problems and gradually move on to more challenging ones. Consider using resources like textbooks, online problem sets, or math apps to practice. \\
3. Check and Review Your Work: After solving a problem, always double-check your work for any errors. Pay attention to details such as signs, decimal points, and units. If time permits, try solving the problem using a different method to verify your initial answer. By reviewing your work, you can catch and correct any mistakes, improving your overall accuracy.}\\};
\node [anchor=east](y1) at ([xshift=-.2cm]output1.west) {$\mathbf{y}_a$};
\node [ynode,anchor=north] (output2) at ([yshift=-.2cm]output1.south) {
\scriptsize{There are three tips for improving accuracy in solving math problems: \\
1.Practice regularly. \\
2.Understand the concepts. \\
3.Double-check your work.\\}
};
\node [anchor=east](y2) at ([xshift=-.2cm]output2.west) {$\mathbf{y}_b$};
\begin{axis}[
compat=1.3,
anchor=north west,
at={(note.north east)},
axis x line=bottom,axis y line=left,
xshift=1.4cm,
yshift=-.5cm,
xtick={0,1,2,3,4},
xticklabels={$\mathbf{y}_a$,$\mathbf{y}_1$,$\mathbf{y}_b$,$\mathbf{y}_2$,$\cdots$},
every tick label/.append style={font=\footnotesize},
every axis label/.append style={font=\footnotesize},
xtick distance=.1cm,
ymin=0,
ymax=.6,
ylabel=$\mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})$,
ylabel shift=-.1cm,
enlarge x limits=0.3,
width=5cm,
height=4cm,
ybar,
bar shift=0pt,]
\addplot[fill=lolred,draw=none] coordinates {
(2, .34)
(3, .24)
};
\addplot[fill=lolblue,draw=none] coordinates {
(0, .19)
(1, .15)
};
\addplot[fill=lightgray,draw=none] coordinates {
(4, .08)
};
\end{axis}
\begin{axis}[
compat=1.3,
anchor=south west,
at={(output2.south east)},
axis x line=bottom,axis y line=left,
xshift=1.4cm,
yshift=.5cm,
xtick={0,1,2,3,4},
xticklabels={$\mathbf{y}_a$,$\mathbf{y}_1$,$\mathbf{y}_b$,$\mathbf{y}_2$,$\cdots$},
every tick label/.append style={font=\footnotesize},
every axis label/.append style={font=\footnotesize},
xtick distance=.1cm,
ymin=0,
ymax=.6,
ylabel=$\mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})$,
ylabel shift=-.1cm,
enlarge x limits=0.3,
width=5cm,
height=4cm,
ybar,
bar shift=0pt,]
\addplot[fill=lolred,draw=none] coordinates {
(2, .14)
(3, .04)
};
\addplot[fill=lolblue,draw=none] coordinates {
(0, .39)
(1, .35)
};
\addplot[fill=lightgray,draw=none] coordinates {
(4, .08)
};
\end{axis}
\node at ([xshift=2.35cm,yshift=-2.00cm]note.north east) {\scriptsize{Preferred}};
\node at ([xshift=3.45cm,yshift=-1.4cm]note.north east) {\scriptsize{Dispreferred}};
\node at ([xshift=2.35cm,yshift=2.25cm]output2.south east) {\scriptsize{Preferred}};
\node at ([xshift=3.45cm,yshift=1.22cm]output2.south east) {\scriptsize{Dispreferred}};
\draw [->,thick] ([xshift=3cm,yshift=-2.8cm]output1.north east) -- node [align=center,fill=white,text width=5.5cm,inner sep=2pt] {
\scriptsize{Minizing the negative probability: \\
$- \log\mathrm{Sigmoid} \big( \beta \log \frac{\mathrm{Pr}_{\theta}(\mathbf{y}_{a}|\mathbf{\mathbf{x}})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_{a}|\mathbf{\mathbf{x}})} - \beta \log \frac{\mathrm{Pr}_{\theta}(\mathbf{y}_{b}|\mathbf{\mathbf{x}})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_{b}|\mathbf{\mathbf{x}})} \big)$}
}([xshift=3cm,yshift=3.2cm]output2.south east);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}[background rectangle/.style={fill=black!5!white,rounded corners=5pt,minimum width=\textwidth}, show background rectangle]
\begin{scope}
\matrix (m) [matrix anchor=west,
matrix of nodes,
nodes={anchor=north west,}
] at (-\textwidth/2,0) {
\color{gray}{Input} &
\begin{varwidth}{18em}\small{Give me three tips to improve my accuracy in solving math problems. Return in JSON format with ``tip1'', ``tip2'', ``tip3'' as the key and the corresponding tip as the value. (\textbf{Expect a JSON output})}
\end{varwidth} \\
\color{gray}{Output 1} &
\begin{varwidth}{18em}
\small{ \{\\
\hspace*{.4cm}"tip1": " Delve into the labyrinth of equations with the focus of a hawk hunting ...", \\
\hspace*{.4cm}"tip2": " Engage in the art of problem-solving like a master alchemist turning ...", \\
\hspace*{.4cm}"tip3": "Practice, practice, practice, until the numbers themselves bow in reverence ..." \\
\} (\textbf{Valid: True})}
\end{varwidth} \\
\color{gray}{Output 2} &
\begin{varwidth}{18em}
\small{tip1: " Delve into the labyrinth of equations with the focus of a hawk hunting ...", \\
tip2: " Engage in the art of problem-solving like a master alchemist turning ...", \\
tip3: "Practice, practice, practice, until the numbers themselves bow in reverence ..."\\
(\textbf{Valid: False})}
\end{varwidth} \\
};
\path let \p1=($(m-2-2.north east)-(m-3-2.south east)$) in node [minimum height=\y1,minimum width=2cm,draw,anchor=north west,rounded corners=2pt,align=center] (counter) at ([xshift=1cm]m-2-2.north east) {Format\\Checker};
\node at ([xshift=4.5cm]m-1-2.east) {Reward};
\node (res1) at ([xshift=4.5cm]m-2-2.east) {\textbf{+1}};
\node (res2) at ([xshift=4.5cm]m-3-2.east) {\textbf{+0}};
\draw [->] (m-2-2.east) -- ([xshift=-.1cm]counter.west |- m-2-2.east);
\draw [->] (m-3-2.east) -- ([xshift=-.1cm]counter.west |- m-3-2.east);
\draw [->] ([xshift=.1cm]counter.east|-m-2-2.east) -- (res1.west);
\draw [->] ([xshift=.1cm]counter.east|-m-3-2.east) -- (res2.west);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}[background rectangle/.style={fill=black!5!white,rounded corners=5pt,minimum width=\textwidth}, show background rectangle]
% rule of math
\begin{scope}[yshift=-6cm]
\matrix (m) [matrix anchor=west,
matrix of nodes,
nodes={anchor=north west,}
] at (-\textwidth/2,0) {
\color{gray}{Input} &
\begin{varwidth}{18em}
\small{Weng earns \$12 an hour for babysitting. Yesterday, she just did 50 minutes of babysitting. How much did she earn? (\textbf{Answer is 10})}
\end{varwidth} \\
\color{gray}{Output 1} &
\begin{varwidth}{18em}
\small{Weng earns 12/60 = \$0.2 per minute. Working 50 minutes, she earned 0.2 x 50 = \$10. \\
\#\#\#\# \boxed{10} (\textbf{Res: 10})
}
\end{varwidth} \\
\color{gray}{Output 2} &
\begin{varwidth}{18em}
\small{To find out how much Weng earned from the 50 minutes of babysitting, we need ... \\
50 minutes x \$12 per hour = \$600 \\
So, Weng earned \$600 from the 50 minutes of babysitting.\\
\#\#\#\# \boxed{600} (\textbf{Res: 600})
}
\end{varwidth} \\
};
\path let \p1=($(m-2-2.north east)-(m-3-2.south east)$) in node [minimum height=\y1,minimum width=2cm,draw,anchor=north west,rounded corners=2pt,align=center] (counter) at ([xshift=1cm]m-2-2.north east) {Answer\\Verification};
\node at ([xshift=4.5cm]m-1-2.east) {Reward};
\node (res1) at ([xshift=4.5cm]m-2-2.east) {\textbf{+1}};
\node (res2) at ([xshift=4.5cm]m-3-2.east) {\textbf{-1}};
\draw [->] (m-2-2.east) -- ([xshift=-.1cm]counter.west |- m-2-2.east);
\draw [->] (m-3-2.east) -- ([xshift=-.1cm]counter.west |- m-3-2.east);
\draw [->] ([xshift=.1cm]counter.east|-m-2-2.east) -- (res1.west);
\draw [->] ([xshift=.1cm]counter.east|-m-3-2.east) -- (res2.west);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\def\ssep{1cm}
\def\nodewidth{3*\ssep}
\def\seg{0.3cm}
% LLMs as RMs
\begin{scope}
\node [anchor=south west,minimum width=3.2*\nodewidth,minimum height=1.3*\ssep,draw,thick,align=center] (llm) at (0,0) {\large{Reward Model (LLM)}};
\node [anchor=north west,minimum width=0.45*\nodewidth,,minimum height=0.45cm,rounded corners=2pt,draw,fill=gray!20] (input) at ([yshift=-\seg]llm.south west) {\footnotesize{$\mathbf{c}$}};
\node [anchor=west,minimum width=0.45*\nodewidth,,minimum height=0.45cm,rounded corners=2pt,draw,fill=gray!20] (input2) at ([xshift=6pt]input.east) {\footnotesize{$\mathbf{x}$}};
\node [anchor=west,minimum width=0.45*\nodewidth,,minimum height=0.45cm,rounded corners=2pt,draw,fill=cyan!20] (input3) at ([xshift=6pt]input2.east) {\footnotesize{$\mathbf{y}_{a}$}};
\node [anchor=west,minimum width=0.45*\nodewidth,,minimum height=0.45cm,rounded corners=2pt,draw,fill=cyan!20] (input4) at ([xshift=6pt]input3.east) {\footnotesize{$\mathbf{y}_{b}$}};
\draw [->] ([yshift=1pt]input.north) -- ([yshift=\seg-1pt]input.north);
\draw [->] ([yshift=1pt]input2.north) -- ([yshift=\seg-1pt]input2.north);
\draw [->] ([yshift=1pt]input3.north) -- ([yshift=\seg-1pt]input3.north);
\draw [->] ([yshift=1pt]input4.north) -- ([yshift=\seg-1pt]input4.north);
\node [anchor=center] (inputlabel) at ([yshift=-0.17cm]input.south) {\scriptsize{prompt}};
\node [anchor=center] (input2label) at ([yshift=-0.17cm]input2.south) {\scriptsize{input}};
\node [anchor=center] (input3label) at ([yshift=-0.18cm]input3.south) {\scriptsize{1st output}};
\node [anchor=center] (input4label) at ([yshift=-0.18cm]input4.south) {\scriptsize{2nd output}};
\node [anchor=south east,minimum width=0.45cm,minimum height=0.45cm,rounded corners=0pt,draw,fill=orange!20] (output) at ([yshift=\seg,xshift=-0.8*\seg]llm.north east) {\footnotesize{$w$}};
\draw [->] ([yshift=-\seg+1pt]output.south) -- ([yshift=-1pt]output.south);
\node [anchor=east] (tokenlabel) at (output.west) {\footnotesize{next token (`A' or `B')}};
\node [anchor=south west,align=left] (loss) at ([yshift=0.9cm]llm.north west) {\scriptsize{Minimizing the negative probability of token prediction:}\\ \scriptsize{$-\log \mathrm{Pr}_{\phi}(w=\text{A}|[\mathbf{c},\mathbf{x},\mathbf{y}_a,\mathbf{y}_b])$}};
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\tikzstyle{model} = [minimum width=2cm, minimum height=1.2cm,align=center,inner sep=2pt,draw,thick,drop shadow={fill=gray,shadow xshift=.6ex,shadow yshift=-.6ex},fill=white]
\tikzstyle{label} = [minimum width=2ex, minimum height=1.5ex,draw,scale=.75,fill=white]
\tikzstyle{circled} = [shape=circle,draw,inner sep=2pt]
\begin{scope}
\node [anchor=west] (x) at (-0.2cm,0) {$\mathbf{x}$};
\node [anchor=west,model] (policy model) at ([xshift=.8cm]x) {Policy\\Model};
\node [anchor=north west,label] at ([xshift=-.1cm,yshift=.15cm]policy model.north west) {To Learn};
\node [anchor=west,align=center] (y) at ([xshift=1.2cm]policy model.east) {$\mathbf{y}_1$\\[1ex]$\mathbf{y}_2$\\[1ex]$\cdots$\\[1ex]$\mathbf{y}_G$};
\node [anchor=west,minimum width=2cm, minimum height=1.2cm] (reward model) at ([xshift=1.4cm]y.east) {};
\node [anchor=south,model] (reference model) at ([yshift=.4cm]reward model.north) {Reference\\Model};
\node [anchor=north west,label] at ([xshift=-.1cm,yshift=.15cm]reference model.north west) {Fixed};
\node [anchor=north,model] (value model) at ([yshift=-.4cm]reward model.south) {Reward\\Model};
\node [anchor=north west,label] at ([xshift=-.1cm,yshift=.15cm]value model.north west) {Fixed};
\node [anchor=west] (reward) at ([xshift=0.1cm]reward model.east) {};
\node [anchor=south west,align=center] (value) at ([xshift=0.1cm,yshift=-.5cm]value model.south east) {$R_\phi(\mathbf{x}, \mathbf{y}_1)$\\[1ex]$R_\phi(\mathbf{x}, \mathbf{y}_2)$\\[1ex]$\cdots$\\[1ex]$R_\phi(\mathbf{x}, \mathbf{y}_G)$};
\node [anchor=west,align=center,rounded corners=5pt,inner sep=12pt,fill=lightgray] (advantage estimation) at ([xshift=1.5cm]reward) {Advantage\\Estimation};
\node [anchor=west] (advantage estimation out) at (advantage estimation.east) {$\{A^{\mathrm{grpo}}_{1,t},\cdots, A^{\mathrm{grpo}}_{G,t}\}$};
\node (left bottom) at (x.west|-value model.south) {};
\node (right bottom) at (advantage estimation out.east|-value model.south) {};
\node [anchor=north,minimum width=13.6cm, minimum height=2cm,rounded corners=8pt,draw=lightgray,thick] (note) at ($(left bottom)!.5!(right bottom)+(0, -1cm)$) {};
\matrix [anchor=center, every even column/.style={column sep=1cm}] at (note.center) {
\node {\circled{1}}; & \node [anchor=west, inner sep=0pt] {sample multiple outputs $\{\mathbf{y}_{1}, \mathbf{y}_{2}, \cdots, \mathbf{y}_{G}\}$} ; & \node {\circled{2}}; & \node [anchor=west, inner sep=0pt] {compute the $\mathrm{Penalty}$} ; \\
\node {\circled{3}}; & \node [anchor=west, inner sep=0pt] {compute the reward for all sampled outputs} ; & \node {\circled{4}}; & \node [anchor=west, inner sep=0pt] {optimize the policy model} ; \\
};
\draw [->,thick] (x.east) -- (policy model.west);
\draw [->,thick] (policy model.east) node [xshift=.5cm,yshift=.0cm,above,align=center] {\circled{1}} -- (y.west);
\draw [->,thick] (y.east) .. controls (reward model.west) and ($(y.east |- reference model.west)$) .. (reference model.west);
\draw [->,thick] (y.east) .. controls (reward model.west) and ($(y.east |- value model.west)$) .. (value model.west);
\draw [->,thick] ([xshift=2cm]value model.east) .. controls (value model.east -| advantage estimation.south) .. node [above,xshift=-.2cm] {\circled{3}}(advantage estimation.south);
\draw [->,thick, dashed] (reference model.north) .. controls ([yshift=.5cm]reference model.north) .. ([xshift=-1cm, yshift=.5cm]reference model.north) node (tmp) {} -- node [midway, above] {$\mathrm{Penalty}$} ([xshift=1cm]policy model.north|-tmp) .. controls (policy model.north|-tmp) .. ([yshift=-1cm]policy model.north|-tmp) -- node [right, solid] {\circled{2}} ([yshift=.2cm]policy model.north);
% update model
\node (policy model bottom) at ($(right bottom -| policy model)+(0,-.6cm)$) {};
\node [anchor=north] (policy model bottom1) at (policy model bottom |- advantage estimation out.north) {};
\node (value model bottom) at ($(right bottom -| value model)+(0,-.6cm)$) {};
\node (advantage estimation out bottom) at ($(advantage estimation out |- right bottom)+(0,-.6cm)$) {};
\draw [->,thick,dashed] (advantage estimation out) -- ($(advantage estimation out)!.5!(advantage estimation out bottom)$) .. controls (advantage estimation out bottom) .. ($(advantage estimation out bottom.center)!.5!(policy model bottom.center)+(5em,0)$) -- ($(advantage estimation out bottom.center)!.5!(policy model bottom.center)-(5em,0)$) .. controls (policy model bottom.center) .. ($(policy model bottom1.center)!.5!(policy model bottom)$) -- node [above,xshift=0.3cm,yshift=-.5cm,solid] {\circled{4}}(policy model.south);
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\begin{scope}
\node [draw, text width=5cm, align=left, rounded corners=2pt] (y lt t) at (0,0) {\footnotesize{Improving math accuracy requires careful attention to detail, strategic problem-solving techniques, and consistent practice. Here are three tips:}};
\def\percent{0.2}
\node [anchor=west,draw, text width=7.5cm, align=left, rounded corners=2pt,minimum width=8cm] (y gt t 1) at ([xshift=2cm]y lt t.east) {\footnotesize{1. Understand the Fundamentals: Building a strong foundation in mathematics starts with a solid understanding of the basics. ...}};
\node [anchor=south] at (y lt t.north) {$\mathbf{y}_{<t}$};
\node [anchor=north west,draw, minimum height=.75ex, minimum width=8cm,inner sep=0] at (y gt t 1.south west) {};
\node [anchor=north west, minimum height=.75ex, minimum width=8cm*\percent,fill=lolblue,inner sep=0] at (y gt t 1.south west) {};
\node [anchor=south east, draw,fill=black!5!white] at (y gt t 1.south east) {\scriptsize{Total Tokens: 60, \textbf{Reward: 6}}};
\def\percent{0.28}
\node [anchor=west,draw, text width=7.5cm, align=left, rounded corners=2pt,minimum width=8cm] (y gt t 2) at ([xshift=2cm,yshift=2cm]y lt t.east) {\footnotesize{1.Practice regularly: \\
- Consistent practice builds fluency and familiarity with different problem ... \\}};
\node [anchor=north west,draw, minimum height=.75ex, minimum width=8cm,inner sep=0] at (y gt t 2.south west) {};
\node [anchor=north west, minimum height=.75ex, minimum width=8cm*\percent,fill=lolblue,inner sep=0] at (y gt t 2.south west) {};
\node [anchor=south east, draw,fill=black!5!white] at (y gt t 2.south east) {\scriptsize{Total Tokens: 80, \textbf{Reward: 8}}};
\def\percent{0.17}
\node [anchor=west,draw, text width=7.5cm, align=left, rounded corners=2pt,minimum width=8cm] (y gt t 3) at ([xshift=2cm,yshift=4cm]y lt t.east) {\footnotesize{1. Understand the fundamentals: Solidify your understanding of basic mathematical principles such as arithmetic, algebra, ...}};
\node [anchor=north west,draw, minimum height=.75ex, minimum width=8cm,inner sep=0] at (y gt t 3.south west) {};
\node [anchor=north west, minimum height=.75ex, minimum width=8cm*\percent,fill=lolblue,inner sep=0] at (y gt t 3.south west) {};
\node [anchor=south east, draw,fill=black!5!white] at (y gt t 3.south east) {\scriptsize{Total Tokens: 50, \textbf{Reward: 5}}};
\node [anchor=west] (y gt t cdots) at ([xshift=2cm,yshift=-2cm]y lt t.east) {\LARGE$\cdots$};
\def\percent{0.78}
\node [anchor=west,draw, text width=7.5cm, align=left, rounded corners=2pt,minimum width=8cm] (y gt t 4) at ([xshift=2cm,yshift=-4cm]y lt t.east) {\footnotesize{Tip 1: Delve into the labyrinth of equations with the focus of a hawk hunting its prey. Leave no variable unturned, ...}};
\node [anchor=north west,draw, minimum height=.75ex, minimum width=8cm,inner sep=0] at (y gt t 4.south west) {};
\node [anchor=north west, minimum height=.75ex, minimum width=8cm*\percent,fill=lolred,inner sep=0] at (y gt t 4.south west) {};
\node [anchor=south east, draw,fill=black!5!white] at (y gt t 4.south east) {\scriptsize{Total Tokens: 300, \textbf{Reward: 30}}};
\draw [->] (y lt t.east) .. controls +(+1cm,0) and +(-2cm,0) .. (y gt t 3.west) node [anchor=south east] {$\mathbf{y}_{1, t:T}$};
\draw [->] (y lt t.east) .. controls +(+1cm,0) and +(-2cm,0) .. (y gt t 2.west) node [anchor=south east] {$\mathbf{y}_{2, t:T}$};
\draw [->] (y lt t.east) -- (y gt t 1.west) node [anchor=south east] {$\mathbf{y}_{3, t:T}$};
\draw [->] (y lt t.east) .. controls +(+1cm,0) and +(-2cm,0) .. (y gt t cdots.west);
\draw [->] (y lt t.east) .. controls +(+1cm,0) and +(-2cm,0) .. (y gt t 4.west) node [anchor=south east,yshift=.5cm] {$\mathbf{y}_{d,t:T}$};
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\def\ssep{.7cm}
\def\nsize{0.3cm}
\tikzstyle{lnode} = [minimum width=7cm,minimum height=1.1cm,inner sep=2pt,draw,thick,fill=white];
\begin{scope}
\node [anchor=west] (x0) at (0,0) {\footnotesize{$x_1$}};
\node [anchor=center] (x1) at ([xshift=\ssep]x0.center) {\footnotesize{$x_2$}};
\node [anchor=center] (x cdots) at ([xshift=\ssep]x1.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (x m) at ([xshift=\ssep]x cdots.center) {\footnotesize{$x_m$}};
\node [anchor=center] (y1) at ([xshift=\ssep]x m.center) {\footnotesize{$y_1$}};
\node [anchor=center] (y cdots) at ([xshift=\ssep]y1.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (y n) at ([xshift=\ssep]y cdots.center) {\footnotesize{$y_T$}};
\node [anchor=north] (y n label) at ([yshift=0.1cm]y n.south) {\scriptsize{(Last Token)}};
\foreach \x in {x0,x1,x m,y1,y n}
\draw [->] ([yshift=0.2cm]\x.center) -- ([yshift=0.55cm]\x.center);
% \node [anchor=center] (ox0) at ([yshift=2.6cm]x0.center) {\footnotesize{$\mathbf{h}_{x_1}$}};
% \node [anchor=center] (ox1) at ([yshift=2.6cm]x1.center) {\footnotesize{$\mathbf{h}_{x_2}$}};
% \node [anchor=center] (ox cdots) at ([yshift=2.6cm]x cdots.center) {\footnotesize{$\cdots$}};
% \node [anchor=center] (ox m) at ([yshift=2.6cm]x m.center) {\footnotesize{$\mathbf{h}_{x_m}$}};
\node [anchor=center] (oy1) at ([yshift=2.6cm]y1.center) {\footnotesize{$\mathbf{h}_{y_1}$}};
\node [anchor=center] (oy cdots) at ([yshift=2.6cm]y cdots.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (oy last) at ([yshift=2.6cm]y n.center) {\footnotesize{$\mathbf{h}_{y_{T}}$}};
\foreach \x in {oy1,oy last}
\draw [<-] ([yshift=-0.25cm]\x.center) -- ([yshift=-0.6cm]\x.center);
\node [anchor=south,draw,thick,minimum width=5.2cm,minimum height=1.3cm,fill=white,drop shadow={fill=gray,shadow xshift=.6ex,shadow yshift=-.6ex}] (llm) at ([yshift=.6cm]x m.center) {Reward Model (LLM)};
\node [anchor=center] (freeze) at ([xshift=-.25cm,yshift=.25cm]llm.south east) {\textcolor{cyan!60}{ \faSnowflake}};
% \node [anchor=east] (representation) at ([xshift=-0.3cm]ox0.west) {\scriptsize{Representation}};
% \node [anchor=north west] (representation2) at ([yshift=0.1cm]representation.south west) {\scriptsize{at Each Position}};
\filldraw [fill=lightgray,draw=white] ([xshift=-0.5cm,yshift=0.6cm]oy last.north west) -- ([xshift=0.5cm,yshift=0.6cm]oy last.north east) -- ([yshift=1.4cm,xshift=0]oy last.north east) -- ([yshift=1.4cm,xshift=0]oy last.north west) -- ([xshift=-0.5cm,yshift=0.6cm]oy last.north west);
\node [anchor=south] (reward) at ([yshift=1.6cm]oy last.north) {\footnotesize{Reward Model Loss}};
\node [anchor=south] (Wr) at ([yshift=0.7cm]oy last.north) {\footnotesize{$\mathbf{W}_r$}};
\node [anchor=center] (not freeze) at ([xshift=1.2em]Wr) {\textcolor{red!60}{\scriptsize \faFire}};
\node [anchor=east] (linear) at ([xshift=-0.5cm]Wr.west) {\scriptsize{Linear Map}};
\draw [->] ([yshift=-0.1cm]oy last.north) -- ([yshift=0.58cm]oy last.north);
\draw [<-] ([yshift=0.1cm]reward.south) -- ([yshift=-0.18cm]reward.south);
\node (rside) at ([xshift=0.5cm,yshift=-.6cm]oy last.north east|-y n label) {};
\node (lside) at ([yshift=-.6cm]llm.west|-y n label) {};
\node at ($(lside)!.5!(rside)$) {\scriptsize{(a)~Parameter Freezing}};
\end{scope}
\begin{scope}[xshift=\textwidth/2]
\node [anchor=west] (x0) at (0,0) {\footnotesize{$x_1$}};
\node [anchor=center] (x1) at ([xshift=\ssep]x0.center) {\footnotesize{$x_2$}};
\node [anchor=center] (x cdots) at ([xshift=\ssep]x1.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (x m) at ([xshift=\ssep]x cdots.center) {\footnotesize{$x_m$}};
\node [anchor=center] (y1) at ([xshift=\ssep]x m.center) {\footnotesize{$y_1$}};
\node [anchor=center] (y cdots) at ([xshift=\ssep]y1.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (y n) at ([xshift=\ssep]y cdots.center) {\footnotesize{$y_T$}};
\node [anchor=north] (y n label) at ([yshift=0.1cm]y n.south) {\scriptsize{(Last Token)}};
\foreach \x in {x0,x1,x m,y1,y n}
\draw [->] ([yshift=0.2cm]\x.center) -- ([yshift=0.55cm]\x.center);
\node [anchor=center,minimum width=2.2cm,minimum height=.45cm,draw,densely dashed,rounded corners=2pt] (bg) at ([yshift=2.6cm]y cdots.center) {};
% \node [anchor=center] (ox0) at ([yshift=2.6cm]x0.center) {\footnotesize{$\mathbf{h}_{x_1}$}};
% \node [anchor=center] (ox1) at ([yshift=2.6cm]x1.center) {\footnotesize{$\mathbf{h}_{x_2}$}};
% \node [anchor=center] (ox cdots) at ([yshift=2.6cm]x cdots.center) {\footnotesize{$\cdots$}};
% \node [anchor=center] (ox m) at ([yshift=2.6cm]x m.center) {\footnotesize{$\mathbf{h}_{x_m}$}};
\node [anchor=center] (oy1) at ([yshift=2.6cm]y1.center) {\footnotesize{$\mathbf{h}_{y_1}$}};
\node [anchor=center] (oy cdots) at ([yshift=2.6cm]y cdots.center) {\footnotesize{$\cdots$}};
\node [anchor=center] (oy last) at ([yshift=2.6cm]y n.center) {\footnotesize{$\mathbf{h}_{y_{T}}$}};
\foreach \x in {oy1,oy last}
\draw [<-] ([yshift=-0.25cm]\x.center) -- ([yshift=-0.6cm]\x.center);
\node [anchor=south,draw,minimum width=5.2cm,minimum height=1.3cm,fill=white,drop shadow={fill=gray,shadow xshift=.6ex,shadow yshift=-.6ex},thick] (llm) at ([yshift=.6cm]x m.center) {Reward Model (LLM)};
% \node [anchor=center] (freeze) at ([xshift=-.25cm,yshift=.25cm]llm.south east) {\textcolor{red!60}{ \faFire}};
% \node [anchor=east] (representation) at ([xshift=-0.3cm]ox0.west) {\scriptsize{Representation}};
% \node [anchor=north west] (representation2) at ([yshift=0.1cm]representation.south west) {\scriptsize{at Each Position}};
\filldraw [fill=lightgray,draw=white] ([xshift=-0.5cm,yshift=0.6cm]oy last.north west) -- ([xshift=0.5cm,yshift=0.6cm]oy last.north east) -- ([yshift=1.4cm,xshift=0]oy last.north east) -- ([yshift=1.4cm,xshift=0]oy last.north west) -- ([xshift=-0.5cm,yshift=0.6cm]oy last.north west);
\node [anchor=south] (reward) at ([yshift=1.6cm]oy last.north) {\footnotesize{Reward Model Loss}};
\node [anchor=south] (Wr) at ([yshift=0.7cm]oy last.north) {\footnotesize{$\mathbf{W}_r$}};
% \node [anchor=center] (not freeze) at ([xshift=1.2em]Wr) {\textcolor{red!60}{\scriptsize \faFire}};
% \node [anchor=east] (linear) at ([xshift=-0.5cm]Wr.west) {\scriptsize{Linear Map}};
\node [anchor=south,fill=lightgray,rounded corners=2pt,minimum height=0.8cm,align=center,minimum width=1.9cm,inner sep=0] (ffn softmax) at ([xshift=-2.5cm,yshift=0.6cm]oy last.north west) {\scriptsize{Softmax}\\\scriptsize{Layer}};
\node [anchor=center,inner sep=0] (sft loss) at (ffn softmax.north|-reward.center) {\footnotesize{SFT Loss}};
% \draw [->,thick,dashed] (oy1.north) .. controls ([yshift=-.2cm]oy1.north|-ffn softmax.east) .. ([yshift=-.2cm]ffn softmax.east);
\draw [->,densely dashed] (bg.north) .. controls (bg.north|-ffn softmax.south) and (ffn softmax.south|-bg.north) .. (ffn softmax.south);
\draw [->] ([yshift=.05cm]ffn softmax.north) -- (sft loss.south);
\draw [->] ([yshift=-0.1cm]oy last.north) -- ([yshift=0.58cm]oy last.north);
\draw [<-] ([yshift=0.1cm]reward.south) -- ([yshift=-0.18cm]reward.south);
\node (rside) at ([xshift=0.5cm,yshift=-.6cm]oy last.north east|-y n label) {};
\node (lside) at ([yshift=-.6cm]llm.west|-y n label) {};
\node at ($(lside)!.5!(rside)$) {\scriptsize{(b)~Regularization}};
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\def\ssep{0.9cm}
\tikzstyle{lnode} = [minimum width=3cm,minimum height=1.0cm,draw,thick,fill=white,inner sep=2pt,drop shadow={fill=gray,shadow xshift=.6ex,shadow yshift=-.6ex}]; % rounded corners=1pt,
%%% RLHF
\begin{scope}
\node [anchor=west,fill=gray!30,inner sep=5pt] (preferencedata) at (0,0) {\Large{$\mathbf{y}_a \succ \mathbf{y}_b$}};
\node [anchor=south west,align=left] (preferencelabel) at ([yshift=0.1cm]preferencedata.north west) {\scriptsize{Preference}\\[-1mm] \scriptsize{Data}};
\node [lnode,anchor=west] (rewardmodel) at ([xshift=2.5cm]preferencedata.east) {Reward Model};
\node [anchor=north west,minimum height=5.8cm,minimum width=3.8cm,rounded corners=5pt,fill=lightgray] (bg) at ([xshift=2.1cm,yshift=2.9cm]rewardmodel.east) {};
\node [lnode,anchor=south west] (valuefunction) at ([xshift=2.5cm,yshift=1.4cm]rewardmodel.east) {Value Model};
\node [lnode,anchor=west] (policy) at ([xshift=2.5cm,yshift=-0.0cm]rewardmodel.east) {Policy Model};
\node [lnode,anchor=north west] (reference) at ([xshift=2.5cm,yshift=-1.4cm]rewardmodel.east) {Reference Model};
\node [align=left] at ([yshift=-1.35cm]reference) {\scriptsize{Training with PPO}};
\draw [->,thick] ([xshift=2pt]preferencedata.east)--([xshift=-2pt]rewardmodel.west) node [pos=0.5,above,align=left,xshift=-0.05cm] {\scriptsize{training with MLE}};
\draw [->,thick] ([xshift=2pt]rewardmodel.east)--([xshift=-2pt]bg.west);
% \draw [->,thick] (rewardmodel.east) .. controls ($(rewardmodel.east)!.5!(policy.west)$) and (rewardmodel.east|-valuefunction.west) .. (valuefunction.west);
% \draw [->,thick] (rewardmodel.east) -- (policy.west);
% \draw [->,thick] (rewardmodel.east) .. controls ($(rewardmodel.east)!.5!(policy.west)$) and (rewardmodel.east|-reference.west) .. (reference.west);
\draw [->,dotted,very thick] ([xshift=-0.6cm,yshift=-2pt]valuefunction.south) .. controls +(-150:0.5cm) and +(150:0.5cm) .. ([xshift=-0.6cm,yshift=2pt]policy.north);
\draw [<-,dotted,very thick] ([xshift=0.6cm,yshift=-2pt]valuefunction.south) .. controls +(-30:0.5cm) and +(30:0.5cm) .. ([xshift=0.6cm,yshift=2pt]policy.north);
\draw [<-,dotted,very thick] ([xshift=-0.6cm,yshift=+2pt]reference.north) .. controls +(150:0.5cm) and +(-150:0.5cm) .. ([xshift=-0.6cm,yshift=-2pt]policy.south);
\draw [->,dotted,very thick] ([xshift=0.6cm,yshift=+2pt]reference.north) .. controls +(30:0.5cm) and +(-30:0.5cm) .. ([xshift=0.6cm,yshift=-2pt]policy.south);
\node (mid) at ($(preferencedata.west)!.5!(bg.east)$) {};
\node [anchor=north] (caption) at ([yshift=-.5cm]mid.center|-bg.south) {\small{(a) Training LLMs with PPO}};
\end{scope}
%%% DPO
\begin{scope}[yshift=-5.5cm]
\node [anchor=west,fill=gray!30,inner sep=5pt] (preferencedata) at (0,0) {\Large{$\mathbf{y}_a \succ \mathbf{y}_b$}};
\node [anchor=south west,align=left] (preferencelabel) at ([yshift=0.1cm]preferencedata.north west) {\scriptsize{Preference}\\ [-1mm] \scriptsize{Data}};
\node [anchor=north west,minimum height=2cm,minimum width=3.8cm,rounded corners=5pt,fill=lightgray] (bg2) at ([xshift=7.6cm,yshift=1cm]preferencedata.east) {};
\node [lnode,anchor=west] (policy) at ([xshift=8cm]preferencedata.east) {Policy Model};
\draw [->,thick] ([xshift=2pt]preferencedata.east)--([xshift=-2pt]bg2.west) node [pos=0.5,above,align=left] {\scriptsize{training with MLE}};
\node (mid2) at ($(preferencedata.west)!.5!(bg2.east)$) {};
\node [anchor=north] (caption b) at ([xshift=-0.0cm,yshift=-.18cm]mid2.center|-bg2.south) {\small{(b) Training LLMs with DPO}};
\end{scope}
\end{tikzpicture}
\end{center}
\begin{tabular}{ll}
\toprule
\textbf{Method} & \textbf{Objective} \\ \midrule
IPO \citep{azar:2024general} & $ \left( \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_a|\mathbf{x})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_a|\mathbf{x})} - \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_b|\mathbf{x})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_b|\mathbf{x})} - \frac{1}{2\tau} \right)^2$ \\ \midrule
CPO \citep{xu:2024contrastive} & $-\log \textrm{Sigmoid} \left(\beta \log \mathrm{Pr}_\theta(\mathbf{y}_a|\mathbf{x}) - \beta \log \mathrm{Pr}_\theta(\mathbf{y}_b|\mathbf{x}) \right) - \lambda \log \mathrm{Pr}_\theta (\mathbf{y}_a|\mathbf{x})$ \\ \midrule
\multirow{2}{*}{KTO \citep{ethayarajh:2024kto}} & $-\lambda_a \textrm{Sigmoid} \left( \beta \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_a|\mathbf{x})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_a|\mathbf{x})} - z_{\theta_\mathrm{ref}} \right) + \lambda_b \textrm{Sigmoid} \left( z_{\theta_\mathrm{ref}} - \beta \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_b|\mathbf{x})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_b|\mathbf{x})} \right)\,$ \\
& $\text{where} \,\, z_{\theta_\mathrm{ref}} = \mathbb{E}_{(\mathbf{x}, \mathbf{y}) \sim \mathcal{S}} \left[\beta \text{KL}\left( \mathrm{Pr}_\theta(\mathbf{y}|\mathbf{x}) || \mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}|\mathbf{x}) \right) \right]$ \\ \midrule
\multirow{2}{*}{ORPO \citep{hong:2024orpo}} & $-\log p_\theta(\mathbf{y}_a|\mathbf{x}) - \lambda \log \textrm{Sigmoid} \left(\log \frac{p_\theta(\mathbf{y}_a|\mathbf{x})}{1 - p_\theta(\mathbf{y}_a|\mathbf{x})} - \log \frac{p_\theta(\mathbf{y}_b|\mathbf{x})}{1 - p_\theta(\mathbf{y}_b|\mathbf{x})} \right)\,$ \\
& $\text{where} \,\, p_\theta(\mathbf{y}|\mathbf{x}) = e^{ \frac{1}{|\mathbf{y}|} \log \mathrm{Pr}_\theta(\mathbf{y}|\mathbf{x})}$ \\ \midrule
R-DPO \citep{gallego:2024refined} & $-\log \textrm{Sigmoid} \left( \beta \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_a|\mathbf{x})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_a|\mathbf{x})} - \beta \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_b|\mathbf{x})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_b|\mathbf{x})} + \left(\alpha |\mathbf{y}_a| - \alpha |\mathbf{y}_b| \right) \right)$ \\ \midrule
\multirow{2}{*}{PCDPO \citep{zhou:2024prior}} & $-\log \textrm{Sigmoid} \left(\Delta^* - \Delta_{\mathrm{Pr}_{\theta}}\right) -\log \textrm{Sigmoid} \Delta_{\mathrm{Pr}_{\theta}} \,,$ \\
& $\text{where} \,\, \Delta_{\mathrm{Pr}_{\theta}}= \beta \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_a|\mathbf{x})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_a|\mathbf{x})} - \beta \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_b|\mathbf{x})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_b|\mathbf{x})},\,\, \Delta^* = \frac{\beta_1}{\textrm{Sim}(\textbf{y}_a,\textbf{y}_b|\textbf{x})+\beta_2}+\beta_3 $ \\ \midrule
SimPO \citep{meng:2025simpo} & $-\log \textrm{Sigmoid} \left( \frac{\beta}{|\mathbf{y}_a|} \log \mathrm{Pr}_\theta(\mathbf{y}_a|\mathbf{x}) - \frac{\beta}{|\mathbf{y}_b|} \log \mathrm{Pr}_\theta(\mathbf{y}_b|\mathbf{x}) - \gamma \right)$ \\ \midrule
D$^2$PO \citep{shao:2025earlier} & $-\log \textrm{Sigmoid} \left( \sum_{t=0}^{T}\gamma^t\beta \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_a^t|\mathbf{x},\mathbf{y}_{a,<t})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_{a,t}|\mathbf{x},\mathbf{y}_{a,<t})} - \sum_{t=0}^{T}\gamma^t\beta \log \frac{\mathrm{Pr}_\theta(\mathbf{y}_{b,t}|\mathbf{x},\mathbf{y}_{b,<t})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_{b,t}|\mathbf{x},\mathbf{y}_{b,<t})}\right)$ \\
\bottomrule
\end{tabular}
\section{Improved Reinforcement Learning for LLMs}
In the previous section, we introduced some improvements to RL, such as importance sampling and reward baseline techniques. However, directly applying them to train LLMs still presents numerous challenges. In this section, we will delve deeper into the improvements for using RL to train LLMs.
\subsection{Advanced Reward Models}
Reward models are a fundamental component in RL, as they define what a policy model optimizes for. Therefore, the quality of the reward model training directly influences the effectiveness of the LLM optimization during RLHF. A poorly trained reward model can lead to suboptimal LLM, while a well-optimized reward model ensures that the LLM is effectively aligned with the desired behaviours and objectives. Here we will introduce several methods for obtaining advanced reward models. Our discussion will be relatively general, and since the reward model is widely used in many RL problems, such as robot planning and control. This broad applicability makes it straightforward to adapt the methods discussed here to other related applications.
\subsubsection{Automatic Preference Data Generation}
\label{sec:automatic-preference-data-generation}
Although learning from human preferences is an effective and popular method for aligning LLMs, annotating preference data is costly. Relying on human feedback not only faces scalability limitations but may also introduce bias, as human feedback is inherently subjective. Consequently, AI-based feedback methods offer a promising solution to address these scalability and consistency issues, avoiding the limitations associated with human annotators.
One simple method is to generate preference data using LLM. Given a set of inputs, we first use an LLM to generate pairs of outputs. Then, we prompt the LLM to label the preference between each pair of outputs, along with its corresponding input. Below is an example of prompting the LLM to generate a preference label for a pair of outputs.
\vspace{0.5em}
\begin{tcolorbox}[frame empty]
\begingroup
\setlength{\leftskip}{2em}
\setlength{\rightskip}{2em}
Consider a homework assistant scenario in which a student types an input. You will review two outputs to this input. Please indicate which output you prefer. A good output should be informative, accurate, and well-articulated. It should directly address the student's question, provide thorough explanations or solutions, and maintain an encouraging tone.
\vspace{0.3cm}
Input:
\vspace{0.1cm}
\textit{Give me three tips to improve my accuracy in solving math problems.}
\vspace{0.3cm}
Output A:
\vspace{0.1cm}
\textit{Improving math accuracy requires careful attention to detail, strategic problem-solving techniques, and consistent practice. Here are three tips: \\
1.Practice regularly: \\
- Consistent practice builds fluency and familiarity with different problem types... \\
... \\
By combining your advice with these additional tips, anyone can significantly enhance their math problem-solving abilities.}
\vspace{0.3cm}
Output B:
\vspace{0.1cm}
\textit{There are three tips for improving accuracy in solving math problems: \\
1.Practice regularly. \\
2.Understand the concepts. \\
3.Double-check your work.}
\vspace{0.3cm}
\underline{Output A is preferred.}
\endgroup
\end{tcolorbox}
\vspace{0.5em}
Once we collect such preference labels, we can use them, along with the output pair and input, to train the reward model. Of course, these labels are not entirely accurate either. Recent studies have shown that AI-based feedback often exhibits a location bias problem, making it more likely to prefer the output at the front \citep{zheng-etal:2023judging}. We can consider demonstrating a few examples or using advanced prompting techniques, such as Chain-of-Though (CoT), to improve the labeling performance \citep{liu-etal:2023GEval}.
For data generation, although it is easy to scale up, it is often necessary to ensure the data is accurate and diverse. Here, the data quality and diversity issues involve not only the labeling of preferences but also the inputs and outputs of the model. Therefore, we often need to use a variety of techniques to obtain large-scale, high-quality data. For example, one can generate diverse model outputs and annotations by using different LLMs, prompts, in-context demonstrations, and so on \citep{cui-etal:2024ultra}. Furthermore, \citet{dubois-etal:2024alpacafarm} report that the variability in pairwise preference data is important for training LLMs from either human or AI feedback.
While learning from AI feedback is highly scalable and generally objective, this method is more suited to well-defined tasks where objective performance metrics are available. By contrast, learning from human feedback is more advantageous when aligning AI systems with human values, preferences, and complex real-world tasks that require an understanding of subtle or subjective context. These methods can be combined to train LLMs that benefit from both human insights and the scalability of AI feedback.
\begin{figure*}[!t]
\centering
\input{section4/Figures/figure-an-overlap-example-dense-reward-model}
\caption{
An example of reward shaping. Here, length-based shaping rewards are implemented, e.g., awarding a +1 intermediate reward for every 10 tokens generated. Additionally, a delayed reward from the reward model evaluates the overall quality of the output at the end.
}
\label{fig:example-shaping-rewards}
\end{figure*}
\subsubsection{Reward Shaping}
\label{sec:reward-shaping}
As discussed in Section \ref{sec:training-reward-models}, while the reward model is effective in capturing human preferences, it provides sparse rewards. These rewards, known as delayed rewards, are only at the end of the output generation process, as opposed to intermediate rewards that would be distributed continuously throughout it.
In fact, dealing with sparse rewards has long been a concern in RL, and has been one of the challenges in many practical applications. For example, robotics often needs to shape the reward function to ease optimization rather than relying solely on end-of-sequence rewards. Various methods have been developed to address this issue. One common approach is reward shaping, where the original function is modified to include intermediate rewards, thereby providing more immediate feedback. Here, the intermediate reward is often an indirect way to be able to improve this delayed reward. For example, as shown in Figure \ref{fig:example-shaping-rewards}, setting a length-based reward as an intermediate reward encourages the generation of more content, potentially increasing the overall quality of the final output as assessed by the delayed reward from the reward model. Additional examples of reward shaping in training LLMs can be found in \citet{kumar-etal:2024training}. In this way, we can obtain
\begin{eqnarray}
r'_{t} & = & r_{t} + f(\mathbf{x}, \mathbf{y}_{<t}, y_{t})
\label{eq:transformed-reward-function}
\end{eqnarray}
where $r'(\cdot)$ is the transformed reward, $r(\cdot)$ is the original delayed reward function, and $f(\cdot)$ is the shaping reward function that provides an intermediate reward. To ensure the optimality of the policy under the transformed reward function, the shaping reward function can be given in the form
\begin{eqnarray}
f(\mathbf{x}, \mathbf{y}_{<t}, y_{t}) & = & \gamma \Phi(\mathbf{x}, \mathbf{y}_{<t+1}, y_{t+1}) - \Phi(\mathbf{x}, \mathbf{y}_{<t}, y_{t}) \label{eq:shaping-reward-function}
\end{eqnarray}
where $\Phi(\cdot)$ is called the potential value function. If we define $\Phi(\cdot)$ as the common value function and substitute Eq. (\ref{eq:shaping-reward-function}) into Eq. (\ref{eq:transformed-reward-function}), we obtain
\begin{eqnarray}
r'_{t} & = & r_{t} + \gamma V_{t+1} - V_{t}
\end{eqnarray}
It is interesting to see that this function is exactly the same as the advantage function used in PPO. This relates advantage-based methods to reward shaping: the advantage is essentially a shaped reward.
In practice, an alternative method for implementing reward shaping in training LLMs involves utilizing the reward model to provide intermediate rewards based on the differences between output rewards \citep{bahdanau-etal:2016actor,wu-etal:2023fine}. For example, at time step $t$, the intermediate reward $r_t$ can be obtained by
\begin{eqnarray}
r_{t} & = & R(\mathbf{x}, \mathbf{y}_{<t+1}) - R(\mathbf{x}, \mathbf{y}_{<t})
\end{eqnarray}
However, in this case, it is crucial that the reward model is capable of accurately assigning rewards for partial outputs, i.e., $\{\mathbf{y}_{<2},\mathbf{y}_{<3},\cdots,\mathbf{y}_{<T}\}$.
In addition to reward shaping, another method to address the sparse reward issue is adopting curriculum learning, where tasks are structured sequentially with gradually increasing complexity. This can help models master simpler tasks first, which prepares them for more complex challenges as their skills develop. There are many methods that can mitigate the impact of sparse rewards, such as Monte Carlo methods and intrinsic motivation. Most of these methods are general, and the discussion of them can be found in the broader literature on RL, such as \citet{Sutton-and-Barto:2018RL}'s book.
\subsubsection{Improved Reward Generalization}
\label{sec:improved-reward-generalization}
\begin{figure*}[!t]
\centering
\input{section4/Figures/figure-methods-improving-reward-generalization}
\caption{
When training reward models, we can use parameter freezing and regularization methods to preserve the features of the LLM and thereby improve reward generalization.
}
\label{fig:improve-reward-generalization}
\end{figure*}
As discussed in Section \ref{sec:training-reward-models}, reward models are trained using preference data and subsequently used to optimize LLMs. However, a problem arises: the data distribution during LLM optimization may differ from the distribution of the preference data, making it challenging for the reward model to generalize to unseen input-output pairs.
A well-known failure mode associated with this problem is commonly referred to as \textit{overoptimization} or \textit{reward hacking}, where the optimization stage improves the reward model but deteriorates the alignment with true rewards \citep{gao-etal:2023scaling,eisenstein-etal:2023helping}. This mode occurs because the reward model may incorrectly assign high rewards to unseen input-output pairs, leading the LLM to learn and optimize for behaviours that do not truly align with the desired behaviours and objectives. For example, consider a scenario from Section \ref{sec:policy-gradient} where the reward model is designed to favor informative and accurate outputs for a homework assistant. However, if the reward model fails to generalize to an overly verbose output and assigns a high reward to it (perhaps due to its length or certain keywords), the LLM may learn to prioritize generating a long output, which is not actually more informative but reward better according to the reward model. Consequently, the LLM becomes misaligned with its true objectives—delivering concise and relevant information—because it optimizes for the incorrect reward signal from the reward model with weak generalization.
Addressing this generalization problem is challenging, and no mature solution exists yet. The ideal approach would be to develop an oracle reward model that perfectly captures the true objectives of the task and generalizes across all input-output pairs during LLM optimization. However, creating such a model is extremely difficult due to the complexity of the real-world environment, as well as the challenge of collecting sufficient preference data. Instead, a more practical approach is to combine multiple reward models, improving generalization and proving more correct rewards \citep{coste-etal:2024reward}.
Given a set of reward models, combining them is straightforward, and in some cases, we can simply treat this problem as an ensemble learning problem. A simple yet common approach is to average the outputs of these models to obtain a more precise reward estimation:
\begin{eqnarray}
R_{\mathrm{combine}}(\mathbf{x}, \mathbf{y}) & = & \frac{1}{K} \sum_{k=1}^{K} w_k \cdot R_k(\mathbf{x},\mathbf{y})
\end{eqnarray}
where $R_k(\cdot)$ is the $k$-th reward model in the ensemble, $w_k$ is the weight of $r_k(\cdot)$, and $K$ is the number of reward models. This combined reward can then be used to supervise the training of a policy. In fact, there are many ways to combine different models; for example, one can make predictions using Bayesian model averaging or develop a fusion network to learn to combine the predictions from different models. Alternatively, one can frame this task as a multi-objective optimization problem and use multiple reward models to train the policy simultaneously. These methods have been intensively discussed in the literature on optimization and machine learning \citep{miettinen:1999nonlinear,Bishop:2006}.
On the other hand, to improve the generalization of reward models, it is important to prevent overfitting to preference data. Reward models are typically trained on an SFT LLM, which has a strong generalization capability. However, training the LLM with a large amount of preference data could lead to overfitting, which may ultimately reduce generalization performance. A common strategy to mitigate this issue is to employ a parameter freezing technique, which helps preserve the original features of the LLM while learning the reward model. Another practical approach is to add a regularization term to the preference learning loss function, which helps regulate the features of the LLM \citep{yang-etal:2024regularizing}. For example, we can use a simple SFT loss as regularization by adding a term that maximizes the probability of the preferred output $\mathbf{y}_a$ to Eq. (\ref{eq:pairwise-reward-loss-expectation}):
\begin{eqnarray}
\mathcal{L}_\mathrm{reg}(\phi) & = & -\mathbb{E}_{(\mathbf{x},\mathbf{y}_a,\mathbf{y}_b) \sim \mathcal{D}_r} \big[ \log \mathrm{Pr}_{\phi}(\mathbf{y}_a \succ \mathbf{y}_b | \mathbf{x}) + \alpha \log(\mathrm{Pr}_{\theta}(\mathbf{y}_{a}|\mathbf{x})) \big]
\end{eqnarray}
where $\alpha$ is a balancing factor, we further illustrate the freezing parameter and regularization methods in Figure \ref{fig:improve-reward-generalization}. Note that the regularization term applies only to the LLM. That is, when optimizing this term, only the parameters of the LLM are updated, while the parameters of the reward linear map remain fixed. Therefore, in this equation, the parameter associated with the regularization term is $\theta$, which specifically refers to the parameters of the LLM.
\begin{figure}[!t]
\centering
\input{section4/Figures/figure-generative-reward-model-architecture}
\caption{In generative reward models, we use an LLM to predict the label token given a prompt, an input, and a pair of outputs. This model can be trained in the same way as standard LLMs.}
\label{fig:generative-reward-model-architecture}
\end{figure}
\subsubsection{Generative Reward Models}
\label{sec:generative-reward-models}
Reward models are typically trained as discriminative models to assign numerical rewards to outputs and classify them as preferred or dispreferred. However, this method does not leverage the text-generation capabilities for which LLMs are fundamentally designed. For example, the discriminative reward model can not perform CoT reasoning. To address this, an LLM can alternatively be employed as a reward model, thus endowing it with the ability to engage in text generation and reasoning, as depicted in Figure \ref{fig:generative-reward-model-architecture}. This model works as follows. First, we input a prompt $\mathbf{c}$, along with the tuple $(\mathbf{x},\mathbf{y}_{a},\mathbf{y}_{b})$, to the LLM. The prompt is a description of the task, as demonstrated in the example below.
\vspace{0.5em}
\begin{tcolorbox}[frame empty]
\begingroup
\setlength{\leftskip}{2em}
\setlength{\rightskip}{2em}
You are given two outputs to an input. Evaluate which output is better based on quality, relevance, and clarity. If the first output is better, return `A'. If the second output is better, return `B'.
\endgroup
\end{tcolorbox}
\vspace{0.5em}
Then, the LLM predicts subsequent tokens based on this input sequence. Let $w$ be the label token predicted by the LLM. If $w=\text{A}$, it indicates a preference for $\mathbf{y}_a$ over $\mathbf{y}_b$; if $w=\text{B}$, then $\mathbf{y}_b$ is preferred. Note that here $\mathbf{y}_a$ and $\mathbf{y}_b$ do not have a pre-defined preference relationship as described in Section \ref{sec:training-reward-models}. Their relationship is instead represented by the label token.
The loss function can be defined as the log-probability of predicting `A':
\begin{eqnarray}
\mathcal{L}_\mathrm{gen}(\theta) & = & -\mathbb{E}_{(\mathbf{c}, \mathbf{x},\mathbf{y}_a,\mathbf{y}_b) \sim \mathcal{D}_r} \big[ \log \mathrm{Pr}_{\theta}(w=\text{A}|\mathbf{s}) \big]
\label{eq:gen-reward-modeling}
\end{eqnarray}
where $\mathbf{s}$ denotes the string $[\mathbf{c},\mathbf{x},\mathbf{y}_a,\mathbf{y}_b]$\footnote{In this work, we will use $\mathbf{s}$ interchangeably to refer to either the tuple of a training sample or a string representing that tuple.}, and $\pi_{\phi}(\cdot)$ denotes the probability of token prediction by the LLM.
Although we discuss methods for training a generative reward model here, an interesting question arises: how can we use the generative reward model in RLHF? We provide some guidance as follows. Specifically, when applying this generative reward model to provide a reward for an input-output pair $(\mathbf{x}',\mathbf{y}')$, we can generate a reference output $\mathbf{y}_{\mathrm{ref}}$ by using the LLM, for example, through greedy search, and concatenate $\mathbf{x}'$, $\mathbf{y}'$ and $\mathbf{y}_{\mathrm{ref}}$ into $\mathbf{s}' = [\mathbf{c}, \mathbf{x}', \mathbf{y}', \mathbf{y}_{\mathrm{ref}}]$.
Additionally, to mitigate the positional bias problem \citep{wang-etal:2023large}, we can introduce an alternative input order by transposing the positions of output, i.e., presenting $\mathbf{y}_{\mathrm{ref}}$ before $\mathbf{y}'$, to construct a secondary input string $\mathbf{s}'_{T} = [\mathbf{c}, \mathbf{x}', \mathbf{y}_{\mathrm{ref}}, \mathbf{y}']$.
The reward for $(\mathbf{x}', \mathbf{y}')$ is thus defined as the log-probability that $\mathbf{y}'$ is preferred over $\mathbf{y}_{\mathrm{ref}}$:
\begin{eqnarray}
r_{\phi}(\mathbf{x}', \mathbf{y}') & = & \frac{\mathrm{Pr}_{\theta}(w=\text{A}|\mathbf{s}')+\mathrm{Pr}_{\theta}(w=\text{B}|\mathbf{s}'_{T})}{2}
\label{eq:apply-generative-rm}
\end{eqnarray}
where the reward ranges from 0 to 1.
To further improve the generative reward model, we can label the relevant explanation to enable the generative reward model to produce a CoT rationale \citep{zhang-etal:2024generative}. In this case, we use a prompt $\mathbf{c}$ with a CoT rationale generation instruction, as shown below.
\vspace{0.5em}
\begin{tcolorbox}[frame empty]
\begingroup
\setlength{\leftskip}{2em}
\setlength{\rightskip}{2em}
You are given two outputs to a user input. Evaluate which output is better based on quality, relevance, and clarity. If the first output is better, return `A'. If the second output is better, return `B'. \uline{Note that, before presenting the evaluation results, you should provide a detailed analytical process and reasoning steps.}
\endgroup
\end{tcolorbox}
\vspace{0.5em}
We can then define a new loss function for training the generative reward model as follows:
\begin{eqnarray}
\mathcal{L}_\mathrm{gen}(\theta) & = & -\mathbb{E}_{(\mathbf{c},\mathbf{x},\mathbf{y}_a,\mathbf{y}_b,\mathbf{rat}) \sim \mathcal{D}_r} \big[ \log \mathrm{Pr}_{\theta}(\mathbf{rat}|\mathbf{s}) + \log \mathrm{Pr}_{\theta}(w=\text{A}|[\mathbf{s},\mathbf{rat}]) \big]
\end{eqnarray}
where $\mathbf{rat}$ is the labeled CoT rationale for generating a preference between $\mathbf{y}_a$ and $\mathbf{y}_b$. In practice, we can obtain the evaluation results by prompting a standard LLM (as known as LLM-as-a-judge), as demonstrated in Section \ref{sec:automatic-preference-data-generation}. However, this approach typically underperforms compared to LLMs trained with preference data \citep{zhang-etal:2024generative,mahan-etal:2024generative}. One reason is that standard LLMs are not fine-tuned specifically for the reward modeling task, and as a result, they may not fully capture the nuanced decision-making process that aligns better with human preferences. Also, LLMs trained with preference data learn to prioritize outputs based on feedback, enhancing their capability to evaluate outputs according to the desired behaviors and objectives.
\subsection{Better Advantage Estimation}
Accurate advantage estimation is critical in RL, particularly when optimizing the policy model to truly reflect the potential benefits of different actions (or tokens in training LLMs). In this subsection, we will delve into methods for improving advantage estimation, providing more accurate advantages in the process of optimizing the policy model.
\subsubsection{Temporal Difference-based Advantage Estimation}
Let us first review the Monte Carlo-based advantage estimation. As discussed in Section \ref{sec:reduce-gradient-variance}, outputs are sampled and their values computed. The advantage can be obtained by comparing the actual received rewards with the predicted values at time step $t$:
\begin{eqnarray}
A_{t} & = & \sum_{t=k}^{T}r_{k} - V_{t}
\label{eq:advantage-estimation-monte-carlo}
\end{eqnarray}
While the Monte Carlo-based estimation of advantage provides a relatively stable training process, there are two main challenges associated with it:
\begin{itemize}
\item Sampling an entire output for each input is necessary. In some cases where immediate rewards are available, it could be more efficient to update the policy model with partial output. Unfortunately, when using Eq. (\ref{eq:advantage-estimation-monte-carlo}) for advantage estimation, this becomes unfeasible. This is because we must compute the term $\sum_{t=k}^T r_{k}$, which requires the rewards of the entire output.
\item This method still results in high variance. Since a single sample operation might be influenced by random factors, the estimated results may not be stable, as illustrated in Figure \ref{fig:monte-carlo-advantage-issue}.
\end{itemize}
\begin{figure*}[!t]
\centering
\input{section4/Figures/figure-issue-monte-carlo}
\caption{
Illustration of the issue of high variance in Monte Carlo-based advantage estimation, using a length-based reward function as described in Eq. (\ref{eq:langth-based-reward-function}).
Typically, if the value model is well-trained, it would predict an average $V_{t}$ of 6 at time step $t$, based on most outputs having around 60 tokens in length. However, for outlier outputs like $\mathbf{y}_{d,t:T}$, which might extend up to 300 tokens, the advantage estimation can exhibit high variance. For instance, the advantage for such an outlier could be computed as $V_{t}(\mathbf{x},\mathbf{y}_{d,<t},y_{d,t}) = \sum_{k=t}^T r_{k} - V_t = 30 - 6 = 24$, a stark contrast to another output, say $\mathbf{y}_{1,t:T}$, which is closer to the average length, where the advantage might be $V_{t}(\mathbf{x},\mathbf{y}_{1,<t},y_{1,t}) = -1$. This discrepancy leads to high gradient variance, potentially destabilizing the learning process.
}
\label{fig:monte-carlo-advantage-issue}
\end{figure*}
To address these challenges, an alternative approach is the temporal difference-based advantage estimation \citep{sutton-and-richard:1988learning}, which allows for estimating the advantage using incomplete outputs. More specifically, this method computes the difference between the predicted values at consecutive time steps (i.e., current and next stapes), adjusting the advantage estimate based on new information as it becomes available, thus reducing the dependency on the entire output. In a practical implementation, the temporal difference-based advantage estimation can be given by
\begin{eqnarray}
A_{t}^{\mathrm{TD}} & = & r_{t} + \gamma V_{t+1} - V_{t}
\label{eq:td-advantage-estimation}
\end{eqnarray}
By focusing on the current and next steps, temporal difference-based advantage estimation minimizes the influence of distant future events. This leads to lower variance in optimizing the policy model and improves the stability of policy training.
\subsubsection{Generalized Advantage Estimation}
While temporal difference-based estimation effectively reduces variance, it can also increase estimation bias due to its heavy reliance on the predictions of the value model. By contrast, one effective method is generalized advantage estimation (GAE), which is one of the most popular advantage estimation methods used in LLMs and other fields \citep{schulman-etal:2015high}. This method mainly refines the advantage estimation by using exponentially weighted averages of temporal difference advantages across multiple future steps. More specifically, consider a bias $\delta_t$ in the temporal difference at time step $t$, defined as
\begin{eqnarray}
\delta_t & = & r_t + \gamma V_{t+1} - V_{t}
\end{eqnarray}
GAE introduces parameters that dynamically balance the trade-off between bias and variance. This basic idea is to use the weighted sum of multi-step temporal difference bias as an estimation of the advantage, combining the advantages of Monte Carlo (high variance and low bias) and temporal difference methods (low variance and high bias). This combination can be expressed as
\begin{eqnarray}
A_t^{\mathrm{gae}} & = & \sum_{n=1}^{T-t} (\gamma \lambda)^{n} \delta_{t+n}
\end{eqnarray}
We can recursively develop the above model:
\begin{eqnarray}
A_t^{\mathrm{gae}} & = & \delta_{t} + \gamma \lambda A_{t+1}^{\mathrm{gae}}
\label{eq:GAE-advantage-estimation}
\end{eqnarray}
This recursive formulation shows that the current advantage estimate incorporates not only the immediate temporal difference error at time $t$ but also the estimated advantages of future time steps, weighted and adjusted by the parameters $\gamma$ and $\lambda$. This method effectively combines the strengths of both Monte Carlo and temporal difference methods, leading to a more stable and efficient training process.
\subsubsection{Group Relative Policy Optimization}
Although the methods discussed in the previous subsections effectively estimate advantage during the learning process, they all rely on a value model that is typically trained concurrently with the policy model. This dependence not only complicates the training process but also requires significant computational resources when applied to training LLMs. Given this, a natural idea arises: could we use more lightweight methods to compute the advantage without relying on the value model? By doing so, we could forego the value model component, thereby improving efficiency.
\begin{figure*}[!t]
\centering
\input{section4/Figures/figure-grpo}
\caption{
Group relative policy optimization. Compared to PPO, it foregoes the value model by estimating advantages directly from the rewards of a group of outputs. This method can simplify the training process and reduce the computational resources.
}
\label{fig:workflow-grpo}
\end{figure*}
In fact, this is entirely feasible, and one example of such an approach is \citet{shao-etal:2024deepseekmath}'s approach, called group relative policy optimization or GRPO for short. In the GRPO approach, the advantage is computed based on the relative rewards of the outputs within each group. More specifically, given an input $\mathbf{x}$, GRPO samples a group of outputs $\mathbf{Y}=\{\mathbf{y}_1, \mathbf{y}_2, \cdots, \mathbf{y}_{G}\}$. Then, the rewards for each output are computed using a reward model or a reward function, denoted as $\mathbf{R} = \{R(\mathbf{x}, \mathbf{y}_{1}), R(\mathbf{x}, \mathbf{y}_{2}), \dots, R(\mathbf{x}, \mathbf{y}_{G})\}$. The advantage of the $i$-th output can be computed through a relative reward in comparison to the other outputs:
\begin{eqnarray}
A^{\mathrm{grpo}}_{i,t} = \frac{R(\mathbf{x}, \mathbf{y}_{i})-\mathrm{Mean}(\mathbf{R})}{\mathrm{Std}(\mathbf{R})}
\label{eq:advantage-grpo}
\end{eqnarray}
where the $\mathrm{Mean}(\cdot)$ and $\mathrm{Std}(\cdot)$ represent the mean and standard deviation functions, respectively. Note that the advantage computation used here differs slightly from that in PPO. In this case, this computed advantage is applied to each time step, whereas in the traditional PPO, we separately compute an advantage for each time step. At this point, we can also use the process reward model to supplement the advantage computation, allowing us to compute an advantage specific to each time step. A simple way to achieve this is by implementing Eq. (\ref{eq:advantage-grpo}) at each time step, where the rewards are computed using the process reward model. As a result, the objective for GRPO can be defined according to Eq. (\ref{eq:ppo-loss}):
\begin{eqnarray}
\mathcal{L}_{g}(\theta) & = & - \mathbb{E}_{\mathbf{x} \sim \mathcal{S}_{x}} \frac{1}{G} \sum_{i=1}^{G} \left[
\sum_{t=1}^{T} \mathrm{Clip}\big(\frac{\mathrm{Pr}_{\theta}(y_{i,t}|\mathbf{x},\mathbf{y}_{i,<t})}{\mathrm{Pr}_{\theta_{\mathrm{ref}}}(y_{i,t}|\mathbf{x},\mathbf{y}_{i,<t})}A_{i,t}^{\mathrm{grpo}}\big) - \beta \mathrm{Penalty}
\right]
\label{eq:grpo-loss}
\end{eqnarray}
In addition to optimizing the advantage computation, GRPO modifies the penalty term to enhance the optimization objective further. However, since our focus here is on advantage estimation, we will not delve into its details. Interested readers can refer to the GRPO paper for further details. The workflow of GRPO is illustrated in Figure \ref{fig:workflow-grpo}, showcasing how these modifications integrate into the training process.
Moreover, there are other methods like GRPO that design advantage estimation without relying on a value model for training LLMs, such as those discussed in \citet{li-etal:2023remax} and \citet{hu:2025reinforce++}.
% 实际上,还有一些类似的工作,尝试消除PPO的组件,。。。
% remax cpo等等reference or value models的工作
\subsection{Efficient RL Methods}
Efficiency is a critical consideration for many practical applications of RL, and it becomes even more significant in the context of LLMs due to their vast number of parameters. For example, we might wish to train LLMs using RL, given memory and time constraints. In practice, the efficiency of RL is not a single issue but encompasses a wide range of challenges. While these challenges can be categorized in various ways in the existing literature, we focus on two fundamental aspects: \textit{time} and \textit{space} efficiency, which are commonly considered in efficiency-related problems. For these two efficiency problems, we aim to achieve the desired RL performance with the least amount of time and memory required.
In this section, we will not discuss all the issues related to the efficiency of RL, which is an extensive topic. Instead, we will focus on the commonly used efficient methods for training LLMs with RL. Some of these methods refine the sampling process, while others aim to eliminate certain components, such as the reward model, and utilize alternative lightweight methods in their place. Nonetheless, although these efficient methods are suggested for training LLMs, they are rather general and can be utilized in other RL scenarios.
\subsubsection{Dynamic Sampling}
As mentioned in Section \ref{sec:policy-gradient}, training LLMs with RL typically requires sampling for each sample $\mathbf{x}$ in $\mathcal{S}_x$, which introduces significant computational time overhead. This becomes particularly challenging in large-scale RL scenarios, where thousands of training steps need to be conducted, such as in DeepSeek-R1 \citep{guo:2025deepseek}.
From the perspective of LLM inference, there are many methods to reduce the time overhead of sampling: since the sampling process is implemented by LLM inference, we have reason to believe that any method that accelerates inference can also be applied to reduce the time required for sampling.
Examples include KV-Cache \citep{pope-etal:2023efficiently}, quantization \citep{zhao-etal:2024atom}, and speculative decoding \citep{chen-etal:2023accelerating}. However, in this section, we will not discuss these methods in detail, as there are numerous approaches, each addressing different aspects of LLM inference. We refer the interested readers to these papers for more details. Instead, we will focus on one particular method that aims to reduce sampling overhead from the perspective of optimizing RL for training LLMs.
Before discussing specific methods, let us first review the purpose of sampling. Given a sample $\mathbf{x}$, the goal of sampling is to explore a better output $\mathbf{y}$ from the policy model. This process allows us to evaluate different possible outputs in order to find those that lead to improved outcomes, enabling the model to make more informed decisions and optimize its performance on the state. Unlike typical RL scenarios, where policy models may start with limited information or random initialization, the policy model in the context of LLMs is usually quite advanced. These LLMs often begin as pre-trained models equipped with substantial knowledge acquired through pre-training and SFT. As a result, for certain samples, we can consider that the policy model may already perform well without the need for further exploration and optimization. This can be viewed through the lens of the classic RL dilemma of exploration versus exploitation \citep{Sutton-and-Barto:2018RL}. In such cases, exploration may be less necessary, and exploiting the knowledge already embedded in the pre-trained model could be more effective. Thus, a balance must be struck between exploring new possibilities and leveraging the pre-existing knowledge of the policy model to maximize efficiency.
To achieve this balance between exploration and exploitation, one simple and straightforward approach to improving the efficiency of sampling is to identify the samples that are required to explore and perform sampling only for those samples. Given an input-only dataset $\mathcal{S}_x = \{\mathbf{x}_1, \mathbf{x}_2, \dots, \mathbf{x}_N\}$, where $N$ denotes the size of the dataset, we can use a reward model to evaluate the need for exploration by the policy model. First, we generate $\hat{\mathbf{y}}$ for each $\mathbf{x}$ using greedy search. Then, by taking $\mathbf{x}$ and $\hat{\mathbf{y}}$ as input to the reward model, the reward for the $k$-th sample $\mathbf{x}_k$ is computed as $R_{\phi}(\mathbf{x}_{k}, \hat{\mathbf{y}}_{k})$.
At this point, we can directly use the reward value to determine whether the policy model requires further exploration and optimization for the sample. Specifically, we set a threshold $r_{\mathrm{upper}}$. If the reward for a sample exceeds $r_{\mathrm{upper}}$, we can consider that no further sampling and optimization is required. Otherwise, exploration and optimization are required. This approach enables selective sampling and optimization, improving the efficiency of the overall process by focusing resources on the most promising samples that are likely to benefit from further exploration \citep{wang-etal:2024esrl}. A challenge in implementation is determining the value of $r_{\mathrm{upper}}$. Typically, we set it as the average reward:
\begin{eqnarray}
r_\mathrm{upper} & = & \frac{1}{N} \sum_{i=1}^{N} R_{\phi}(\mathbf{x}_{i}, \hat{\mathbf{y}}_{i})
\end{eqnarray}
A similar dynamic sampling strategy, known as prioritized sampling, has proven effective in large-scale RL scenarios \citep{kimi-team:2025kimi}. However, we must consider another important factor in dynamic sampling. For samples with very low rewards, the policy model may lack the capacity to learn and optimize effectively. Excessive exploration of such samples can be inefficient. Consequently, we can introduce an additional threshold $r_{\mathrm{lower}}$ to eliminate unnecessary exploration of these low-reward samples, thereby enhancing efficiency. In this case, the policy model can concentrate its exploration and optimization efforts on a select group of samples, avoiding wasted sampling time on those that are unlikely to explore a more optimal output \citep{li-etal:2025limr}.
A concept closely related to the discussion here is sample efficiency. This topic has been widely discussed in recent literature. For example, some research on knowledge distillation seeks to select a smaller subset of samples to achieve more optimal performance \citep{wang-etal:2021selective}. More recently, much work on instruction sample selection in SFT aims to improve SFT efficiency by learning from a small number of instruction samples \citep{chen-etal:2023alpagasus}.
\subsubsection{Lightweight Reward Methods}
\label{sec:lightweight-reward-methods}
An additional computational overhead in training LLMs with RL is the frequent need to call the reward model to compute rewards. Furthermore, to ensure the reliability and generalizability of the reward model, we often scale it to a large size \citep{gao-etal:2023scaling}, which increases the computational time for reward computations.
One approach for mitigating this issue is to explore rule-based rewards, for example, as discussed in Section \ref{sec:policy-gradient}, the length of the outputs could be utilized as a reward. Instead of relying on large-scale, resource-intensive models, rule-based rewards assign rewards using predefined and task-specific rules. These rules, often derived from expert knowledge or task-specific requirements, are computationally inexpensive and provide fast, effective feedback to the policy model. By replacing or complementing traditional reward models with rule-based approaches, we can reduce both the time and computational cost associated with reward computations while still offering meaningful guidance for the learning process. This approach is also based on the understanding that, in some tasks, human preferences are simple to describe and do not require the complexity of training a reward model. Instead, we can use rules to express these preferences efficiently.
Here, in addition to the previously mentioned length-based rewards, we further demonstrate two examples of rule-based rewards. First, considering the output format, we can design rewards that encourage the policy model to generate outputs adhering to specific formats. For example, in response to the input "Give me three tips to improve my accuracy in solving math problems," we can assign rewards based on whether the output can be parsed as JSON correctly and include ``tip1'', ``tip2'', and ``tip3'' as the keys.
% Example 1
\input{section4/Figures/figure-examples-of-rule-based-reward-length}
Similarly, in mathematical reasoning tasks, where the desired outcome is a correct final answer, we can design a rule to verify the correctness of the answer and provide rewards based on whether the result is correct \citep{havrilla-etal:2024teaching,shao-etal:2024deepseekmath}.
% Example 2
\input{section4/Figures/figure-examples-of-rule-based-reward-math}
In addition to improving computational efficiency, another benefit of using rule-based rewards is their stability. Compared to reward models, rule-based rewards are less prone to issues such as reward hacking. As the rules are predefined and not subject to the same complex training process as traditional models, the risk of misalignment between the optimization objectives of the policy model and the oracle rewards is reduced. This stability makes rule-based rewards a more predictable and reliable approach, ensuring that the learning process of the policy model is less influenced by prediction errors or biases that may arise from the reward model. Furthermore, recent research has shown that employing rule-based rewards can significantly enhance the reasoning capabilities of LLMs through RL, underscoring its practical effectiveness \citep{guo:2025deepseek,xie-etal:2025logic}.
Another lightweight reward method is to consider achieving a lightweight reward model with fewer parameters, which would reduce both computational time and resource usage. There are several methods to achieve this. For example, given that reward models fundamentally utilize an LLM, many efficient methods, such as Mixture of Experts (MoE) \citep{masoudnia-etal:2014mixture} and model pruning \citep{sun-etal:2023simple}, originally developed for LLMs can be adapted to the reward model to improve its efficiency. Furthermore, knowledge distillation techniques offer a pathway to construct a smaller reward model that learns to emulate the behaviour of a larger one \citep{wang-etal:2023learning}. These approaches not only improve the efficiency of the reward model but also maintain its ability to deliver accurate rewards.
\subsection{Direct Preference Optimization}
\begin{figure*}[!t]
\centering
\input{section4/Figures/figure-rlhf-vs-dpo}
\caption{Standard PPO vs. DPO in training LLMs. In PPO, the human preference data is used to train a reward model, which is then employed to train the policy and the value function. In DPO, the use of human preference data is more direct, and the policy is trained on this data without the need for reward model training.}
\label{fig:comparsion-rl-dpo}
\end{figure*}
Although learning reward models is a standard step in RL, it makes the entire training process much more complex than supervised training. Training a reliable reward model is itself not an easy task and a poorly trained reward model can greatly affect the outcome of policy learning. We now consider an alternative method, called direct preference optimization (DPO), which simplifies the training framework by eliminating the need to explicitly model rewards \citep{rafailov:2023direct}. This method directly optimizes the policy model based on user preferences rather than developing a separate reward model. As a result, we can achieve human preference alignment in a supervised learning-like fashion. Figure \ref{fig:comparsion-rl-dpo} shows a comparison of the standard PPO method and the DPO method.
DPO simplifies the RL process for LLMs by utilizing a straightforward cross-entropy loss, which streamlines the learning process and enhances its stability. Rather than delving into the detailed derivation of the DPO loss function, we will discuss its optimization objective from the perspective of optimizing an LLM. The loss function derived from a Bradley-Terry reward model can be given by
\begin{eqnarray}
\mathcal{L}_{\mathrm{dpo}}(\theta) & = & - \mathbb{E}_{(\mathbf{x}, \mathbf{y}_{a}, \mathbf{y}_{b}) \sim \mathcal{D}_{r}} \big[ \log\mathrm{Sigmoid} \big( \beta \log \frac{\mathrm{Pr}_{\theta}(\mathbf{y}_{a}|\mathbf{\mathbf{x}})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_{a}|\mathbf{\mathbf{x}})} - \beta \log \frac{\mathrm{Pr}_{\theta}(\mathbf{y}_{b}|\mathbf{\mathbf{x}})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}_{b}|\mathbf{\mathbf{x}})} \big) \big]
\label{eq:dpo-loss}
\end{eqnarray}
This loss function utilizes the implicit reward for DPO training, which replaces the need for an external reward model. The implicit reward is computed as the log-ratio of probabilities:
\begin{eqnarray}
R(\mathbf{x}, \mathbf{y}) = \beta \frac{\mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{\mathbf{x}})}{\mathrm{Pr}_{\theta_\mathrm{ref}}(\mathbf{y}|\mathbf{\mathbf{x}})}
\end{eqnarray}
\begin{figure*}[!t]
\centering
\input{section4/Figures/figure-dpo}
\caption{
Illustration of the optimization objective of DPO from a hypothesis space perspective. The right-top chart shows the probability distributions over different hypotheses before optimization, where $\mathbf{y}_{a}$ is the preferred output and $\mathbf{y}_{b}$ is the dispreferred output. The right-bottom chart shows the probabilities after applying DPO, where the likelihood of the preferred output increases and that of dispreferred outputs decreases. Note that during this optimization process, outputs similar to the preferred output $\mathbf{y}_{a}$ (such as $\mathbf{y}_{1}$) will obtain an increased probability, while those similar to the less preferred output $\mathbf{y}_{b}$ (such as $\mathbf{y}_{2}$) will experience a decreased probability, due to the generalization of the model.
}
\label{fig:understand-dpo-optimization-objective}
\end{figure*}
Here, we can consider that the DPO directly models human feedback through the relative likelihood of outputs under the policy and reference models. This method provides a more straightforward approach than traditional methods that use a separate reward model and then apply RL algorithms like PPO to adjust the probabilities of sampled outputs. In this way, we can further understand the optimization objective of the DPO by leveraging the hypothesis space as mentioned in Section \ref{sec:policy-gradient}. In practice, as illustrated in Figure \ref{fig:understand-dpo-optimization-objective}, we can see the DPO as reranking in the hypothesis space through the Bradley-Terry model approach. Specifically, this method increases the probabilities of preferred outputs in preference data while decreasing those of dispreferred ones. Similar to TRPO, the DPO also incorporates a penalty derived from the reference model, which ensures that updates remain within a trusted behaviour space for the policy model.
However, there are two sides to every coin. While DPO simplifies the RL process, it also introduces limitations. Notably, since DPO eliminates the exploration phase during training, its potential peak performance is generally considered lower compared to RL-based methods like PPO, which incorporate exploration to discover more effective policies. In other words, the performance of DPO is inherently bounded by the labeled preferred outputs in the preference data. To address this limitation, current approaches focus on ensuring high-quality preferred outputs, either by employing advanced LLMs like GPT-4 or through human labeling \citep{cui-etal:2023ultrafeedback,morimura-etal:2024filtered}.
\begin{table*}[!t]
\centering
\resizebox{\textwidth}{!}{
\input{section4/Tables/dpo_variants}}
\caption{DPO variants and their optimization objectives.}
\label{tab:dpo_varients}
\end{table*}
Another limitation associated with DPO is over-optimization. Since DPO employs the Bradley-Terry model for modeling preferences, it can suffer from over-optimization, such as length exploitation, where a longer output might be mistakenly deemed as more aligned with human preferences \citep{singhal-etal:2023long,wang-etal:2023far}. Many efforts have been made to address this issue and propose different variants of DPO, as detailed in Table \ref{tab:dpo_varients}. Note that here we only provide an introduction to their optimization objectives. For more discussions on these variants, interested readers can refer to the related papers.
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\begin{scope}
\def\sep{2.8cm}
\node [draw, text width=5cm, align=left, rounded corners=2pt] (y lt t) at (0,0) {\footnotesize{Karen's students are about to take a standardized test. Karen gets a \$500 bonus if their average score is above 75, plus an extra \$10 bonus for every additional point the average score increases above 75. So far, Karen has graded 8 tests, and the average is 70. Given that each student can have a maximum score of 150, what combined score do the last two tests need to have for Karen to earn a \$600 bonus?}};
\node [anchor=west,draw, text width=7.5cm, align=left, rounded corners=2pt,minimum width=8cm] (y gt t 1) at ([xshift=1.5cm,yshift=\sep/2]y lt t.east) {\footnotesize{First subtract \$500 from Karen's goal bonus amount to find how much she makes from the extra ... \\ Then divide the extra bonus by the extra rate: ...\\ Then ... : 850 points - 560 points = 290 points \\ \#\#\#\# $\boxed{290}$\\}};
\node [anchor=south] at (y lt t.north) {$\mathbf{x}$};
\node [anchor=south east, draw,fill=black!5!white] (reward) at (y gt t 1.south east) {\footnotesize{\textbf{Reward: 6}}};
\node [anchor=west,draw, text width=7.5cm, align=left, rounded corners=2pt,minimum width=8cm] (y gt t 2) at ([xshift=1.5cm,yshift=\sep*3/2]y lt t.east) {\footnotesize{To determine the combined score required for the last two tests in order for Karen to earn a \$600 ... \\ Since she wants to earn a total of \$600 ... \\ ... Thus, to summarize, the combined score needed on the last two tests for Karen to earn a \$600 bonus is: $\boxed{290}$\\}};
\node [anchor=south east, draw,fill=black!5!white] (reward) at (y gt t 2.south east) {\footnotesize{\textbf{Reward: 8}}};
\node at ([xshift=-.1cm,yshift=.1cm]reward.south east) {\textcolor{ugreen}{\faCheck}};
\node [anchor=west] (y gt t cdots) at ([xshift=1.5cm,yshift=-\sep/2]y lt t.east) {\LARGE$\cdots$};
\node [anchor=west,draw, text width=7.5cm, align=left, rounded corners=2pt,minimum width=8cm] (y gt t 4) at ([xshift=1.5cm,yshift=-\sep*3/2]y lt t.east) {\footnotesize{The total bonus from the first 8 tests is \$500 x 8 = \$4000. \\ The average score needs to be at least 75 ... \\ So the minimum possible score for these 2 tests is 85 x 2 = 170. \\ To find the combined score needed for all 10 tests ... \\ Therefore, the combined score needed for Karen to earn a \$600 bonus is 1370. \\ The answer is $\boxed{1370}$}\\};
\node [anchor=south east, draw,fill=black!5!white] (reward) at (y gt t 4.south east) {\footnotesize{\textbf{Reward: -3}}};
\draw [->] (y lt t.east) .. controls +(+1cm,0) and +(-1.5cm,0) .. (y gt t 2.west) node [anchor=south east] {$\mathbf{y}_{1}$};
\draw [->] (y lt t.east) .. controls +(+1cm,0) and +(-1.5cm,0) .. (y gt t 1.west) node [anchor=south east] {$\mathbf{y}_{2}$};
\draw [->] (y lt t.east) .. controls +(+1cm,0) and +(-1.5cm,0) .. (y gt t cdots.west);
\draw [->] (y lt t.east) .. controls +(+1cm,0) and +(-1.5cm,0) .. (y gt t 4.west) node [anchor=south east,yshift=.5cm,xshift=.1cm] {$\mathbf{y}_{N}$};
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}[remember picture]
\tikzset {
block/.style={draw,inner sep=0pt,fill=white, minimum width=4cm, minimum height=1.5cm, align=center},
miniblock/.style={draw,inner sep=0pt,fill=white, minimum width=1.5cm, minimum height=.8cm, align=center, rounded corners=2pt, text width=1.4cm},
linetext/.style={fill=#1, minimum height=1.5pt, minimum width=.8cm, inner sep=0},
}
\def\sep{1.5cm}
\def\ssep{1cm}
\def\vsep{.2cm}
\begin{scope}
\node [anchor=center, block] (b1) at (0,0) {Pre-training an LLM};
\node [anchor=south, block] (b2) at ([yshift=\sep]b1.north) {\footnotesize{Cold Start} \\ \footnotesize{for Reasoning}};
\node [anchor=south, block] (b3) at ([yshift=\sep]b2.north) {\footnotesize{Large-scale Reinforcement}\\ \footnotesize{Learning for Reasoning}};
\node [anchor=south, block] (b4) at ([yshift=\sep]b3.north) {\footnotesize{Rejection Sampling} \\ \footnotesize{for Multi-task}};
\node [anchor=south, block] (b5) at ([yshift=\sep]b4.north) {\footnotesize{Reinforcement Learning} \\ \footnotesize{for Multi-task}};
% \node [anchor=south] (b6) at ([yshift=\sep]b5.north) {\LARGE{$\cdots$}};
\draw [->] (b1.north) -- (b2.south);
\draw [->] (b2.north) -- (b3.south);
\draw [->] (b3.north) -- (b4.south);
\draw [->] (b4.north) -- (b5.south);
% \draw [->] (b5.north) -- (b6.south);
\def\notesep{10cm}
%%
\node [anchor=north west, text width=\notesep] (n1) at ([xshift=\ssep]b1.north east) {\scriptsize{Starting with a pre-trained model, the iterative process incorporates a series of strategically designed phases to enhance the reasoning capabilities of the LLM.\\}};
%%
\def\hsep{2.4cm}
\node [anchor=north west, text width=\notesep] (n2) at ([xshift=\ssep]b2.north east) {\scriptsize{This phase involves collecting high-quality data from various reasoning tasks, which is then used to fine-tune the pre-trained LLM, providing a cold start.\\}};
\node [anchor=north] at (n2.south) {
\tikz{
\node [anchor=north west] (m1-1) at (0,0) {\scriptsize{$\textbf{x}$}};
\node [anchor=north west] (m1-2) at ([yshift=-\vsep]m1-1.south west) {\scriptsize{$\textbf{y}$}};
\node [anchor=west, linetext=lolred] (m1-3) at (m1-1.east) {};
\node [anchor=west, linetext=lolblue] (m1-4) at (m1-2.east) {};
\node [anchor=west, miniblock] (m1-5) at ([xshift=\hsep]$(m1-3.east)!.5!(m1-4.east)$) {\scriptsize{Pre-trained \\ LLM \\}};
\draw [->] ([xshift=-\hsep]m1-5.west) -- (m1-5.west);
\node [anchor=west, miniblock] (m1-6) at ([xshift=\hsep]m1-5.east) {\scriptsize{SFT\\LLM\\}};
\draw [->] ([xshift=-\hsep]m1-6.west) -- node [midway, above] {\scriptsize{SFT}} (m1-6.west);
}
};
%%
\def\hsep{.95cm}
\node [anchor=north west, text width=\notesep] (n3) at ([xshift=\ssep]b3.north east) {\scriptsize{Following the cold start, the model undergoes large-scale RL, leveraging rule-based rewards to enhance reasoning capabilities.\\}};
\node [anchor=north] at (n3.south){
\tikz{
\node [anchor=north west] (m2-1) at (0,0) {\scriptsize{$\textbf{x}$}};
\node [anchor=west, linetext=lolred] (m2-2) at (m2-1.east) {};
\node [anchor=west,miniblock] (m2-3) at ([xshift=\hsep]m2-2.east) {\scriptsize{SFT\\LLM\\}};
\matrix [matrix anchor=west,nodes={inner sep=0,anchor=center},column sep={.1cm,between borders},row sep={.2cm,between borders}] (m2-4) at ([xshift=\hsep]m2-3.east) {
\node{\scriptsize{$\mathbf{y}_1$}}; & \node [linetext=lolblue] {}; & \node (t1){\scriptsize{$6$}}; \\
\node{\scriptsize{$\cdots$}}; & \node{\scriptsize{$\cdots$}}; & \node{\scriptsize{$\cdots$}}; \\
\node{\scriptsize{$\mathbf{y}_G$}}; & \node [linetext=lolblue] {}; & \node{\scriptsize{$12$}}; \\};
\node [anchor=south] at ([yshift=.1cm]t1.north) {\scriptsize{Rewards}};
\node [anchor=west,miniblock] (m2-5) at ([xshift=\hsep]m2-4.east) {\scriptsize{Reinforced \\ LLM\\}};
\draw [->] ([xshift=.1cm]m2-2.east) -- (m2-3.west);
\draw [->] (m2-3.east) -- (m2-4.west);
\draw [->] (m2-4.east) -- node [midway, above] {\scriptsize{GRPO}} (m2-5.west);
}
};
%%
\def\hsep{1cm}
\node [anchor=north west, text width=\notesep] (n4) at ([xshift=\ssep]b4.north east) {\scriptsize{This phase performs multi-task rejection sampling to further enhance the readability of outputs. \\}};
\node[anchor=north] at (n4.south) {
\tikz{
\node [anchor=north west] (m3-1) at (0,0) {\scriptsize{$\textbf{x}$}};
\node [anchor=west, linetext=lolred] (m3-2) at (m3-1.east) {};
\node [anchor=west,miniblock] (m3-3) at ([xshift=\hsep]m3-2.east) {\scriptsize{Reinforced\\LLM\\}};
\matrix [matrix anchor=west,nodes={inner sep=0,anchor=center},column sep={.1cm,between borders},row sep={.2cm,between borders}] (m3-4) at ([xshift=\hsep]m3-3.east) {
\node{\scriptsize{$\mathbf{y}_1$}}; & \node [linetext=lolblue] {}; & \\
\node{\scriptsize{$\cdots$}}; & \node{\scriptsize{$\cdots$}}; & \\
\node{\scriptsize{$\mathbf{y}_N$}}; & \node [linetext=lolblue] {}; & \node {\textcolor{ugreen}{\scriptsize{\faCheck}}}; \\ };
\node [anchor=west,miniblock] (m3-5) at ([xshift=\hsep]m3-4.east) {\scriptsize{SFT\\LLM\\}};
\draw [->] ([xshift=.1cm]m3-2.east) -- (m3-3.west);
\draw [->] (m3-3.east) -- (m3-4.west);
\draw [->] (m3-4.east) -- node [midway, above] {\scriptsize{SFT}} (m3-5.west);
}
};
%%
\def\hsep{.95cm}
\node [anchor=north west, text width=\notesep] (n5) at ([xshift=\ssep]b5.north east) {\scriptsize{In the final phase, multi-task RL is implemented to ensure broader generalization and maintain high performance across tasks beyond reasoning.\\}};
\node [anchor=north] at (n5.south){
\tikz{
\node [anchor=north west] (m2-1) at (0,0) {\scriptsize{$\textbf{x}$}};
\node [anchor=west, linetext=lolred] (m2-2) at (m2-1.east) {};
\node [anchor=west,miniblock] (m2-3) at ([xshift=\hsep]m2-2.east) {\scriptsize{SFT\\LLM\\}};
\matrix [matrix anchor=west,nodes={inner sep=0,anchor=center},column sep={.1cm,between borders},row sep={.2cm,between borders}] (m2-4) at ([xshift=\hsep]m2-3.east) {
\node{\scriptsize{$\mathbf{y}_1$}}; & \node [linetext=lolblue] {}; & \node (t1){\scriptsize{$0.6$}}; \\
\node{\scriptsize{$\cdots$}}; & \node{\scriptsize{$\cdots$}}; & \node{\scriptsize{$\cdots$}}; \\
\node{\scriptsize{$\mathbf{y}_G$}}; & \node [linetext=lolblue] {}; & \node{\scriptsize{$0.7$}}; \\};
\node [anchor=south] at ([yshift=.1cm]t1.north) {\scriptsize{Rewards}};
\node [anchor=west,miniblock] (m2-5) at ([xshift=\hsep]m2-4.east) {\scriptsize{Reinforced\\LLM\\}};
\draw [->] ([xshift=.1cm]m2-2.east) -- (m2-3.west);
\draw [->] (m2-3.east) -- (m2-4.west);
\draw [->] (m2-4.east) -- node [midway, above] {\scriptsize{GRPO}} (m2-5.west);
}
};
%%
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\tikzset {
block/.style={,draw,inner sep=2pt,rounded corners=2pt,fill=white},
emphasize/.style={draw=lolred,line width=.04cm}
}
\begin{scope}
\draw [-] (-\textwidth/2,0) -- (\textwidth/2,0);
\draw [-] (0,-\textwidth/2-.6cm) -- (0,\textwidth/2-2cm);
\begin{scope}[xshift=-\textwidth/4, yshift=\textwidth/2-3.2cm]
\node (x) at (0,0) {$\mathbf{x}$};
\node [anchor=north,fill=black!5!white,minimum width=\textwidth/2-.5cm,minimum height=.7cm] at (0,1cm) {(1) Selection};
\node [anchor=north,text width=.4\textwidth,block] (x) at (x.south) {\scriptsize{Karen's students are about to take a standardized test. Karen ... what combined score do the last two tests need to have for Karen to earn a \$600 bonus?\\}};
\def\blockheight{1.0cm}
\node [anchor=north west,text width=.15\textwidth,block,minimum height=\blockheight,emphasize] (y part1) at ([yshift=-.5cm]x.south west) {\scriptsize{First subtract \$500 from Karen's goal ... \$600 - \$500 = \$100\\}};
\node [anchor=north east,text width=.15\textwidth,block,minimum height=\blockheight] (y part2) at ([yshift=-.5cm]x.south east) {\scriptsize{Let's denote the combined score of the ... $8 \times 70 = 560$.\\}};
\node (cdots) at ($(y part1)!.5!(y part2)$) {$\cdots$};
\draw [->] (x.south) -- (y part1.north);
\draw [->] (x.south) -- (y part2.north);
\draw [->] (x.south) -- ($(y part1.north)!.5!(y part2.north)$);
\end{scope}
\begin{scope}[xshift=\textwidth/4, yshift=\textwidth/2-3.2cm]
\node (x) at (0,0) {$\mathbf{x}$};
\node [anchor=north,fill=black!5!white,minimum width=\textwidth/2-.5cm,minimum height=.7cm] at (0,1cm) {(2) Expansion};
\node [anchor=north,text width=.4\textwidth,block] (x) at (x.south) {\scriptsize{Karen's students are about to take a standardized test. Karen ... what combined score do the last two tests need to have for Karen to earn a \$600 bonus?\\}};
\def\blockheight{1.0cm}
\node [anchor=north west,text width=.15\textwidth,block,minimum height=\blockheight] (y part1) at ([yshift=-.5cm]x.south west) {\scriptsize{First subtract \$500 from Karen's goal ... \$600 - \$500 = \$100\\}};
\node [anchor=north east,text width=.15\textwidth,block,minimum height=\blockheight] (y part2) at ([yshift=-.5cm]x.south east) {\scriptsize{Let's denote the combined score of the ... $8 \times 70 = 560$.\\}};
\def\blockheight{1.0cm}
\node [anchor=north,text width=.15\textwidth,block,minimum height=\blockheight,emphasize] (y part1 1) at ([yshift=-.5cm]y part1.south) {\scriptsize{Then divide the extra bonus by the extra ... = 10 points\\}};
\node (cdots) at ($(y part1)!.5!(y part2)$) {$\cdots$};
\draw [->] (x.south) -- (y part1.north);
\draw [->] (x.south) -- (y part2.north);
\draw [->] (x.south) -- ($(y part1.north)!.5!(y part2.north)$);
\draw [->] (y part1.south) -- (y part1 1.north);
\end{scope}
\begin{scope}[xshift=-\textwidth/4, yshift=-1.2cm]
\node (x) at (0,0) {$\mathbf{x}$};
\node [anchor=north,fill=black!5!white,minimum width=\textwidth/2-.5cm,minimum height=.7cm] at (0,1cm) {(3) Simulation};
\node [anchor=north,text width=.4\textwidth,block] (x) at (x.south) {\scriptsize{Karen's students are about to take a standardized test. Karen ... what combined score do the last two tests need to have for Karen to earn a \$600 bonus?\\}};
\def\blockheight{1.0cm}
\node [anchor=north west,text width=.15\textwidth,block,minimum height=\blockheight] (y part1) at ([yshift=-.5cm]x.south west) {\scriptsize{First subtract \$500 from Karen's goal ... \$600 - \$500 = \$100\\}};
\node [anchor=north east,text width=.15\textwidth,block,minimum height=\blockheight] (y part2) at ([yshift=-.5cm]x.south east) {\scriptsize{Let's denote the combined score of the ... $8 \times 70 = 560$.\\}};
\def\blockheight{1.0cm}
\node [anchor=north,text width=.15\textwidth,block,minimum height=\blockheight] (y part1 1) at ([yshift=-.5cm]y part1.south) {\scriptsize{Then divide the extra bonus by the extra ... = 10 points\\}};
\node [anchor=north] (cdots 1) at ([yshift=-.2cm]y part1 1.south) {\textcolor{gray}{$\cdots$}};
\node [anchor=north,text width=.15\textwidth,block,minimum height=.8cm,align=center,emphasize,dashed] (y part1 n-1) at ([yshift=-.8cm]y part1 1.south) {\textcolor{gray}{\scriptsize{Then multiply the current average by the ... = 560 points}\\}};
\node [anchor=north,text width=.15\textwidth,block,minimum height=.8cm,align=center,emphasize,dashed] (y part1 n) at ([yshift=-.5cm]y part1 n-1.south) {\textcolor{gray}{\scriptsize{Then subtract the number of points ... = 290 points}\\}};
\node (cdots) at ($(y part1)!.5!(y part2)$) {$\cdots$};
\draw [->] (x.south) -- (y part1.north);
\draw [->] (x.south) -- (y part2.north);
\draw [->] (x.south) -- ($(y part1.north)!.5!(y part2.north)$);
\draw [->] (y part1.south) -- (y part1 1.north);
\draw [->,dashed] (y part1 1.south) -- (cdots 1.north);
\draw [->,dashed] (cdots 1.south) -- (y part1 n-1.north);
\draw [->,dashed] (y part1 n-1) -- (y part1 n);
\end{scope}
\begin{scope}[xshift=\textwidth/4, yshift=-1.2cm]
\node (x) at (0,0) {$\mathbf{x}$};
\node [anchor=north,fill=black!5!white,minimum width=\textwidth/2-.5cm,minimum height=.7cm] at (0,1cm) {(4) Backpropagation};
\node [anchor=north,text width=.4\textwidth,block] (x) at (x.south) {\scriptsize{Karen's students are about to take a standardized test. Karen ... what combined score do the last two tests need to have for Karen to earn a \$600 bonus?\\}};
\def\blockheight{1.0cm}
\node [anchor=north west,text width=.15\textwidth,block,minimum height=\blockheight] (y part1) at ([yshift=-.5cm]x.south west) {\scriptsize{First subtract \$500 from Karen's goal ... \$600 - \$500 = \$100\\}};
\node [anchor=north east,text width=.15\textwidth,block,minimum height=\blockheight] (y part2) at ([yshift=-.5cm]x.south east) {\scriptsize{Let's denote the combined score of the ... $8 \times 70 = 560$.\\}};
\def\blockheight{1.0cm}
\node [anchor=north,text width=.15\textwidth,block,minimum height=\blockheight] (y part1 1) at ([yshift=-.5cm]y part1.south) {\scriptsize{Then divide the extra bonus by the extra ... = 10 points\\}};
\node (cdots) at ($(y part1)!.5!(y part2)$) {$\cdots$};
\node [anchor=north] (cdots 1) at ([yshift=-.2cm]y part1 1.south) {\textcolor{gray}{$\cdots$}};
\node [anchor=north,text width=.15\textwidth,block,minimum height=.8cm,align=center,draw=gray,opacity=.6] (y part1 n-1) at ([yshift=-.8cm]y part1 1.south) {\textcolor{gray}{\scriptsize{Then multiply the current average by the ... = 560 points}\\}};
\node [anchor=north,text width=.15\textwidth,block,minimum height=.8cm,align=center] (y part1 n) at ([yshift=-.5cm]y part1 n-1.south) {\scriptsize{Then subtract the number of points ... = 290 points}\\};
\draw [->] (x.south) -- (y part1.north);
\draw [->] (x.south) -- (y part2.north);
\draw [->] (x.south) -- ($(y part1.north)!.5!(y part2.north)$);
\draw [->] (y part1.south) -- (y part1 1.north);
\draw [->,gray,opacity=.6] (y part1 1.south) -- (cdots 1.north);
\draw [->,gray,opacity=.6] (cdots 1.south) -- (y part1 n-1.north);
\draw [->,gray,opacity=.6] (y part1 n-1) -- (y part1 n);
\draw [->,emphasize] ([xshift=-.5cm]y part1 n.north) .. controls +(-.5cm,+.5cm) and +(-.5cm,-.5cm) .. ([xshift=-.5cm]y part1 1.south);
\draw [->,emphasize] ([xshift=-.5cm]y part1 1.north) .. controls +(-.5cm,+.1cm) and +(-.5cm,-.1cm) .. ([xshift=-.5cm]y part1.south);
\draw [->,emphasize] ([xshift=-.5cm]y part1.north) .. controls +(-.5cm,+.1cm) and +(-.5cm,-.1cm) .. ([xshift=-.5cm]y part1.north|-x.south);
\end{scope}
\end{scope}
\end{tikzpicture}
\end{center}
% !Mode:: "TeX:UTF-8"
% !TEX encoding = UTF-8 Unicode
\begin{center}
\begin{tikzpicture}
\tikzset {
block/.style={,draw,inner sep=2pt,rounded corners=2pt,fill=white}
}
\begin{scope}
\node (x) at (0,0) {$\mathbf{x}$};
\node [anchor=north,text width=.75\textwidth,block] (x) at (x.south) {\footnotesize{Karen's students are about to take a standardized test. Karen gets a \$500 bonus if their average score is above 75, plus an extra \$10 bonus for every additional point the average score increases above 75. So far, Karen has graded 8 tests, and the average is 70. Given that each student can have a maximum score of 150, what combined score do the last two tests need to have for Karen to earn a \$600 bonus?\\}};
\def\blockheight{2.5cm}
\node [anchor=north,text width=.25\textwidth,block,minimum height=\blockheight] (y part1) at ([yshift=-1cm]x.south) {\footnotesize{Let's denote the combined score of the last two tests as $x$. Given that Karen has graded 8 tests and the average is 70, the sum of the scores of the first 8 tests is $8 \times 70 = 560$.\\}};
\node [anchor=south east, draw,fill=black!5!white] at (y part1.south east) {\scriptsize{\textbf{Reward: 7}}};
\node [anchor=east,text width=.25\textwidth,block,minimum height=\blockheight] (y part2) at ([xshift=-.5cm]y part1.west) {\footnotesize{The first step is to calculate how much the average score needs to be above 75 for that extra \$100:\\ Extra points required = $\frac{100}{10}$ = 10\\}};
\node [anchor=south east, draw,fill=black!5!white] at (y part2.south east) {\scriptsize{\textbf{Reward: 1}}};
\node [anchor=east,inner sep=0] (lcdots) at ([xshift=-.1cm]y part2.west) {\LARGE{$\cdots$}};
\node [anchor=west,text width=.25\textwidth,block,minimum height=\blockheight,draw=lolred,line width=.05cm] (y part3) at ([xshift=.5cm]y part1.east) {\footnotesize{First subtract \$500 from Karen's goal bonus amount to find how much she makes from the extra \$10/point bonus: \$600 - \$500 = \$100\\}};
\node [anchor=south east, draw,fill=black!5!white] at (y part3.south east) {\scriptsize{\textbf{Reward: 8}}};
\node [anchor=west,inner sep=0] (rcdots) at ([xshift=.1cm]y part3.east) {\LARGE{$\cdots$}};
\begin{scope}[on background layer]
\node [anchor=north,fill=black!5!white, rounded corners=2pt, inner sep=8pt, fit=(lcdots)(y part1)(y part2)(y part3)(rcdots)] (y) {};
\end{scope}
\def\blockheight{2.5cm}
\node [anchor=north,text width=.25\textwidth,block,minimum height=\blockheight,draw=lolred,line width=.05cm] (y part1 1) at ([yshift=-1cm]y part1.south) {\footnotesize{Then divide the extra bonus by the extra rate: \$100 / \$10/point = 10 points\\}};
\node [anchor=south east, draw,fill=black!5!white] at (y part1 1.south east) {\scriptsize{\textbf{Reward: 9}}};
\node [anchor=east,text width=.25\textwidth,block,minimum height=\blockheight] (y part2 1) at ([xshift=-.5cm]y part1 1.west) {\footnotesize{Then, divide the remaining bonus amount by the extra bonus amount per point to find out how many additional points Karen needs: \$100 / \$10 = \$10\\}};
\node [anchor=south east, draw,fill=black!5!white] at (y part2 1.south east) {\scriptsize{\textbf{Reward: 2}}};
\node [anchor=east,inner sep=0] (lcdots) at ([xshift=-.1cm]y part2 1.west) {\LARGE{$\cdots$}};
\node [anchor=west,text width=.25\textwidth,block,minimum height=\blockheight] (y part3 1) at ([xshift=.5cm]y part1 1.east) {\footnotesize{Then, divide the \$100 bonus by \$10/point to find the additional average score Karen needs: \$100 / \$10/point = 10 additional points.\\}};
\node [anchor=south east, draw,fill=black!5!white] at (y part3 1.south east) {\scriptsize{\textbf{Reward: -1}}};
\node [anchor=west,inner sep=0] (rcdots) at ([xshift=.1cm]y part3 1.east) {\LARGE{$\cdots$}};
\begin{scope}[on background layer]
\node [anchor=north,fill=black!5!white, rounded corners=2pt, inner sep=8pt, fit=(lcdots)(y part1 1)(y part2 1)(y part3 1)(rcdots)] (y 1) {};
\end{scope}
\def\blockheight{3.5cm}
\node [anchor=north,text width=.25\textwidth,block,minimum height=\blockheight] (y part1 2) at ([yshift=-1cm]y part1 1.south) {\footnotesize{Then, multiply the number of additional points needed by the number of tests Karen has yet to grade, which is 2: 10 points * 2 tests = 20 points.\\}};
\node [anchor=south east, draw,fill=black!5!white] at (y part1 2.south east) {\scriptsize{\textbf{Reward: -4}}};
\node [anchor=east,text width=.25\textwidth,block,minimum height=\blockheight] (y part2 2) at ([xshift=-.5cm]y part1 2.west) {\footnotesize{Then, Karen needs the combined score of the last two tests to be 10 points higher than the average of 75. Since there are two tests, the combined score for the last two tests needs to be 20 points higher than the average.\\}};
\node [anchor=south east, draw,fill=black!5!white] at (y part2 2.south east) {\scriptsize{\textbf{Reward: 3}}};
\node [anchor=east,inner sep=0] (lcdots) at ([xshift=-.1cm]y part2 2.west) {\LARGE{$\cdots$}};
\node [anchor=west,text width=.25\textwidth,block,minimum height=\blockheight,draw=lolred,line width=.05cm] (y part3 2) at ([xshift=.5cm]y part1 2.east) {\footnotesize{Then add the 10 extra points to the baseline 75 point goal to find the students' average test score: 10 points + 75 points = 85 points\\}};
\node [anchor=south east, draw,fill=black!5!white] at (y part3 2.south east) {\scriptsize{\textbf{Reward: 10}}};
\node [anchor=west,inner sep=0] (rcdots) at ([xshift=.1cm]y part3 2.east) {\LARGE{$\cdots$}};
\begin{scope}[on background layer]
\node [anchor=north,fill=black!5!white, rounded corners=2pt, inner sep=8pt, fit=(lcdots)(y part1 2)(y part2 2)(y part3 2)(rcdots)] (y 2) {};
\end{scope}
\draw [->] (x.south) -- ([yshift=.28cm]y part1.north);
\draw [->] (y part3.south) .. controls +(0,-.8cm) and +(0,+.8cm) .. ([yshift=.28cm]y part1 1.north);
\draw [->] (y part1 1.south) -- ([yshift=.28cm]y part1 2.north);
\draw [->,lolblue] (y part3.south) node [xshift=1.0cm,yshift=-.5cm,right] {\footnotesize{\textbf{Self-refinement}}} arc [start angle=225, end angle=495, radius=1.95cm];
\node [anchor=north] (cdots) at ([yshift=-1cm]y part1 2.south) {\LARGE{$\cdots$}};
\draw [->] (y part3 2.south) .. controls +(0,-.8cm) and +(0,+.8cm) .. (cdots.north);
\end{scope}
\end{tikzpicture}
\end{center}
\section{Reinforcement Learning for LLM Reasoning}
So far, our discussion has mainly focused on various aspects of using and improving RL for training LLMs. The methods mentioned can be easily adapted to a wild of scenarios where the correctness of an output can be examined by checking whether the desired result is included. For example, in the task of calculating a mathematical expression, a reward model can provide positive feedback if the answer is correct and negative feedback if the answer is wrong. However, in many problems that require complex reasoning, simply examining the correctness of the final answer is insufficient for learning. Imagine a student who is only given the final answer to a challenging math problem. Knowing whether the final answer is right or wrong does not help the student figure out where they went wrong and how to calculate the correct answer. A better approach would be to guide the student with a step-by-step breakdown of the problem-solving process and encourage understanding of the underlying concepts and logic behind these steps. To address this, researchers also explore the potential of RL to teach LLMs not just to generate correct answers but to develop and present coherent, step-by-step reasoning that aligns with human cognitive processes. In this regard, RL has previously demonstrated its effectiveness in training neural networks for complex planning and reasoning within game environments, as evidenced by notable successes such as AlphaGo \citep{silver-etal:2016mastering} and AlphaStar \citep{vinyals-rtal:2019grandmaster}. Given these advancements and the inherently interactive nature of problem-solving, it is also natural to consider the application of RL to LLM reasoning.
In this section, we delve deeper into the application of RL to enhance the reasoning capabilities of the LLM, a topic that has recently garnered significant attention. We begin by giving a general introduction to test-time scaling that fully unleashes the reasoning potential of LLMs, including best-of-$N$ sampling, step-by-step verification, and Monte Carlo Tree Search. We then discuss two critical issues: how to scale RL effectively, and how to iterate the RL process to enhance the reasoning capabilities of LLMs. Note that the following methods are primarily illustrated through the mathematical reasoning problem, they are applicable to a broad range of decision-making problems.
\subsection{Test-time Scaling}
Initially, we review the fundamental objective of RL: to maximize the rewards obtained during output generation. This goal has been extensively achieved through various training-time optimization techniques, such as policy gradient or PPO algorithms. Beyond training, recent research has illuminated the benefits of scaling up test-time computation, a process also known as test-time scaling, which can significantly enhance the maximization of rewards, especially in the reasoning task. In a practical implementation, one simple approach to achieve test-time scaling is the use of prompting techniques. For example, appending the phrase ``Let's think step-by-step.'' to the input can effectively stimulate the LLM to engage in a more detailed reasoning process during the generation. While this approach can improve reasoning accuracy, it often leads to suboptimal performance because the model is not inherently trained to understand the most beneficial reasoning processes. To address this issue, we can perform the test-time scaling with guidance from a reward model, as discussed in the following subsections.
\subsubsection{Best-of-N Sampling}
\label{sec:BoN-sampling}
\begin{figure*}[!t]
\centering
\input{section5/Figures/figure-bon-sampling}
\caption{
Best-of-n sampling. Given an input, multiple candidate outputs are sampled, and a reward model is used to select the best output. A majority vote can also determine the final output in scenarios without a reward model. For example, in this figure, most of the outputs suggest an answer of 290. Therefore, we can consider any one of the outputs resulting in 290 to be the final output.
}
\label{fig:bon-sampling}
\end{figure*}
One approach to test-time scaling using a reward model involves sampling multiple reasoning paths given input and then using the reward model to select the best one from $N$ alternative outputs generated by the LLM, called best-of-$N$ sampling (BoN sampling). We can consider BoN sampling a reranking technique. In fact, reranking methods have been a prevalent technique in NLP, particularly in machine translation, where they have been employed for a long time to enhance output quality by selecting the most appropriate translation from a set of candidates. Additionally, this method often functions as a simple model ensemble approach. In such cases, different outputs generated by various models can be reranked according to a specified metric.
As illustrated in Figure \ref{fig:bon-sampling}, in the BoN sampling, we first sample $N$ different outputs $\{\mathbf{y}_{1}, \mathbf{y}_{2}, \cdots, \mathbf{y}_{N}\}$ for the input $\mathbf{x}$:
\begin{eqnarray}
\{\mathbf{y}_1,...,\mathbf{y}_N\} & = & \mathop{\mathrm{argTopN}}_{\mathbf{y}} \left[ \mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x}) \right]
\end{eqnarray}
where the $\mathrm{argTopN}$ operation returns the top-$N$ outputs that maximize the function $\Pr(\mathbf{y}|\mathbf{x})$. These outputs can be sampled in various ways, depending on the search algorithm used by the model (e.g., nucleus sampling or beam search). Once the $N$-best output candidates are sampled, the reward model is used to evaluate and select the best one:
\begin{eqnarray}
\mathbf{y}_{\mathrm{best}} & = & \max\{R_{\phi}(\mathbf{x},\mathbf{y}_1),...,R_{\phi}(\mathbf{x},\mathbf{y}_N)\}
\end{eqnarray}
This process not only identifies the output with the highest reward but also makes it a direct evaluation of the effectiveness of the reward model. For example, the agreement between the reward model and human judgments can be assessed by judging whether $\mathbf{y}_{\mathrm{best}}$ matches the best output selected by humans. Given its speed and cost-efficiency compared to more complex evaluation methods such as those used in PPO, this approach is widely used for evaluating the performance of reward models \citep{rafailov:2023direct,gao-etal:2023scaling}. Nevertheless, in scenarios without a reward model, we can also employ the majority vote approach to determine the final output of the reasoning task. The basic steps involve first identifying the answer that most reasoning paths converge upon and then considering any of the outputs leading to this answer as the final output.
It is worth noting that the result of BoN sampling is also influenced by the diversity of the $N$-best list. This is a common issue with most reranking methods. Typically, we wish the $N$-best output candidates to be relatively high quality but sufficiently different from each other. In many text generation systems, the $N$-best outputs are very similar, often differing by just one or two words. The diversity issue is even more challenging in LLMs, as the $N$-best outputs sampled by an LLM can differ in their wordings. Yet, their semantic meanings are often quite similar. In practice, one can adjust the model hyperparameters and/or adopt different LLMs to generate more diverse output candidates for reranking. Nevertheless, as with many practical systems, we need to make a trade-off between selecting high-quality candidates and ensuring sufficient variation in the sampled outputs.
BoN sampling can also be used to train LLMs. A closely related method is rejection sampling. In this method, we first select the ``best'' outputs from the $N$-best lists via the reward model and then take these selected outputs to fine-tune the LLM. In this way, we can introduce human preferences into the training of LLMs via a much simpler approach than PPO. Many LLMs have adopted rejection sampling for fine-tuning \citep{nakano-etal:2021webgpt,touvron-etal:2023llama2}.
\subsubsection{Step-by-step Verification}
\label{sec:step-by-step-verification}
While BoN sampling is effective for scaling test-time computation, it is inefficient because the model must generate the entire reasoning path before evaluating its quality. Addressing this inefficiency, one approach is step-by-step verification, which evaluates the quality of each step while generating a reasoning path\footnote{We typically evaluate each reasoning step from multiple dimensions. A primary consideration is the presence of errors, assessing whether the intermediate step contains any inaccuracies. Furthermore, we evaluate the advantage conferred by the step, which examines the potential of the reasoning step to facilitate superior outcomes in subsequent steps \citep{setlur-etal:2024rewarding}.}. This allows us to identify the potential of a path at intermediate steps early and further reduce computational resources by abandoning unpromising paths. Note that in this approach, traditional reward models, which are designed to evaluate the entire output, are inadequate. Instead, we need to develop a process reward model to evaluate each step in the reasoning path.
We can collect or generate reasoning paths corresponding to problems from existing datasets to train a process reward model. Human experts then annotate each step in these paths for correctness. These annotations can be used to train LLMs or as rewards in reward modeling directly. However, in practice, richer annotations are often introduced \citep{lightman-etal:2024lets}. In addition to the \textit{correct} and \textit{incorrect} labels, a step can also be labeled as \textit{neutral} to indicate that while the step may be technically correct, it might still be problematic within the overall reasoning process. Additionally, an automatic process annotation framework can be utilized, which relies on generating multiple entire reasoning paths based on the current step and using the accuracy of these paths to serve as the quality annotation of the step \citep{wang-etal:2023math}.
Given a set of step-level annotated reasoning paths and corresponding inputs, we can train a reward model to provide a reward for each step in the reasoning process. The reward model can be treated as a classification model. So its architecture can be an LLM with a Softmax layer stacked on top, akin to the architecture depicted in Figure \ref{fig:reward-model}. Here, consider an reasoning path including $n_{s}$ steps, represented as $\mathbf{y} = \{\bar{\mathbf{y}}_{1},\cdots ,\bar{\mathbf{y}}_{n_{s}}\}$. At each step $k$, the process reward model takes both the problem description, denoted by $\mathbf{x}$, and the reasoning steps generated so far, denoted by $\bar{\mathbf{y}}$, as inputs. It then outputs a probability distribution over the set of labels ${\text{\textit{correct}}, \text{\textit{incorrect}}}$, or ${\text{\textit{correct}}, \text{\textit{incorrect}}, \text{\textit{neutral}}}$, to evaluate the reasoning at that point. This model can trained by a casual classification loss, e.g., Sigmoid Cross-entropy. Once trained, the reward model can be used to evaluate reasoning paths by assessing the correctness of each step. A simple method to use log-probabilities of classification to define the reward of each reasoning step, for example, the reward of the $k$-th reasoning step can be given by
\begin{eqnarray}
R_{\phi}(\mathbf{x},\bar{\mathbf{y}}_{\le k}) & = & \mathrm{Pr}_{\phi}(\textit{correct}|\mathbf{x},\bar{\mathbf{y}}_{\le k})
\end{eqnarray}
where $\Pr_{\phi}(\textit{correct}|\mathbf{x},\bar{\mathbf{y}}_{\le k})$ denotes the probability of the \textit{correct} label generated by the reward model. The reward score $R_{\phi}(\mathbf{x},\mathbf{y})$ can then be used to select the best step while generating a reasoning path. Additionally, as discussed in Section \ref{sec:generative-reward-models}, there is the option to train a generative process reward model that further improves this performance on evaluating the reasoning step, also called generative verifier in the literature \citep{zhang-etal:2024generative}. Note that in practice, the process reward model serves not only to provide rewards for test-time scaling but also to train the model using RL, e.g., the rewards from this model can be employed as shaping rewards.
\begin{figure*}[!t]
\centering
\input{section5/Figures/figure-step-by-step-with-prm.tex}
\caption{
Step-by-step verification with greedy search. This process can be readily extended to beam search by retaining the top steps at each step from all candidate paths for further scaling up test-time computation. As the blue line indicates, the reasoning steps can be refined using verification feedback, such as rewards, to enhance accuracy.
}
\label{fig:step-by-step-with-prm}
\end{figure*}
As illustrated in Figure \ref{fig:step-by-step-with-prm}, step-by-step verification can be conducted using a simple greedy search method. Specifically, multiple candidate reasoning steps are sampled at each step, and the process reward model is employed to select the best one. This selected step then serves as the foundation for generating the subsequent step, continuing this process until the entire reasoning path is obtained. In this way, any search method can be applied to improve test-time scaling. For example, we can expand the search space with beam search, retaining multiple promising reasoning steps at each step. Furthermore, the reasoning steps can be refined through the self-refinement technique at each step using verification feedback, such as rewards, to enhance accuracy \citep{yao-etal:2023tree}.
\subsubsection{Monte Carlo Tree Search}
In this subsection, we discuss the use of Monte Carlo Tree Search (MCTS), a popular search method in step-by-step verification. In practice, MCTS is not a new method but is well-established across various domains. Notably, it typically serves as an effective alternative to traditional RL methods, particularly in environments characterized by large or intricate state spaces, where conventional RL algorithms may falter. For example, within an RL framework, MCTS can aid in planning by simulating diverse actions to maximize rewards, as demonstrated by its success in complex game environments \citep{silver-etal:2016mastering}. In LLM reasoning, MCTS leverages randomness and structured tree search to probe potential reasoning paths, thereby expanding the search space in an efficient way. More specifically, as illustrated in Figure \ref{fig:monte-carlo-tree-search}, MCTS repeatedly cycles through the following four stages to explore potential reasoning paths:
\begin{figure*}[!t]
\centering
\input{section5/Figures/figure-mcts}
\caption{
Illustration of step-by-step verification with MCTS. The process consists of four phases: (1) Selection, where the algorithm navigates the decision tree to select the most promising node based on strategies like UCT; (2) Expansion, where new potential nodes are added to the tree for further exploration; (3) Simulation, where each new node is evaluated by simulating possible outcomes (i.e., reasoning paths); and (4) Backpropagation, where the outcomes from the simulations are used to refine node values, such as UCT values, thus enhancing the decision-making for subsequent explorations.
}
\label{fig:monte-carlo-tree-search}
\end{figure*}
\begin{itemize}
\item \textbf{Selection}.
This process begins at the root node (i.e., an input), where the algorithm selects promising child nodes (i.e., reasoning steps) based on specific selection strategies, such as the Upper Confidence Bound (UCT). More specifically, at each step $k$, among the sampled candidate reasoning steps, we aim to select the one that maximizes the UCT objective:
\begin{eqnarray}
\bar{\mathbf{y}}_{k}^{*} & = & \argmax_{\bar{\mathbf{y}}_{k}} \big(\bar{R}(\bar{\mathbf{y}}_{k}) + c \sqrt{\frac{\ln N_p}{N_{\bar{\mathbf{y}}_{k}}}} \big)
\label{eq:uct}
\end{eqnarray}
where $N_p$ denotes the total number of times the parent node has been visited, $N_{\bar{\mathbf{y}}_{k}}$ is the number of visits to the child node $\bar{\mathbf{y}}_{k}$, and $c$ is a constant that balances the trade-off between exploitation of known good paths and exploration of new paths. $\bar{R}(\cdot)$ denotes the reward of a reasoning step. This reward can be static, set by a process reward model, or dynamic, continuously updated based on the delayed rewards obtained from simulations.
\item \textbf{Expansion}. Once a leaf node (i.e., a newly generated reasoning step based on previously selected steps) is encountered and it does not represent a terminal state of the reasoning process, the algorithm expands this node by adding it as a new child node.
\item \textbf{Simulation}. The process performs simulations (or rollouts) for each new node added during the expansion stage. These simulated paths help evaluate the effectiveness of the reasoning steps initiated from the node. For example, the final answer obtained from the simulation can be compared with the correct answer to derive a delayed reward.
\item \textbf{Backpropagation}. The process updates the UCT values of prior nodes using the results of these simulations. Key updates include the visitation counts, $N_p$ and $N_{\bar{\mathbf{y}}_{k}}$ in Eq. (\ref{eq:uct}), reflecting how frequently each node has been explored. Additionally, this stage allows for the incorporation of delayed rewards to adjust the $\bar{R}(\bar{\mathbf{y}}_k)$ values.
\end{itemize}
\subsection{Large-scale Reinforcement Learning}
\label{sec:large-sclae-rl}
While test-time scaling is instrumental in exploring more effective reasoning paths, its potential is restricted if the model has not learned to generate high-quality reasoning paths. In other words, in practice, test-time scaling struggles to achieve desired outcomes without foundational reasoning ability, no matter how extensive it is \citep{yuan-etal:2023scaling,zeng-etal:2025revisiting}. Consequently, there has been a growing research interest in training LLMs specifically to develop reasoning abilities that mimic human-like processes. A straightforward approach is to use annotated reasoning data to train the LLM through SFT. However, a significant challenge in this approach is the high cost of annotations, especially since annotating detailed step-by-step reasoning paths is significantly more cost-intensive than annotating SFT data in other tasks like machine translation and summarization.
Another more advanced approach is to utilize large-scale RL, transitioning from human-annotated to self-reinforced learning processes. Unlike the RL used as a fine-tuning method discussed in Section \ref{sec:example-using-rl-training-llms}, which typically involves training for a few dozen or hundreds of steps at a small scale, this approach employs it on a larger scale with rewards to break free from the limitation of supervised reasoning data. This approach has enabled the development of robust reasoning models, such as OpenAI-o1 \citep{openai:2024learning}, DeepSeek-R1 \citep{guo:2025deepseek}, and Kimi-1.5 \citep{kimi-team:2025kimi}. Notably, DeepSeek-R1-Zero was able to develop a robust reasoning model by applying large-scale RL to a pre-trained LLM without the need for any annotated reasoning data. However, despite its successes, large-scale RL is not easy and introduces unique challenges that are not present at smaller scales, as follows.
One is that large-scale RL requires a highly generalizable reward model. As the scale of training increases, the model is trained on broader data, necessitating a reward model capable of effectively generalizing across this varied data. On the other hand, the extensive scope of training introduces significant variability in the model, leading to considerable changes in sampling behaviours. Consequently, it is crucial to ensure that the reward model possesses robust generalization capabilities to prevent overfitting and maintain the effectiveness of the learning process. There are several methods to achieve this. For example, \citet{guo:2025deepseek} incorporated rule-based rewards, as discussed in Section \ref{sec:lightweight-reward-methods}, such as format checking and answer verification in reasoning scenarios, which can provide a stable reward throughout the learning process. This suggests that in certain RL scenarios, prioritizing rule-based rewards may be beneficial if they can effectively describe human preferences. \citet{yuan-etal:2024selfrewarding} introduced a self-rewarding framework that dynamically updates the reward model based on the currently optimized policy model so that this reward model can effectively evaluate the behaviours from the current policy model.
Another is that large-scale RL might be constrained by the reference model. To prevent the policy model from deviating too far from its initial state, an SFT LLM is typically employed as the reference model, adding a penalty that restricts updates to ensure the policy model operates within a desirable behaviour region. However, reliance on a static reference model can limit the adaptability and innovation potential of the policy model in large-scale RL. A common strategy to mitigate this issue is to periodically update the reference model to better align with the improving capabilities and knowledge of the policy model, thus striking a balance between keeping stability and fostering adaptiveness during the training process \citep{gorbatovski-etal:2024learn}.
\subsection{Iterative Reinforcement Learning}
Training LLMs with RL usually follows a two-phase approach: training a pre-trained LLM with SFT and further training with an RL algorithm applied to the SFT LLM. However, using large-scale RL in such a two-phase approach to train LLMs in reasoning may lead to significant knowledge forgetting as the model continuously adjusts to fit the reward model. For example, as mentioned in DeepSeek-R1 \citep{guo:2025deepseek}, directly applying large-scale RL can achieve the desired reasoning outcomes, but often at the cost of reduced readability in the reasoning process. To address these challenges, we can utilize an iterative RL approach to continuously enhance the various capabilities of the LLM. Here we consider DeepSeek-R1 as an example to illustrate how to perform an iterative RL. The idea is to split the RL process into multiple phases, each designed to enhance different capabilities using varied rewards, to develop a robust reasoning model that is able to generate clearer and more comprehensible reasoning paths. Figure \ref{fig:iterative-rl} provides a schematic illustration of the iterative RL process. Here we give a brief outline of each phase involved.
\begin{figure*}[!t]
\centering
\input{section5/Figures/figure-iterative-rl}
\caption{
Illustration of iterative RL in DeepSeek-R1 \citep{guo:2025deepseek}. This method maintains multi-phase RL to develop a robust reasoning LLM with a GRPO algorithm. Initially, a pre-trained LLM is fine-tuned using a small set of high-quality labeled reasoning data with SFT as a cold start. Then, large-scale RL is applied specifically to enhance reasoning capabilities. After that, the model undergoes reject sampling across multiple tasks to refine output quality. In the final phase, the model is further trained using RL to enhance generalization across various tasks.
}
\label{fig:iterative-rl}
\end{figure*}
\begin{itemize}
\item The iterative RL aims to train an LLM that generates a human-like reasoning process. Initially, it involves collecting high-quality data from various reasoning tasks, such as mathematical problem-solving and code generation. This data is then used to fine-tune a pre-trained LLM as a cold start. This phase is designed to equip the LLM with basic reasoning capabilities, preventing linguistically confused or nonsensical reasoning paths from being sampled during the following RL phase.
\item After the cold start phase, large-scale RL is employed on the reasoning task. During this phase, rule-based rewards (i.e., format checking and answer verification) are utilized to optimize the reasoning process continuously. By doing so, the model learns longer and more reasonable reasoning paths, progressively enhancing its capacity to tackle complex reasoning tasks.
\item While the model has learned how to generate reasoning paths through large-scale RL, it often produces outputs with poor readability. This can mainly be attributed to the large-scale RL phase, which does not focus on optimizing the readability of the reasoning paths but instead solely on format correctness and answer accuracy. To address this issue, the third phase involves using rejection sampling to further enhance the readability of outputs. More specifically, multiple outputs are sampled from the model trained in the previous phase across various tasks such as mathematical reasoning, code generation, and question answering. A reward model then selects the best outputs, as discussed in Section \ref{sec:BoN-sampling}. These selected outputs are used to train the model using SFT, which aims to refine its performance by aligning it more closely with human-like reasoning.
\item In the final phase, multi-task RL is implemented to ensure broader generalization and maintain high performance across tasks beyond reasoning. The process involves training the model using RL on various tasks against a general reward model.
\end{itemize}
An interesting issue arises with this design of iterative RL: why is RL aimed at enhancing reasoning capabilities placed at the initial phase rather than at the end? This is partly because the rewards for the reasoning task, which are inherently more stable as described in Section \ref{sec:lightweight-reward-methods}, are well-suited for large-scale RL compared to traditional reward models. Furthermore, the reasoning capabilities could be applicable across many tasks, allowing for large-scale learning to boost performance in subsequent multi-task learning. We consider that this foundational enhancement in reasoning capabilities may contribute to the observed robust performance across various tasks, even without the use of labeled data in later phases. In this case, large-scale RL could be viewed as an approach to continued pre-training.
Since the optimization objective of each phase is different, the design of iterative RL can also be analyzed and understood from the perspective of multi-objective optimization \citep{wang-etal:2024hybrid}. In the context of multi-objective optimization, iterative RL for LLMs can be likened to interactive methods where the solution process is iterative, and preferences are actively defined and refined by the decision-maker during the search for the most preferred solutions \citep{miettinen-etal:2008introduction,deb-etal:2016multi}. More specifically, in this scenario, RL can be seen as a decision-maker, iteratively refining and enhancing the different capabilities of the LLM across different phases.
\section{Reinforcement Learning for Multimodal Models}
Multimodal learning involves models that process and relate information from multiple data modalities (e.g. vision, text, speech), enabling more comprehensive understanding than single-modality systems. By combining modalities, models can make more robust predictions and capture complementary information that one modality alone might miss. Such multimodal approaches have benefits across various applications, such as image caption and image generation. Given these advantages, multimodal learning has become increasingly significant as AI systems aim to perceive and reason more like humans, who naturally integrate sight, sound, and language \citep{baltruvsaitis-etal:2018multimodal}.
In the era of LLM, we can extend their capabilities into the multimodal domain by training them with diverse data types. For example, we can develop a visual language model by training an LLM with image-text pairs using SFT. Here, we return to the issue of aligning models with human preferences. Ideally, once aligned with human preferences through RL, the LLM would also generalize the target modality well by the SFT. However, the reality does not always meet expectations. Although an LLM may align well with human preferences in one aspect, it often struggles to generalize this alignment across a different modality. Thus, to align multimodal models with human preferences, we perform RL to enhance their performance further within the specific modality.
In this section, we use the visual language, speech generation, and diffusion models to discuss the application of RL in multimodal models.
\subsection{Visual Language Models}
While RL is commonly used to train LLMs, its application to other domains has been a prominent research topic. In multimodal language models\footnote{A multimodal language model is defined as a model that integrates an LLM with a multimodal encoder, such as CLIP \citep{radford-etal:2021learning}, allowing the LLM to process inputs beyond text, such as images. Recent literature has also introduced the use of LLMs for generating outputs in non-text modalities, such as images and speech \citep{xu-etal:2025qwen2,zhang-etal:2024mm}. However, in this section, we focus on the former definition of multimodal language models.}, for example, a notable trend is to perform RL training to improve their trustworthiness and helpfulness. This section considers Visual Language Models (VLMs), which connect a visual encoder to an LLM through a linear projector, facilitating general-purpose visual and language understanding. VLMs are currently the most explored extension in multimodal language research, and they form the foundation for many open-source multimodal language models, such as Qwen2.5-VL \citep{qwenTeam:2025qwen2.5-VL} and LLaMA-3.2-11B-Vision \citep{grattafiori-etal:2024llama}.
Training LLMs and VLMs with RL exhibits only minimal differences, primarily related to the input content. Unlike the textual input used for LLMs, the input for VLMs typically comprises a combination of one or multiple images and an instruction, denoted by $(\mathrm{\mathbf{I}}, \mathrm{\mathbf{x}})$, where $\mathrm{\mathbf{I}}$ represents the input images. These images are encoded into representations that are either concatenated with instruction embeddings or integrated through cross-attention mechanisms into the LLM. In practice, this subtle difference does not significantly affect the applicability of RL algorithms to VLMs. As a result, the RL training process for LLMs, as described in Section \ref{sec:example-using-rl-training-llms}, can be seamlessly adapted to train VLMs without major improvements \citep{yu-etal:2024rlhf,wang-etal:2024rovrm,zang2025internlm,ji2025safe}.
\begin{figure*}[!t]
\centering
\caption{RoVRM}
\label{fig:preference-transfer}
\end{figure*}
However, training VLMs with RL is not a low-hanging fruit in practical applications. This is because it typically encounters the challenge of training a visual reward model due to the scarcity of high-quality visual preference data. One straightforward approach to address this issue is to generate visual preference data through the automatic preference data generation method described in Section \ref{sec:automatic-preference-data-generation} \citep{yu-etal:2024rlaif}. Another alternative is a multi-stage training approach for the visual reward model, motivated by a simple idea: human preferences are well captured in text, and these preferences can be transferred across modalities \citep{wang-etal:2024rovrm}. By leveraging textual preference data, this approach reduces the dependence on visual preference data in training a visual reward model. More specifically, as illustrated in Figure \ref{fig:preference-transfer}, we can train a visual reward model in the following three stages:
\begin{itemize}
\item Stage 1: pre-training with large-scale textual preference data. Given the transferability of human preferences across different modalities, we can begin by using large-scale textual preference data to pre-train the visual reward model. Note that the projector parameters are frozen without images. This stage can be considered as providing a stronger starting point for training the visual reward model, as it enables the model to pre-learn general human preferences.
\item Stage 2: fine-tuning with image caption-based preference data. Pre-learned human preferences cannot be directly applied to vision tasks due to both \textit{task gap} and \textit{modality gap}. To bridge the task gap, we fine-tune the reward model using image caption-based preference data in this stage. The rationale behind this approach is that general textual preference data does not cover vision-specific tasks, such as ``Please describe the content in this image''. We use such data to fine-tune the model, adapting it to vision-specific tasks. The projector parameters are frozen at this stage.
\item Stage 3: fine-tuning with small-scale visual preference data. To further bridge the modality gap, we use visual preference data in the final stage to train the model. Note that in this phase, we also train the projector parameters.
\end{itemize}
In this process, not all preference data may align with the preferences used in subsequent phases, potentially leading to preference conflicts. We can enhance the visual reward model through data selection techniques, such as LESS \citep{xia-etal:2024less} to address this. In fact, preference transfer is effective across modalities and has also been shown to work across different tasks and languages \citep{cheng-etal:2023everyone,wu-etal:2024reuse}. Interested readers can refer to these papers for more detailed discussions of these topics.
Apart from preference data, another approach to improving the visual model is to integrate additional image content, such as image captions, into the reward model \citep{sun-etal:2023aligning}. This approach aims to achieve factually augmented reward prediction, addressing reward hacking. Specifically, in the original setup, the reward model predicts a reward based solely on the image, input, and output; that is, the reward model’s input is $[\mathbf{I}, \mathbf{x}, \mathbf{y}]$. In the factually augmented setup, the reward model also receives additional input in the form of the textual image caption $\mathbf{C}$, resulting in an input of $[\mathbf{I}, \mathbf{C}, \mathbf{x}, \mathbf{y}]$. The basic idea is that the backbone of the visual reward model remains a well-trained LLM with an in-context learning ability. With this ability, we can provide additional content to help the model predict rewards more accurately.
While our discussion primarily focuses on models for visual inputs in the context of visual language models, the RL techniques are adaptable across inputs of various modalities, such as video and audio.
% visual language model as reward model
% 视觉奖励模型训练大致思路
% image reward
% 融合文本数据进行提升
% 视觉奖励模型应用场景---对齐图生文模型 文生图评估模型
% 视觉语言模型的定义
% 框架
% Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
% Aligning Large Multimodal Models with Factually Augmented RLHF
% 偏好数据
% RLAIF-V
% Silkie: Preference Distillation for Large Visual Language Models
% RoVRM
\subsection{Speech Generation Models}
% SpeechAlign: Aligning Speech Generation to Human Preferences
\subsection{Diffusion Models}
% Image Reward
% 为什么要进行强化学习
% 大体训练思路和流程
% DDPO (denoising diffusion preference optimization model)
\section{Summary}
% future work
% 强化学习作为一种预训练方式
% 高效强化学习方法
% 全模态RL
\ No newline at end of file
\clearpage
\section{Systems and Datasets}
\begin{table}[h]
\centering
\scalebox{0.88}{
\input{section8/tables/dataset}}
\caption{datasets}
\label{tab:dataset}
\end{table}
\begin{table}[h]
\centering
\scalebox{0.88}{
\input{section8/tables/systems}}
\caption{systems}
\label{tab:system}
\end{table}
\ No newline at end of file
\begin{tabular}{lccccc}
\toprule[1.1pt]
\multirow{2}{*}{Dataset Name} & \multirow{2}{*}{\begin{tabular}[c]{@{}c@{}}Sample\\Size\end{tabular}} & \multirow{2}{*}{\begin{tabular}[c]{@{}c@{}}Response\\Size\end{tabular}} & \multirow{2}{*}{Modality} & \multicolumn{2}{c}{Feedback} \\ \cmidrule(l){5-6}
& & & & Source & Category \\ \midrule
\href{https://url}{HuggingFaceH4/stack-exchange-preferences}
&10,000 &1 & \faIcon{file-alt} & Human & Score \\
\href{https://url}{Skywork/Skywork-Reward-Preference-80K-v0.2} &10,000 &2 & \faIcon{images} & AI & Ranking \\
\href{https://huggingface.co/datasets/openbmb/UltraFeedback}{openbmb/UltraFeedback} &10,000 &3 & \faIcon{video} & Rule & \\
&10,000 &4 & \faIcon{volume-up} & & \\
& & &\faIcon{file-alt} \faIcon{images} & & \\
& & & & & \\
& & & & & \\
& & & & & \\
\bottomrule[1.1pt]
\end{tabular}
\ No newline at end of file
\begin{tabular}{lccccc}
\toprule[1.1pt]
\multirow{2}{*}{System Name} & \multicolumn{4}{c}{Supported Modality} & \multirow{2}{*}{Supported Training Approaches} \\ \cmidrule(r){2-5}
& Text & Image & Video & Audio & \\ \midrule
\href{https://github.com/huggingface/trl}{TRL} & \CheckmarkBold & \CheckmarkBold &\XSolidBrush & \CheckmarkBold & Reward Modeling, PPO, GRPO, DPO, Online-DPO, etc. \\
\href{https://github.com/OpenRLHF/OpenRLHF}{OpenRLHF}& \CheckmarkBold & \XSolidBrush & \XSolidBrush & \XSolidBrush & Sft, Reject Sampling, PPO, GRPO, DPO, KTO, etc.\\
\href{https://github.com/hiyouga/EasyR1}{EasyR1}& \CheckmarkBold & \CheckmarkBold & \XSolidBrush & \XSolidBrush & GRPO, Reinforce++, Remax, RLOO, etc. \\
\href{https://github.com/volcengine/verl}{veRL}& \CheckmarkBold & \XSolidBrush & \XSolidBrush & \XSolidBrush & GRPO, PPO, Remax, RLOO, SFT, etc. \\
\href{https://github.com/OpenRLHF/OpenRLHF-M}{OpenRLHF-M}& \XSolidBrush & \CheckmarkBold & \XSolidBrush & \XSolidBrush & PPO, GRPO, RLOO, Online-RLHF, Reject-Sampling, etc. \\
\href{https://github.com/PKU-Alignment/align-anything}{Align-Anything}& \CheckmarkBold & \CheckmarkBold & \CheckmarkBold & \CheckmarkBold &PPO, GRPO, DPO,KTO, ORPO, etc. \\
\bottomrule[1.1pt]
\end{tabular}
\ No newline at end of file
Markdown 格式
0%
您添加了 0 到此讨论。请谨慎行事。
请先完成此评论的编辑!
注册 或者 后发表评论