\subsection{Fundamentals of Large Language Models}
In this subsection, we introduce the fundamentals of supervised fine-tuning for LLMs and focus on presenting the key concepts and notations necessary for the discussions in the following sections.
In this subsection, we introduce the fundamentals of LLMs and focus on presenting the key concepts and notations necessary for the discussions in the following sections.
Although pre-trained LLMs possess a vast amount of general knowledge, they are typically limited to tasks related to language modeling. Our ultimate goal is to enable them to perform various specific tasks, such as summarization, question answering, and machine translation.
One straightforward method to achieve this goal is supervised fine-tuning (SFT), in which the pre-trained LLM is further trained on a dataset comprising task-specific inputs (i.e., instruction + user input)\footnote{In this paper, the \textit{instruction} represents the description of a specific task, e.g., ``Summarize the following article''; the \textit{user input} represents the extra content added to the instruction, e.g., ``Article: In recent years, solar energy has seen a significant increase in the amount of energy available to the public. Solar energy has seen unprecedented growth, becoming the fastest-growing ...''} paired with their expected outputs \citep{ouyang:2022training,wei-etal:2022finetuned}.
\node [ynode,anchor=north west] (output1) at ([yshift=-.2cm]note.south west) {
\scriptsize{Improving accuracy in solving math problems is crucial for success in mathematics. Here are three tips that can help: \\
1.Practice regularly... \\
2.Understand the concepts... \\
\ctext[RGB]{255,204,204}{3.Double-check your work...any-one can significantly enhance their math problem-solving abilities.}\\
1.Practice regularly... \\
2.Understand the concepts... \\
\ctext[RGB]{255,204,204}{3.Double-check your work...any-one can significantly enhance their math problem-solving abilities.}\\
\textbf{\{Total Tokens: 160, Reward: 16\}}}
};
\node [anchor=east](y1) at ([xshift=-.2cm]output1.west) {$\mathbf{y}_1$};
\node [ynode,anchor=north] (output2) at ([yshift=-.2cm]output1.south) {
\scriptsize{Improving math accuracy requires careful attention to detail, strategic problem-solving techniques, and consistent practice. Here are three tips: \\
1.Avoid rushing and read carefully... \\
2.Use estimation to validate answers... \\
\ctext[RGB]{255,204,204}{3.Double-check your work...any-one can significantly enhance their math problem-solving abilities.}\\
1.Avoid rushing and read carefully... \\
2.Use estimation to validate answers... \\
\ctext[RGB]{255,204,204}{3.Double-check your work...any-one can significantly enhance their math problem-solving abilities.}\\
\textbf{\{Total Tokens: 750, Reward: 75\}}}
};
\node [anchor=east](y2) at ([xshift=-.2cm]output2.west) {$\mathbf{y}_2$};
\node [anchor=north] (x) at ([yshift=-.5cm]llm.south) {\large{$\mathbf{x}$}};
\node [draw, dashed, rounded corners=2pt,anchor=south, align=left, text width=\textwid, minimum height=2.6cm] (output) at ([yshift=.5cm]llm.north) {\footnotesize{There are three tips }\footnotesize{\sethlcolor{red!20}\hl{for improving for improving for improving}\footnotesize{accuracy in solving math problems:}}\\
\footnotesize{1.Practice regularly. \\
2.Understand the concepts. \\
3.Double-check your work.}};
\footnotesize{1.Practice regularly. \\
2.Understand the concepts. \\
3.Double-check your work.}};
\node [draw, dashed, rounded corners=2pt,anchor=north, align=left, text width=\textwid] (input) at ([yshift=0cm]x.south) {\footnotesize{Give me three tips to improve my accuracy in solving math problems.}};
\node [anchor=south] (y) at ([yshift=0cm]output.north) {\large{$\mathbf{y}$}};
...
...
@@ -33,7 +33,7 @@
\node [block=lolred,anchor=west] (llm1) at ([xshift=\sep*3]llm.east) {Reference Policy (LLM)};
\node [anchor=north] (x) at ([yshift=-.5cm]llm1.south) {\large{$\mathbf{x}$}};
\node [draw, dashed, rounded corners=2pt,anchor=south, align=left, text width=\textwid, minimum height=2.6cm] (output) at ([yshift=.5cm]llm1.north) {\footnotesize{There are three tips }\footnotesize{\sethlcolor{red!20}\hl{for improving improving improving improving improving improving improving improving }\footnotesize{accuracy in solving math problems:}}\\
\node [draw, dashed, rounded corners=2pt,anchor=north, align=left, text width=\textwid] (input) at ([yshift=0cm]x.south) {\footnotesize{Give me three tips to improve my accuracy in solving math problems.}};
\node [anchor=south] (y) at ([yshift=0cm]output.north) {\large{$\mathbf{y}$}};
\node [ynode,anchor=north west] (output1) at ([yshift=-.2cm]note.south west) {
\scriptsize{There are three tips for improving accuracy in solving math problems: \\
1.Practice regularly. \\
2.Understand the concepts. \\
3.Double-check your work. \\
1.Practice regularly. \\
2.Understand the concepts. \\
3.Double-check your work. \\
\textbf{\{Total Tokens: 20\}}}
};
\node [anchor=east](y1) at ([xshift=-.2cm]output1.west) {$\mathbf{y}_1$};
\node [ynode,anchor=north] (output2) at ([yshift=-.2cm]output1.south) {
\scriptsize{Improving accuracy in solving math problems is crucial for success in mathematics. Here are three tips that can help: \\
1.Practice regularly: \\
1.Practice regularly: \\
~~~- Consistent practice builds fluency and familiarity with different problem types. \\
... \\
By combining your advice with these additional tips, anyone can significantly enhance their math problem-solving abilities. \\
...
...
@@ -28,7 +28,7 @@
\node [anchor=east](y2) at ([xshift=-.2cm]output2.west) {$\mathbf{y}_2$};
\node [ynode,anchor=north] (output3) at ([yshift=-.2cm]output2.south) {
\scriptsize{Improving math accuracy requires careful attention to detail, strategic problem-solving techniques, and consistent practice. Here are three tips: \\
1.Practice regularly: \\
1.Practice regularly: \\
~~~- Mathematics is a skill, and like any skill, it improves with consistent practice.... \\
... \\
This is a great approach for anyone looking to improve their math skills, from students to professionals. \\
\node [draw,rounded corners=2pt,inner sep=2pt,text width=5.5em, anchor=north west,font=\linespread{0.8}\selectfont, minimum height=9ex] (x) at ([yshift=-1cm,xshift=-1cm]llm.south west) {\scriptsize{Give me three tips to improve my accuracy in solving math problems.\\}};
\node [draw,rounded corners=2pt,inner sep=2pt,text width=13.5em, anchor=north east,font=\linespread{0.8}\selectfont, minimum height=9ex] (y) at ([yshift=-1cm,xshift=1cm]llm.south east) {\scriptsize{Improving math accuracy requires careful attention to detail, ... \\
1.Practice regularly: ... \\
1.Practice regularly: ... \\
... , anyone can significantly enhance their math problem-solving abilities.\\}};
\node [anchor=north] at ([yshift=-.2cm]x.south) {$\mathbf{x}$};
\node [anchor=north] at ([yshift=-.2cm]y.south) {$\mathbf{y}_a$};
...
...
@@ -33,9 +33,9 @@
\node [draw,rounded corners=2pt,inner sep=2pt,text width=5.5em, anchor=north west,font=\linespread{0.8}\selectfont, minimum height=9ex] (x) at ([yshift=-1cm,xshift=-1cm]llm.south west) {\scriptsize{Give me three tips to improve my accuracy in solving math problems.\\}};
\node [draw,rounded corners=2pt,inner sep=2pt,text width=13.5em, anchor=north east,font=\linespread{0.8}\selectfont, minimum height=9ex] (y) at ([yshift=-1cm,xshift=1cm]llm.south east) {\scriptsize{There are three tips for improving accuracy in solving math problems: \\
1.Practice regularly. \\
2.Understand the concepts.\\
3.Double-check your work.\\}};
1.Practice regularly. \\
2.Understand the concepts.\\
3.Double-check your work.\\}};
\node [anchor=north] at ([yshift=-.2cm]x.south) {$\mathbf{x}$};
\node [anchor=north] at ([yshift=-.2cm]y.south) {$\mathbf{y}_b$};
% An example of REINFORCE and optimization objective
\section{An Example of Using Reinforcement Learning to Train LLMs}
\section{A Simple Example}
\label{sec:example-using-rl-training-llms}
We begin by considering a practical scenario: using an LLM as a homework assistant. In this context, we aim to improve the capability of an SFT LLM to handle education-related inputs more effectively. Suppose we have a homework assistant powered by an SFT LLM, and a student types the input ``Give me three tips to improve my accuracy in solving math problems.'' A typical SFT LLM might generate a short output, such as
...
...
@@ -12,9 +12,9 @@ We begin by considering a practical scenario: using an LLM as a homework assista
\setlength{\leftskip}{2em}
\setlength{\rightskip}{2em}
There are three tips for improving accuracy in solving math problems: \\ [1mm]
1.Practice regularly. \\ [1mm]
2.Understand the concepts. \\ [1mm]
3.Double-check your work.
1.Practice regularly. \\ [1mm]
2.Understand the concepts. \\ [1mm]
3.Double-check your work.
\endgroup
...
...
@@ -97,7 +97,7 @@ We can understand this optimization process from the perspective of hypothesis s
\subsection{Temporal Decomposition}
\label{sec:temporal-decomposition}
While optimizing using Eq. (\ref{eq:rl-j-theta-gradient-simplified}) is intuitive, a potential issue arises: in some cases, the sampled outputs may significantly overlap, yet the rewards for the overlapping parts differ. For example, as illustrated in Figure \ref{fig:an-overlap-example}, if outputs $\mathbf{y}_1$ and $\mathbf{y}_2$ both include the same tokens in the third tip ``3.Double-check your work...anyone can significantly enhance their math problem-solving abilities.'', but due to variations in the entire outputs, they receive markedly different rewards (e.g., $R(\mathbf{y}_1)=50$ vs. $R(\mathbf{y}_2)=10$). Ideally, we would want similar generation behaviors to be rewarded consistently, without significant variation, as their contributions are equivalent to the sum of rewards.
While optimizing using Eq. (\ref{eq:rl-j-theta-gradient-simplified}) is intuitive, a potential issue arises: in some cases, the sampled outputs may significantly overlap, yet the rewards for the overlapping parts differ. For example, as illustrated in Figure \ref{fig:an-overlap-example}, if outputs $\mathbf{y}_1$ and $\mathbf{y}_2$ both include the same tokens in the third tip ``3.Double-check your work...anyone can significantly enhance their math problem-solving abilities.'', but due to variations in the entire outputs, they receive markedly different rewards (e.g., $R(\mathbf{y}_1)=50$ vs. $R(\mathbf{y}_2)=10$). Ideally, we would want similar generation behaviors to be rewarded consistently, without significant variation, as their contributions are equivalent to the sum of rewards.
Before discussing an approach to solve this issue, we first perform temporal decomposition for the term $\frac{\partial\log\mathrm{Pr}_{\theta}(\mathbf{y}|\mathbf{x})}{\partial\theta} R(\mathbf{y})$ in Eq. (\ref{eq:rl-j-theta-gradient-simplified}), and obtain
\textit{Improving math accuracy requires careful attention to detail, strategic problem-solving techniques, and consistent practice. Here are three tips: \\
1.Practice regularly: \\
1.Practice regularly: \\
- Consistent practice builds fluency and familiarity with different problem types... \\
... \\
By combining your advice with these additional tips, anyone can significantly enhance their math problem-solving abilities.}
...
...
@@ -48,9 +48,9 @@ Output B:
\vspace{0.1cm}
\textit{There are three tips for improving accuracy in solving math problems: \\