Commit 2da46525 by wangchenglong

update.

parent c3b29726
...@@ -163,3 +163,4 @@ Based on these estimated advantages, we formulate the GRPO objective as follows: ...@@ -163,3 +163,4 @@ Based on these estimated advantages, we formulate the GRPO objective as follows:
\mathrm{clip}(\frac{\mathrm{Pr}_\theta(a_t^i|s_t^i)}{\mathrm{Pr}_{\theta_{\mathrm{old}}}(a_t^i|s_t^i)},1-\epsilon,1+\epsilon)\hat{A}_i)\right] \mathrm{clip}(\frac{\mathrm{Pr}_\theta(a_t^i|s_t^i)}{\mathrm{Pr}_{\theta_{\mathrm{old}}}(a_t^i|s_t^i)},1-\epsilon,1+\epsilon)\hat{A}_i)\right]
\end{eqnarray} \end{eqnarray}
During training, one challenge is trajectory sampling
\ No newline at end of file
Markdown 格式
0%
您添加了 0 到此讨论。请谨慎行事。
请先完成此评论的编辑!
注册 或者 后发表评论