Commit 66a643f6 by wangchenglong

update.

parent 5d7cc1b9
......@@ -6,7 +6,7 @@ Reinforcement learning (RL) has become an important training paradigm for large
\max_{\pi} \; \mathbb{E}_{\tau \sim \pi}
\left[
\sum_{t=0}^{T} r_t
\right],
\right]
````
where $\pi$ denotes the policy, $\tau$ denotes a trajectory of interactions, and $r_t$ is the reward received at step $t$.
......
Markdown 格式
0%
您添加了 0 到此讨论。请谨慎行事。
请先完成此评论的编辑!
注册 或者 后发表评论