Skip to content
项目
群组
代码片段
帮助
当前项目
正在载入...
登录 / 注册
切换导航面板
R
rl-introduction
概览
Overview
Details
Activity
Cycle Analytics
版本库
Repository
Files
Commits
Branches
Tags
Contributors
Graph
Compare
Charts
问题
0
Issues
0
列表
Board
标记
里程碑
合并请求
0
Merge Requests
0
CI / CD
CI / CD
流水线
作业
日程表
图表
维基
Wiki
代码片段
Snippets
成员
Collapse sidebar
Close sidebar
活动
图像
聊天
创建新问题
作业
提交
Issue Boards
Open sidebar
wangchenglong
rl-introduction
Commits
66a643f6
Commit
66a643f6
authored
Sep 11, 2026
by
wangchenglong
Browse files
Options
Browse Files
Download
Email Patches
Plain Diff
update.
parent
5d7cc1b9
隐藏空白字符变更
内嵌
并排
正在显示
1 个修改的文件
包含
1 行增加
和
1 行删除
+1
-1
README.md
+1
-1
没有找到文件。
README.md
查看文件 @
66a643f6
...
@@ -6,7 +6,7 @@ Reinforcement learning (RL) has become an important training paradigm for large
...
@@ -6,7 +6,7 @@ Reinforcement learning (RL) has become an important training paradigm for large
\max_{\pi} \; \mathbb{E}_{\tau \sim \pi}
\max_{\pi} \; \mathbb{E}_{\tau \sim \pi}
\left[
\left[
\sum_{t=0}^{T} r_t
\sum_{t=0}^{T} r_t
\right]
,
\right]
```
`
```
`
where $\pi$ denotes the policy, $\tau$ denotes a trajectory of interactions, and $r_t$ is the reward received at step $t$.
where $\pi$ denotes the policy, $\tau$ denotes a trajectory of interactions, and $r_t$ is the reward received at step $t$.
...
...
编写
预览
Markdown
格式
0%
重试
或
添加新文件
添加附件
取消
您添加了
0
人
到此讨论。请谨慎行事。
请先完成此评论的编辑!
取消
请
注册
或者
登录
后发表评论