Commit c872babb by wangchenglong

update.

parent 71f5dbb2
\begin{thebibliography}{104} \begin{thebibliography}{109}
\providecommand{\natexlab}[1]{#1} \providecommand{\natexlab}[1]{#1}
\providecommand{\url}[1]{\texttt{#1}} \providecommand{\url}[1]{\texttt{#1}}
\expandafter\ifx\csname urlstyle\endcsname\relax \expandafter\ifx\csname urlstyle\endcsname\relax
...@@ -158,6 +158,11 @@ Jian Hu. ...@@ -158,6 +158,11 @@ Jian Hu.
\newblock \emph{ArXiv preprint}, abs/2501.03262, 2025. \newblock \emph{ArXiv preprint}, abs/2501.03262, 2025.
\newblock URL \url{https://arxiv.org/abs/2501.03262}. \newblock URL \url{https://arxiv.org/abs/2501.03262}.
\bibitem[Hu et~al.(2025)Hu, Zhao, Xu, Sun, Lou, Lin, Luo, and Rajmohan]{hu-etal:agentgen}
Mengkang Hu, Pu~Zhao, Can Xu, Qingfeng Sun, Jian-Guang Lou, Qingwei Lin, Ping Luo, and Saravan Rajmohan.
\newblock Agentgen: Enhancing planning abilities for large language model based agent via environment and task generation.
\newblock In \emph{Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1}, pp.\ 496--507, 2025.
\bibitem[Ji et~al.(2025)Ji, Chen, Pan, Zhu, Zhang, Li, Hong, Chen, Zhou, Wang, et~al.]{ji2025safe} \bibitem[Ji et~al.(2025)Ji, Chen, Pan, Zhu, Zhang, Li, Hong, Chen, Zhou, Wang, et~al.]{ji2025safe}
Jiaming Ji, Xinyu Chen, Rui Pan, Han Zhu, Conghui Zhang, Jiahao Li, Donghai Hong, Boyuan Chen, Jiayi Zhou, Kaile Wang, et~al. Jiaming Ji, Xinyu Chen, Rui Pan, Han Zhu, Conghui Zhang, Jiahao Li, Donghai Hong, Boyuan Chen, Jiayi Zhou, Kaile Wang, et~al.
\newblock Safe rlhf-v: Safe reinforcement learning from human feedback in multimodal large language models. \newblock Safe rlhf-v: Safe reinforcement learning from human feedback in multimodal large language models.
...@@ -205,6 +210,11 @@ Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. ...@@ -205,6 +210,11 @@ Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu.
\newblock \doi{10.18653/v1/2023.emnlp-main.153}. \newblock \doi{10.18653/v1/2023.emnlp-main.153}.
\newblock URL \url{https://aclanthology.org/2023.emnlp-main.153}. \newblock URL \url{https://aclanthology.org/2023.emnlp-main.153}.
\bibitem[Longpre et~al.(2023)Longpre, Hou, Vu, Webson, Chung, Tay, Zhou, Le, Zoph, Wei, et~al.]{longpre-etal:flan}
Shayne Longpre, Le~Hou, Tu~Vu, Albert Webson, Hyung~Won Chung, Yi~Tay, Denny Zhou, Quoc~V Le, Barret Zoph, Jason Wei, et~al.
\newblock The flan collection: Designing data and methods for effective instruction tuning.
\newblock In \emph{International conference on machine learning}, pp.\ 22631--22648. PMLR, 2023.
\bibitem[Mahan et~al.(2024)Mahan, Van~Phung, Rafailov, Blagden, Lile, Castricato, Fr{\"a}nken, Finn, and Albalak]{mahan-etal:2024generative} \bibitem[Mahan et~al.(2024)Mahan, Van~Phung, Rafailov, Blagden, Lile, Castricato, Fr{\"a}nken, Finn, and Albalak]{mahan-etal:2024generative}
Dakota Mahan, Duy Van~Phung, Rafael Rafailov, Chase Blagden, Nathan Lile, Louis Castricato, Jan-Philipp Fr{\"a}nken, Chelsea Finn, and Alon Albalak. Dakota Mahan, Duy Van~Phung, Rafael Rafailov, Chase Blagden, Nathan Lile, Louis Castricato, Jan-Philipp Fr{\"a}nken, Chelsea Finn, and Alon Albalak.
\newblock Generative reward models. \newblock Generative reward models.
...@@ -519,6 +529,11 @@ Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and K ...@@ -519,6 +529,11 @@ Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and K
\newblock In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), \emph{Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023}, 2023. \newblock In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), \emph{Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023}, 2023.
\newblock URL \url{http://papers.nips.cc/paper\_files/paper/2023/hash/271db9922b8d1f4dd7aaef84ed5ac703-Abstract-Conference.html}. \newblock URL \url{http://papers.nips.cc/paper\_files/paper/2023/hash/271db9922b8d1f4dd7aaef84ed5ac703-Abstract-Conference.html}.
\bibitem[Yin et~al.(2024)Yin, Brahman, Ravichander, Chandu, Chang, Choi, and Lin]{yin-etal:agentlumos}
Da~Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, and Bill~Yuchen Lin.
\newblock Agent lumos: Unified and modular training for open-source language agents.
\newblock In \emph{Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, pp.\ 12380--12403, 2024.
\bibitem[Yu et~al.(2024{\natexlab{a}})Yu, Yao, Zhang, He, Han, Cui, Hu, Liu, Zheng, Sun, et~al.]{yu-etal:2024rlhf} \bibitem[Yu et~al.(2024{\natexlab{a}})Yu, Yao, Zhang, He, Han, Cui, Hu, Liu, Zheng, Sun, et~al.]{yu-etal:2024rlhf}
Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, et~al. Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, Maosong Sun, et~al.
\newblock Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. \newblock Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.
...@@ -546,6 +561,11 @@ Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenx ...@@ -546,6 +561,11 @@ Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenx
\newblock Internlm-xcomposer2. 5-reward: A simple yet effective multi-modal reward model. \newblock Internlm-xcomposer2. 5-reward: A simple yet effective multi-modal reward model.
\newblock \emph{arXiv preprint arXiv:2501.12368}, 2025. \newblock \emph{arXiv preprint arXiv:2501.12368}, 2025.
\bibitem[Zeng et~al.(2024)Zeng, Liu, Lu, Wang, Liu, Dong, and Tang]{zeng-etal:agenttuning}
Aohan Zeng, Mingdao Liu, Rui Lu, Bowen Wang, Xiao Liu, Yuxiao Dong, and Jie Tang.
\newblock Agenttuning: Enabling generalized agent abilities for llms.
\newblock In \emph{Findings of the Association for Computational Linguistics: ACL 2024}, pp.\ 3053--3077, 2024.
\bibitem[Zeng et~al.(2025)Zeng, Cheng, Yin, Zhou, and Qiu]{zeng-etal:2025revisiting} \bibitem[Zeng et~al.(2025)Zeng, Cheng, Yin, Zhou, and Qiu]{zeng-etal:2025revisiting}
Zhiyuan Zeng, Qinyuan Cheng, Zhangyue Yin, Yunhua Zhou, and Xipeng Qiu. Zhiyuan Zeng, Qinyuan Cheng, Zhangyue Yin, Yunhua Zhou, and Xipeng Qiu.
\newblock Revisiting the test-time scaling of o1-like models: Do they truly possess test-time scaling capabilities? \newblock Revisiting the test-time scaling of o1-like models: Do they truly possess test-time scaling capabilities?
...@@ -584,6 +604,11 @@ Lianmin Zheng, Wei{-}Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao ...@@ -584,6 +604,11 @@ Lianmin Zheng, Wei{-}Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao
\newblock In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), \emph{Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023}, 2023. \newblock In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), \emph{Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023}, 2023.
\newblock URL \url{http://papers.nips.cc/paper\_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets\_and\_Benchmarks.html}. \newblock URL \url{http://papers.nips.cc/paper\_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets\_and\_Benchmarks.html}.
\bibitem[Zhou et~al.(2023)Zhou, Liu, Xu, Iyer, Sun, Mao, Ma, Efrat, Yu, Yu, et~al.]{zhou-etal:lima}
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et~al.
\newblock Lima: Less is more for alignment.
\newblock \emph{Advances in Neural Information Processing Systems}, 36:\penalty0 55006--55021, 2023.
\bibitem[Zhou et~al.(2024)Zhou, Wang, Hu, Xiao, Zhang, and Zhu]{zhou:2024prior} \bibitem[Zhou et~al.(2024)Zhou, Wang, Hu, Xiao, Zhang, and Zhu]{zhou:2024prior}
Hang Zhou, Chenglong Wang, Yimin Hu, Tong Xiao, Chunliang Zhang, and Jingbo Zhu. Hang Zhou, Chenglong Wang, Yimin Hu, Tong Xiao, Chunliang Zhang, and Jingbo Zhu.
\newblock Prior constraints-based reward model training for aligning large language models. \newblock Prior constraints-based reward model training for aligning large language models.
......
@inproceedings{hu-etal:agentgen,
title={Agentgen: Enhancing planning abilities for large language model based agent via environment and task generation},
author={Hu, Mengkang and Zhao, Pu and Xu, Can and Sun, Qingfeng and Lou, Jian-Guang and Lin, Qingwei and Luo, Ping and Rajmohan, Saravan},
booktitle={Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1},
pages={496--507},
year={2025}
}
@inproceedings{yin-etal:agentlumos,
title={Agent lumos: Unified and modular training for open-source language agents},
author={Yin, Da and Brahman, Faeze and Ravichander, Abhilasha and Chandu, Khyathi and Chang, Kai-Wei and Choi, Yejin and Lin, Bill Yuchen},
booktitle={Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
pages={12380--12403},
year={2024}
}
@article{zhou-etal:lima,
title={Lima: Less is more for alignment},
author={Zhou, Chunting and Liu, Pengfei and Xu, Puxin and Iyer, Srinivasan and Sun, Jiao and Mao, Yuning and Ma, Xuezhe and Efrat, Avia and Yu, Ping and Yu, Lili and others},
journal={Advances in Neural Information Processing Systems},
volume={36},
pages={55006--55021},
year={2023}
}
@inproceedings{longpre-etal:flan,
title={The flan collection: Designing data and methods for effective instruction tuning},
author={Longpre, Shayne and Hou, Le and Vu, Tu and Webson, Albert and Chung, Hyung Won and Tay, Yi and Zhou, Denny and Le, Quoc V and Zoph, Barret and Wei, Jason and others},
booktitle={International conference on machine learning},
pages={22631--22648},
year={2023},
organization={PMLR}
}
@inproceedings{zeng-etal:agenttuning,
title={Agenttuning: Enabling generalized agent abilities for llms},
author={Zeng, Aohan and Liu, Mingdao and Lu, Rui and Wang, Bowen and Liu, Xiao and Dong, Yuxiao and Tang, Jie},
booktitle={Findings of the Association for Computational Linguistics: ACL 2024},
pages={3053--3077},
year={2024}
}
@article{zhang-etal:agentohana, @article{zhang-etal:agentohana,
title={Agentohana: Design unified data and training pipeline for effective agent learning}, title={Agentohana: Design unified data and training pipeline for effective agent learning},
......
No preview for this file type
...@@ -157,12 +157,38 @@ minus0.2ex}{\large\bf\raggedright}} ...@@ -157,12 +157,38 @@ minus0.2ex}{\large\bf\raggedright}}
\def\subsubsection{\@startsection{subsubsection}{3}{\z@}{-1.5ex \def\subsubsection{\@startsection{subsubsection}{3}{\z@}{-1.5ex
plus -0.5ex minus -.2ex}{0.5ex plus plus -0.5ex minus -.2ex}{0.5ex plus
.2ex}{\normalsize\bf\itshape\raggedright}} .2ex}{\normalsize\bf\itshape\raggedright}}
\def\paragraph{\@startsection{paragraph}{4}{\z@}{1.5ex plus \setcounter{secnumdepth}{4}
\setcounter{tocdepth}{4}
\newcounter{subsubsubsection}[subsubsection]
\renewcommand\thesubsubsubsection{\thesubsubsection.\arabic{subsubsubsection}}
\def\l@subsubsubsection{\@dottedtocline{4}{7.0em}{4.1em}}
\def\toclevel@subsubsubsection{4}
\def\l@paragraph{\@dottedtocline{5}{10.0em}{5.0em}}
\def\toclevel@paragraph{5}
\def\l@subparagraph{\@dottedtocline{6}{12.0em}{6.0em}}
\def\toclevel@subparagraph{6}
\def\subsubsubsection{\@ifstar\subsubsubsection@star\subsubsubsection@nostar}
\def\subsubsubsection@nostar{\@dblarg\subsubsubsection@@nostar}
\def\subsubsubsection@@nostar[#1]#2{%
\par
\addpenalty\@secpenalty
\addvspace{1.2ex plus 0.4ex minus 0.2ex}%
\refstepcounter{subsubsubsection}%
\addcontentsline{toc}{subsubsubsection}{\protect\numberline{\thesubsubsubsection}#1}%
{\normalsize\bf\itshape\raggedright \thesubsubsubsection\quad #2\par}%
\nobreak\vspace{0.4ex plus 0.2ex}%
\@afterheading}
\def\subsubsubsection@star#1{%
\par
\addpenalty\@secpenalty
\addvspace{1.2ex plus 0.4ex minus 0.2ex}%
{\normalsize\bf\itshape\raggedright #1\par}%
\nobreak\vspace{0.4ex plus 0.2ex}%
\@afterheading}
\def\paragraph{\@startsection{paragraph}{5}{\z@}{1.5ex plus
0.5ex minus .2ex}{-1em}{\normalsize\bf}} 0.5ex minus .2ex}{-1em}{\normalsize\bf}}
\def\subparagraph{\@startsection{subparagraph}{5}{\z@}{1.5ex plus \def\subparagraph{\@startsection{subparagraph}{6}{\z@}{1.5ex plus
0.5ex minus .2ex}{-1em}{\normalsize\it}} 0.5ex minus .2ex}{-1em}{\normalsize\it}}
\def\subsubsubsection{\vskip
5pt{\noindent\normalsize\raggedright}}
% Footnotes % Footnotes
...@@ -216,5 +242,3 @@ plus -0.5ex minus -.2ex}{0.5ex plus ...@@ -216,5 +242,3 @@ plus -0.5ex minus -.2ex}{0.5ex plus
\def\bottomtitlebar{\vskip .29in\vskip-\parskip\hrule height1pt\vskip \def\bottomtitlebar{\vskip .29in\vskip-\parskip\hrule height1pt\vskip
.09in} % .09in} %
%Reduced second vskip to compensate for adding the strut in \@author %Reduced second vskip to compensate for adding the strut in \@author
...@@ -53,16 +53,37 @@ $\cdots \cdots$ ...@@ -53,16 +53,37 @@ $\cdots \cdots$
\label{fig:agent-planning-sft} \label{fig:agent-planning-sft}
\end{figure*} \end{figure*}
One commonly used approach to improve the planning capability of LLM-based agents is SFT \citep{chen-etal:fireact,zhang-etal:agentohana}. As illustrated in Figure \ref{fig:agent-planning-sft}, this process typically consists of three stages: 1) preparing environments and planning tasks; 2) synthesizing expert-level trajectories, which consist of sequences of action-observation pairs, on these tasks; and 3) instruction-tuning LLMs using the synthesized trajectory data. Specifically, expert trajectories can be generated by leveraging state-of-the-art LLMs as agent policies and selecting high-quality trajectories based on predefined reward signals or task success criteria. Through this process, LLMs can acquire planning behaviors from demonstrations and improve their ability to generate structured execution plans. Based on this procedure, we can construct SFT data for enhancing the planning capability of LLM-based agents. Specifically, each training instance consists of an input-output pair: % 一段话来去引出强化学习对planning能力的关注,然后使用引出具体下面的内容
\subsubsubsection{Supervised Fine-Tuning}
Before applying RL, SFT is commonly adopted to provide a cold start for LLM-based agents by teaching them basic planning and interaction behaviors \citep{chen-etal:fireact,zhang-etal:agentohana,zeng-etal:agenttuning}. As illustrated in Figure \ref{fig:agent-planning-sft}, constructing SFT data for agent planning typically involves three stages: (1) preparing diverse environments and planning tasks; (2) synthesizing expert-level trajectories, consisting of action-observation sequences, through agent-environment interactions; and (3) selecting high-quality trajectories based on predefined evaluation criteria. Specifically, expert trajectories can be generated by leveraging state-of-the-art LLMs as expert agents and retaining reliable demonstrations according to reward signals or task success criteria. Through this process, LLMs can learn planning behaviors from demonstrations and improve their ability to generate structured execution plans.
\begin{eqnarray} \begin{eqnarray}
\mathbf{x} &=& \{e,q\} \\ \mathbf{x} &=& [e,q] \\
\mathbf{y} &=& \{p,u_1,o_1,\cdots,u_T,o_T\} \mathbf{y} &=& [p,u_1,o_1,\cdots,u_T,o_T]
\end{eqnarray} \end{eqnarray}
where $e$ denotes the environment information, including available tools and external resources, and $q$ represents the task description provided by users. $p$ denotes the generated plan that decomposes the task into intermediate steps. $u_t$ and $o_t$ represent the agent's behavior and the corresponding observation from the environment at the $t$-th interaction step, respectively. Specifically, $u_t$ denotes the executable behavior performed by the agent, such as tool invocation, while $o_t$ denotes the feedback returned by the environment after executing $u_t$. Therefore, $\{u_1,o_1,\cdots,u_T,o_T\}$ forms a sequential interaction trajectory between the agent and the environment. where $e$ denotes the environment information, including available tools and external resources, and $q$ represents the task description provided by users. $p$ denotes the generated plan that decomposes the task into intermediate steps. $u_t$ and $o_t$ represent the agent's behavior and the corresponding observation from the environment at the $t$-th interaction step, respectively. Specifically, $u_t$ denotes the executable behavior performed by the agent, such as tool invocation, while $o_t$ denotes the feedback returned by the environment after executing $u_t$. Therefore, $\{u_1,o_1,\cdots,u_T,o_T\}$ forms a sequential interaction trajectory between the agent and the environment. It is worth noting that although $u_t$ and $a_t$ in Section \ref{sec:preliminary} both represent ``action'', they are defined at different levels of abstraction. Specifically, $u_t$ represents a high-level interaction step in agent planning, such as calling a specific tool, whereas $a_t$ represents the low-level token-level action in the standard LLM-based RL formulation, corresponding to the generation of an individual token by the language model.
Similar to LLM \citep{longpre-etal:flan,zhou-etal:lima}, the scale and quality of trajectory data significantly influence the effectiveness of agent training. High-quality trajectories provide explicit supervision signals for agents to learn how to decompose complex tasks, select appropriate tools, and interact with dynamic environments. Therefore, recent studies have explored constructing trajectory-based instruction tuning datasets to enhance the planning capabilities of LLM-based agents.
A straightforward approach is to utilize generated trajectories for instruction tuning. However, trajectories collected from LLM-based agents may contain unsuccessful plans, ineffective tool usage, or execution failures, which can introduce noisy supervision. Therefore, existing methods typically employ filtering strategies based on task success criteria or heuristic rules to select high-quality demonstrations. For example, AgentTuning constructs AgentInstruct by collecting interaction trajectories and filtering unsuccessful samples, while FireAct investigates the impact of diverse trajectory data from multiple tasks and reasoning paradigms for improving agent fine-tuning \citep{zeng-etal:agenttuning,chen-etal:fireact}.
Beyond selecting high-quality trajectories, another challenge lies in effectively representing the complex interaction processes within trajectories. Unlike conventional instruction tuning, agent trajectories contain not only final responses but also intermediate planning processes and multi-step interactions with environments. Therefore, how to organize and annotate different components of trajectories becomes crucial for enabling agents to acquire planning and execution capabilities. To address this challenge, AgentLUMOS introduces a modular training framework that explicitly separates high-level planning and low-level execution \citep{yin-etal:agentlumos}. Specifically, its planning module learns to generate high-level subgoals, while its grounding module learns to translate these subgoals into executable actions through external tools. Such a modular design enables agents to acquire different capabilities from structured trajectory annotations.
Furthermore, even with effective trajectory selection and representation, scaling the generation of diverse and high-quality trajectory data remains challenging. Existing approaches often rely on manually designed environments and planning tasks, which limits the diversity and coverage of collected trajectories. Designing new environments requires substantial human expertise, while constructing tasks with appropriate difficulty levels remains difficult. To address this limitation, AGENTGEN explores automatically generating diverse environments and planning tasks with LLMs, enabling large-scale synthesis of trajectory data with varying task complexity for agent training \citep{hu-etal:agentgen}.
\subsubsubsection{Reinforcement Learning}
\subsubsubsection{}
\subsubsubsection{}
It is worth noting that although $u_t$ and $a_t$ in Section \ref{sec:preliminary} both represent ``action'', they are defined at different levels of abstraction. Specifically, $u_t$ represents a high-level interaction step in agent planning, such as calling a specific tool, whereas $a_t$ represents the low-level token-level action in the standard LLM-based RL formulation, corresponding to the generation of an individual token by the language model.
% 先讲述使用SFT来去提升、然后再讲述与环境进行交互,之后再讲述搜索(这一部分可以重点讲述一下如何构建reward function) % 先讲述使用SFT来去提升、然后再讲述与环境进行交互,之后再讲述搜索(这一部分可以重点讲述一下如何构建reward function)
......
Markdown 格式
0%
您添加了 0 到此讨论。请谨慎行事。
请先完成此评论的编辑!
注册 或者 后发表评论