Commit 9dc12e9a by wangchenglong

update.

parent 1852b590
...@@ -6,4 +6,4 @@ ...@@ -6,4 +6,4 @@
*.out *.out
*.synctex.gz *.synctex.gz
*.toc *.toc
*.png # *.png
\begin{thebibliography}{147} \begin{thebibliography}{152}
\providecommand{\natexlab}[1]{#1} \providecommand{\natexlab}[1]{#1}
\providecommand{\url}[1]{\texttt{#1}} \providecommand{\url}[1]{\texttt{#1}}
\expandafter\ifx\csname urlstyle\endcsname\relax \expandafter\ifx\csname urlstyle\endcsname\relax
...@@ -153,6 +153,11 @@ Jiazhan Feng, Shijue Huang, Xingwei Qu, Ge~Zhang, Yujia Qin, Baoquan Zhong, Chen ...@@ -153,6 +153,11 @@ Jiazhan Feng, Shijue Huang, Xingwei Qu, Ge~Zhang, Yujia Qin, Baoquan Zhong, Chen
\newblock Retool: Reinforcement learning for strategic tool use in llms. \newblock Retool: Reinforcement learning for strategic tool use in llms.
\newblock \emph{arXiv preprint arXiv:2504.11536}, 2025. \newblock \emph{arXiv preprint arXiv:2504.11536}, 2025.
\bibitem[Fu et~al.(2025)Fu, He, Wang, Hong, Gongque, Zeng, Wang, Wang, Cai, and Xu]{fu-etal:agentrefine}
Dayuan Fu, Keqing He, Yejie Wang, Wentao Hong, Zhuoma Gongque, Weihao Zeng, Wei Wang, Jingang Wang, Xunliang Cai, and Weiran Xu.
\newblock Agentrefine: Enhancing agent generalization through refinement tuning.
\newblock In \emph{International Conference on Learning Representations}, volume 2025, pp.\ 65185--65204, 2025.
\bibitem[Fu et~al.(2026)Fu, Li, Ai, Wu, Wu, Zhao, Wang, He, and Wang]{fu-etal:self} \bibitem[Fu et~al.(2026)Fu, Li, Ai, Wu, Wu, Zhao, Wang, He, and Wang]{fu-etal:self}
Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Haoyu Wu, Minghui Wu, Chenxu Zhao, Ante Wang, Guannan He, and Changwei Wang. Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Haoyu Wu, Minghui Wu, Chenxu Zhao, Ante Wang, Guannan He, and Changwei Wang.
\newblock Self-play meets skill evolution: Self-evolving search agents that pose, solve, and remember. \newblock Self-play meets skill evolution: Self-evolving search agents that pose, solve, and remember.
...@@ -282,6 +287,11 @@ Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. ...@@ -282,6 +287,11 @@ Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu.
\newblock \doi{10.18653/v1/2023.emnlp-main.153}. \newblock \doi{10.18653/v1/2023.emnlp-main.153}.
\newblock URL \url{https://aclanthology.org/2023.emnlp-main.153}. \newblock URL \url{https://aclanthology.org/2023.emnlp-main.153}.
\bibitem[Liu et~al.(2024)Liu, Yao, Zhang, Liu, Yang, RN, Lan, Zhu, Tan, Kokane, et~al.]{liu-etal:pract}
Zhiwei Liu, Weiran Yao, Jianguo Zhang, Zuxin Liu, Liangwei Yang, Rithesh RN, Tian Lan, Ming Zhu, Juntao Tan, Shirley Kokane, et~al.
\newblock Pract: Optimizing principled reasoning and acting of llm agent.
\newblock In \emph{Proceedings of the 28th Conference on Computational Natural Language Learning}, pp.\ 442--446, 2024.
\bibitem[Longpre et~al.(2023)Longpre, Hou, Vu, Webson, Chung, Tay, Zhou, Le, Zoph, Wei, et~al.]{longpre-etal:flan} \bibitem[Longpre et~al.(2023)Longpre, Hou, Vu, Webson, Chung, Tay, Zhou, Le, Zoph, Wei, et~al.]{longpre-etal:flan}
Shayne Longpre, Le~Hou, Tu~Vu, Albert Webson, Hyung~Won Chung, Yi~Tay, Denny Zhou, Quoc~V Le, Barret Zoph, Jason Wei, et~al. Shayne Longpre, Le~Hou, Tu~Vu, Albert Webson, Hyung~Won Chung, Yi~Tay, Denny Zhou, Quoc~V Le, Barret Zoph, Jason Wei, et~al.
\newblock The flan collection: Designing data and methods for effective instruction tuning. \newblock The flan collection: Designing data and methods for effective instruction tuning.
...@@ -471,6 +481,11 @@ Xiaoshuai Song, Haofei Chang, Guanting Dong, Yutao Zhu, Zhicheng Dou, and Ji-Ron ...@@ -471,6 +481,11 @@ Xiaoshuai Song, Haofei Chang, Guanting Dong, Yutao Zhu, Zhicheng Dou, and Ji-Ron
\newblock Envscaler: Scaling tool-interactive environments for llm agent via programmatic synthesis. \newblock Envscaler: Scaling tool-interactive environments for llm agent via programmatic synthesis.
\newblock \emph{arXiv preprint arXiv:2601.05808}, 2026. \newblock \emph{arXiv preprint arXiv:2601.05808}, 2026.
\bibitem[Song et~al.(2024)Song, Yin, Yue, Huang, Li, and Lin]{song-etal:trial}
Yifan Song, Da~Yin, Xiang Yue, Jie Huang, Sujian Li, and Bill~Yuchen Lin.
\newblock Trial and error: Exploration-based trajectory optimization of llm agents.
\newblock In \emph{Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, pp.\ 7584--7600, 2024.
\bibitem[Sullivan et~al.(2025)Sullivan, Hartmann, and Koller]{sullivan-etal:randomworld} \bibitem[Sullivan et~al.(2025)Sullivan, Hartmann, and Koller]{sullivan-etal:randomworld}
Michael Sullivan, Mareike Hartmann, and Alexander Koller. Michael Sullivan, Mareike Hartmann, and Alexander Koller.
\newblock Procedural environment generation for tool-use agents. \newblock Procedural environment generation for tool-use agents.
...@@ -581,6 +596,11 @@ Fusheng Wang, Jianhao Yan, Fandong Meng, and Jie Zhou. ...@@ -581,6 +596,11 @@ Fusheng Wang, Jianhao Yan, Fandong Meng, and Jie Zhou.
\newblock \doi{10.18653/v1/2021.acl-long.504}. \newblock \doi{10.18653/v1/2021.acl-long.504}.
\newblock URL \url{https://aclanthology.org/2021.acl-long.504}. \newblock URL \url{https://aclanthology.org/2021.acl-long.504}.
\bibitem[Wang et~al.(2025{\natexlab{b}})Wang, Wang, Leong, and Li]{wang-etal:steca}
Hanlin Wang, Jian Wang, Chak~Tou Leong, and Wenjie Li.
\newblock Steca: Step-level trajectory calibration for llm agent learning.
\newblock In \emph{Findings of the Association for Computational Linguistics: ACL 2025}, pp.\ 11597--11614, 2025{\natexlab{b}}.
\bibitem[Wang et~al.(2026{\natexlab{d}})Wang, Yan, Wang, Tian, Mishra, Xu, Gandhi, Xu, and Cheong]{wang-etal:reinforcement} \bibitem[Wang et~al.(2026{\natexlab{d}})Wang, Yan, Wang, Tian, Mishra, Xu, Gandhi, Xu, and Cheong]{wang-etal:reinforcement}
Jiongxiao Wang, Qiaojing Yan, Yawei Wang, Yijun Tian, Soumya~Smruti Mishra, Zhichao Xu, Megha Gandhi, Panpan Xu, and Lin~Lee Cheong. Jiongxiao Wang, Qiaojing Yan, Yawei Wang, Yijun Tian, Soumya~Smruti Mishra, Zhichao Xu, Megha Gandhi, Panpan Xu, and Lin~Lee Cheong.
\newblock Reinforcement learning for self-improving agent with skill library. \newblock Reinforcement learning for self-improving agent with skill library.
...@@ -609,10 +629,10 @@ Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi C ...@@ -609,10 +629,10 @@ Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi C
\newblock In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), \emph{Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023}, 2023{\natexlab{d}}. \newblock In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), \emph{Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023}, 2023{\natexlab{d}}.
\newblock URL \url{http://papers.nips.cc/paper\_files/paper/2023/hash/ec6413875e4ab08d7bc4d8e225263398-Abstract-Datasets\_and\_Benchmarks.html}. \newblock URL \url{http://papers.nips.cc/paper\_files/paper/2023/hash/ec6413875e4ab08d7bc4d8e225263398-Abstract-Datasets\_and\_Benchmarks.html}.
\bibitem[Wang et~al.(2025{\natexlab{b}})Wang, Takanobu, Liang, Mao, Hu, McAuley, and Wu]{wang-etal:mem} \bibitem[Wang et~al.(2025{\natexlab{c}})Wang, Takanobu, Liang, Mao, Hu, McAuley, and Wu]{wang-etal:mem}
Yu~Wang, Ryuichi Takanobu, Zhiqi Liang, Yuzhen Mao, Yuanzhe Hu, Julian McAuley, and Xiaojian Wu. Yu~Wang, Ryuichi Takanobu, Zhiqi Liang, Yuzhen Mao, Yuanzhe Hu, Julian McAuley, and Xiaojian Wu.
\newblock Mem-$\{$$\backslash$alpha$\}$: Learning memory construction via reinforcement learning. \newblock Mem-$\{$$\backslash$alpha$\}$: Learning memory construction via reinforcement learning.
\newblock \emph{arXiv preprint arXiv:2509.25911}, 2025{\natexlab{b}}. \newblock \emph{arXiv preprint arXiv:2509.25911}, 2025{\natexlab{c}}.
\bibitem[Wang et~al.(2026{\natexlab{e}})Wang, Xu, Liu, Wang, Han, Yao, Yao, and He]{wang-etal:awm} \bibitem[Wang et~al.(2026{\natexlab{e}})Wang, Xu, Liu, Wang, Han, Yao, Yao, and He]{wang-etal:awm}
Zhaoyang Wang, Canwen Xu, Boyi Liu, Yite Wang, Siwei Han, Zhewei Yao, Huaxiu Yao, and Yuxiong He. Zhaoyang Wang, Canwen Xu, Boyi Liu, Yite Wang, Siwei Han, Zhewei Yao, Huaxiu Yao, and Yuxiong He.
...@@ -732,6 +752,11 @@ Tianyu Yu, Haoye Zhang, Yuan Yao, Yunkai Dang, Da~Chen, Xiaoman Lu, Ganqu Cui, T ...@@ -732,6 +752,11 @@ Tianyu Yu, Haoye Zhang, Yuan Yao, Yunkai Dang, Da~Chen, Xiaoman Lu, Ganqu Cui, T
\newblock Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness. \newblock Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.
\newblock \emph{arXiv preprint arXiv:2405.17220}, 2024{\natexlab{b}}. \newblock \emph{arXiv preprint arXiv:2405.17220}, 2024{\natexlab{b}}.
\bibitem[Yuan et~al.(2025)Yuan, Chen, Xi, Ye, Du, and Chen]{yuan-etal:agent}
Siyu Yuan, Zehui Chen, Zhiheng Xi, Junjie Ye, Zhengyin Du, and Jiecao Chen.
\newblock Agent-r: Training language model agents to reflect via iterative self-training.
\newblock \emph{arXiv preprint arXiv:2501.11425}, 2025.
\bibitem[Yuan et~al.(2024)Yuan, Pang, Cho, Li, Sukhbaatar, Xu, and Weston]{yuan-etal:2024selfrewarding} \bibitem[Yuan et~al.(2024)Yuan, Pang, Cho, Li, Sukhbaatar, Xu, and Weston]{yuan-etal:2024selfrewarding}
Weizhe Yuan, Richard~Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. Weizhe Yuan, Richard~Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston.
\newblock Self-rewarding language models. \newblock Self-rewarding language models.
......
...@@ -2,8 +2,45 @@ ...@@ -2,8 +2,45 @@
@inproceedings{song-etal:trial,
title={Trial and error: Exploration-based trajectory optimization of LLM agents},
author={Song, Yifan and Yin, Da and Yue, Xiang and Huang, Jie and Li, Sujian and Lin, Bill Yuchen},
booktitle={Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
pages={7584--7600},
year={2024}
}
@inproceedings{fu-etal:agentrefine,
title={Agentrefine: Enhancing agent generalization through refinement tuning},
author={Fu, Dayuan and He, Keqing and Wang, Yejie and Hong, Wentao and Gongque, Zhuoma and Zeng, Weihao and Wang, Wei and Wang, Jingang and Cai, Xunliang and Xu, Weiran},
booktitle={International Conference on Learning Representations},
volume={2025},
pages={65185--65204},
year={2025}
}
@article{yuan-etal:agent,
title={Agent-r: Training language model agents to reflect via iterative self-training},
author={Yuan, Siyu and Chen, Zehui and Xi, Zhiheng and Ye, Junjie and Du, Zhengyin and Chen, Jiecao},
journal={arXiv preprint arXiv:2501.11425},
year={2025}
}
@inproceedings{wang-etal:steca,
title={Steca: Step-level trajectory calibration for llm agent learning},
author={Wang, Hanlin and Wang, Jian and Leong, Chak Tou and Li, Wenjie},
booktitle={Findings of the Association for Computational Linguistics: ACL 2025},
pages={11597--11614},
year={2025}
}
@inproceedings{liu-etal:pract,
title={Pract: Optimizing principled reasoning and acting of llm agent},
author={Liu, Zhiwei and Yao, Weiran and Zhang, Jianguo and Liu, Zuxin and Yang, Liangwei and RN, Rithesh and Lan, Tian and Zhu, Ming and Tan, Juntao and Kokane, Shirley and others},
booktitle={Proceedings of the 28th Conference on Computational Natural Language Learning},
pages={442--446},
year={2024}
}
@article{shinn-etal:reflexion, @article{shinn-etal:reflexion,
title={Reflexion: Language agents with verbal reinforcement learning}, title={Reflexion: Language agents with verbal reinforcement learning},
......
No preview for this file type
...@@ -317,6 +317,7 @@ The three approaches mentioned above can be implemented through various techniqu ...@@ -317,6 +317,7 @@ The three approaches mentioned above can be implemented through various techniqu
\subsubsection{Memory Management} \subsubsection{Memory Management}
\label{sec:memory-management}
Memory management is a direct way for agents to learn from agentic experience. During interaction with an environment, an agent may observe useful information. If this information is discarded after the current task, the agent must solve similar problems from scratch in the future. Therefore, the goal of agent memory is to store useful information from previous interactions and retrieve it when needed. As illustrated in Figure~\ref{fig:memory-and-retrieve}, a typical memory system usually consists of three main stages \citep{chhikara-etal:mem0}: Memory management is a direct way for agents to learn from agentic experience. During interaction with an environment, an agent may observe useful information. If this information is discarded after the current task, the agent must solve similar problems from scratch in the future. Therefore, the goal of agent memory is to store useful information from previous interactions and retrieve it when needed. As illustrated in Figure~\ref{fig:memory-and-retrieve}, a typical memory system usually consists of three main stages \citep{chhikara-etal:mem0}:
\begin{itemize} \begin{itemize}
...@@ -369,6 +370,7 @@ Beyond optimizing memory update operations, another important question is how to ...@@ -369,6 +370,7 @@ Beyond optimizing memory update operations, another important question is how to
\subsubsection{Skill Optimization} \subsubsection{Skill Optimization}
\label{sec:skill-optimization}
Skill represents a higher-level abstraction of agentic experience. Different from memory, which focuses on storing useful information extracted from previous interactions, skill aims to summarize recurring patterns across multiple experiences and transform them into reusable capabilities. Specifically, a skill can be viewed as a structured instruction package that describes when the skill should be invoked, what workflow it should follow, which tools are required, and how the final outcome should be verified. For example, a reasoning-oriented skill may be represented as follows: Skill represents a higher-level abstraction of agentic experience. Different from memory, which focuses on storing useful information extracted from previous interactions, skill aims to summarize recurring patterns across multiple experiences and transform them into reusable capabilities. Specifically, a skill can be viewed as a structured instruction package that describes when the skill should be invoked, what workflow it should follow, which tools are required, and how the final outcome should be verified. For example, a reasoning-oriented skill may be represented as follows:
...@@ -544,28 +546,10 @@ Verify the solution: ...@@ -544,28 +546,10 @@ Verify the solution:
$x=5$. $x=5$.
\end{tcolorbox} \end{tcolorbox}
It is worth noting that we can obtain feedback from either internal evaluation or external evaluators. For example, we can use rule-based verifiers, LLM-based evaluators, or binary reward signals to provide feedback to the agent. Regardless of the feedback source or format, we need to convert it into a language-based representation, enabling the agent to understand errors and refine future behaviors. Here, inspired by the step-level feedback discussed in Section~\ref{sec:step-by-step-verification}, we can further consider whether providing fine-grained feedback at each decision step can help agents generate better-refined trajectories. Compared with outcome-level feedback, step-level feedback provides more precise guidance about where and why the agent makes mistakes \citep{liu-etal:pract,wang-etal:steca}. For example, when an agent fails to solve a mathematical reasoning problem, a binary reward only indicates that the final answer is incorrect, but it does not reveal which reasoning step contains the error or how the agent should revise its solution. However, obtaining step-level feedback is often expensive and challenging, as it requires evaluating intermediate decisions throughout the trajectory. To address this challenge, recent studies explore using methods such as MCTS to automatically construct step-level feedback signals \citep{yuan-etal:agent}.
Training-free: \\ The above approaches mainly rely on prompting to enable trajectory refinement. Such approaches still face several limitations. First, their performance is bounded by the intrinsic refinement ability of the LLM. A pre-trained LLM may not naturally know how to interpret feedback and effectively revise its trajectory without additional optimization. Second, prompt-based refinement mainly improves the current trajectory at inference time, which limits its ability to accumulate and transfer experience across tasks. As discussed in Sections~\ref{sec:memory-management} and~\ref{sec:skill-optimization}, failed trajectories contain valuable experiences about agent weaknesses and can provide useful learning signals for improving future behaviors. However, prompting-based approaches only leverage these experiences within the current interaction context. To address these limitations, recent studies explore training agents to acquire trajectory refinement abilities \citep{fu-etal:agentrefine}, either by updating model parameters or optimizing refinement behaviors based on collected experiences \citep{song-etal:trial}. Interested readers can refer to these works for more information.
Reflexion: Language Agents with Verbal Reinforcement Learning \\
PRACT: Optimizing Principled Reasoning and Acting of LLM Agent \\
Training-based: \\
Trial and Error: Exploration-Based Trajectory Optimization of LLM Agents (DPO training) \\
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training (Self-Training) \\
STeCa: Step-level Trajectory Calibration for LLM Agent Learning (step-level Trajectories Calibration) \\
AgentRefine: Enhancing Agent Generalization through Refinement Tuning (use refinement tuning to improve generalization) \\
Feedback is very important. How to obtain feedback from convention?
OpenClaw-RL: Train Any Agent Simply by Talking
Although trajectory refinement does not explicitly update model parameters like traditional RL, it shares the fundamental principle of RL: improving agent behaviors based on feedback collected from interactions. From the perspective of in-context learning, the feedback obtained during trajectory refinement can be viewed as a temporary reward signal that modifies the agent's subsequent decision-making process. Instead of updating policy parameters, the agent updates its context with reflections, which potentially modifies the policy used for future actions within the current interaction. Therefore, trajectory refinement can be regarded as an in-context form of policy improvement.
\ No newline at end of file
Markdown 格式
0%
您添加了 0 到此讨论。请谨慎行事。
请先完成此评论的编辑!
注册 或者 后发表评论