David Silver, Aja Huang, Chris~J Maddison, Arthur Guez, Laurent Sifre, George Van Den~Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et~al.
\newblock Mastering the game of go with deep neural networks and tree search.
@@ -302,7 +302,9 @@ Scaling environments also introduces substantial heterogeneity. Different enviro
\subsection{Learning from Agentic Experience}
\label{sec:learning_from_agentic_experience}
% AI 正在从依赖人类数据的"人类数据时代",转向主要依靠智能体自身交互经验的"经验时代". As mentioned in Section~\ref{sec:building-agent-capabilities}, 构建人类文本、示范和偏好数据,虽然仍然有用,但很难单独支撑贴近人类或者超越人类的智能。这其中主要的原因一方面是,人类的智能无法使用大量的样例数据来进行枚举的。另外一个问题,
AI systems are gradually moving from an era dominated by human-generated data toward an era in which agents increasingly learn from their own experience \citep{sutton-etal:welcome}. As discussed in Section~\ref{sec:building-agent-capabilities}, supervised data remains useful for teaching agents basic behaviors, such as planning, tool invocation, and response formatting. However, human demonstrations alone are unlikely to cover the full range of situations that an agent may encounter in complex environments. The space of possible tasks, observations, tool responses, and execution failures is too large to be exhaustively annotated in advance.