Reflexion:让 Agent 用「言语」做强化学习一、为什么需要 Reflexion? 传统强化学习的修复路径是: 失败 → 算梯度 → 更新权重 → 上千次试验 → 收敛但生产...admin2026-06-11服务器技术 阅读(35)评论(0)赞(0)