欢迎光临
我们一直在努力

Spring AI多Agent架构三大生产级踩坑:从Demo到生产的避坑指南

一、背景介绍

2026年以来,Spring AI + Multi-Agent 是Java圈最热的技术方向之一。不少团队开始在真实项目里落地多智能体架构,用来做智能客服、工单分诊、代码审查助手等场景。

但落地过程中,踩坑的团队比跑通的多。Demo跑通只需要一天,但生产级稳定运行可能需要一个月的踩坑积累。

有开发团队在上线Spring AI多Agent系统后,排查了整整两天的异常,最终定位到三个高频坑。这三个坑在官方文档里几乎没有人系统总结,但每一个都可能让项目组花大量时间排查,甚至造成线上事故。

本文将带你深入剖析这三个生产级大坑,从现象到根因,再到完整的解决方案,附带可直接复制的生产级代码。


二、核心技术讲解

2.1 多Agent架构的协作模式

在Spring AI多Agent架构中,通常采用"调度Agent + 多个Worker Agent"的协作模式:

用户请求 → Orchestrator Agent(调度)

┌─────────┼─────────┐
↓ ↓ ↓
TechAgent SalesAgent SupportAgent
(技术) (销售) (客服)

每个Agent都有自己的专业领域和Prompt设定,调度Agent负责将用户请求分发给最合适的Worker Agent处理。

2.2 Spring AI ChatMemory 机制

Spring AI 通过 ChatMemory 组件管理对话上下文,核心依赖 conversationId 来区分不同的对话会话。

// ChatMemory 核心接口
public interface ChatMemory {
// 添加消息
void add(String conversationId, ChatMessage message);

// 获取历史消息
List<ChatMessage> get(String conversationId);

// 清除会话
void clear(String conversationId);
}

问题就出在这个看似简单的机制上…


三、坑一:ConversationId串话——多Agent上下文互相污染

3.1 现象描述

这是最多团队第一个踩到的坑。

生产环境现象:

  • 用户问技术问题,得到的却是销售话术
  • Agent回复内容"乱七八糟",前后逻辑不一致
  • 不同用户的对话历史互相窜台

典型错误日志:

用户A:"Java内存溢出怎么排查?"
系统返回:"您好,您咨询的资费问题是这样的…"

3.2 根因分析

Spring AI的ChatMemory机制,依赖conversationId来区分不同对话会话。如果多个Worker Agent共用了同一个conversationId,问题就出现了:不同Agent的对话历史会混在一起,导致后续输出结果不可预期。

// ❌ 错误写法:所有Agent共用同一个conversationId
public String handleMultiAgent(String userInput, String conversationId) {
// 调度Agent分析用户意图
String intent = orchestratorAgent.handle(userInput, conversationId);

// 技术支撑Agent处理
if ("TECH".equals(intent)) {
return techAgent.handle(userInput, conversationId); // 串话根源
}

// 销售Agent处理
if ("SALES".equals(intent)) {
return salesAgent.handle(userInput, conversationId); // 串话根源
}

return supportAgent.handle(userInput, conversationId); // 串话根源
}

问题本质:

  • 技术Agent处理完后,对话历史被写入 conversationId
  • 销售Agent再用同一个 conversationId 处理时,读到了技术Agent的历史
  • 结果就是销售Agent的回答被技术Agent的上下文"污染"

隐蔽性:本地测试时,单轮对话几乎发现不了,只有多轮对话、多Agent协作场景下才会暴露。上线后用户反馈"回复乱七八糟",项目组才开始排查。

3.3 完整解决方案

方案A:每个Agent使用独立的ConversationId(推荐)

// ✅ 正确写法:每个Agent使用独立的conversationId
@Service
public class MultiAgentService {

private final Agent orchestratorAgent;
private final Agent techAgent;
private final Agent salesAgent;
private final Agent supportAgent;

public String handleMultiAgent(String userInput, String baseConversationId) {
// 给每个Agent的conversationId加唯一后缀
String orchId = baseConversationId + "-orch";
String techId = baseConversationId + "-tech";
String salesId = baseConversationId + "-sales";
String supportId = baseConversationId + "-support";

// 1. 调度Agent分析意图(使用独立ID)
String intent = orchestratorAgent.handle(userInput, orchId);

// 2. 根据意图分发到对应的Worker Agent
switch (intent) {
case "TECH_SUPPORT":
return techAgent.handle(userInput, techId);
case "SALES_CONSULT":
return salesAgent.handle(userInput, salesId);
case "CUSTOMER_SERVICE":
return supportAgent.handle(userInput, supportId);
default:
return supportAgent.handle(userInput, supportId);
}
}
}

方案B:完整的上下文隔离工具类

/**
* Agent会话ID生成器
* 确保每个Agent拥有独立的对话上下文
*/

public final class AgentConversationIdGenerator {

private AgentConversationIdGenerator() {}

// Agent类型枚举
public enum AgentType {
ORCHESTRATOR("orch"),
TECHNICAL("tech"),
SALES("sales"),
CUSTOMER_SERVICE("support"),
CODE_REVIEW("code"),
DOCUMENT("doc");

private final String suffix;

AgentType(String suffix) {
this.suffix = suffix;
}

public String getSuffix() {
return suffix;
}
}

/**
* 为特定Agent生成独立的会话ID
* @param baseId 基础会话ID(通常是用户ID或请求ID)
* @param agentType Agent类型
* @return 独立的会话ID
*/

public static String generate(String baseId, AgentType agentType) {
return String.format("%s-%s", baseId, agentType.getSuffix());
}

/**
* 批量生成所有Agent的会话ID
*/

public static Map<AgentType, String> generateAll(String baseId) {
Map<AgentType, String> ids = new EnumMap<>(AgentType.class);
for (AgentType type : AgentType.values()) {
ids.put(type, generate(baseId, type));
}
return ids;
}
}

使用示例

@Service
public class ImprovedMultiAgentService {

public String handleRequest(String userInput, String userId) {
// 批量生成所有Agent的独立会话ID
Map<AgentConversationIdGenerator.AgentType, String> agentIds =
AgentConversationIdGenerator.generateAll(userId);

// 使用独立ID调用各Agent
String orchId = agentIds.get(AgentConversationIdGenerator.AgentType.ORCHESTRATOR);
String techId = agentIds.get(AgentConversationIdGenerator.AgentType.TECHNICAL);

// … 后续逻辑
return processWithIsolatedContext(userInput, orchId, techId);
}
}


四、坑二:JVM内存记忆——服务重启后上下文全部丢失

4.1 现象描述

生产环境现象:

  • 用户正在跟智能客服聊到第三步,刷新页面后"一切归零"
  • 服务重启、容器重新调度、Pod滚动更新后,所有用户的对话上下文全部清空
  • 用户投诉:“我刚才说了三遍订单号,一刷新又让我重新报一遍”

典型用户反馈:

“客服机器人根本记不住我说的话,每次刷新都要重新开始,体验太差了!”

4.2 根因分析

Spring AI默认使用 MessageWindowChatMemory,它将对话历史存在JVM内存中:

// Spring AI 默认实现:内存存储
public class InMemoryChatMemory implements ChatMemory {
private final Map<String, List<ChatMessage>> memory = new ConcurrentHashMap<>();

@Override
public void add(String conversationId, ChatMessage message) {
memory.computeIfAbsent(conversationId, k -> new ArrayList<>())
.add(message);
}
// …
}

问题本质:

  • 开发环境:服务一直运行,内存不会清空,一切正常
  • 生产环境:K8s滚动更新、服务重启、Pod漂移时,所有内存数据丢失
  • 这是一个"开发环境永远发现不了,生产环境必中招"的典型陷阱

隐蔽性:本地调试时服务一直开着,这个问题在开发环境几乎不可能发现。只有上到生产环境、经历第一次重启后,才会暴露。

4.3 完整解决方案

方案:Redis持久化ChatMemory(生产环境标配)

/**
* Redis持久化的ChatMemory实现
* 支持自动过期、会话隔离、性能优化
*/

@Slf4j
public class RedisChatMemory implements ChatMemory {

private final RedisTemplate<String, String> redisTemplate;
private final ObjectMapper objectMapper;

// Redis Key前缀
private static final String KEY_PREFIX = "chat:memory:";
// 默认过期时间:7天
private static final long DEFAULT_EXPIRE_DAYS = 7;
// 每个会话最大消息数
private static final int MAX_MESSAGES_PER_SESSION = 50;

public RedisChatMemory(RedisTemplate<String, String> redisTemplate,
ObjectMapper objectMapper) {
this.redisTemplate = redisTemplate;
this.objectMapper = objectMapper;
// 配置ObjectMapper支持Jackson模块
this.objectMapper.registerModule(new ParameterNamesModule());
this.objectMapper.registerModule(new Jdk8Module());
this.objectMapper.registerModule(new JavaTimeModule());
}

@Override
public void add(String conversationId, ChatMessage message) {
String key = buildKey(conversationId);

try {
// 序列化消息
String messageJson = objectMapper.writeValueAsString(message);

// LPUSH: 新消息添加到列表头部
redisTemplate.opsForList().leftPush(key, messageJson);

// LTRIM: 保留最新的N条消息,防止列表无限增长
redisTemplate.opsForList().trim(key, 0, MAX_MESSAGES_PER_SESSION 1);

// 设置过期时间
redisTemplate.expire(key, DEFAULT_EXPIRE_DAYS, TimeUnit.DAYS);

log.debug("消息已添加到Redis会话: {}", conversationId);
} catch (JsonProcessingException e) {
log.error("消息序列化失败,conversationId: {}", conversationId, e);
throw new RuntimeException("消息序列化失败", e);
}
}

@Override
public List<ChatMessage> get(String conversationId) {
String key = buildKey(conversationId);

// LRANGE: 获取列表所有元素(0到-1表示全部)
List<String> messageJsons = redisTemplate.opsForList().range(key, 0, 1);

if (CollectionUtils.isEmpty(messageJsons)) {
return new ArrayList<>();
}

List<ChatMessage> messages = new ArrayList<>();
for (String messageJson : messageJsons) {
try {
ChatMessage message = objectMapper.readValue(messageJson, ChatMessage.class);
messages.add(message);
} catch (JsonProcessingException e) {
log.error("消息反序列化失败,跳过该消息: {}", messageJson, e);
}
}

// 注意:Redis存储是最新的在前面,需要反转以保持时间顺序
Collections.reverse(messages);

log.debug("从Redis获取会话消息: {}, 共{}条", conversationId, messages.size());
return messages;
}

@Override
public void clear(String conversationId) {
String key = buildKey(conversationId);
redisTemplate.delete(key);
log.info("会话已清除: {}", conversationId);
}

private String buildKey(String conversationId) {
return KEY_PREFIX + conversationId;
}
}

Spring Boot配置类

@Configuration
public class ChatMemoryConfig {

@Bean
@ConditionalOnProperty(name = "chat.memory.type", havingValue = "redis")
public ChatMemory redisChatMemory(RedisTemplate<String, String> redisTemplate,
ObjectMapper objectMapper) {
log.info("使用Redis持久化ChatMemory(生产环境)");
return new RedisChatMemory(redisTemplate, objectMapper);
}

@Bean
@ConditionalOnProperty(name = "chat.memory.type", havingValue = "memory", matchIfMissing = true)
public ChatMemory inMemoryChatMemory() {
log.warn("使用内存ChatMemory(仅开发环境!生产环境请使用redis)");
return new InMemoryChatMemory();
}
}

application.yml配置

chat:
memory:
type: redis # 生产环境必须用redis!开发环境可临时用memory

spring:
redis:
host: yourredishost
port: 6379
password: yourpassword
database: 1 # 建议使用独立的DB存储对话数据
timeout: 3000ms
lettuce:
pool:
max-active: 50
max-idle: 10
min-idle: 5


五、坑三:LLM输出JSON格式不稳定——解析直接抛500

5.1 现象描述

生产环境现象:

  • 接口偶发性返回500错误
  • 错误日志:JsonProcessingException: Unrecognized token
  • 问题不是必现,有的请求正常,有的报错,测试难以复现
  • 高峰期报错概率明显上升

典型错误堆栈:

com.fasterxml.jackson.core.JsonParseException:
Unrecognized token '```json': was expecting (JSON String, Number, Array, Object or token 'null', 'true' or 'false')
at [Source: (String)"```json
{
"intent": "TECH_SUPPORT",
"confidence": 0.92
}
```"; line: 1, column: 8]

5.2 根因分析

多Agent架构里,调度Agent(Orchestrator)通常需要将用户意图解析为结构化JSON,再分发给对应Worker Agent。

但LLM的输出格式并不稳定:

// 开发者期望LLM返回的纯JSON
{
"intent": "TECH_SUPPORT",
"confidence": 0.92
}

// LLM实际返回的可能是Markdown包裹的JSON
```json
{
"intent": "TECH_SUPPORT",
"confidence": 0.92
}

**问题本质**:
– 即使Prompt里明确要求"只输出JSON,不要包裹Markdown代码块"
– LLM仍然有一定概率返回Markdown格式的JSON
– 后端直接用ObjectMapper解析时,会直接抛出JsonProcessingException
– 前端看到的就是500错误

**隐蔽性**:这个错误不是必现的。有的query返回纯JSON,有的返回Markdown包裹的JSON,测试阶段很容易漏掉。

### 5.3 完整解决方案

#### 方案A:JSON解析前清理 + 降级处理

```java
@Slf4j
@Service
public class JsonParsingService {

private final ObjectMapper objectMapper;

// Markdown代码块标记正则
private static final Pattern JSON_CODE_BLOCK_PATTERN =
Pattern.compile("```(json)?\\\\s*|```\\\\s*$", Pattern.MULTILINE);

public JsonParsingService(ObjectMapper objectMapper) {
this.objectMapper = objectMapper;
this.objectMapper.configure(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES, false);
}

/**
* 安全解析LLM返回的JSON
* 支持清理Markdown标记、自动降级
*/
public <T> T safeParse(String rawJson, Class<T> targetClass, T fallbackValue) {
if (StringUtils.isBlank(rawJson)) {
log.warn("JSON内容为空,返回降级值");
return fallbackValue;
}

// 第一步:清理Markdown代码块标记
String cleanedJson = cleanJsonMarkdown(rawJson);

// 第二步:提取真正的JSON部分(处理前置说明文本)
cleanedJson = extractJsonContent(cleanedJson);

try {
T result = objectMapper.readValue(cleanedJson, targetClass);
log.debug("JSON解析成功: {}", targetClass.getSimpleName());
return result;
} catch (JsonProcessingException e) {
log.error("JSON解析失败,原始内容:\\n{}\\n清理后:\\n{}",
rawJson, cleanedJson, e);

// 第三步:尝试更宽松的解析(修复常见格式错误)
try {
String repairedJson = repairCommonJsonIssues(cleanedJson);
T result = objectMapper.readValue(repairedJson, targetClass);
log.info("JSON修复后解析成功");
return result;
} catch (JsonProcessingException e2) {
log.error("JSON修复后仍解析失败,返回降级值", e2);
return fallbackValue;
}
}
}

/**
* 清理Markdown代码块标记
*/
private String cleanJsonMarkdown(String rawJson) {
return JSON_CODE_BLOCK_PATTERN.matcher(rawJson).replaceAll("").trim();
}

/**
* 从文本中提取JSON内容
* 处理LLM可能返回"好的,这是解析结果:{…}"这种情况
*/
private String extractJsonContent(String text) {
int firstBrace = text.indexOf('{');
int firstBracket = text.indexOf('[');

// 找到第一个JSON开始标记
int startIndex = -1;
if (firstBrace >= 0 && firstBracket >= 0) {
startIndex = Math.min(firstBrace, firstBracket);
} else if (firstBrace >= 0) {
startIndex = firstBrace;
} else if (firstBracket >= 0) {
startIndex = firstBracket;
}

if (startIndex < 0) {
return text; // 没找到,原样返回
}

// 找到对应的结束标记
char startChar = text.charAt(startIndex);
char endChar = (startChar == '{') ? '}' : ']';

int depth = 0;
int endIndex = -1;
for (int i = startIndex; i < text.length(); i++) {
char c = text.charAt(i);
if (c == startChar) depth++;
if (c == endChar) depth–;
if (depth == 0) {
endIndex = i;
break;
}
}

if (endIndex > 0) {
return text.substring(startIndex, endIndex + 1);
}

return text.substring(startIndex);
}

/**
* 修复常见的JSON格式错误
*/
private String repairCommonJsonIssues(String json) {
String repaired = json;

// 修复:尾部多余的逗号
repaired = repaired.replaceAll(",\\\\s*([}\\\\]])", "$1");

// 修复:单引号改为双引号
repaired = repaired.replaceAll("'", "\\"");

// 修复:键名缺少引号
repaired = repaired.replaceAll("(\\\\w+)(\\\\s*:)", "\\"$1\\"$2");

return repaired;
}
}

在多Agent服务中使用

@Service
public class OrchestratorAgent {

private final JsonParsingService jsonParsingService;
private final ChatClient chatClient;

public TaskPlan analyzeIntent(String userInput, String conversationId) {
// 调用LLM获取结构化意图
String rawResponse = chatClient.prompt()
.user(buildIntentPrompt(userInput))
.call()
.content();

// 使用安全解析 + 降级策略
TaskPlan fallbackPlan = TaskPlan.createFallback(userInput);

return jsonParsingService.safeParse(
rawResponse,
TaskPlan.class,
fallbackPlan // 解析失败时的降级值
);
}

private String buildIntentPrompt(String userInput) {
return """
请分析用户问题,返回JSON格式的意图分析结果。

用户问题:%s

【重要要求】
1. 只返回纯JSON,不要任何解释说明
2. 不要使用Markdown代码块包裹
3. 严格按照以下格式:
{
"intent": "TECH_SUPPORT|SALES|GENERAL",
"confidence": 0.95,
"suggestedAgent": "tech_agent"
}
""".formatted(userInput);
}
}


六、实际应用场景与踩坑总结

6.1 三大坑的共性特征

坑点开发环境表现生产环境表现发现难度
ConversationId串话 单轮对话正常 多轮对话回复混乱 ⭐⭐⭐⭐⭐ 极难发现
内存记忆丢失 一切正常 服务重启后上下文归零 ⭐⭐⭐⭐⭐ 极难发现
JSON格式不稳定 大部分请求正常 偶发性500错误 ⭐⭐⭐⭐ 难以复现

核心启示:

Demo跑通 ≠ 生产可用。Spring AI多Agent落地最难的地方,就是这些"开发环境永远发现不了"的坑。

6.2 生产环境检查清单

上线前务必检查以下项目:

生产环境检查清单:
✅ 会话隔离: 每个Agent使用独立的conversationId
✅ 持久化: ChatMemory使用Redis而非内存
✅ 降级策略: JSON解析失败有降级逻辑
✅ 超时处理: LLM调用超时有熔断机制
✅ 限流保护: 防止LLM API被打爆
✅ 监控告警: 异常率、响应时间监控
✅ 日志审计: 完整的请求响应日志链

6.3 适合多Agent架构的场景

✅ 推荐场景:

  • 智能客服(多领域专家协作)
  • 工单自动分诊系统
  • 代码审查助手(不同Agent检查不同维度)
  • 智能文档生成(规划→撰写→校对)

❌ 不推荐场景:

  • 简单的问答系统(单Agent足够)
  • 对响应时间要求极高(<1s)的场景
  • 资源受限的边缘设备

七、结尾总结

Spring AI多Agent架构是2026年Java AI开发的核心方向,但从Demo到生产,需要跨过无数坑。本文总结的三大生产级大坑,每一个都有团队踩过,每一个都可能让项目延期或造成线上事故。

核心要点回顾:

  • 会话隔离:每个Agent必须使用独立的conversationId,避免上下文污染
  • 持久化存储:生产环境必须用Redis替代内存存储对话历史
  • 容错降级:JSON解析要做多层清理+降级,不能让LLM的不稳定性直接抛给用户
  • 技术演进的真相:

    新技术的Demo永远美好,但生产级落地需要的是:对细节的极致把控、对异常的充分预案、对边界情况的反复验证。

    作为Java开发者,我们正站在AI与后端融合的黄金时代。Spring AI让我们能快速构建AI应用,但真正的挑战在于——如何让这些应用稳定、可靠地运行在生产环境中。

    希望这篇踩坑指南能帮你少走弯路,让你的多Agent系统顺利从Demo走向生产。


    本文代码仓库:https://github.com/example/spring-ai-multiagent-best-practices

    推荐阅读:

    • Spring AI官方文档
    • Spring AI Alibaba企业级实践
    • AgentScope Java多Agent框架

    赞(0)
    未经允许不得转载:171主机测评 » Spring AI多Agent架构三大生产级踩坑:从Demo到生产的避坑指南
    分享到: 更多 (0)

    评论 抢沙发

    • 昵称 (必填)
    • 邮箱 (必填)
    • 网址