欢迎光临
我们一直在努力

【推理与部署篇03】SGLang深度解析

🚀 【推理与部署篇03】SGLang深度解析:RadixAttention与结构化生成

2026年最新版 | 覆盖 SGLang v0.5.12+ | 400,000+ GPU 生产验证


📑 目录

  • 1. SGLang核心架构与设计理念
  • 2. RadixAttention原理深度解析
  • 3. XGrammar结构化生成
  • 4. Overlap Scheduling调度优化
  • 5. PD分离与HiCache层级缓存
  • 6. 多模态与Embedding统一服务
  • 7. 性能基准测试
  • 8. 生产部署最佳实践
  • 9. 关键参数调优指南
  • 10. SGLang vs vLLM选型决策
  • 面试加分点

1. SGLang核心架构与设计理念

1.1 什么是SGLang?

SGLang(Structured Generation Language)是由伯克利大学 LMSYS 实验室(LLM System)研发的新一代大模型推理引擎,定位不仅是推理加速器,更是一套完整的"大模型程序运行时"。

截至2026年5月,SGLang已在全球超过40万张GPU上部署,被 xAI、AMD、NVIDIA、Intel、LinkedIn、Cursor、Oracle Cloud 等顶级企业采用,GitHub 28K+ Stars,2026年已发布14个版本 [1]。

1.2 三层架构

┌──────────────────────────────────────────────┐
│ 前端DSL层(sglang.xxx) │
│ – @sgl.function 装饰器 + sgl.gen / sgl.select │
│ – 结构化生成流、约束解码、工具调用 │
├──────────────────────────────────────────────┤
│ 运行时层(Runtime) │
│ – RadixAttention 缓存管理 │
│ – Overlap Scheduling 调度引擎 │
│ – XGrammar 结构化输出后端 │
│ – HiCache 三级缓存卸载 │
├──────────────────────────────────────────────┤
│ 内核层(Kernels) │
│ – FlashInfer Attention │
│ – Triton FP8 GEMM / DeepGemm MLA │
│ – CUTLASS MoE Kernels │
│ – PD分离分布式协同 │
└──────────────────────────────────────────────┘

1.3 设计哲学

与 vLLM 的"纯推理引擎"定位不同,SGLang 将编译器概念引入推理系统:

维度传统推理引擎SGLang
定位 HTTP 服务 + 推理加速 程序运行时 + 推理加速
前端 OpenAI API 兼容 DSL 编程模型 + OpenAI API
缓存粒度 Page/Block 级 Token 级基数树
约束解码 外部工具(Outlines) 原生 XGrammar(零开销)
服务整合 需独立部署 LLM + VLM + Embedding 同一进程

2. RadixAttention原理深度解析

2.1 问题:传统前缀缓存的局限性

vLLM 的 Prefix Caching 基于哈希精确匹配:只有前缀完全一致才能复用缓存。在 RAG 或多轮对话中,不同请求的前缀往往部分相似(共享系统提示词但检索文档不同),哈希方案只能丢弃无法精确匹配的部分。

传统方案(vLLM Hash-based):
请求1: [系统提示词] + [文档A] + [问题]
请求2: [系统提示词] + [文档B] + [问题]

请求1缓存: [系统提示词||文档A||问题] → 只能完整匹配
请求2: 系统提示词部分需重新计算 ❌

RadixAttention(SGLang):
请求1缓存: ┌系统提示词┐┌文档A┐┌问题┐
请求2: 重用 ↑ 计算新↑ 重用↑

系统提示词部分自动复用 ✅

2.2 基数树(Radix Tree)结构

RadixAttention 的核心数据结构是基数树,每个节点存储一段连续的 Token 序列 [2]:

Root
├── "System: You are a helpful assistant."
│ ├── "What is the capital of France?" → "Paris"
│ ├── "What is the capital of Japan?" → "Tokyo"
│ └── "Translate to French: Hello" → "Bonjour"
├── "System: You are a coding assistant."
│ ├── "Write a Python function to…"
│ └── "Explain the time complexity of…"
└── "System: You are a data analyst."
└── …

关键特性:

  • 最长公共前缀自动检测:新请求到达时,自动匹配已有树节点
  • Token 级粒度:节点按连续 token 序列组织,比 Page 级更精细
  • 部分复用:即使只有前缀部分匹配,也能复用共享节点
  • LRU 驱逐:缓存满时淘汰最近最少使用节点

2.3 完整实现

import threading
from typing import Dict, List, Optional, Tuple
import heapq

class RadixTreeNode:
"""基数树节点"""
def __init__(self, token_ids: List[int]):
self.token_ids = token_ids # 节点存储的 token 序列
self.children: Dict[int, 'RadixTreeNode'] = {} # 子节点(首 token -> 节点)
self.kv_cache: Optional[Tuple] = None # 此节点对应的 KV Cache
self.last_access_time: int = 0 # 最后访问时间(LRU)
self.access_count: int = 0 # 访问次数
self.is_leaf: bool = True

class RadixAttention:
"""
RadixAttention 缓存管理器
支持最长公共前缀匹配、部分复用、LRU驱逐
"""

def __init__(self, max_cache_size: int = 100000):
self.root = RadixTreeNode([])
self.max_cache_size = max_cache_size
self.current_size = 0
self.lock = threading.Lock()
self.access_counter = 0

def match_prefix(self, token_ids: List[int]) > Tuple[int, RadixTreeNode, List[RadixTreeNode]]:
"""
匹配最长公共前缀
返回:(匹配长度, 匹配到的节点, 路径上的节点列表)
"""

node = self.root
matched_len = 0
path = [node]
i = 0

while i < len(token_ids):
# 查找匹配的子节点
first_token = token_ids[i]
if first_token not in node.children:
break

child = node.children[first_token]
# 检查子节点内的序列是否能匹配
child_tokens = child.token_ids
j = 0
while j < len(child_tokens) and i < len(token_ids):
if child_tokens[j] != token_ids[i]:
break
j += 1
i += 1

if j > 0:
matched_len += j
path.append(child)
node = child

# 如果子节点的 token 没有完全匹配完,需要分裂
if j < len(child_tokens):
self._split_node(child, j)
# 分裂后重新获取节点引用
child = node.children.get(token_ids[matched_len j] if matched_len j < len(token_ids) else None)
if child:
path[1] = child
continue
break

return matched_len, node, path

def _split_node(self, node: RadixTreeNode, split_pos: int):
"""分裂节点:从 split_pos 处切开"""
if split_pos <= 0 or split_pos >= len(node.token_ids):
return

# 新节点:保留后半部分
new_node = RadixTreeNode(node.token_ids[split_pos:])
new_node.kv_cache = None # 分裂后需要重新计算
new_node.children = node.children
new_node.is_leaf = node.is_leaf

# 原节点变为前半部分
node.token_ids = node.token_ids[:split_pos]
node.children = {new_node.token_ids[0]: new_node}
node.is_leaf = False

# 调整缓存大小
self.current_size -= len(new_node.token_ids)

def insert(self, token_ids: List[int], kv_cache: Tuple):
"""插入新的缓存序列"""
with self.lock:
# 先匹配已有前缀
matched_len, node, path = self.match_prefix(token_ids)

# 计算需要新插入的 token
remaining = token_ids[matched_len:]
if not remaining:
# 完全命中已有节点,更新 KV Cache
node.kv_cache = kv_cache
node.last_access_time = self.access_counter
node.access_count += 1
self.access_counter += 1
return

# 逐段创建新节点
current = node
start = 0
while start < len(remaining):
# 尝试找到最大的可合并段
end = start + 1
while end <= len(remaining):
if self._is_mergeable(remaining[start:end]):
end += 1
else:
break
end -= 1

segment = remaining[start:end]
new_node = RadixTreeNode(segment)
new_node.last_access_time = self.access_counter
self.access_counter += 1

current.children[segment[0]] = new_node
current.is_leaf = False
current = new_node
start = end

# 最后的节点保存 KV Cache
current.kv_cache = kv_cache
current.is_leaf = True
self.current_size += len(remaining)

# LRU 驱逐
while self.current_size > self.max_cache_size:
self._evict_lru()

def _is_mergeable(self, tokens: List[int]) > bool:
"""判断一段 token 是否可以合并为单一节点(启发式)"""
return len(tokens) <= 64 # 最大节点长度

def _evict_lru(self):
"""LRU 驱逐:找到最久未访问的叶子节点并移除"""
# BFS 找到最久未访问的叶子
stack = [(self.root, 0)]
oldest_node = None
oldest_time = float('inf')

while stack:
node, depth = stack.pop()
if node.is_leaf and node != self.root:
if node.last_access_time < oldest_time:
oldest_time = node.last_access_time
oldest_node = node
for child in node.children.values():
stack.append((child, depth + 1))

if oldest_node and oldest_node.kv_cache:
self.current_size -= len(oldest_node.token_ids)
oldest_node.kv_cache = None
oldest_node.last_access_time = 0

def get_cache(self, token_ids: List[int]) > Optional[Tuple]:
"""获取匹配的 KV Cache"""
matched_len, node, _ = self.match_prefix(token_ids)
if node and node.kv_cache:
node.last_access_time = self.access_counter
node.access_count += 1
self.access_counter += 1
return node.kv_cache, matched_len
return None, matched_len

def get_cache_hit_rate(self) > float:
"""计算缓存命中率"""
total_nodes = 0
cached_nodes = 0
stack = [self.root]
while stack:
node = stack.pop()
if node != self.root:
total_nodes += 1
if node.kv_cache is not None:
cached_nodes += 1
for child in node.children.values():
stack.append(child)
return cached_nodes / max(total_nodes, 1)

2.4 实测效果

场景缓存命中率首Token延迟降低吞吐提升
单用户 RAG (c=1) 92% 3.5× 2.8×
多轮对话 5轮 83% 28% 29%
Agent工具调用 78% 2.1× 1.8×
纯随机查询 12% 基准 基准

在共享前缀明显的场景(RAG、Agent 系统提示词、多轮对话),RadixAttention 相比 hash-based 方案有显著优势 [3]。

2.5 UnifiedRadixTree

SGLang v0.5.0+ 引入 UnifiedRadixTree,将传统 Radix Tree 与层级缓存结合:

GPU VRAM (最快, 最小)
└── Radix Tree 热节点
↓ LRU 驱逐
CPU RAM (中等, 大)
└── Radix Tree 温节点
↓ LRU 驱逐
SSD (最慢, 最大)
└── 冷节点持久化

允许缓存远超 GPU 显存限制的 KV Cache,对长上下文服务至关重要。


3. XGrammar结构化生成

3.1 为什么需要结构化生成?

Agent 场景的核心需求:模型输出必须严格遵守 JSON Schema、正则表达式或 EBNF 语法。传统方案(先生成后验证):

传统方案:生成文本 → JSON解析 → 验证schema → 失败重试 ❌
→ 30%+ 重试率
XGrammar:约束解码 → 每一步只生成合法token ✅
→ 0 重试率

3.2 XGrammar 工作原理

XGrammar 是 SGLang 默认的结构化输出后端(2024.11 发布),实现零开销约束解码:

用户定义 Schema

XGrammar 编译为上下文无关文法 (CFG)

在采样过程中实时计算合法 token 集合

LM Head 输出概率 → 屏蔽非法 token → 采样

核心特性:

  • 编译时优化:Schema 预编译为 CFG,运行时零开销
  • token 级约束:每一步只采样合法 token
  • 支持格式:JSON Schema、正则表达式、EBNF、structural_tag
  • 多工具注册:一个请求内同时注册多个函数

import sglang as sgl

# === 方式1:@sgl.function 装饰器 DSL ===
@sgl.function
def extract_city_info(s, document):
s += "从以下文档提取城市信息:\\n"
s += document + "\\n"
s += "输出 JSON:\\n"
s += "{\\n"
s += ' "name": ' + sgl.gen("name", max_tokens=8,
regex=r'"[^"\\\\]*(?:\\\\.[^"\\\\]*)*"') + ",\\n"
s += ' "country": ' + sgl.gen("country", max_tokens=8,
regex=r'"[^"\\\\]*(?:\\\\.[^"\\\\]*)*"') + ",\\n"
s += ' "latitude": ' + sgl.gen("latitude", max_tokens=8,
regex=r'-?\\d+\\.\\d+') + ",\\n"
s += ' "population": ' + sgl.gen("population", max_tokens=8,
regex=r'-?\\d+') + "\\n"
s += "}\\n"

# === 方式2:OpenAI API 兼容(XGrammar structural_tag)===
response = client.chat.completions.create(
model="qwen3-32b",
messages=[{"role": "user", "content": "北京的天气怎么样?"}],
response_format={
"type": "structural_tag",
"format": {
"type": "triggered_tags",
"triggers": ["<function="],
"tags": [
{
"begin": "<function=get_current_weather>",
"content": {
"type": "json_schema",
"json_schema": {
"type": "object",
"properties": {
"location": {"type": "string"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["location", "unit"]
}
},
"end": "</function>"
},
{
"begin": "<function=get_forecast>",
"content": {
"type": "json_schema",
"json_schema": {
"type": "object",
"properties": {
"location": {"type": "string"},
"days": {"type": "integer", "minimum": 1, "maximum": 7}
},
"required": ["location", "days"]
}
},
"end": "</function>"
}
]
}
}
)

3.3 性能对比

指标SGLang (XGrammar)vLLM + OutlinesGuidance提升倍数
吞吐量 (req/s) 36 12 8 3.0×
延迟 (ms) 28 83 125 4.5×
内存占用 (GB) 4.2 6.8 7.5 1.8×
准确率 (%) 98.7 95.2 93.8 +3.5%

4. Overlap Scheduling调度优化

4.1 传统调度瓶颈

LLM 推理的典型流程:

CPU: 接收请求 → 预处理(Tokenize) → 调度决策 → …
↓ 等待
GPU: ← 执行推理 →

CPU 预处理和调度决策是同步阻塞的,GPU 在等待期间空闲。

4.2 Overlap Scheduling

SGLang 实现 CPU 和 GPU 的流水线重叠:

传统调度:
GPU: ██████推理██████ ██████推理██████ ██████推理██████
CPU: 准备下一批 准备下一批
↑ GPU 空闲等待 ↑ ↑ GPU 空闲等待 ↑

Overlap Scheduling:
GPU: ██████推理██████ ██████推理██████ ██████推理██████
CPU: 准备下一批 准备下一批 准备下一批
时间 →
↑ 零等待 ↑ ↑ GPU 始终满载 ↑

4.3 Zero-Overhead Batch Scheduler

SGLang v0.4+ 的 Zero-Overhead Batch Scheduler 进一步消除调度开销:

class OverlapScheduler:
"""
CPU-GPU 重叠调度器
在 GPU 执行当前批次时,CPU 预计算下一批
"""

def __init__(self, max_batch_size: int = 256):
self.max_batch_size = max_batch_size
self.pending_queue = [] # 等待调度的请求
self.running_batch = None # 当前 GPU 执行批次
self.next_batch = None # 预计算的下一批

def submit(self, request):
"""提交请求"""
self.pending_queue.append(request)

def schedule_async(self):
"""
异步调度:在 GPU 空闲时预计算下一批
此方法在 CPU 线程中异步执行
"""

if self.next_batch is not None:
return # 已经预计算了下一批

if not self.pending_queue:
return

# 批处理策略:优先相同长度的请求
batch = []
remaining = []
ref_len = None

for req in self.pending_queue:
if ref_len is None:
ref_len = len(req.input_ids)
batch.append(req)
elif abs(len(req.input_ids) ref_len) / ref_len < 0.3:
batch.append(req)
else:
remaining.append(req)

self.pending_queue = remaining

# 预计算:Tokenize、KV Cache 分配、调度决策
self.next_batch = self._prepare_batch(batch)

# 利用 RadixAttention 匹配缓存
for req in self.next_batch:
cache, matched = self.radix_cache.get_cache(req.input_ids)
req.cached_prefix_len = matched

def _prepare_batch(self, batch):
"""预计算批次元数据"""
# 最大序列长度
max_len = max(len(r.input_ids) for r in batch)

processed = []
for req in batch:
processed.append({
"input_ids": req.input_ids,
"cached_prefix_len": 0,
"prefill_len": len(req.input_ids),
"decode_len": 0,
"max_tokens": req.max_tokens,
"sampling_params": req.sampling_params,
"block_table": self._allocate_blocks(max_len + req.max_tokens),
})

return processed

def step(self):
"""
执行一步调度
返回需要 GPU 执行的批次
"""

if self.running_batch is None and self.next_batch is None:
self.schedule_async()
return None

# 交换批次
self.running_batch = self.next_batch
self.next_batch = None

# 异步预计算下一批
self.schedule_async()

return self.running_batch

实测效果:在低并发(c=1)场景下 Overlap Scheduling 可将 GPU 利用率从 65% 提升至 92%+ [4]。


5. PD分离与HiCache层级缓存

5.1 PD分离(Prefill-Decode Disaggregation)

预填充和解码阶段对资源需求完全不同:

维度预填充(Prefill)解码(Decode)
计算密集度 Compute-bound Memory-bound
并行度 高(一次性处理所有 token) 低(逐 token)
KV Cache 生成 写密集 读密集
理想硬件 高算力 GPU 高带宽 GPU

SGLang 支持将 Prefill 和 Decode 分配到不同 GPU:

请求到达

Prefill GPU(高算力:H100/B300)
├── 计算完整 KV Cache
├── 通过 NCCL 传输

Decode GPU(高带宽:H200/B300)
├── 复用已计算 KV Cache
├── 逐 token 生成

返回结果

配置示例:

# Prefill 节点(2 GPU)
python -m sglang.launch_server \\
–model-path Qwen/Qwen3-32B \\
–disaggregation-mode prefill \\
–tp 2 \\
–host 0.0.0.0 –port 30001

# Decode 节点(4 GPU)
python -m sglang.launch_server \\
–model-path Qwen/Qwen3-32B \\
–disaggregation-mode decode \\
–tp 4 \\
–host 0.0.0.0 –port 30002

# 路由节点
python -m sglang.launch_server \\
–model-path Qwen/Qwen3-32B \\
–disaggregation-mode router \\
–prefill-server http://prefill:30001 \\
–decode-server http://decode:30002 \\
–host 0.0.0.0 –port 8000

5.2 HiCache层级KV Cache卸载

HiCache 实现 GPU → CPU → SSD 三级缓存:

class HiCacheManager:
"""
HiCache 层级缓存管理器
支持 GPU VRAM → CPU RAM → SSD 三级卸载
"""

def __init__(self, gpu_cache_size: int, cpu_cache_size: int,
ssd_cache_path: str):
self.gpu_cache = {} # GPU VRAM 缓存
self.cpu_cache = {} # CPU RAM 缓存
self.ssd_cache_path = ssd_cache_path # SSD 持久化
self.gpu_limit = gpu_cache_size
self.cpu_limit = cpu_cache_size

# 访问频率跟踪
self.access_stats = defaultdict(lambda: {"count": 0, "last_access": 0})

def get(self, key: str) > Optional[bytes]:
"""按优先级查找缓存"""
if key in self.gpu_cache:
self._update_stats(key)
return self.gpu_cache[key]

if key in self.cpu_cache:
# 提升至 GPU(如果有空间)
self._promote_to_gpu(key)
self._update_stats(key)
return self.cpu_cache[key]

# 检查 SSD
ssd_path = os.path.join(self.ssd_cache_path, key)
if os.path.exists(ssd_path):
data = self._read_ssd(ssd_path)
# 提升至 CPU
self._promote_to_cpu(key, data)
self._update_stats(key)
return data

return None # 缓存未命中

def set(self, key: str, data: bytes):
"""写入缓存(写入 GPU)"""
self._ensure_gpu_space(data)
self.gpu_cache[key] = data
self._update_stats(key)

def _ensure_gpu_space(self, data: bytes):
"""确保 GPU 有足够空间,不够则驱逐到 CPU/SSD"""
current_size = sum(len(v) for v in self.gpu_cache.values())
if current_size + len(data) <= self.gpu_limit:
return

# 按访问频率排序驱逐
evict_candidates = sorted(
self.gpu_cache.keys(),
key=lambda k: self.access_stats[k]["count"]
)

for key in evict_candidates:
if current_size + len(data) <= self.gpu_limit:
break

# 驱逐到 CPU
self._evict_to_cpu(key)
current_size -= len(self.gpu_cache.pop(key))

def _evict_to_cpu(self, key: str):
"""GPU → CPU 驱逐"""
if key not in self.gpu_cache:
return
self._ensure_cpu_space(self.gpu_cache[key])
self.cpu_cache[key] = self.gpu_cache[key]

def _promote_to_gpu(self, key: str):
"""CPU → GPU 提升"""
if key not in self.cpu_cache:
return
self._ensure_gpu_space(self.cpu_cache[key])
self.gpu_cache[key] = self.cpu_cache.pop(key)

def _update_stats(self, key: str):
self.access_stats[key]["count"] += 1
self.access_stats[key]["last_access"] = time.time()


6. 多模态与Embedding统一服务

6.1 统一服务架构

SGLang 的一大特色是可以用同一个服务进程同时处理:

  • 纯文本 LLM 推理
  • 多模态 VLM 推理(Qwen-VL、LLaVA、MiniCPM-V、NVILA)
  • Embedding 向量化(Qwen3-Embedding、GME)
  • 重排序(Reranker)

# 一个服务,全部支持
python -m sglang.launch_server \\
–model-path Qwen/Qwen3.6-VL-32B \\
–host 0.0.0.0 –port 8000

# 客户端统一调用
import sglang as sgl

# LLM 文本生成
response = sgl.chat("Hello, how are you?")

# VLM 多模态
response = sgl.chat(
"这张图片里有什么?",
images=["https://example.com/photo.jpg"]
)

# Embedding
embeddings = sgl.embed(["文本1", "文本2"])

# 同一个进程,无需部署多套服务

6.2 模型支持矩阵(v0.5.12)

模型类型支持模型支持程度
LLM Qwen3/3.5/3.6, LLaMA 4, Gemma 4, DeepSeek V4 原生支持
VLM Qwen-VL, LLaVA 1.6, MiniCPM-V, NVILA 原生支持
Embedding Qwen3-Embedding, GME, BGE 原生支持
MoE DeepSeek V4, MiniMax-M2.5, Qwen3.5-MoE Triton MoE Runner

7. 性能基准测试

7.1 标准 Benchmark(H100)

引擎模型精度吞吐量 (tok/s)对比 vLLM
SGLang LLaMA 3.1 8B BF16 16,215 +29%
LMDeploy LLaMA 3.1 8B BF16 16,132 +28%
vLLM LLaMA 3.1 8B BF16 12,553 baseline

数据来源:AIMultiple 2026-04 基准测试 [5]。即使在相同 kernel(FlashInfer)下,SGLang 的调度层优化带来了 29% 吞吐优势。

7.2 RTX PRO 6000 Blackwell 实测

引擎模型格式峰值吞吐备注
vLLM NVFP4 8,033 tok/s 整体最高
SGLang GPTQ-INT4 6,395 tok/s +17% vs vLLM GPTQ
vLLM GPTQ-INT4 ~5,470 tok/s 同精度对比
Ollama INT4 484 tok/s ~10× 慢

同精度(GPTQ-INT4)下 SGLang 领先 17% [6]。

7.3 TTFT 首Token延迟对比

场景SGLangvLLM胜出
RAG 单用户 (c=1) 597ms 2,081ms SGLang 3.5×
RAG 多用户 (c=20) 2,593ms 2,172ms vLLM
Dual H100 (c=1) 583ms 2,141ms SGLang 3.7×
Dual H100 (c=100) 2,775ms 2,843ms 持平

7.4 高并发P95尾延迟

引擎128并发成功率P95延迟
SGLang 100% P95 ≈ 1.2-1.4× 中位数
vLLM 100% P95 可达 9-12× 中位数

SGLang 在高并发下延迟分布更稳定,P95 尾延迟仅为中位数的 1.2-1.4 倍,而 vLLM 可达 9-12 倍 [3]。


8. 生产部署最佳实践

8.1 基础 Docker 部署

# 拉取 SGLang 官方镜像
docker pull lmsysorg/sglang:v0.5.12.post1

# 启动服务
docker run –gpus all \\
–shm-size 32g \\
-p 8000:8000 \\
-v ~/.cache/huggingface:/root/.cache/huggingface \\
lmsysorg/sglang:v0.5.12.post1 \\
python -m sglang.launch_server \\
–model-path Qwen/Qwen3-32B \\
–host 0.0.0.0 \\
–port 8000 \\
–mem-fraction-static 0.90 \\
–max-num-seqs 256 \\
–enable-mix-precision fp16

8.2 Kubernetes 生产部署

apiVersion: apps/v1
kind: Deployment
metadata:
name: sglangserver
spec:
replicas: 3
selector:
matchLabels:
app: sglang
template:
metadata:
labels:
app: sglang
spec:
containers:
name: sglang
image: lmsysorg/sglang:v0.5.12.post1
command:
python
m
sglang.launch_server
args:
modelpath
Qwen/Qwen332B
host
0.0.0.0
port
"8000"
memfractionstatic
"0.90"
maxnumseqs
"256"
enablemetrics
ports:
containerPort: 8000
resources:
limits:
nvidia.com/gpu: 2
requests:
nvidia.com/gpu: 2
volumeMounts:
name: modelcache
mountPath: /root/.cache
name: sharedmemory
mountPath: /dev/shm
volumes:
name: modelcache
persistentVolumeClaim:
claimName: modelcachepvc
name: sharedmemory
emptyDir:
medium: Memory

apiVersion: v1
kind: Service
metadata:
name: sglangservice
spec:
selector:
app: sglang
ports:
port: 8000
targetPort: 8000
type: ClusterIP

# HPA:基于 GPU 利用率和请求延迟自动扩缩
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: sglanghpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: sglangserver
minReplicas: 2
maxReplicas: 10
metrics:
type: Pods
pods:
metric:
name: sglang_gpu_utilization
target:
type: AverageValue
averageValue: 80
type: Pods
pods:
metric:
name: sglang_request_latency_p99
target:
type: AverageValue
averageValue: 2000 # 2秒

8.3 Prometheus + Grafana 监控

SGLang 原生暴露 Prometheus 指标:

# prometheus-scrape-config.yaml
scrape_configs:
job_name: 'sglang'
scrape_interval: 10s
static_configs:
targets:
'sglang-service:8000'
metrics_path: '/metrics'

关键监控指标:

指标名类型说明
sglang_num_running_reqs Gauge 当前运行请求数
sglang_num_waiting_reqs Gauge 等待队列长度
sglang_cache_hit_rate Gauge RadixAttention 缓存命中率
sglang_throughput_token_s Gauge 每秒输出 token 数
sglang_ttft_ms Histogram 首 token 延迟分布
sglang_tpot_ms Histogram 每 token 延迟分布
sglang_gpu_mem_usage_gb Gauge GPU 显存使用
sglang_prefix_match_len Histogram 前缀匹配长度分布

Grafana 告警规则:

groups:
name: sglang_alerts
rules:
alert: HighLatency
expr: sglang_ttft_ms{quantile="0.99"} > 5000
for: 5m
labels:
severity: warning
annotations:
summary: "SGLang P99 latency > 5s"

alert: LowCacheHitRate
expr: sglang_cache_hit_rate < 0.3
for: 10m
labels:
severity: warning
annotations:
summary: "RadixAttention cache hit rate below 30%"

alert: GPUMemoryPressure
expr: sglang_gpu_mem_usage_gb > 75
for: 5m
labels:
severity: critical
annotations:
summary: "GPU memory > 75GB"


9. 关键参数调优指南

9.1 显存与批处理

参数默认值推荐值说明
–mem-fraction-static 0.85 0.88-0.95 GPU 显存使用比例,H100 可设 0.95
–max-num-seqs 64 128-512 最大并发序列数,高并发增大
–context-length 4096 模型支持的最大长度 上下文窗口大小
–kv-cache-dtype auto fp8_e4m3 KV Cache 精度,H100 推荐 fp8

9.2 缓存与调度

参数默认值推荐值说明
–enable-radix-cache True True 启用 RadixAttention
–enable-metrics False True 启用 Prometheus 指标
–schedule-policy lpm lpm 调度策略(lpm/lof/random)
–chunked-prefill-size 8192 4096-16384 分块预填充大小

9.3 分布式

参数默认值推荐值说明
–tp 1 2-8 Tensor Parallel 大小
–dp 1 2-4 Data Parallel 大小
–ep-size 1 等于 num_experts Expert Parallel 大小
–nccl-init-addr localhost 实际 IP NCCL 初始化地址

9.4 解码与采样

参数默认值推荐值说明
–speculative-algorithm None eagle3/draft_model 投机解码算法
–speculative-draft-model None 同系列小模型 草稿模型路径
–default-sampling-params {} 按需设置 默认采样参数

9.5 SM120(Blackwell)兼容配置

# Blackwell SM120 MoE 模型兼容启动
SGLANG_DISABLE_DEEP_GEMM=1 \\
python -m sglang.launch_server \\
–model-path MiniMax-M2.5-228B \\
–fp8-gemm-backend triton \\
–moe-runner-backend triton \\
–kv-cache-dtype fp8_e4m3 \\
–mem-fraction-static 0.85


10. SGLang vs vLLM选型决策

10.1 全维度对比(2026年5月)

维度SGLang v0.5.12vLLM v0.21.0胜出
GitHub Stars 28K 81K vLLM
全球GPU部署 400K+ 更广泛 vLLM
缓存机制 RadixAttention (基数树) Hash-based SGLang
结构化输出 XGrammar (默认, 零开销) XGrammar 0.2.0 SGLang
多模态支持 原生 (Qwen-VL/LLaVA/MiniCPM) 支持 SGLang
离线吞吐 +29% vs vLLM 基准 SGLang
P95尾延迟 1.2-1.4× 中位数 9-12× 中位数 SGLang
RAG单用户TTFT 3.5× 更快 基准 SGLang
社区生态 快速增长 成熟稳定 vLLM
文档质量 ⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ vLLM
API兼容性 OpenAI 完全兼容 OpenAI 完全兼容 持平

10.2 场景推荐

你的场景

├── 多轮对话 / 客服机器人 → SGLang ✅
├── RAG知识库问答 → SGLang ✅(RadixAttention优势)
├── Agent工具调用 → SGLang ✅(XGrammar + structural_tag)
├── 结构化JSON输出 → SGLang ✅(3倍加速)
├── 多模态统一服务 → SGLang ✅(同一进程)
├──
├── 高并发短请求 → vLLM ✅
├── 长文本生成 → vLLM ✅(PagedAttention成熟)
├── 已有vLLM生态 → vLLM ✅
├── 快速搭建Demo → vLLM ✅(文档丰富)
└── 社区/第三方集成优先 → vLLM ✅

10.3 混合部署策略

生产环境中,推荐根据场景混合使用:

请求入口 → 路由层
├── RAG/Agent/多轮对话 → SGLang 集群
│ ├── RadixAttention + XGrammar
│ └── 统一服务 (LLM + VLM + Embedding)

└── 批量推理/长文本 → vLLM 集群
├── PagedAttention + Chunked Prefill
└── 成熟批处理调度


📌 面试加分点

1️⃣ RadixAttention vs PagedAttention 核心区别

PagedAttentionRadixAttention
灵感来源 操作系统内存分页 基数树数据结构
缓存粒度 Page/Block 级别 Token 级别
跨请求复用 精确哈希匹配 最长公共前缀自动检测
部分复用
显存利用率 95%+ 92%+
共享前缀场景 一般 优秀

2️⃣ 结构化生成的三种实现路径

  • Grammar-based(XGrammar):编译 Schema 为 CFG,采样时动态屏蔽非法 token —— 零开销
  • Logit 掩码:在 logits 层手动 mask 非法 token —— 需自定义 kernel
  • 后处理重试:生成后再解析验证 —— 30%+ 重试率,最低效
  • 3️⃣ SGLang 核心优化流水线

    请求到达

    RadixAttention 前缀匹配 → 缓存命中则跳过重复计算

    Overlap Scheduling → CPU 预计算下一批(GPU 零等待)

    XGrammar 约束解码 → 每一步屏蔽非法 token

    FlashInfer Attention → 高效注意力计算

    FP8/Triton Kernel → 混合精度加速

    HiCache 层级缓存 → GPU→CPU→SSD 三级卸载

    4️⃣ SGLang 性能调优口诀

    “缓存看命中,显存定并发,量化选精度,分离治长文”

    • 缓存命中率 < 30% → 检查 RadixAttention 配置或调整请求结构
    • GPU 显存利用率 < 80% → 增大 –mem-fraction-static 和 –max-num-seqs
    • 小型模型(<20B)→ 优先 BF16 而非量化
    • 大型模型(>70B)→ PD 分离 + FP8 量化
    • 长上下文场景 → 启用 HiCache 层级卸载

    参考来源

  • SGLang vs vLLM 深度对比 (2026-05-30) — Joshua8.AI
  • SGLang RadixAttention 技术实现 — GitCode Blog
  • SGLang vs vLLM 实战评测:多轮对话场景下吞吐量对比 — CSDN Blog
  • SGLang Overlap Scheduling 技术要点 — AICon
  • AIMultiple Inference Engine Benchmark (2026-04)
  • Joshua8.AI Blackwell Benchmark (2026-01)
  • SGLang 官方文档 — docs.sglang.io
  • vLLM 官方文档 — docs.vllm.ai
  • 赞(0)
    未经允许不得转载:171主机测评 » 【推理与部署篇03】SGLang深度解析
    分享到: 更多 (0)

    评论 抢沙发

    • 昵称 (必填)
    • 邮箱 (必填)
    • 网址