写 LLM Agent 或者接 Function Calling / Tool Use 的时候,role 字段是消息路由的核心。四种 role 分工明确,工具调用要按固定生命周期走。
四种 role
1. system — 全局规则
设定模型的行为准则、可用工具、输出格式约束:
{
"role": "system",
"content": "你是一个只回答天气问题的助手。工具调用格式必须严格 JSON。"
}
优先级最高。一段对话通常只有一条 system 消息(放最前面)。
2. user — 用户输入
真实用户的问题、指令、上下文:
{ "role": "user", "content": "上海明天下雨吗?" }
3. assistant — 模型响应
模型的回复。可以是普通文本、也可以是 tool_call 请求:
普通回复:
{ "role": "assistant", "content": "上海明天多云,温度 15-22°C。" }
发起工具调用:
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"city\":\"上海\",\"date\":\"tomorrow\"}"
}
}
]
}
content 为 null 表示模型选择用工具而不是直接回答。
4. tool — 工具执行结果
外部函数运行完,把结果塞回上下文让模型继续:
{
"role": "tool",
"tool_call_id": "call_abc123",
"content": "{\"weather\":\"多云\",\"temp\":\"15-22\"}"
}
tool_call_id 必须和上一条 assistant.tool_calls[].id 对上——多工具并行时靠这个匹配。
完整生命周期
user → "上海明天天气"
↓
assistant → tool_calls: [ get_weather({city:"上海"}) ]
↓
[外部执行 get_weather,返回 "多云 15-22°C"]
↓
tool → "多云 15-22°C" (tool_call_id = call_abc123)
↓
assistant → "上海明天多云,15-22°C,建议带件外套"
四轮消息、四种 role。每条 tool 消息必须对应上一轮某个 tool_call,不能凭空出现。
“不需要工具”的返回约定
有些 Agent 框架要求 assistant 明确表达”这一轮我不需要工具”。约定俗成的写法是空数组:
{
"role": "assistant",
"content": "你好,我是助手。",
"tool_calls": []
}
严格规范里:
- 不需要工具 →
tool_calls: [](或者干脆不带该字段) - 需要工具 →
tool_calls: [{...}]
[] 明确表示”模型已经判断过、决定不用工具”,比 null 或缺字段更清晰。写自己的 tool routing 时,这样解析更省心:
if (Array.isArray(msg.tool_calls) && msg.tool_calls.length > 0) {
executeToolCalls(msg.tool_calls);
} else {
displayText(msg.content);
}
并行工具调用
现代模型(GPT-4、Claude 3.5+)支持一次返回多个 tool_calls:
{
"role": "assistant",
"tool_calls": [
{ "id": "call_1", "function": {"name": "get_weather", "arguments": "{\"city\":\"上海\"}"} },
{ "id": "call_2", "function": {"name": "get_weather", "arguments": "{\"city\":\"北京\"}"} }
]
}
Agent 应该并发执行两个 get_weather,然后按顺序 append 两条 tool 消息回上下文:
[
{ "role": "tool", "tool_call_id": "call_1", "content": "上海:多云" },
{ "role": "tool", "tool_call_id": "call_2", "content": "北京:晴" }
]
再让模型继续。
常见坑
1. tool_call_id 忘了对齐
{ "role": "assistant", "tool_calls": [{"id": "call_1", ...}] },
{ "role": "tool", "tool_call_id": "call_2", ... } // 对不上
模型下一轮会困惑——大概率报错或者胡说。每个 tool_call 必须对应恰好一个 tool 消息。
2. 直接输出工具调用当文本
有些开发者 prompt 里让模型 “输出 <tool>...</tool> 格式”,然后自己解析。能用原生 tool_calls 就用,稳定性和生态好得多(错误处理、并行、streaming 都有官方支持)。
3. 工具报错没处理
工具执行失败,应该把错误信息作为 tool 消息返回,让模型知道并决定重试或换策略:
{
"role": "tool",
"tool_call_id": "call_1",
"content": "{\"error\":\"API rate limit exceeded, retry after 60s\"}"
}
而不是抛异常终止对话。
4. 工具太多导致 token 爆炸
每次请求都要把所有 tool schema 塞进去。只给模型看当前场景需要的工具子集——按用户意图动态选。20 个以上工具建议做工具路由。
OpenAI / Anthropic 差异
| 字段 | OpenAI | Anthropic Claude |
|---|---|---|
role: assistant 里的工具调用 | tool_calls: [] | content 里是 array,含 type: "tool_use" 项 |
| 工具结果 role | tool | user(但 content 里是 type: "tool_result") |
| 工具定义位置 | tools: [] 参数 | tools: [] 参数(结构略不同) |
Anthropic 把工具结果放 user role 是历史原因——本质数据一样,只是包装略不同。
一句话总结
四种 role:system 全局规则、user 用户输入、assistant 模型输出(可含 tool_calls)、tool 外部执行结果。“不用工具”的规范返回是 tool_calls: []。每个 tool 消息必须靠 tool_call_id 对齐前一轮的调用。
