LLM Tool Calling 的四种 role 和 "不用调工具" 的空数组约定

写 LLM Agent 或者接 Function Calling / Tool Use 的时候,role 字段是消息路由的核心。四种 role 分工明确,工具调用要按固定生命周期走。

四种 role

1. system — 全局规则

设定模型的行为准则、可用工具、输出格式约束:

{
  "role": "system",
  "content": "你是一个只回答天气问题的助手。工具调用格式必须严格 JSON。"
}

优先级最高。一段对话通常只有一条 system 消息(放最前面)。

2. user — 用户输入

真实用户的问题、指令、上下文:

{ "role": "user", "content": "上海明天下雨吗?" }

3. assistant — 模型响应

模型的回复。可以是普通文本、也可以是 tool_call 请求

普通回复

{ "role": "assistant", "content": "上海明天多云,温度 15-22°C。" }

发起工具调用

{
  "role": "assistant",
  "content": null,
  "tool_calls": [
    {
      "id": "call_abc123",
      "type": "function",
      "function": {
        "name": "get_weather",
        "arguments": "{\"city\":\"上海\",\"date\":\"tomorrow\"}"
      }
    }
  ]
}

content 为 null 表示模型选择用工具而不是直接回答。

4. tool — 工具执行结果

外部函数运行完,把结果塞回上下文让模型继续

{
  "role": "tool",
  "tool_call_id": "call_abc123",
  "content": "{\"weather\":\"多云\",\"temp\":\"15-22\"}"
}

tool_call_id 必须和上一条 assistant.tool_calls[].id 对上——多工具并行时靠这个匹配。

完整生命周期

user  → "上海明天天气"

assistant → tool_calls: [ get_weather({city:"上海"}) ]

[外部执行 get_weather,返回 "多云 15-22°C"]

tool → "多云 15-22°C"   (tool_call_id = call_abc123)

assistant → "上海明天多云,15-22°C,建议带件外套"

四轮消息、四种 role。每条 tool 消息必须对应上一轮某个 tool_call,不能凭空出现。

“不需要工具”的返回约定

有些 Agent 框架要求 assistant 明确表达”这一轮我不需要工具”。约定俗成的写法是空数组

{
  "role": "assistant",
  "content": "你好,我是助手。",
  "tool_calls": []
}

严格规范里

  • 不需要工具tool_calls: [](或者干脆不带该字段)
  • 需要工具tool_calls: [{...}]

[] 明确表示”模型已经判断过、决定不用工具”,比 null 或缺字段更清晰。写自己的 tool routing 时,这样解析更省心:

if (Array.isArray(msg.tool_calls) && msg.tool_calls.length > 0) {
    executeToolCalls(msg.tool_calls);
} else {
    displayText(msg.content);
}

并行工具调用

现代模型(GPT-4、Claude 3.5+)支持一次返回多个 tool_calls:

{
  "role": "assistant",
  "tool_calls": [
    { "id": "call_1", "function": {"name": "get_weather", "arguments": "{\"city\":\"上海\"}"} },
    { "id": "call_2", "function": {"name": "get_weather", "arguments": "{\"city\":\"北京\"}"} }
  ]
}

Agent 应该并发执行两个 get_weather,然后按顺序 append 两条 tool 消息回上下文:

[
  { "role": "tool", "tool_call_id": "call_1", "content": "上海:多云" },
  { "role": "tool", "tool_call_id": "call_2", "content": "北京:晴" }
]

再让模型继续。

常见坑

1. tool_call_id 忘了对齐

{ "role": "assistant", "tool_calls": [{"id": "call_1", ...}] },
{ "role": "tool", "tool_call_id": "call_2", ... }   // 对不上

模型下一轮会困惑——大概率报错或者胡说。每个 tool_call 必须对应恰好一个 tool 消息

2. 直接输出工具调用当文本

有些开发者 prompt 里让模型 “输出 <tool>...</tool> 格式”,然后自己解析。能用原生 tool_calls 就用,稳定性和生态好得多(错误处理、并行、streaming 都有官方支持)。

3. 工具报错没处理

工具执行失败,应该把错误信息作为 tool 消息返回,让模型知道并决定重试或换策略:

{
  "role": "tool",
  "tool_call_id": "call_1",
  "content": "{\"error\":\"API rate limit exceeded, retry after 60s\"}"
}

而不是抛异常终止对话。

4. 工具太多导致 token 爆炸

每次请求都要把所有 tool schema 塞进去。只给模型看当前场景需要的工具子集——按用户意图动态选。20 个以上工具建议做工具路由。

OpenAI / Anthropic 差异

字段OpenAIAnthropic Claude
role: assistant 里的工具调用tool_calls: []content 里是 array,含 type: "tool_use"
工具结果 roletooluser(但 content 里是 type: "tool_result"
工具定义位置tools: [] 参数tools: [] 参数(结构略不同)

Anthropic 把工具结果放 user role 是历史原因——本质数据一样,只是包装略不同。

一句话总结

四种 rolesystem 全局规则、user 用户输入、assistant 模型输出(可含 tool_calls)、tool 外部执行结果。“不用工具”的规范返回是 tool_calls: []。每个 tool 消息必须靠 tool_call_id 对齐前一轮的调用。