Llm nodes cn
Jianmu 提供两个内置 LLM 节点来驱动模型推理:SimpleLLMNode 处理无需工具调用的单次文本生成,AgentLLMNode 在其基础上扩展了工具 Schema 暴露、结构化 ToolCall 提取和跨轮次用量累积。两者共享同一套上下文构建管线——从状态中拉取对话历史,通过可组合的 Provider→Filter 流水线装配成最终发送给模型的消息序列。本章将沿着"消息准备 → 上下文装配 → 模型调用 → 结果回写"这条链路,逐步拆解两个节点的设计差异与协作关系。
节点继承与运行时基座¶
SimpleLLMNode 和 AgentLLMNode 最终都继承自 AsyncBehaviour,后者是 Jianmu 在 py_trees 行为树基础上的异步扩展。AsyncBehaviour 在每个 tick 周期内启动一个 asyncio.Task 执行 update_async() 协程,并在任务完成时通过 _wake_up 回调唤醒 ReactiveRunner,从而将异步 IO 无缝嵌入行为树的同步 tick 循环。与此同时,JianmuNodeMixin 为节点注入了 state_manager、ctx(RunContext)、端口绑定解析(read_port / write_port / write_ports)和消息追加(append_port_messages)等运行时能力。这意味着 LLM 节点无需关心底层状态存储细节——它们通过声明式端口与 StateManager 交互。
以下 Mermaid 图展示了两个 LLM 节点在类型层级中的定位,以及它们与上下文构建器和模型客户端之间的核心依赖关系:
classDiagram
class AsyncBehaviour {
+update_async() Status
+inject(payload) None
+read_port(name, state_key) Any
+write_port(name, value, state_key) None
+write_ports(values, state_keys) None
+append_port_messages(name, messages, state_key) None
+model_client ModelClient
}
class SimpleLLMNode {
+_prepare_messages() list~Message~
+_resolve_context_builder() ContextBuilderProtocol
+_build_model_config() ModelConfig
+_emit_runtime_event(type, payload) None
+update_async() Status
}
class AgentLLMNode {
+_tools_schema list~dict~
+_tools_description str
+_persist_usage(response_msg) None
+_persist_result(response_msg, content) None
+_resolve_context_builder() ContextBuilderProtocol
+_build_model_config(tools_schema) ModelConfig
+update_async() Status
}
class ContextBuilder {
+providers list~MessageProvider~
+filters list~ContextFilter~
+build(local_state, global_state, ctx) Sequence~Message~
+resolve_for_llm(...) ContextBuilderProtocol
+for_chat(...) ContextBuilder
+for_react(...) ContextBuilder
+for_skill(...) ContextBuilder
}
class ModelClient {
+invoke(messages, config, stream) tuple~Message,str~
+resolve(...) ModelClient
}
class AgentLLMConfig {
+model str
+temperature float
+max_tokens int
+stream bool
+keys StateKeys
+resilience ModelResilienceConfig
}
AsyncBehaviour <|-- SimpleLLMNode
SimpleLLMNode <|-- AgentLLMNode
SimpleLLMNode ..> ContextBuilder : _resolve_context_builder
SimpleLLMNode ..> ModelClient : invoke
SimpleLLMNode ..> AgentLLMConfig : config
AgentLLMNode ..> ContextBuilder : prefer_react=True
SimpleLLMNode:单轮文本生成的最小契约¶
SimpleLLMNode 的核心语义是一次模型调用 → 一次回答回写。它的构造函数接收三个必需/可选依赖:model_client(必需,会被 _require_model_client 校验为 Jianmu 的 ModelClient 类型)、context_builder(可选,用于覆盖默认的上下文装配逻辑)和 config(可选,AgentLLMConfig 实例,承载模型参数与状态键路由)。
消息准备:StateMessageStore 归一化¶
_prepare_messages() 不直接读取原始状态列表,而是通过 StateMessageStore.read_input() 进行归一化——无论状态中存储的是 Message 对象、字典还是裸字符串,都会被强制转换为标准的 Message 列表,fallback_role 默认为 "user"。这一层抽象使节点代码免于分支处理各种消息存储格式。
上下文装配:ContextBuilder 的三层解析¶
_resolve_context_builder() 调用 ContextBuilder.resolve_for_llm(),按优先级依次尝试:显式传入的 builder → 运行时注入的 ctx.prompt_runtime.context_builder → 默认的 for_chat(...) 预设。SimpleLLMNode 不传 prefer_react(默认为 False)且 tools_desc 为空,因此走 chat 预设——仅装配 persona 提示词、可选的 bootstrap 文件内容和对话历史,不注入 ReAct 协议或工具描述。
模型配置构建¶
_build_model_config() 忽略 tools_schema 参数(它始终为空列表),将 AgentLLMConfig 的 model、temperature、max_tokens、top_p、timeout、resilience 配置等映射到 ModelConfig dataclass 上。注意 SimpleLLMNode 不设置 top_k,这与其"不需要工具调用"的定位一致。
update_async 主流程¶
update_async() 的执行分为七个阶段:
- 重置调用字段:
state_manager.reset_call_fields()清理上一次调用的临时标记 - 消息准备:调用
_prepare_messages()获取归一化消息列表,若为空则直接返回 FAILURE - 模型客户端解析:优先使用注入的
self.model_client,否则回退到ModelClient.resolve(...)全局解析 - 上下文构建:调用
_resolve_context_builder().build()将本地消息(local_state={"messages": messages})和全局状态送入 Provider→Filter 流水线,产出full_messages - 事件发射:依次发射
reply.started和model.call.started运行时事件 - 模型调用:通过
model_client.invoke()发起实际调用,支持流式与非流式,流式模式下通过on_text_update回调同步更新streaming_output端口并发射text.delta事件 - 结果回写:将 assistant 消息追加到
messages端口,将文本内容写入final_answer,将done标记设为True
运行时事件体系¶
SimpleLLMNode 在单次调用中发射四类事件,构建完整的可观测链路:
| 事件类型 | 发射时机 | 关键 Payload 字段 |
|---|---|---|
reply.started |
上下文构建完成后,模型调用前 | message_count, full_message_count, stream |
model.call.started |
同上(紧跟 reply.started) | model, stream, tool_schema_count |
text.delta |
流式模式下每次文本增量 | text, length |
model.call.completed |
模型返回后 | model, tool_call_count, usage |
reply.completed |
结果回写后 | final_text, tool_call_count |
所有事件通过 emit_runtime_event() 统一发射,该函数在 RuntimeEventBus 存在时才实际推送,确保在无总线配置的场景下零开销。
AgentLLMNode:工具感知的 Agent 推理步¶
AgentLLMNode 继承自 SimpleLLMNode,在单轮模型调用的基础上附加了三项关键能力:工具 Schema 暴露(将结构化 tool definitions 传给模型)、结构化 ToolCall 提取(从响应中解析 ToolCall 列表并回写到 actions 状态字段)和跨轮次用量累积(将每次调用的 token 消耗累加到 llm.usage)。这些差异使 AgentLLMNode 成为 ReAct 循环中 Agent 推理步的标准执行单元。
构造函数差异¶
相比父类,AgentLLMNode 新增两个参数:
tools_schema: Optional[List[Dict[str, Any]]]— 符合 OpenAI function calling 格式的工具定义列表,由 ReAct 预设通过ToolSet.schemas()预计算tools_description: str— 人类可读的工具描述文本块,注入系统提示词供模型理解可用工具
上下文装配:偏好 ReAct 预设¶
AgentLLMNode 重写了 _resolve_context_builder(),调用 ContextBuilder.resolve_for_llm() 时显式传入 prefer_react=True 和 tools_desc=self._tools_description。这导致解析逻辑优先选择 for_react(...) 预设,在 persona 提示词和对话历史之间插入 ReAct 协议提示词和工具描述块。
ReAct 协议的核心指令是:
当需要使用工具时,系统会自动处理工具调用。你不需要以文本格式输出工具调用——直接决定使用哪个工具,系统会通过结构化 function calling 为你调用它。当你得出最终答案时,以 "Final Answer: [你的回答]" 格式输出。
此协议提示词通过 get_react_protocol_prompt() 解析,支持用户在 jianmu.yaml 中覆盖。
模型配置:注入工具 Schema¶
AgentLLMNode 的 _build_model_config() 接受 tools_schema 参数并将其传入 ModelConfig 的 tools 字段。此外增加了 top_k 默认值(40),并使 max_tokens 的默认值通过 AgentLLMConfig 的配置链路解析。
结果持久化:_persist_result 与 _persist_usage¶
_persist_result() 是 AgentLLMNode 与 SimpleLLMNode 最关键的差异点。它不仅回写 messages 和 final_answer,还:
- 从
response_msg.tool_calls中通过extract_tool_call_from_dict()提取标准化的ToolCall列表 - 写入
actions端口(供 ToolExecutor 消费) - 递增
rounds计数器 - 仅在 没有 tool calls 时将
done设为True(有工具调用意味着还需要继续循环) - 仅在 没有 tool calls 时填充
final_answer(有工具调用时 final_answer 留空等待后续轮次)
_persist_usage() 则从响应消息的 metadata 中提取 usage 字典,按 key 将数值累加到状态的 llm.usage 字段中,实现跨轮次的 token 统计。
两节点关键差异对比¶
| 维度 | SimpleLLMNode | AgentLLMNode |
|---|---|---|
| 工具 Schema | 不暴露 | 通过 tools_schema 暴露 |
| 上下文预设 | for_chat() |
for_react()(含协议+工具描述) |
done 信号 |
始终 True |
仅在无 tool_calls 时为 True |
actions 回写 |
无 | 解析 tool_calls → ToolCall 列表 |
rounds 递增 |
无 | 每次调用 +1 |
| Token 用量累积 | 无 | 累加到 llm.usage |
top_k 配置 |
不设置 | 默认 40 |
final_answer 策略 |
始终写入 | 仅无工具调用时写入 |
ContextBuilder:可组合上下文装配流水线¶
两个 LLM 节点都依赖 ContextBuilder 完成消息装配。ContextBuilder 采用 Provider → Filter 两阶段流水线架构:Provider 负责从不同来源产生消息片段,Filter 对聚合后的消息序列进行后处理(截断、token 预算控制等)。
Provider 类型¶
| Provider | 职责 | 使用的预设 |
|---|---|---|
StaticPromptProvider |
注入静态系统提示词(persona、ReAct 协议) | chat / react / skill |
ToolsDescProvider |
注入格式化后的工具描述块 | react |
StateHistoryProvider |
从状态中加载对话历史并归一化语义标注 | chat / react / skill |
BootstrapFilesProvider |
加载工作空间引导文件(如 CLAUDE.md) | chat / react / skill |
SkillSetPromptProvider |
渲染已选技能集的提示词(摘要/完整模式) | skill |
StateHistoryProvider 在加载历史消息时还执行语义标注:它将之前的 assistant 工具调用计划标注为 semantic_kind=self_tool_plan,将 tool 消息标注为 semantic_kind=tool_observation,并前置一条 get_history_semantics_prompt() 的系统消息来告知模型这些标注的含义,防止模型将历史中的工具交互误认为新的用户消息。
Filter 类型¶
| Filter | 职责 | 关键参数 |
|---|---|---|
MaxMessagesFilter |
限制消息总数,保留前导 pinned 角色 | max_messages, pinned_roles(默认 ("system",)) |
TokenBudgetFilter |
按 token 预算截断,支持三种计数器后端 | max_tokens, token_counter, trim_from_start |
TokenBudgetFilter 支持三种 token 计数器:SimpleTokenCounter(基于字符数 / 4 估算)、TiktokenTokenCounter(使用 OpenAI tiktoken 库精确计算)和 HFTokenCounter(使用 HuggingFace tokenizer)。默认使用 SimpleTokenCounter,可在构建时替换。
resolve_for_llm:上下文构建器的五级解析¶
ContextBuilder.resolve_for_llm() 是连接 LLM 节点与上下文装配的核心枢纽。它的解析优先级如下:
flowchart TD
A[resolve_for_llm] --> B{explicit_builder?}
B -->|是| C[返回 explicit_builder]
B -->|否| D{runtime_prompt.context_builder?}
D -->|是| E[返回 runtime builder]
D -->|否| F{allow_skill_context AND 请求 skill-aware prompt?}
F -->|是| G[for_skill(...)]
F -->|否| H{prefer_react OR tools_desc?}
H -->|是| I[for_react(...)]
H -->|否| J[for_chat(...)]
SimpleLLMNode 调用时不传 prefer_react 且 tools_desc 为空,因此落到 for_chat;AgentLLMNode 传入 prefer_react=True 和 tools_desc,因此匹配 for_react。如果运行时 PromptRuntimeContext 中包含技能目录配置且调用方启用了 allow_skill_context,则可能走到 for_skill 分支——这为 SkillNode 等高级场景提供了扩展入口。
ModelClient:统一调用外观与弹性策略¶
两个 LLM 节点都通过 ModelClient.invoke() 发起实际的模型调用,而非直接与 Provider 交互。ModelClient 作为外观层提供了:
- Provider 自动解析:按
["openai", "litellm"]优先级检测环境变量中的 API Key,自动选择可用 Provider - 重试与 Fallback:可重试错误(超时、限流、5xx)自动重试,重试耗尽后尝试同 Provider 家族的 fallback 模型
- 流式/非流式统一:
invoke()内部根据stream参数分发到consume_stream()或provider.generate(),返回统一的(Message, str)元组
调用流程图¶
sequenceDiagram
participant Node as SimpleLLMNode / AgentLLMNode
participant MC as ModelClient
participant IM as invoke_model
participant P as Provider (OpenAI/LiteLLM)
participant CS as consume_stream
Node->>Node: _prepare_messages()
Node->>Node: _resolve_context_builder().build()
Node->>MC: invoke(messages, config, stream, on_text_update)
MC->>MC: 重试循环 (max_retries + 1)
alt stream = True
MC->>IM: invoke_model(stream=True)
IM->>P: stream(messages, config)
P-->>CS: async for chunk
CS-->>Node: on_text_update(accumulated_text)
CS-->>IM: (response_msg, content)
else stream = False
MC->>IM: invoke_model(stream=False)
IM->>P: generate(messages, config)
P-->>IM: response_msg
end
IM-->>MC: (response_msg, content)
MC-->>MC: 失败? → fallback 模型重试
MC-->>Node: (response_msg, content)
Node->>Node: _persist_result / write_port
重试分类与 Fallback 策略¶
ModelClient 将错误分为可重试和不可重试两类。可重试错误包括 TimeoutError、ConnectionError 以及错误消息中包含 timeout、rate limit、429、5xx 等关键词的异常。不可重试错误(如 4xx 参数错误、模型不存在)直接抛出,不浪费重试次数。Fallback 机制仅在主模型与 fallback 模型属于同一 Provider 家族时触发(通过 _model_provider_family() 判断),避免跨 Provider 的配置不兼容。
流式处理:tool call 增量合并¶
流式模式下,consume_stream() 不仅累积文本增量,还处理流式工具调用:通过 merge_stream_tool_calls() 按 index 将分散在多个 chunk 中的 tool name 和 arguments 片段合并,最终由 finalize_stream_tool_calls() 去重并产出规范的 tool_calls 列表。
配置体系:AgentLLMConfig 与 StateKeys¶
AgentLLMConfig 是 Pydantic BaseModel,其每个字段的默认值都通过 get_config() 从 jianmu.yaml 中延迟解析。这意味着用户无需在代码中显式设置 model、temperature 等参数——只需在 jianmu.yaml 中配置一次即可全局生效。
配置字段解析表¶
| 字段 | 默认值来源 | 说明 |
|---|---|---|
model |
get_config().models.default |
默认模型标识符 |
temperature |
get_config().llm.temperature |
采样温度 |
max_tokens |
get_config().llm.max_tokens |
最大生成 token 数 |
top_p |
get_config().llm.top_p |
Nucleus 采样阈值 |
top_k |
get_config().llm.top_k |
Top-k 采样(仅 AgentLLMNode 使用) |
timeout |
get_config().llm.timeout |
请求超时(秒) |
stream |
False(硬编码) |
是否启用流式输出 |
max_budget_tokens |
get_config().llm.max_budget_tokens |
跨轮次 token 预算上限 |
resilience.max_retries |
get_config().llm.max_retries |
模型调用最大重试次数 |
resilience.fallback_model |
get_config().models.fallback |
重试耗尽后的 fallback 模型 |
system_prompt |
None |
覆盖默认 persona 提示词 |
keys |
StateKeys() |
状态字段名路由配置 |
StateKeys 定义了 LLM 节点、ToolExecutor、SkillNode 与 ReAct 预设之间共享的状态字段命名约定:
| Key | 默认值 | 存储内容 |
|---|---|---|
messages |
"messages" |
对话历史(Message 列表) |
rounds |
"rounds" |
当前 Agent 轮次计数 |
actions |
"actions" |
待执行的 ToolCall 列表 |
done |
"done" |
循环终止标志 |
text_output |
"text_output" |
最近一次纯文本节点输出 |
final_answer |
"final_answer" |
最终回答文本 |
skill_result |
"skill_result" |
结构化 skill 输出桥接槽 |
usage |
"llm.usage" |
累积 token 用量 |
streaming_output |
"streaming_output" |
实时流式输出文本 |
tool_effects |
"tool_effects" |
工具副作用记录 |
用户可通过自定义 StateKeys 实例来重命名这些字段。当前公开 API 里,这通常通过 AgentLLMConfig.keys、ReActConfig.keys、SkillNodeConfig.keys、ToolExecutorConfig.keys 这类每节点配置对象完成,使状态路由契约集中且显式。
ReAct 预设中 AgentLLMNode 的组装模式¶
create_react_node() 展示了 AgentLLMNode 的标准使用模式。它将 AgentLLMNode、ToolExecutor 和 StateCondition 组合为一个 LoopUntilSuccess 子树:
flowchart LR
subgraph LoopUntilSuccess["LoopUntilSuccess (max_iterations)"]
direction TB
A[AgentLLMNode] --> B[ToolExecutor]
B --> C[StateCondition: done?]
C -->|否| A
end
C -->|是| D[返回 final_answer]
关键设计细节:
- AgentLLMNode 将 tools_schema 和 tools_description 传入,确保模型知晓可用工具
- AgentLLMNode 与 ToolExecutor 共享同一套 StateKeys,通过状态字段隐式通信
- LoopUntilSuccess 在每次循环前检查 token 预算(abort_condition),超出则提前终止
- StateCondition 检查 done 字段:当 AgentLLMNode 判定无工具调用时设为 True,循环终止
关键设计要点¶
上下文构建器的可替换性。两个 LLM 节点都通过 _resolve_context_builder() 在每次 update_async() 调用时动态解析上下文构建器。这意味着即使节点已创建,后续通过 inject() 注入新的 RunContext(包含不同的 PromptRuntimeContext)也能切换上下文装配策略——测试中验证了这一行为:同一节点在两次 inject() 后分别使用了不同的 builder。
技能上下文的隐式忽略。SimpleLLMNode 和 AgentLLMNode 默认不启用技能感知上下文(allow_skill_context=False)。即使 PromptRuntimeContext 中配置了技能目录和 include_skills_summary=True,也不会触发 for_skill 预设。技能上下文仅由 SkillNode 等明确开启 allow_skill_context=True 的调用方触发。
ModelClient 的类型守卫。_require_model_client() 确保传入的 model_client 是 Jianmu 的 ModelClient 包装类型而非裸 Provider。这是架构分层的关键约束——LLM 节点不应直接依赖 Provider 接口,所有调用必须经过 ModelClient.invoke() 以获得重试、fallback 和可观测性。
阅读下一步¶
- 了解 AgentLLMNode 产出的 ToolCall 如何被消费:工具与技能节点:ToolExecutor、SkillNode 与约束联动
- 深入上下文构建的过滤与 token 预算:上下文构建器:消息过滤、Token 预算控制与多源 Prompt 装配
- 理解 ModelClient 的 Provider 架构:ModelClient 外观:统一 OpenAI 与 LiteLLM Provider 的可观测调用
- 查看 ReAct 预设的完整组装逻辑:ReAct 节点工厂:LLM 调用 → 工具执行 → 完成的循环回路