跳转至

Llm nodes cn

Jianmu 提供两个内置 LLM 节点来驱动模型推理:SimpleLLMNode 处理无需工具调用的单次文本生成,AgentLLMNode 在其基础上扩展了工具 Schema 暴露、结构化 ToolCall 提取和跨轮次用量累积。两者共享同一套上下文构建管线——从状态中拉取对话历史,通过可组合的 Provider→Filter 流水线装配成最终发送给模型的消息序列。本章将沿着"消息准备 → 上下文装配 → 模型调用 → 结果回写"这条链路,逐步拆解两个节点的设计差异与协作关系。

节点继承与运行时基座

SimpleLLMNode 和 AgentLLMNode 最终都继承自 AsyncBehaviour,后者是 Jianmu 在 py_trees 行为树基础上的异步扩展。AsyncBehaviour 在每个 tick 周期内启动一个 asyncio.Task 执行 update_async() 协程,并在任务完成时通过 _wake_up 回调唤醒 ReactiveRunner,从而将异步 IO 无缝嵌入行为树的同步 tick 循环。与此同时,JianmuNodeMixin 为节点注入了 state_manager、ctx(RunContext)、端口绑定解析(read_port / write_port / write_ports)和消息追加(append_port_messages)等运行时能力。这意味着 LLM 节点无需关心底层状态存储细节——它们通过声明式端口与 StateManager 交互。

以下 Mermaid 图展示了两个 LLM 节点在类型层级中的定位,以及它们与上下文构建器和模型客户端之间的核心依赖关系:

classDiagram
    class AsyncBehaviour {
        +update_async() Status
        +inject(payload) None
        +read_port(name, state_key) Any
        +write_port(name, value, state_key) None
        +write_ports(values, state_keys) None
        +append_port_messages(name, messages, state_key) None
        +model_client ModelClient
    }
    class SimpleLLMNode {
        +_prepare_messages() list~Message~
        +_resolve_context_builder() ContextBuilderProtocol
        +_build_model_config() ModelConfig
        +_emit_runtime_event(type, payload) None
        +update_async() Status
    }
    class AgentLLMNode {
        +_tools_schema list~dict~
        +_tools_description str
        +_persist_usage(response_msg) None
        +_persist_result(response_msg, content) None
        +_resolve_context_builder() ContextBuilderProtocol
        +_build_model_config(tools_schema) ModelConfig
        +update_async() Status
    }
    class ContextBuilder {
        +providers list~MessageProvider~
        +filters list~ContextFilter~
        +build(local_state, global_state, ctx) Sequence~Message~
        +resolve_for_llm(...) ContextBuilderProtocol
        +for_chat(...) ContextBuilder
        +for_react(...) ContextBuilder
        +for_skill(...) ContextBuilder
    }
    class ModelClient {
        +invoke(messages, config, stream) tuple~Message,str~
        +resolve(...) ModelClient
    }
    class AgentLLMConfig {
        +model str
        +temperature float
        +max_tokens int
        +stream bool
        +keys StateKeys
        +resilience ModelResilienceConfig
    }

    AsyncBehaviour <|-- SimpleLLMNode
    SimpleLLMNode <|-- AgentLLMNode
    SimpleLLMNode ..> ContextBuilder : _resolve_context_builder
    SimpleLLMNode ..> ModelClient : invoke
    SimpleLLMNode ..> AgentLLMConfig : config
    AgentLLMNode ..> ContextBuilder : prefer_react=True

SimpleLLMNode:单轮文本生成的最小契约

SimpleLLMNode 的核心语义是一次模型调用 → 一次回答回写。它的构造函数接收三个必需/可选依赖:model_client(必需,会被 _require_model_client 校验为 Jianmu 的 ModelClient 类型)、context_builder(可选,用于覆盖默认的上下文装配逻辑)和 config(可选,AgentLLMConfig 实例,承载模型参数与状态键路由)。

消息准备:StateMessageStore 归一化

_prepare_messages() 不直接读取原始状态列表,而是通过 StateMessageStore.read_input() 进行归一化——无论状态中存储的是 Message 对象、字典还是裸字符串,都会被强制转换为标准的 Message 列表,fallback_role 默认为 "user"。这一层抽象使节点代码免于分支处理各种消息存储格式。

上下文装配:ContextBuilder 的三层解析

_resolve_context_builder() 调用 ContextBuilder.resolve_for_llm(),按优先级依次尝试:显式传入的 builder → 运行时注入的 ctx.prompt_runtime.context_builder → 默认的 for_chat(...) 预设。SimpleLLMNode 不传 prefer_react(默认为 False)且 tools_desc 为空,因此走 chat 预设——仅装配 persona 提示词、可选的 bootstrap 文件内容和对话历史,不注入 ReAct 协议或工具描述。

模型配置构建

_build_model_config() 忽略 tools_schema 参数(它始终为空列表),将 AgentLLMConfig 的 model、temperature、max_tokens、top_p、timeout、resilience 配置等映射到 ModelConfig dataclass 上。注意 SimpleLLMNode 不设置 top_k,这与其"不需要工具调用"的定位一致。

update_async 主流程

update_async() 的执行分为七个阶段:

  1. 重置调用字段:state_manager.reset_call_fields() 清理上一次调用的临时标记
  2. 消息准备:调用 _prepare_messages() 获取归一化消息列表,若为空则直接返回 FAILURE
  3. 模型客户端解析:优先使用注入的 self.model_client,否则回退到 ModelClient.resolve(...) 全局解析
  4. 上下文构建:调用 _resolve_context_builder().build() 将本地消息(local_state={"messages": messages})和全局状态送入 Provider→Filter 流水线,产出 full_messages
  5. 事件发射:依次发射 reply.started 和 model.call.started 运行时事件
  6. 模型调用:通过 model_client.invoke() 发起实际调用,支持流式与非流式,流式模式下通过 on_text_update 回调同步更新 streaming_output 端口并发射 text.delta 事件
  7. 结果回写:将 assistant 消息追加到 messages 端口,将文本内容写入 final_answer,将 done 标记设为 True

运行时事件体系

SimpleLLMNode 在单次调用中发射四类事件,构建完整的可观测链路:

事件类型 发射时机 关键 Payload 字段
reply.started 上下文构建完成后,模型调用前 message_count, full_message_count, stream
model.call.started 同上(紧跟 reply.started) model, stream, tool_schema_count
text.delta 流式模式下每次文本增量 text, length
model.call.completed 模型返回后 model, tool_call_count, usage
reply.completed 结果回写后 final_text, tool_call_count

所有事件通过 emit_runtime_event() 统一发射,该函数在 RuntimeEventBus 存在时才实际推送,确保在无总线配置的场景下零开销。

AgentLLMNode:工具感知的 Agent 推理步

AgentLLMNode 继承自 SimpleLLMNode,在单轮模型调用的基础上附加了三项关键能力:工具 Schema 暴露(将结构化 tool definitions 传给模型)、结构化 ToolCall 提取(从响应中解析 ToolCall 列表并回写到 actions 状态字段)和跨轮次用量累积(将每次调用的 token 消耗累加到 llm.usage)。这些差异使 AgentLLMNode 成为 ReAct 循环中 Agent 推理步的标准执行单元。

构造函数差异

相比父类,AgentLLMNode 新增两个参数:

  • tools_schema: Optional[List[Dict[str, Any]]] — 符合 OpenAI function calling 格式的工具定义列表,由 ReAct 预设通过 ToolSet.schemas() 预计算
  • tools_description: str — 人类可读的工具描述文本块,注入系统提示词供模型理解可用工具

上下文装配:偏好 ReAct 预设

AgentLLMNode 重写了 _resolve_context_builder(),调用 ContextBuilder.resolve_for_llm() 时显式传入 prefer_react=True 和 tools_desc=self._tools_description。这导致解析逻辑优先选择 for_react(...) 预设,在 persona 提示词和对话历史之间插入 ReAct 协议提示词和工具描述块。

ReAct 协议的核心指令是:

当需要使用工具时,系统会自动处理工具调用。你不需要以文本格式输出工具调用——直接决定使用哪个工具,系统会通过结构化 function calling 为你调用它。当你得出最终答案时,以 "Final Answer: [你的回答]" 格式输出。

此协议提示词通过 get_react_protocol_prompt() 解析,支持用户在 jianmu.yaml 中覆盖。

模型配置:注入工具 Schema

AgentLLMNode 的 _build_model_config() 接受 tools_schema 参数并将其传入 ModelConfig 的 tools 字段。此外增加了 top_k 默认值(40),并使 max_tokens 的默认值通过 AgentLLMConfig 的配置链路解析。

结果持久化:_persist_result 与 _persist_usage

_persist_result() 是 AgentLLMNode 与 SimpleLLMNode 最关键的差异点。它不仅回写 messages 和 final_answer,还:

  1. 从 response_msg.tool_calls 中通过 extract_tool_call_from_dict() 提取标准化的 ToolCall 列表
  2. 写入 actions 端口(供 ToolExecutor 消费)
  3. 递增 rounds 计数器
  4. 仅在 没有 tool calls 时将 done 设为 True(有工具调用意味着还需要继续循环)
  5. 仅在 没有 tool calls 时填充 final_answer(有工具调用时 final_answer 留空等待后续轮次)

_persist_usage() 则从响应消息的 metadata 中提取 usage 字典,按 key 将数值累加到状态的 llm.usage 字段中,实现跨轮次的 token 统计。

两节点关键差异对比

维度 SimpleLLMNode AgentLLMNode
工具 Schema 不暴露 通过 tools_schema 暴露
上下文预设 for_chat() for_react()(含协议+工具描述)
done 信号 始终 True 仅在无 tool_calls 时为 True
actions 回写 无 解析 tool_calls → ToolCall 列表
rounds 递增 无 每次调用 +1
Token 用量累积 无 累加到 llm.usage
top_k 配置 不设置 默认 40
final_answer 策略 始终写入 仅无工具调用时写入

ContextBuilder:可组合上下文装配流水线

两个 LLM 节点都依赖 ContextBuilder 完成消息装配。ContextBuilder 采用 Provider → Filter 两阶段流水线架构:Provider 负责从不同来源产生消息片段,Filter 对聚合后的消息序列进行后处理(截断、token 预算控制等)。

Provider 类型

Provider 职责 使用的预设
StaticPromptProvider 注入静态系统提示词(persona、ReAct 协议) chat / react / skill
ToolsDescProvider 注入格式化后的工具描述块 react
StateHistoryProvider 从状态中加载对话历史并归一化语义标注 chat / react / skill
BootstrapFilesProvider 加载工作空间引导文件(如 CLAUDE.md) chat / react / skill
SkillSetPromptProvider 渲染已选技能集的提示词(摘要/完整模式) skill

StateHistoryProvider 在加载历史消息时还执行语义标注:它将之前的 assistant 工具调用计划标注为 semantic_kind=self_tool_plan,将 tool 消息标注为 semantic_kind=tool_observation,并前置一条 get_history_semantics_prompt() 的系统消息来告知模型这些标注的含义,防止模型将历史中的工具交互误认为新的用户消息。

Filter 类型

Filter 职责 关键参数
MaxMessagesFilter 限制消息总数,保留前导 pinned 角色 max_messages, pinned_roles(默认 ("system",))
TokenBudgetFilter 按 token 预算截断,支持三种计数器后端 max_tokens, token_counter, trim_from_start

TokenBudgetFilter 支持三种 token 计数器:SimpleTokenCounter(基于字符数 / 4 估算)、TiktokenTokenCounter(使用 OpenAI tiktoken 库精确计算)和 HFTokenCounter(使用 HuggingFace tokenizer)。默认使用 SimpleTokenCounter,可在构建时替换。

resolve_for_llm:上下文构建器的五级解析

ContextBuilder.resolve_for_llm() 是连接 LLM 节点与上下文装配的核心枢纽。它的解析优先级如下:

flowchart TD
    A[resolve_for_llm] --> B{explicit_builder?}
    B -->|是| C[返回 explicit_builder]
    B -->|否| D{runtime_prompt.context_builder?}
    D -->|是| E[返回 runtime builder]
    D -->|否| F{allow_skill_context AND 请求 skill-aware prompt?}
    F -->|是| G[for_skill(...)]
    F -->|否| H{prefer_react OR tools_desc?}
    H -->|是| I[for_react(...)]
    H -->|否| J[for_chat(...)]

SimpleLLMNode 调用时不传 prefer_react 且 tools_desc 为空,因此落到 for_chat;AgentLLMNode 传入 prefer_react=True 和 tools_desc,因此匹配 for_react。如果运行时 PromptRuntimeContext 中包含技能目录配置且调用方启用了 allow_skill_context,则可能走到 for_skill 分支——这为 SkillNode 等高级场景提供了扩展入口。

ModelClient:统一调用外观与弹性策略

两个 LLM 节点都通过 ModelClient.invoke() 发起实际的模型调用,而非直接与 Provider 交互。ModelClient 作为外观层提供了:

  1. Provider 自动解析:按 ["openai", "litellm"] 优先级检测环境变量中的 API Key,自动选择可用 Provider
  2. 重试与 Fallback:可重试错误(超时、限流、5xx)自动重试,重试耗尽后尝试同 Provider 家族的 fallback 模型
  3. 流式/非流式统一:invoke() 内部根据 stream 参数分发到 consume_stream() 或 provider.generate(),返回统一的 (Message, str) 元组

调用流程图

sequenceDiagram
    participant Node as SimpleLLMNode / AgentLLMNode
    participant MC as ModelClient
    participant IM as invoke_model
    participant P as Provider (OpenAI/LiteLLM)
    participant CS as consume_stream

    Node->>Node: _prepare_messages()
    Node->>Node: _resolve_context_builder().build()
    Node->>MC: invoke(messages, config, stream, on_text_update)
    MC->>MC: 重试循环 (max_retries + 1)
    alt stream = True
        MC->>IM: invoke_model(stream=True)
        IM->>P: stream(messages, config)
        P-->>CS: async for chunk
        CS-->>Node: on_text_update(accumulated_text)
        CS-->>IM: (response_msg, content)
    else stream = False
        MC->>IM: invoke_model(stream=False)
        IM->>P: generate(messages, config)
        P-->>IM: response_msg
    end
    IM-->>MC: (response_msg, content)
    MC-->>MC: 失败? → fallback 模型重试
    MC-->>Node: (response_msg, content)
    Node->>Node: _persist_result / write_port

重试分类与 Fallback 策略

ModelClient 将错误分为可重试和不可重试两类。可重试错误包括 TimeoutError、ConnectionError 以及错误消息中包含 timeout、rate limit、429、5xx 等关键词的异常。不可重试错误(如 4xx 参数错误、模型不存在)直接抛出,不浪费重试次数。Fallback 机制仅在主模型与 fallback 模型属于同一 Provider 家族时触发(通过 _model_provider_family() 判断),避免跨 Provider 的配置不兼容。

流式处理:tool call 增量合并

流式模式下,consume_stream() 不仅累积文本增量,还处理流式工具调用:通过 merge_stream_tool_calls() 按 index 将分散在多个 chunk 中的 tool name 和 arguments 片段合并,最终由 finalize_stream_tool_calls() 去重并产出规范的 tool_calls 列表。

配置体系:AgentLLMConfig 与 StateKeys

AgentLLMConfig 是 Pydantic BaseModel,其每个字段的默认值都通过 get_config() 从 jianmu.yaml 中延迟解析。这意味着用户无需在代码中显式设置 model、temperature 等参数——只需在 jianmu.yaml 中配置一次即可全局生效。

配置字段解析表

字段 默认值来源 说明
model get_config().models.default 默认模型标识符
temperature get_config().llm.temperature 采样温度
max_tokens get_config().llm.max_tokens 最大生成 token 数
top_p get_config().llm.top_p Nucleus 采样阈值
top_k get_config().llm.top_k Top-k 采样(仅 AgentLLMNode 使用)
timeout get_config().llm.timeout 请求超时(秒)
stream False(硬编码) 是否启用流式输出
max_budget_tokens get_config().llm.max_budget_tokens 跨轮次 token 预算上限
resilience.max_retries get_config().llm.max_retries 模型调用最大重试次数
resilience.fallback_model get_config().models.fallback 重试耗尽后的 fallback 模型
system_prompt None 覆盖默认 persona 提示词
keys StateKeys() 状态字段名路由配置

StateKeys 定义了 LLM 节点、ToolExecutor、SkillNode 与 ReAct 预设之间共享的状态字段命名约定:

Key 默认值 存储内容
messages "messages" 对话历史(Message 列表)
rounds "rounds" 当前 Agent 轮次计数
actions "actions" 待执行的 ToolCall 列表
done "done" 循环终止标志
text_output "text_output" 最近一次纯文本节点输出
final_answer "final_answer" 最终回答文本
skill_result "skill_result" 结构化 skill 输出桥接槽
usage "llm.usage" 累积 token 用量
streaming_output "streaming_output" 实时流式输出文本
tool_effects "tool_effects" 工具副作用记录

用户可通过自定义 StateKeys 实例来重命名这些字段。当前公开 API 里,这通常通过 AgentLLMConfig.keys、ReActConfig.keys、SkillNodeConfig.keys、ToolExecutorConfig.keys 这类每节点配置对象完成,使状态路由契约集中且显式。

ReAct 预设中 AgentLLMNode 的组装模式

create_react_node() 展示了 AgentLLMNode 的标准使用模式。它将 AgentLLMNode、ToolExecutor 和 StateCondition 组合为一个 LoopUntilSuccess 子树:

flowchart LR
    subgraph LoopUntilSuccess["LoopUntilSuccess (max_iterations)"]
        direction TB
        A[AgentLLMNode] --> B[ToolExecutor]
        B --> C[StateCondition: done?]
        C -->|否| A
    end
    C -->|是| D[返回 final_answer]

关键设计细节: - AgentLLMNode 将 tools_schema 和 tools_description 传入,确保模型知晓可用工具 - AgentLLMNode 与 ToolExecutor 共享同一套 StateKeys,通过状态字段隐式通信 - LoopUntilSuccess 在每次循环前检查 token 预算(abort_condition),超出则提前终止 - StateCondition 检查 done 字段:当 AgentLLMNode 判定无工具调用时设为 True,循环终止

关键设计要点

上下文构建器的可替换性。两个 LLM 节点都通过 _resolve_context_builder() 在每次 update_async() 调用时动态解析上下文构建器。这意味着即使节点已创建,后续通过 inject() 注入新的 RunContext(包含不同的 PromptRuntimeContext)也能切换上下文装配策略——测试中验证了这一行为:同一节点在两次 inject() 后分别使用了不同的 builder。

技能上下文的隐式忽略。SimpleLLMNode 和 AgentLLMNode 默认不启用技能感知上下文(allow_skill_context=False)。即使 PromptRuntimeContext 中配置了技能目录和 include_skills_summary=True,也不会触发 for_skill 预设。技能上下文仅由 SkillNode 等明确开启 allow_skill_context=True 的调用方触发。

ModelClient 的类型守卫。_require_model_client() 确保传入的 model_client 是 Jianmu 的 ModelClient 包装类型而非裸 Provider。这是架构分层的关键约束——LLM 节点不应直接依赖 Provider 接口,所有调用必须经过 ModelClient.invoke() 以获得重试、fallback 和可观测性。

阅读下一步