已执行调用追踪
模型发出工具调用后,Codex 除了把工具结果写回历史,还可以在下一次模型请求中附带一份 executed_tool_calls 元数据。它回答的是“运行时尝试调度过哪些调用”,不是“哪些调用已经成功完成”。记录动作发生在 handler 之前,所以未知工具、参数错误或被 handler 拒绝的调用也可能出现;反过来,取消、压缩或 Code Mode cell 在 yielded 后没有再次 wait,也可能让尚未绑定的记录消失。
本文面向已经读过ToolOrchestrator执行流程、 并行结果排序和ToolLifecycle事件的读者。前两篇解释工具如何 执行以及输出怎样进入历史,后一篇解释 extension lifecycle;本文只研究 session 内的 attempted-tool metadata:谁记录、 怎样与 output 配对、重试时如何复用、超出预算后保留什么。读完后,读者应能从一个 output 的 call_id 追到元数据来源, 并判断某条记录能否作为“工具成功”的依据。
1. 元数据边界
1.1 三种事实
同一次工具调用会留下三类容易混淆的事实。Tool output 是模型可读的业务结果;lifecycle event 面向 extension,描述 start、finish 或 aborted;executed_tool_calls 则是下一次 Responses 请求中的内部附加字段。三者的生产时机和可靠性 不同,不能互相替代。
| 事实 | 产生位置 | 主要消费者 | 能否证明成功 |
|---|---|---|---|
| Tool output | handler 或错误适配层结束后 | 模型与会话历史 | 需结合 output success/错误类型判断 |
| Lifecycle event | Registry 接受调用后的 start/finish 路径 | extension contributor | finish outcome 才能表达终态 |
| Executed calls | dispatch runtime 入口 | 下一次模型请求 | 不能,只表达尝试过的调用 |
图中的关键分叉在 Dispatch入口:记录器与 handler 从这里走向不同路径。即使 handler 随后失败,记录器也已经见过 调用;即使 handler 成功,pending 记录也必须等相应 output 出现在 prompt 输入中才能被附加。
1.2 Feature门槛
记录器是 Session service。Session 创建时只有 ExecutedToolCallMetadata feature 已启用才构造它;默认关闭时,后续 dispatch 和 prompt 构造都拿不到 recorder,因此不会产生该字段。
源码位置:codex-rs/core/src/session/session.rs :: Session::new
let executed_tool_calls = config
.features
.enabled(Feature::ExecutedToolCallMetadata)
.then(|| Arc::new(crate::state::ExecutedToolCallRecorder::default()));
let services = SessionServices {
// ...
executed_tool_calls,
code_mode_service: crate::tools::code_mode::CodeModeService::new(
Arc::clone(&code_mode_session_provider),
&config.features,
),
// ...
};所有者是 Session,而不是单个 Turn。这样同一 session 的后续 prompt 可以重新绑定历史 output;代价是记录器只提供 best-effort 元数据,不承担持久化执行账本的职责。
2. 状态所有权
2.1 记录器结构
ExecutedToolCallRecorder 用一个 std::sync::Mutex 保护五组状态。锁内没有异步等待,临界区只做 HashMap、Vec 与字节 计数,因此使用同步 mutex;调用者不会在持锁时进入 handler 或发网络请求。
源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorderState
pub(crate) struct ExecutedToolCallRecorder {
state: std::sync::Mutex<ExecutedToolCallRecorderState>,
}
struct ExecutedToolCallRecorderState {
direct_calls: HashMap<String, ExecutedToolCall>,
cells: HashMap<CellId, RecordedCell>,
output_cells: HashMap<String, CellId>,
retained_calls: HashMap<
(std::mem::Discriminant<ResponseItem>, String),
Vec<ExecutedToolCall>,
>,
pending_nested_calls: usize,
}
struct RecordedCell {
pending_calls: Vec<ExecutedToolCall>,
pending_full_argument_bytes: usize,
}direct_calls 用 output call_id 直接寻址;Code Mode 的 nested 调用没有自己的外层 Responses output,所以先按 CellId 聚合到 cells,再用 output_cells 把外层 exec/wait output 映射到该 cell。retained_calls 和调用方 提供的 retry cache 则负责重复构造 prompt 时恢复已经消费过的记录。
2.2 协议形状
单条记录只有工具名与参数。参数采用 untagged enum:正常情况保存原始 JSON,超限时换成本地生成的 truncation 对象。executed_tool_calls 位于 internal_chat_message_metadata_passthrough,并被排除在公开 app-server schema 和 TypeScript 类型之外。
相关源码:
codex-rs/protocol/src/models/executed_tool_calls.rs :: ExecutedToolCallcodex-rs/protocol/src/models.rs :: InternalChatMessageMetadataPassthrough
#[serde(untagged)]
pub enum ExecutedToolCallArguments {
Raw(serde_json::Value),
#[serde(skip_deserializing)]
Truncated {
#[serde(rename = "_codex_executed_tool_call_truncated")]
truncation: ExecutedToolCallTruncation,
},
}
pub struct ExecutedToolCall {
pub name: String,
arguments: ExecutedToolCallArguments,
}
pub struct ExecutedToolCallTruncation {
original_bytes: usize,
max_bytes: usize,
omitted_calls: Option<usize>,
original_name_bytes: Option<usize>,
}源码位置:codex-rs/protocol/src/models.rs :: InternalChatMessageMetadataPassthrough
pub struct InternalChatMessageMetadataPassthrough {
pub turn_id: Option<String>,
#[serde(default, skip_deserializing, skip_serializing_if = "Option::is_none")]
#[schemars(skip)]
#[ts(skip)]
pub executed_tool_calls: Option<Vec<ExecutedToolCall>>,
}skip_deserializing 建立了信任边界:provider 或 rollout 输入中的同名字段不会被当成本地 attempted-tool metadata 恢复。记录只能由当前进程的运行时生成,再序列化进发往支持方的请求。
3. 调度记录
3.1 记录时机
ToolCallRuntime::handle_tool_call_with_source 在查询 router、等待 runtime readiness、取得并行锁和执行 handler 之前 调用 recorder。feature 与 recorder 必须同时存在;这段顺序决定了“executed”在这里更接近 attempted,而不是 completed。
源码位置:codex-rs/core/src/tools/parallel.rs :: ToolCallRuntime::handle_tool_call_with_source
if self
.step_context
.turn
.config
.features
.enabled(codex_features::Feature::ExecutedToolCallMetadata)
&& let Some(executed_tool_calls) = self.session.services.executed_tool_calls.as_ref()
{
executed_tool_calls.record_tool_call(
&call,
&source,
super::effective_tool_mode(&self.step_context.turn),
);
}
let router = &self.step_context.tool_router;
let supports_parallel = router.tool_supports_parallel(&call);
let tool_runtime = router.tool_runtime(&call);这解释了一个看似反常的测试结果:nested exec_command 参数无效并抛出异常时,外层脚本可以捕获错误,但下一次 请求仍包含该调用元数据。元数据说明 dispatch runtime 收到过它,不说明 handler 返回了成功结果。
3.2 参数归一化
记录器先按 payload 类型计算序列化长度。Function 参数是 JSON 文本,未超过 8 KiB 时尝试解析;非法 JSON 不会让 记录失败,而是作为普通字符串保存。Custom input 始终是字符串,ToolSearch 参数则序列化为 JSON 值。
源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::record_tool_call
let original_bytes = match &call.payload {
ToolPayload::Function { arguments } => arguments.len(),
ToolPayload::Custom { input } => serialized_json_bytes(input),
ToolPayload::ToolSearch { arguments } => serialized_json_bytes(arguments),
};
let name = codex_tools::code_mode_name_for_tool_name(&call.tool_name);
let recorded_call = if original_bytes > MAX_EXECUTED_TOOL_CALL_ARGUMENT_BYTES {
ExecutedToolCall::truncated(
name,
original_bytes,
MAX_EXECUTED_TOOL_CALL_ARGUMENT_BYTES,
)
} else {
let arguments = match &call.payload {
ToolPayload::Function { arguments } => serde_json::from_str(arguments)
.unwrap_or_else(|_| JsonValue::String(arguments.clone())),
ToolPayload::Custom { input } => JsonValue::String(input.clone()),
ToolPayload::ToolSearch { arguments } => {
serde_json::to_value(arguments).unwrap_or_default()
}
};
ExecutedToolCall::new(name, arguments)
};工具名还会经过 code_mode_name_for_tool_name,把 namespace 形式转换为 Code Mode 可见名称。因此测试中的 test_namespace::unsupported_tool 被记录为 test_namespace__unsupported_tool,它与原始 Responses call 的 namespace/name 字段不是同一种 wire shape。
3.3 保留字防伪
模型参数本身可能恰好包含 _codex_executed_tool_call_truncated。若直接保存,消费者无法区分“模型给出的普通参数”与 “本地生成的截断标记”。构造函数因此把这类原始对象再包一层,只有 set_truncation 能产生受信任的 truncation 变体。
源码位置:codex-rs/protocol/src/models/executed_tool_calls.rs :: ExecutedToolCall::new
pub fn new(name: String, arguments: serde_json::Value) -> Self {
let arguments = if arguments
.as_object()
.is_some_and(|object| object.contains_key("_codex_executed_tool_call_truncated"))
{
serde_json::json!({ "_codex_executed_tool_call_raw": arguments })
} else {
arguments
};
Self {
name,
arguments: ExecutedToolCallArguments::Raw(arguments),
}
}4. 直接调用
Direct 与 DirectPlaintextMessage 都按 call_id 存入 direct_calls。相同 id 只保留首次记录;达到 256 条后, 第 257 个不同 id 不再保存完整参数,而是保存 max_bytes: 0 的 overflow marker,之后的新增 id 被忽略。
源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::record_tool_call
match source {
ToolCallSource::Direct | ToolCallSource::DirectPlaintextMessage => {
let mut state = self
.state
.lock()
.unwrap_or_else(std::sync::PoisonError::into_inner);
if state.direct_calls.len() < MAX_PENDING_EXECUTED_TOOL_CALLS {
state
.direct_calls
.entry(call.call_id.clone())
.or_insert(recorded_call);
} else if state.direct_calls.len() == MAX_PENDING_EXECUTED_TOOL_CALLS
&& !state.direct_calls.contains_key(&call.call_id)
{
state.direct_calls.insert(
call.call_id.clone(),
ExecutedToolCall::truncated(
recorded_call.name,
original_bytes,
0,
),
);
}
}
ToolCallSource::CodeMode { cell_id, .. } => {
self.record_nested_tool_call(
CellId::new(cell_id.clone()),
recorded_call,
original_bytes,
);
}
}这个上限控制的是尚未与 output 配对的 session 状态,避免模型连续生成大量无法闭合的调用时无限占用内存。overflow marker 仍保留工具名和原始参数字节数,让消费者知道“还有一次调用,但参数内容未保存”。
5. Code Mode配对
5.1 外壳过滤
在有效工具模式确实是 Code Mode 或 CodeModeOnly 时,模型直接调用的 exec 与 wait 是承载脚本 cell 的外壳。真正需要记录的是脚本 内部通过 tools.* 发起的 nested calls;若连外壳也记录,外层 output 会同时出现“exec 调用了自己”和 nested calls。 因此 recorder 在特定模式下跳过默认 namespace 的直接 exec/wait。
源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::record_tool_call
if matches!(source, ToolCallSource::Direct)
&& matches!(tool_mode, ToolMode::CodeMode | ToolMode::CodeModeOnly)
&& call.tool_name.is_default_namespace()
&& matches!(
(call.tool_name.name.as_str(), &call.payload),
(
crate::tools::code_mode::PUBLIC_TOOL_NAME,
ToolPayload::Custom { .. }
) | (
crate::tools::code_mode::WAIT_TOOL_NAME,
ToolPayload::Function { .. }
)
)
{
return;
}这条过滤只针对 Code Mode 生效。Direct 模式下名为 exec 的 custom call 仍会记录,因为此时它不是 cell 外壳,而是 一次普通的直接调用尝试。
工具 feature 开启并不自动保证有效模式仍是 Code Mode。effective_tool_mode 会在 Code Mode provider 不可用且允许 in-process fallback 时返回 Direct;此时外层 exec 不会走上面的过滤分支,后续 metadata 看到的是直接调用形状。阅读或 调试 Code Mode 记录时,必须先确认这个模式选择。
源码位置:codex-rs/core/src/tools/mod.rs :: effective_tool_mode
let requested_tool_mode = requested_tool_mode(turn_context);
if !turn_context.code_mode_available
&& requested_tool_mode == ToolMode::CodeMode
&& !turn_context.config.code_mode.disable_in_process_fallback
{
ToolMode::Direct
} else {
requested_tool_mode
}5.2 Cell映射
Code Mode nested call 先按 CellId 进入 cells。外层 exec 启动 cell 后立即登记 cell_id → call_id;后续 wait 收到 live cell 响应时也登记当前 wait output,使最后一次可见 output 能取回该 cell 的 pending calls。
相关源码:
codex-rs/core/src/tools/code_mode/execute_handler.rs :: ExecuteHandler::handlecodex-rs/core/src/tools/code_mode/wait_handler.rs :: WaitHandler::handle
if let Some(executed_tool_calls) = exec.session.services.executed_tool_calls.as_ref() {
executed_tool_calls.register_cell(&cell_id, &call_id);
}源码位置:codex-rs/core/src/tools/code_mode/wait_handler.rs :: WaitHandler::handle
if let codex_code_mode::WaitOutcome::LiveCell(response) = &wait_response {
let runtime_cell_id = match response {
codex_code_mode::RuntimeResponse::Yielded { cell_id, .. }
| codex_code_mode::RuntimeResponse::Terminated { cell_id, .. }
| codex_code_mode::RuntimeResponse::Result { cell_id, .. } => cell_id,
};
if let Some(executed_tool_calls) =
exec.session.services.executed_tool_calls.as_ref()
{
executed_tool_calls.register_cell(runtime_cell_id, &call_id);
}
}5.3 Cell预算
每个 cell 最多保留 32 KiB 完整参数,单次调用最多 8 KiB;全 session 最多有 256 个 pending nested calls 和 256 个 cell/output mappings。达到调用数上限时仍写入一个零预算 marker,之后的新 cell 或新调用则被丢弃。
源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::record_nested_tool_call
let at_pending_call_limit =
state.pending_nested_calls == MAX_PENDING_EXECUTED_TOOL_CALLS;
let cell = state.cells.entry(cell_id).or_default();
let max_bytes = MAX_EXECUTED_TOOL_CALL_ARGUMENT_BYTES.min(
MAX_EXECUTED_TOOL_CALL_FULL_ARGUMENT_BYTES_PER_OUTPUT
.saturating_sub(cell.pending_full_argument_bytes),
);
let call = if at_pending_call_limit {
ExecutedToolCall::truncated(call.name, original_bytes, 0)
} else if original_bytes <= max_bytes {
cell.pending_full_argument_bytes = cell
.pending_full_argument_bytes
.saturating_add(original_bytes);
call
} else {
ExecutedToolCall::truncated(call.name, original_bytes, max_bytes)
};
cell.pending_calls.push(call);
state.pending_nested_calls += 1;6. Prompt绑定
6.1 绑定入口
每次 sampling request 构造 prompt 前,Turn loop 从历史生成 prompt_input,再让 recorder 附加 pending calls。只有 确实附加过元数据时才运行请求级预算函数。retry cache 在整个 sampling 重试循环之外创建,因此网络重试重新构造 prompt 时仍能取回同一份 calls。
源码位置:codex-rs/core/src/session/turn.rs :: run_sampling_request
let mut executed_tool_calls_by_output = HashMap::new();
loop {
let prompt_input = if let Some(input) = initial_input.take() {
input
} else {
sess.clone_history()
.await
.for_prompt(&turn_context.model_info.input_modalities)
};
let mut prompt_input = prompt_input;
if let Some(executed_tool_calls) = sess.services.executed_tool_calls.as_ref()
&& executed_tool_calls.attach_pending_to_prompt(
&mut prompt_input,
&mut executed_tool_calls_by_output,
)
{
codex_protocol::models::bound_executed_tool_calls_for_prompt(
&mut prompt_input,
);
}
let prompt = build_prompt(
prompt_input,
router.as_ref(),
turn_context.as_ref(),
base_instructions.clone(),
);
// ...
}生效时机因此是“下一次模型采样请求构造时”,不是工具调用刚发生时。元数据附在历史中的 output item 上,再随 build_prompt 进入请求;它不修改 handler 已经生成的 output body。
6.2 Output键
绑定器从 items 尾部向前扫描,只接受 Function、Custom 和带 call id 的 ToolSearch output。key 同时包含 ResponseItem discriminant 与 call_id,避免不同 output 类型恰好复用同一字符串 id 时误配。
源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::attach_pending_to_prompt
for item in items.iter_mut().rev() {
let call_id = match &*item {
ResponseItem::FunctionCallOutput { call_id, .. }
| ResponseItem::CustomToolCallOutput { call_id, .. }
| ResponseItem::ToolSearchOutput {
call_id: Some(call_id),
..
} => call_id,
_ => continue,
};
let key = (std::mem::discriminant(&*item), call_id.clone());
// ...
}反向扫描让重复出现的同类 output 优先绑定最近一项。pending_retry_outputs 与 pending_retained_outputs 还保证同一个 key 在一次 prompt 构造中只附加一次。
6.3 来源优先级
同一个 output key 的调用记录按四级来源选择:本轮 retry cache、session retained history、尚未消费的 direct call、 通过 output-cell 映射取得的 nested calls。新取出的 direct/nested calls 会同时写入 retry cache 和 retained map。
源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::attach_pending_to_prompt
let calls = if let Some(cached) = retry_cache.get(&key) {
if !pending_retry_outputs.remove(&key) {
continue;
}
pending_retained_outputs.remove(&key);
cached.clone()
} else if let Some(retained) = state.retained_calls.get(&key) {
if !pending_retained_outputs.remove(&key) {
continue;
}
retained.clone()
} else {
let mut calls = state
.direct_calls
.remove(call_id)
.into_iter()
.collect::<Vec<_>>();
if let Some(cell_id) = state.output_cells.remove(call_id)
&& let Some(mut cell) = state.cells.remove(&cell_id)
{
state.pending_nested_calls = state
.pending_nested_calls
.saturating_sub(cell.pending_calls.len());
state
.output_cells
.retain(|_, output_cell_id| output_cell_id != &cell_id);
calls.append(&mut cell.pending_calls);
}
if calls.is_empty() {
continue;
}
retry_cache.insert(key.clone(), calls.clone());
state.retained_calls.insert(key, calls.clone());
calls
};
item.append_executed_tool_calls(calls);成功绑定 cell 后,recorder 会减少 pending_nested_calls、删除 cell,并清掉指向同一 cell 的其他 output mapping。 这是一种一次性归属:nested calls 被投影到匹配到的 output 后,不应再从 pending 状态重复取出。
7. 预算与省略
7.1 三层预算
这套元数据有三个不同粒度的上限,不能把它们统称为“8 KiB 截断”。
| 层级 | 上限 | 超限行为 |
|---|---|---|
| 单个调用参数 | 8 KiB | 参数替换为 truncation marker |
| 单个 Code Mode cell 的完整参数 | 32 KiB | 后续调用按剩余空间截断 |
| 整个请求的 attempted-tool metadata | 32 KiB | 公平分配或优先保留最近记录,并报告 omissions |
| Session pending 调用/映射 | 256 | 保留一个零预算 overflow marker,之后丢弃新增项 |
请求级函数先计算 metadata 在 wire 上的精确 JSON 字节数,包括字段名和外围对象开销,而不是只累加参数字符串长度。
源码位置:codex-rs/protocol/src/models/executed_tool_calls.rs :: executed_tool_call_metadata_bytes
pub fn executed_tool_call_metadata_bytes(item: &ResponseItem) -> usize {
let Some(metadata) = item.executed_tool_call_metadata() else {
return 0;
};
let Some(calls) = metadata
.executed_tool_calls
.as_ref()
.filter(|calls| !calls.is_empty())
else {
return 0;
};
serde_json::to_vec(calls)
.map(|bytes| {
bytes
.len()
.saturating_add(executed_tool_call_metadata_field_bytes(metadata))
})
.unwrap_or(usize::MAX)
}7.2 公平与最近优先
发送请求前调用 bound_executed_tool_calls_for_prompt,按剩余 item 数公平分配预算;recorder 整理 retained history 时 使用 bound_executed_tool_calls_for_prompt_prioritizing_recent,先反转 items,让新记录先消耗预算,再恢复顺序。
源码位置:codex-rs/protocol/src/models/executed_tool_calls.rs :: bound_executed_tool_calls_for_prompt
pub fn bound_executed_tool_calls_for_prompt(items: &mut [ResponseItem]) {
bound_executed_tool_calls_for_prompt_with_priority(items, false);
}
pub fn bound_executed_tool_calls_for_prompt_prioritizing_recent(
items: &mut [ResponseItem],
) {
items.reverse();
bound_executed_tool_calls_for_prompt_with_priority(items, true);
items.reverse();
}当记录被整体移除时,算法把缺失数量合并进某条仍保留记录的 omitted_calls。若预算小到一条都放不下,则尽量保留 一个截断后的 fallback call,并把其余数量写入 marker。消费者因此可以区分“历史原本只有这些调用”和“还有若干调用 因预算未展开”。
8. 失败与恢复
8.1 失败调用
记录发生在 handler 之前,所以“调用失败”不是清除 metadata 的条件。对于真正进入 ToolCallSource::CodeMode 的 nested 调用,调用来源会把记录放入 cell;对于 fallback 到 Direct 的 turn,外层 exec 本身会按 direct call 记录。两种结果都 符合“dispatch 已尝试”的定义,但不能从字段名推导 handler 成功。
源码位置:codex-rs/core/src/tools/code_mode/mod.rs :: call_nested_tool
let result = tool_runtime
.handle_tool_call_with_source(
call,
ToolCallSource::CodeMode {
cell_id: cell_id.to_string(),
runtime_tool_call_id,
},
cancellation_token,
)
.await?;
Ok(result.code_mode_result())端到端测试还会受到 Code Mode host 是否可用、effective_tool_mode 是否回退的影响;因此测试输入必须同时说明工具模式, 否则看到外层 exec 记录时,不能误判为 nested 绑定器失效。无论哪条路径,副作用是否回滚仍需检查 handler 结果和真实 外部状态。
8.2 Retry恢复
第一次绑定时,direct/cell pending state 被消费,但同一 calls 会写入 sampling loop 的 retry cache。若 provider 请求失败并 重试,新的 prompt_input 可以从 cache 再附加同一记录。Session 的 retained map 解决更长周期的重复 prompt 构造: 只要相应 output 仍在输入中,就能恢复 metadata;output 因压缩不再出现时,未匹配 retained key 会被清理。
源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::attach_pending_to_prompt
if !pending_retained_outputs.is_empty() {
state
.retained_calls
.retain(|key, _| !pending_retained_outputs.contains(key));
}
if retained_bytes > MAX_EXECUTED_TOOL_CALL_FULL_ARGUMENT_BYTES_PER_OUTPUT {
bound_executed_tool_calls_for_prompt_prioritizing_recent(items);
state.retained_calls.clear();
for item in items {
// 从已经完成预算收束的 items 重建 retained_calls
// ...
}
}8.3 Best-effort边界
源码注释明确把 recorder 定义为 session-scoped best-effort attempted-tool metadata。以下情况不能外推为可靠日志:
- 调用在 output 出现前被取消,pending state 可能没有机会绑定;
- 历史压缩移除了 output,retained metadata 会随未匹配 key 清理;
- Code Mode cell yield 后没有再次 wait,nested calls 可能一直没有对应 output;
- 超过调用数或字节预算的内容只留下 marker,甚至在极端情况下被省略;
- 进程退出后 recorder 内存状态不会成为持久化恢复依据。
因此排查“工具究竟有没有成功”时,应回到 Tool output、lifecycle outcome、handler 日志或真实外部状态;本字段适合告诉 下一次模型请求“运行时近期尝试过什么”,不适合作为可靠执行日志或 exactly-once 证明。
9. 测试路径
9.1 Pending上限
单元测试分别制造 258 个 direct calls、258 个 nested calls 和 258 组 cell mapping。关键断言是 direct 与 nested 都 保留 256 条完整记录加 1 条 overflow marker,marker 的 max_bytes 为 0;绑定 cell output 后,pending 计数归零、 cell 被删除,retry cache 留下可重放副本。
源码位置:codex-rs/core/src/tools/executed_tool_calls_tests.rs :: executed_tool_call_recorder_bounds_pending_calls_and_preserves_overflow
assert_eq!(
state.direct_calls.len(),
MAX_PENDING_EXECUTED_TOOL_CALLS + 1,
);
assert_eq!(
state.pending_nested_calls,
MAX_PENDING_EXECUTED_TOOL_CALLS + 1,
);
assert_eq!(state.cells.len(), MAX_PENDING_EXECUTED_TOOL_CALLS);
assert_eq!(state.output_cells.len(), MAX_PENDING_EXECUTED_TOOL_CALLS);
recorder.attach_pending_to_prompt(&mut items, &mut retry_cache);
let calls = items[0]
.executed_tool_call_metadata()
.and_then(|metadata| metadata.executed_tool_calls.as_ref())
.expect("bounded nested calls must attach to their own output");
assert_eq!(calls.len(), MAX_PENDING_EXECUTED_TOOL_CALLS + 1);
assert_eq!(retry_cache.len(), 1);测试还用空 prompt 再次调用绑定器,断言没有匹配 output 后 retained_calls 被清空。它覆盖容量、绑定、重放和清理, 但没有模拟 Tokio task 取消或真实进程重启。
9.2 Retained省略
另一个单元测试连续建立 512 个约 1 KiB 参数的历史 output。每次把最新调用附到 prompt 并执行请求预算收束,最后断言 retained serialized bytes 不超过 32 KiB,最新调用仍在,所有保留条目数与 marker 中的 omitted_calls 之和等于 原始 512 次调用。
源码位置:codex-rs/core/src/tools/executed_tool_calls_tests.rs :: executed_tool_call_recorder_bounds_retained_history_and_reports_omissions
assert!(
retained_bytes
<= MAX_EXECUTED_TOOL_CALL_FULL_ARGUMENT_BYTES_PER_OUTPUT
);
let omitted_calls = metadata
.iter()
.filter_map(|call| {
call["arguments"]
["_codex_executed_tool_call_truncated"]
["omitted_calls"]
.as_u64()
})
.sum::<u64>();
assert!(omitted_calls > 0);
assert_eq!(metadata.len() as u64 + omitted_calls, 512);这里证明 recent prioritization 与 omissions 计数守恒,不证明模型会怎样解释 marker,也不证明 32 KiB 是任意 provider 都接受的公共协议上限。
9.3 请求投影
集成测试先关闭 feature,断言 custom tool output 中没有 executed_tool_calls;再开启 feature,模拟 namespaced custom call、超长转义字符串和 direct exec。测试检查下一次 mock Responses 请求的完整 JSON,确认 namespace 转换、8 KiB truncation 与跨请求 retained metadata。Code Mode nested 端到端测试还需要可用的 Code Mode host;若 effective_tool_mode 回退到 Direct,应按 Direct 记录路径解释结果。
源码位置:codex-rs/core/tests/suite/tools.rs :: namespaced_custom_tool_call_preserves_namespace_through_dispatch_and_replay
assert_eq!(
escaped_request.custom_tool_call_output(escaped_call_id)
["internal_chat_message_metadata_passthrough"]
["executed_tool_calls"],
json!([{
"name": format!("{namespace}__{tool_name}"),
"arguments": {
"_codex_executed_tool_call_truncated": {
"original_bytes": serde_json::to_vec(&escaped_input)?.len(),
"max_bytes": 8 * 1024,
},
},
}]),
);这个测试把 recorder、历史 output、prompt 构造和请求序列化连成一条真实路径;它使用 mock server,不覆盖远端模型如何 利用该字段。
10. 阅读实践
在源码仓库根目录运行下面的定向测试,可以先观察容量与 omissions,再观察真实请求 JSON。若要运行 Code Mode nested 端到端测试,还需要先准备可用的 Code Mode host,并确认 effective_tool_mode 没有回退到 Direct。
RUST_MIN_STACK=16777216 cargo test -p codex-core executed_tool_call_recorder
RUST_MIN_STACK=16777216 cargo test -p codex-core namespaced_custom_tool_call_preserves_namespace_through_dispatch_and_replay
RUST_MIN_STACK=16777216 cargo test -p codex-protocol executed_tool_call然后用三个场景检查自己的判断:
- 一个 Direct function call 参数合法,但 handler 返回错误。它会在何时进入
direct_calls,最终附在哪类 output 上? - 一个 Code Mode cell 调用两个 nested tools 后 yield,用户没有再调用 wait。哪些状态可能仍留在
cells,为什么不能 据此声称模型已经看到两次调用? - 历史中有数百个大参数调用。哪些函数负责单调用 8 KiB、cell 32 KiB 和整请求 32 KiB,
omitted_calls应满足什么 数量关系?
能够从 handle_tool_call_with_source 复述到 record_tool_call、从 output key 复述到 attach_pending_to_prompt,再走到 bound_executed_tool_calls_for_prompt,就建立了这套元数据的完整心智模型。若要 继续理解工具结果本身如何转成模型输入,可回看ToolOutput与错误模型;若要区分 attempted metadata 与 extension 可见终态,则回看ToolLifecycle事件。
