Skip to content

已执行调用追踪

追踪工具调用如何在调度入口被记录、与直接输出或 Code Mode cell 配对,并在重试和大小预算约束下投影到下一次模型请求。

基于rust-v0.150.0
CodexRustToolsMetadata

已执行调用追踪 ​

模型发出工具调用后,Codex 除了把工具结果写回历史,还可以在下一次模型请求中附带一份 executed_tool_calls 元数据。它回答的是“运行时尝试调度过哪些调用”,不是“哪些调用已经成功完成”。记录动作发生在 handler 之前,所以未知工具、参数错误或被 handler 拒绝的调用也可能出现;反过来,取消、压缩或 Code Mode cell 在 yielded 后没有再次 wait,也可能让尚未绑定的记录消失。

本文面向已经读过ToolOrchestrator执行流程、 并行结果排序和ToolLifecycle事件的读者。前两篇解释工具如何 执行以及输出怎样进入历史,后一篇解释 extension lifecycle;本文只研究 session 内的 attempted-tool metadata:谁记录、 怎样与 output 配对、重试时如何复用、超出预算后保留什么。读完后,读者应能从一个 output 的 call_id 追到元数据来源, 并判断某条记录能否作为“工具成功”的依据。

1. 元数据边界 ​

1.1 三种事实 ​

同一次工具调用会留下三类容易混淆的事实。Tool output 是模型可读的业务结果;lifecycle event 面向 extension,描述 start、finish 或 aborted;executed_tool_calls 则是下一次 Responses 请求中的内部附加字段。三者的生产时机和可靠性 不同,不能互相替代。

事实产生位置主要消费者能否证明成功
Tool outputhandler 或错误适配层结束后模型与会话历史需结合 output success/错误类型判断
Lifecycle eventRegistry 接受调用后的 start/finish 路径extension contributorfinish outcome 才能表达终态
Executed callsdispatch runtime 入口下一次模型请求不能,只表达尝试过的调用

图中的关键分叉在 Dispatch入口:记录器与 handler 从这里走向不同路径。即使 handler 随后失败,记录器也已经见过 调用;即使 handler 成功,pending 记录也必须等相应 output 出现在 prompt 输入中才能被附加。

1.2 Feature门槛 ​

记录器是 Session service。Session 创建时只有 ExecutedToolCallMetadata feature 已启用才构造它;默认关闭时,后续 dispatch 和 prompt 构造都拿不到 recorder,因此不会产生该字段。

源码位置:codex-rs/core/src/session/session.rs :: Session::new

rust
let executed_tool_calls = config
    .features
    .enabled(Feature::ExecutedToolCallMetadata)
    .then(|| Arc::new(crate::state::ExecutedToolCallRecorder::default()));

let services = SessionServices {
    // ...
    executed_tool_calls,
    code_mode_service: crate::tools::code_mode::CodeModeService::new(
        Arc::clone(&code_mode_session_provider),
        &config.features,
    ),
    // ...
};

所有者是 Session,而不是单个 Turn。这样同一 session 的后续 prompt 可以重新绑定历史 output;代价是记录器只提供 best-effort 元数据,不承担持久化执行账本的职责。

2. 状态所有权 ​

2.1 记录器结构 ​

ExecutedToolCallRecorder 用一个 std::sync::Mutex 保护五组状态。锁内没有异步等待,临界区只做 HashMap、Vec 与字节 计数,因此使用同步 mutex;调用者不会在持锁时进入 handler 或发网络请求。

源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorderState

rust
pub(crate) struct ExecutedToolCallRecorder {
    state: std::sync::Mutex<ExecutedToolCallRecorderState>,
}

struct ExecutedToolCallRecorderState {
    direct_calls: HashMap<String, ExecutedToolCall>,
    cells: HashMap<CellId, RecordedCell>,
    output_cells: HashMap<String, CellId>,
    retained_calls: HashMap<
        (std::mem::Discriminant<ResponseItem>, String),
        Vec<ExecutedToolCall>,
    >,
    pending_nested_calls: usize,
}

struct RecordedCell {
    pending_calls: Vec<ExecutedToolCall>,
    pending_full_argument_bytes: usize,
}

direct_calls 用 output call_id 直接寻址;Code Mode 的 nested 调用没有自己的外层 Responses output,所以先按 CellId 聚合到 cells,再用 output_cells 把外层 exec/wait output 映射到该 cell。retained_calls 和调用方 提供的 retry cache 则负责重复构造 prompt 时恢复已经消费过的记录。

2.2 协议形状 ​

单条记录只有工具名与参数。参数采用 untagged enum:正常情况保存原始 JSON,超限时换成本地生成的 truncation 对象。executed_tool_calls 位于 internal_chat_message_metadata_passthrough,并被排除在公开 app-server schema 和 TypeScript 类型之外。

相关源码:

  • codex-rs/protocol/src/models/executed_tool_calls.rs :: ExecutedToolCall
  • codex-rs/protocol/src/models.rs :: InternalChatMessageMetadataPassthrough
rust
#[serde(untagged)]
pub enum ExecutedToolCallArguments {
    Raw(serde_json::Value),
    #[serde(skip_deserializing)]
    Truncated {
        #[serde(rename = "_codex_executed_tool_call_truncated")]
        truncation: ExecutedToolCallTruncation,
    },
}

pub struct ExecutedToolCall {
    pub name: String,
    arguments: ExecutedToolCallArguments,
}

pub struct ExecutedToolCallTruncation {
    original_bytes: usize,
    max_bytes: usize,
    omitted_calls: Option<usize>,
    original_name_bytes: Option<usize>,
}

源码位置:codex-rs/protocol/src/models.rs :: InternalChatMessageMetadataPassthrough

rust
pub struct InternalChatMessageMetadataPassthrough {
    pub turn_id: Option<String>,
    #[serde(default, skip_deserializing, skip_serializing_if = "Option::is_none")]
    #[schemars(skip)]
    #[ts(skip)]
    pub executed_tool_calls: Option<Vec<ExecutedToolCall>>,
}

skip_deserializing 建立了信任边界:provider 或 rollout 输入中的同名字段不会被当成本地 attempted-tool metadata 恢复。记录只能由当前进程的运行时生成,再序列化进发往支持方的请求。

3. 调度记录 ​

3.1 记录时机 ​

ToolCallRuntime::handle_tool_call_with_source 在查询 router、等待 runtime readiness、取得并行锁和执行 handler 之前 调用 recorder。feature 与 recorder 必须同时存在;这段顺序决定了“executed”在这里更接近 attempted,而不是 completed。

源码位置:codex-rs/core/src/tools/parallel.rs :: ToolCallRuntime::handle_tool_call_with_source

rust
if self
    .step_context
    .turn
    .config
    .features
    .enabled(codex_features::Feature::ExecutedToolCallMetadata)
    && let Some(executed_tool_calls) = self.session.services.executed_tool_calls.as_ref()
{
    executed_tool_calls.record_tool_call(
        &call,
        &source,
        super::effective_tool_mode(&self.step_context.turn),
    );
}
let router = &self.step_context.tool_router;
let supports_parallel = router.tool_supports_parallel(&call);
let tool_runtime = router.tool_runtime(&call);

这解释了一个看似反常的测试结果:nested exec_command 参数无效并抛出异常时,外层脚本可以捕获错误,但下一次 请求仍包含该调用元数据。元数据说明 dispatch runtime 收到过它,不说明 handler 返回了成功结果。

3.2 参数归一化 ​

记录器先按 payload 类型计算序列化长度。Function 参数是 JSON 文本,未超过 8 KiB 时尝试解析;非法 JSON 不会让 记录失败,而是作为普通字符串保存。Custom input 始终是字符串,ToolSearch 参数则序列化为 JSON 值。

源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::record_tool_call

rust
let original_bytes = match &call.payload {
    ToolPayload::Function { arguments } => arguments.len(),
    ToolPayload::Custom { input } => serialized_json_bytes(input),
    ToolPayload::ToolSearch { arguments } => serialized_json_bytes(arguments),
};
let name = codex_tools::code_mode_name_for_tool_name(&call.tool_name);
let recorded_call = if original_bytes > MAX_EXECUTED_TOOL_CALL_ARGUMENT_BYTES {
    ExecutedToolCall::truncated(
        name,
        original_bytes,
        MAX_EXECUTED_TOOL_CALL_ARGUMENT_BYTES,
    )
} else {
    let arguments = match &call.payload {
        ToolPayload::Function { arguments } => serde_json::from_str(arguments)
            .unwrap_or_else(|_| JsonValue::String(arguments.clone())),
        ToolPayload::Custom { input } => JsonValue::String(input.clone()),
        ToolPayload::ToolSearch { arguments } => {
            serde_json::to_value(arguments).unwrap_or_default()
        }
    };
    ExecutedToolCall::new(name, arguments)
};

工具名还会经过 code_mode_name_for_tool_name,把 namespace 形式转换为 Code Mode 可见名称。因此测试中的 test_namespace::unsupported_tool 被记录为 test_namespace__unsupported_tool,它与原始 Responses call 的 namespace/name 字段不是同一种 wire shape。

3.3 保留字防伪 ​

模型参数本身可能恰好包含 _codex_executed_tool_call_truncated。若直接保存,消费者无法区分“模型给出的普通参数”与 “本地生成的截断标记”。构造函数因此把这类原始对象再包一层,只有 set_truncation 能产生受信任的 truncation 变体。

源码位置:codex-rs/protocol/src/models/executed_tool_calls.rs :: ExecutedToolCall::new

rust
pub fn new(name: String, arguments: serde_json::Value) -> Self {
    let arguments = if arguments
        .as_object()
        .is_some_and(|object| object.contains_key("_codex_executed_tool_call_truncated"))
    {
        serde_json::json!({ "_codex_executed_tool_call_raw": arguments })
    } else {
        arguments
    };
    Self {
        name,
        arguments: ExecutedToolCallArguments::Raw(arguments),
    }
}

4. 直接调用 ​

Direct 与 DirectPlaintextMessage 都按 call_id 存入 direct_calls。相同 id 只保留首次记录;达到 256 条后, 第 257 个不同 id 不再保存完整参数,而是保存 max_bytes: 0 的 overflow marker,之后的新增 id 被忽略。

源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::record_tool_call

rust
match source {
    ToolCallSource::Direct | ToolCallSource::DirectPlaintextMessage => {
        let mut state = self
            .state
            .lock()
            .unwrap_or_else(std::sync::PoisonError::into_inner);
        if state.direct_calls.len() < MAX_PENDING_EXECUTED_TOOL_CALLS {
            state
                .direct_calls
                .entry(call.call_id.clone())
                .or_insert(recorded_call);
        } else if state.direct_calls.len() == MAX_PENDING_EXECUTED_TOOL_CALLS
            && !state.direct_calls.contains_key(&call.call_id)
        {
            state.direct_calls.insert(
                call.call_id.clone(),
                ExecutedToolCall::truncated(
                    recorded_call.name,
                    original_bytes,
                    0,
                ),
            );
        }
    }
    ToolCallSource::CodeMode { cell_id, .. } => {
        self.record_nested_tool_call(
            CellId::new(cell_id.clone()),
            recorded_call,
            original_bytes,
        );
    }
}

这个上限控制的是尚未与 output 配对的 session 状态,避免模型连续生成大量无法闭合的调用时无限占用内存。overflow marker 仍保留工具名和原始参数字节数,让消费者知道“还有一次调用,但参数内容未保存”。

5. Code Mode配对 ​

5.1 外壳过滤 ​

在有效工具模式确实是 Code Mode 或 CodeModeOnly 时,模型直接调用的 exec 与 wait 是承载脚本 cell 的外壳。真正需要记录的是脚本 内部通过 tools.* 发起的 nested calls;若连外壳也记录,外层 output 会同时出现“exec 调用了自己”和 nested calls。 因此 recorder 在特定模式下跳过默认 namespace 的直接 exec/wait。

源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::record_tool_call

rust
if matches!(source, ToolCallSource::Direct)
    && matches!(tool_mode, ToolMode::CodeMode | ToolMode::CodeModeOnly)
    && call.tool_name.is_default_namespace()
    && matches!(
        (call.tool_name.name.as_str(), &call.payload),
        (
            crate::tools::code_mode::PUBLIC_TOOL_NAME,
            ToolPayload::Custom { .. }
        ) | (
            crate::tools::code_mode::WAIT_TOOL_NAME,
            ToolPayload::Function { .. }
        )
    )
{
    return;
}

这条过滤只针对 Code Mode 生效。Direct 模式下名为 exec 的 custom call 仍会记录,因为此时它不是 cell 外壳,而是 一次普通的直接调用尝试。

工具 feature 开启并不自动保证有效模式仍是 Code Mode。effective_tool_mode 会在 Code Mode provider 不可用且允许 in-process fallback 时返回 Direct;此时外层 exec 不会走上面的过滤分支,后续 metadata 看到的是直接调用形状。阅读或 调试 Code Mode 记录时,必须先确认这个模式选择。

源码位置:codex-rs/core/src/tools/mod.rs :: effective_tool_mode

rust
let requested_tool_mode = requested_tool_mode(turn_context);
if !turn_context.code_mode_available
    && requested_tool_mode == ToolMode::CodeMode
    && !turn_context.config.code_mode.disable_in_process_fallback
{
    ToolMode::Direct
} else {
    requested_tool_mode
}

5.2 Cell映射 ​

Code Mode nested call 先按 CellId 进入 cells。外层 exec 启动 cell 后立即登记 cell_id → call_id;后续 wait 收到 live cell 响应时也登记当前 wait output,使最后一次可见 output 能取回该 cell 的 pending calls。

相关源码:

  • codex-rs/core/src/tools/code_mode/execute_handler.rs :: ExecuteHandler::handle
  • codex-rs/core/src/tools/code_mode/wait_handler.rs :: WaitHandler::handle
rust
if let Some(executed_tool_calls) = exec.session.services.executed_tool_calls.as_ref() {
    executed_tool_calls.register_cell(&cell_id, &call_id);
}

源码位置:codex-rs/core/src/tools/code_mode/wait_handler.rs :: WaitHandler::handle

rust
if let codex_code_mode::WaitOutcome::LiveCell(response) = &wait_response {
    let runtime_cell_id = match response {
        codex_code_mode::RuntimeResponse::Yielded { cell_id, .. }
        | codex_code_mode::RuntimeResponse::Terminated { cell_id, .. }
        | codex_code_mode::RuntimeResponse::Result { cell_id, .. } => cell_id,
    };
    if let Some(executed_tool_calls) =
        exec.session.services.executed_tool_calls.as_ref()
    {
        executed_tool_calls.register_cell(runtime_cell_id, &call_id);
    }
}

5.3 Cell预算 ​

每个 cell 最多保留 32 KiB 完整参数,单次调用最多 8 KiB;全 session 最多有 256 个 pending nested calls 和 256 个 cell/output mappings。达到调用数上限时仍写入一个零预算 marker,之后的新 cell 或新调用则被丢弃。

源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::record_nested_tool_call

rust
let at_pending_call_limit =
    state.pending_nested_calls == MAX_PENDING_EXECUTED_TOOL_CALLS;
let cell = state.cells.entry(cell_id).or_default();
let max_bytes = MAX_EXECUTED_TOOL_CALL_ARGUMENT_BYTES.min(
    MAX_EXECUTED_TOOL_CALL_FULL_ARGUMENT_BYTES_PER_OUTPUT
        .saturating_sub(cell.pending_full_argument_bytes),
);
let call = if at_pending_call_limit {
    ExecutedToolCall::truncated(call.name, original_bytes, 0)
} else if original_bytes <= max_bytes {
    cell.pending_full_argument_bytes = cell
        .pending_full_argument_bytes
        .saturating_add(original_bytes);
    call
} else {
    ExecutedToolCall::truncated(call.name, original_bytes, max_bytes)
};
cell.pending_calls.push(call);
state.pending_nested_calls += 1;

6. Prompt绑定 ​

6.1 绑定入口 ​

每次 sampling request 构造 prompt 前,Turn loop 从历史生成 prompt_input,再让 recorder 附加 pending calls。只有 确实附加过元数据时才运行请求级预算函数。retry cache 在整个 sampling 重试循环之外创建,因此网络重试重新构造 prompt 时仍能取回同一份 calls。

源码位置:codex-rs/core/src/session/turn.rs :: run_sampling_request

rust
let mut executed_tool_calls_by_output = HashMap::new();
loop {
    let prompt_input = if let Some(input) = initial_input.take() {
        input
    } else {
        sess.clone_history()
            .await
            .for_prompt(&turn_context.model_info.input_modalities)
    };
    let mut prompt_input = prompt_input;
    if let Some(executed_tool_calls) = sess.services.executed_tool_calls.as_ref()
        && executed_tool_calls.attach_pending_to_prompt(
            &mut prompt_input,
            &mut executed_tool_calls_by_output,
        )
    {
        codex_protocol::models::bound_executed_tool_calls_for_prompt(
            &mut prompt_input,
        );
    }
    let prompt = build_prompt(
        prompt_input,
        router.as_ref(),
        turn_context.as_ref(),
        base_instructions.clone(),
    );
    // ...
}

生效时机因此是“下一次模型采样请求构造时”,不是工具调用刚发生时。元数据附在历史中的 output item 上,再随 build_prompt 进入请求;它不修改 handler 已经生成的 output body。

6.2 Output键 ​

绑定器从 items 尾部向前扫描,只接受 Function、Custom 和带 call id 的 ToolSearch output。key 同时包含 ResponseItem discriminant 与 call_id,避免不同 output 类型恰好复用同一字符串 id 时误配。

源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::attach_pending_to_prompt

rust
for item in items.iter_mut().rev() {
    let call_id = match &*item {
        ResponseItem::FunctionCallOutput { call_id, .. }
        | ResponseItem::CustomToolCallOutput { call_id, .. }
        | ResponseItem::ToolSearchOutput {
            call_id: Some(call_id),
            ..
        } => call_id,
        _ => continue,
    };
    let key = (std::mem::discriminant(&*item), call_id.clone());
    // ...
}

反向扫描让重复出现的同类 output 优先绑定最近一项。pending_retry_outputs 与 pending_retained_outputs 还保证同一个 key 在一次 prompt 构造中只附加一次。

6.3 来源优先级 ​

同一个 output key 的调用记录按四级来源选择:本轮 retry cache、session retained history、尚未消费的 direct call、 通过 output-cell 映射取得的 nested calls。新取出的 direct/nested calls 会同时写入 retry cache 和 retained map。

源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::attach_pending_to_prompt

rust
let calls = if let Some(cached) = retry_cache.get(&key) {
    if !pending_retry_outputs.remove(&key) {
        continue;
    }
    pending_retained_outputs.remove(&key);
    cached.clone()
} else if let Some(retained) = state.retained_calls.get(&key) {
    if !pending_retained_outputs.remove(&key) {
        continue;
    }
    retained.clone()
} else {
    let mut calls = state
        .direct_calls
        .remove(call_id)
        .into_iter()
        .collect::<Vec<_>>();
    if let Some(cell_id) = state.output_cells.remove(call_id)
        && let Some(mut cell) = state.cells.remove(&cell_id)
    {
        state.pending_nested_calls = state
            .pending_nested_calls
            .saturating_sub(cell.pending_calls.len());
        state
            .output_cells
            .retain(|_, output_cell_id| output_cell_id != &cell_id);
        calls.append(&mut cell.pending_calls);
    }
    if calls.is_empty() {
        continue;
    }
    retry_cache.insert(key.clone(), calls.clone());
    state.retained_calls.insert(key, calls.clone());
    calls
};
item.append_executed_tool_calls(calls);

成功绑定 cell 后,recorder 会减少 pending_nested_calls、删除 cell,并清掉指向同一 cell 的其他 output mapping。 这是一种一次性归属:nested calls 被投影到匹配到的 output 后,不应再从 pending 状态重复取出。

7. 预算与省略 ​

7.1 三层预算 ​

这套元数据有三个不同粒度的上限,不能把它们统称为“8 KiB 截断”。

层级上限超限行为
单个调用参数8 KiB参数替换为 truncation marker
单个 Code Mode cell 的完整参数32 KiB后续调用按剩余空间截断
整个请求的 attempted-tool metadata32 KiB公平分配或优先保留最近记录,并报告 omissions
Session pending 调用/映射256保留一个零预算 overflow marker,之后丢弃新增项

请求级函数先计算 metadata 在 wire 上的精确 JSON 字节数,包括字段名和外围对象开销,而不是只累加参数字符串长度。

源码位置:codex-rs/protocol/src/models/executed_tool_calls.rs :: executed_tool_call_metadata_bytes

rust
pub fn executed_tool_call_metadata_bytes(item: &ResponseItem) -> usize {
    let Some(metadata) = item.executed_tool_call_metadata() else {
        return 0;
    };
    let Some(calls) = metadata
        .executed_tool_calls
        .as_ref()
        .filter(|calls| !calls.is_empty())
    else {
        return 0;
    };
    serde_json::to_vec(calls)
        .map(|bytes| {
            bytes
                .len()
                .saturating_add(executed_tool_call_metadata_field_bytes(metadata))
        })
        .unwrap_or(usize::MAX)
}

7.2 公平与最近优先 ​

发送请求前调用 bound_executed_tool_calls_for_prompt,按剩余 item 数公平分配预算;recorder 整理 retained history 时 使用 bound_executed_tool_calls_for_prompt_prioritizing_recent,先反转 items,让新记录先消耗预算,再恢复顺序。

源码位置:codex-rs/protocol/src/models/executed_tool_calls.rs :: bound_executed_tool_calls_for_prompt

rust
pub fn bound_executed_tool_calls_for_prompt(items: &mut [ResponseItem]) {
    bound_executed_tool_calls_for_prompt_with_priority(items, false);
}

pub fn bound_executed_tool_calls_for_prompt_prioritizing_recent(
    items: &mut [ResponseItem],
) {
    items.reverse();
    bound_executed_tool_calls_for_prompt_with_priority(items, true);
    items.reverse();
}

当记录被整体移除时,算法把缺失数量合并进某条仍保留记录的 omitted_calls。若预算小到一条都放不下,则尽量保留 一个截断后的 fallback call,并把其余数量写入 marker。消费者因此可以区分“历史原本只有这些调用”和“还有若干调用 因预算未展开”。

8. 失败与恢复 ​

8.1 失败调用 ​

记录发生在 handler 之前,所以“调用失败”不是清除 metadata 的条件。对于真正进入 ToolCallSource::CodeMode 的 nested 调用,调用来源会把记录放入 cell;对于 fallback 到 Direct 的 turn,外层 exec 本身会按 direct call 记录。两种结果都 符合“dispatch 已尝试”的定义,但不能从字段名推导 handler 成功。

源码位置:codex-rs/core/src/tools/code_mode/mod.rs :: call_nested_tool

rust
let result = tool_runtime
    .handle_tool_call_with_source(
        call,
        ToolCallSource::CodeMode {
            cell_id: cell_id.to_string(),
            runtime_tool_call_id,
        },
        cancellation_token,
    )
    .await?;
Ok(result.code_mode_result())

端到端测试还会受到 Code Mode host 是否可用、effective_tool_mode 是否回退的影响;因此测试输入必须同时说明工具模式, 否则看到外层 exec 记录时,不能误判为 nested 绑定器失效。无论哪条路径,副作用是否回滚仍需检查 handler 结果和真实 外部状态。

8.2 Retry恢复 ​

第一次绑定时,direct/cell pending state 被消费,但同一 calls 会写入 sampling loop 的 retry cache。若 provider 请求失败并 重试,新的 prompt_input 可以从 cache 再附加同一记录。Session 的 retained map 解决更长周期的重复 prompt 构造: 只要相应 output 仍在输入中,就能恢复 metadata;output 因压缩不再出现时,未匹配 retained key 会被清理。

源码位置:codex-rs/core/src/tools/executed_tool_calls.rs :: ExecutedToolCallRecorder::attach_pending_to_prompt

rust
if !pending_retained_outputs.is_empty() {
    state
        .retained_calls
        .retain(|key, _| !pending_retained_outputs.contains(key));
}
if retained_bytes > MAX_EXECUTED_TOOL_CALL_FULL_ARGUMENT_BYTES_PER_OUTPUT {
    bound_executed_tool_calls_for_prompt_prioritizing_recent(items);
    state.retained_calls.clear();
    for item in items {
        // 从已经完成预算收束的 items 重建 retained_calls
        // ...
    }
}

8.3 Best-effort边界 ​

源码注释明确把 recorder 定义为 session-scoped best-effort attempted-tool metadata。以下情况不能外推为可靠日志:

  • 调用在 output 出现前被取消,pending state 可能没有机会绑定;
  • 历史压缩移除了 output,retained metadata 会随未匹配 key 清理;
  • Code Mode cell yield 后没有再次 wait,nested calls 可能一直没有对应 output;
  • 超过调用数或字节预算的内容只留下 marker,甚至在极端情况下被省略;
  • 进程退出后 recorder 内存状态不会成为持久化恢复依据。

因此排查“工具究竟有没有成功”时,应回到 Tool output、lifecycle outcome、handler 日志或真实外部状态;本字段适合告诉 下一次模型请求“运行时近期尝试过什么”,不适合作为可靠执行日志或 exactly-once 证明。

9. 测试路径 ​

9.1 Pending上限 ​

单元测试分别制造 258 个 direct calls、258 个 nested calls 和 258 组 cell mapping。关键断言是 direct 与 nested 都 保留 256 条完整记录加 1 条 overflow marker,marker 的 max_bytes 为 0;绑定 cell output 后,pending 计数归零、 cell 被删除,retry cache 留下可重放副本。

源码位置:codex-rs/core/src/tools/executed_tool_calls_tests.rs :: executed_tool_call_recorder_bounds_pending_calls_and_preserves_overflow

rust
assert_eq!(
    state.direct_calls.len(),
    MAX_PENDING_EXECUTED_TOOL_CALLS + 1,
);
assert_eq!(
    state.pending_nested_calls,
    MAX_PENDING_EXECUTED_TOOL_CALLS + 1,
);
assert_eq!(state.cells.len(), MAX_PENDING_EXECUTED_TOOL_CALLS);
assert_eq!(state.output_cells.len(), MAX_PENDING_EXECUTED_TOOL_CALLS);

recorder.attach_pending_to_prompt(&mut items, &mut retry_cache);
let calls = items[0]
    .executed_tool_call_metadata()
    .and_then(|metadata| metadata.executed_tool_calls.as_ref())
    .expect("bounded nested calls must attach to their own output");
assert_eq!(calls.len(), MAX_PENDING_EXECUTED_TOOL_CALLS + 1);
assert_eq!(retry_cache.len(), 1);

测试还用空 prompt 再次调用绑定器,断言没有匹配 output 后 retained_calls 被清空。它覆盖容量、绑定、重放和清理, 但没有模拟 Tokio task 取消或真实进程重启。

9.2 Retained省略 ​

另一个单元测试连续建立 512 个约 1 KiB 参数的历史 output。每次把最新调用附到 prompt 并执行请求预算收束,最后断言 retained serialized bytes 不超过 32 KiB,最新调用仍在,所有保留条目数与 marker 中的 omitted_calls 之和等于 原始 512 次调用。

源码位置:codex-rs/core/src/tools/executed_tool_calls_tests.rs :: executed_tool_call_recorder_bounds_retained_history_and_reports_omissions

rust
assert!(
    retained_bytes
        <= MAX_EXECUTED_TOOL_CALL_FULL_ARGUMENT_BYTES_PER_OUTPUT
);

let omitted_calls = metadata
    .iter()
    .filter_map(|call| {
        call["arguments"]
            ["_codex_executed_tool_call_truncated"]
            ["omitted_calls"]
            .as_u64()
    })
    .sum::<u64>();
assert!(omitted_calls > 0);
assert_eq!(metadata.len() as u64 + omitted_calls, 512);

这里证明 recent prioritization 与 omissions 计数守恒,不证明模型会怎样解释 marker,也不证明 32 KiB 是任意 provider 都接受的公共协议上限。

9.3 请求投影 ​

集成测试先关闭 feature,断言 custom tool output 中没有 executed_tool_calls;再开启 feature,模拟 namespaced custom call、超长转义字符串和 direct exec。测试检查下一次 mock Responses 请求的完整 JSON,确认 namespace 转换、8 KiB truncation 与跨请求 retained metadata。Code Mode nested 端到端测试还需要可用的 Code Mode host;若 effective_tool_mode 回退到 Direct,应按 Direct 记录路径解释结果。

源码位置:codex-rs/core/tests/suite/tools.rs :: namespaced_custom_tool_call_preserves_namespace_through_dispatch_and_replay

rust
assert_eq!(
    escaped_request.custom_tool_call_output(escaped_call_id)
        ["internal_chat_message_metadata_passthrough"]
        ["executed_tool_calls"],
    json!([{
        "name": format!("{namespace}__{tool_name}"),
        "arguments": {
            "_codex_executed_tool_call_truncated": {
                "original_bytes": serde_json::to_vec(&escaped_input)?.len(),
                "max_bytes": 8 * 1024,
            },
        },
    }]),
);

这个测试把 recorder、历史 output、prompt 构造和请求序列化连成一条真实路径;它使用 mock server,不覆盖远端模型如何 利用该字段。

10. 阅读实践 ​

在源码仓库根目录运行下面的定向测试,可以先观察容量与 omissions,再观察真实请求 JSON。若要运行 Code Mode nested 端到端测试,还需要先准备可用的 Code Mode host,并确认 effective_tool_mode 没有回退到 Direct。

bash
RUST_MIN_STACK=16777216 cargo test -p codex-core executed_tool_call_recorder
RUST_MIN_STACK=16777216 cargo test -p codex-core namespaced_custom_tool_call_preserves_namespace_through_dispatch_and_replay
RUST_MIN_STACK=16777216 cargo test -p codex-protocol executed_tool_call

然后用三个场景检查自己的判断:

  1. 一个 Direct function call 参数合法,但 handler 返回错误。它会在何时进入 direct_calls,最终附在哪类 output 上?
  2. 一个 Code Mode cell 调用两个 nested tools 后 yield,用户没有再调用 wait。哪些状态可能仍留在 cells,为什么不能 据此声称模型已经看到两次调用?
  3. 历史中有数百个大参数调用。哪些函数负责单调用 8 KiB、cell 32 KiB 和整请求 32 KiB,omitted_calls 应满足什么 数量关系?

能够从 handle_tool_call_with_source 复述到 record_tool_call、从 output key 复述到 attach_pending_to_prompt,再走到 bound_executed_tool_calls_for_prompt,就建立了这套元数据的完整心智模型。若要 继续理解工具结果本身如何转成模型输入,可回看ToolOutput与错误模型;若要区分 attempted metadata 与 extension 可见终态,则回看ToolLifecycle事件。