Skip to content

CodeMode等待取消与遥测

从 cell 观察者、V1 与 gRPC 取消退休走到 Turn 中断、终态输出和 analytics,解释 Code Mode wait 与 terminate 的完整控制链。

基于rust-v0.150.0
CodexRustExecutionCodeMode

CodeMode等待取消与遥测 ​

Code Mode 的 wait 不是“暂停十秒再看结果”,而是为一个仍存活的 cell 注册下一位观察者。观察者等待三类边界之一:指定的 yield 时间到达、脚本自然完成、cell 被终止。调用方中途放弃等待时,transport 还必须先让旧观察请求退休,才能安全开始同一 cell 的下一次观察。

这条控制链比普通工具调用多出一个长期存活对象:cell 可以跨多个 Turn 返回 Yielded,但每次 wait 都是新的 Core 工具调用,拥有新的 call ID、输出预算、耗时和 telemetry。因而必须同时分清 cell 生命周期、单次 wait 生命周期与 transport request 生命周期。

本文承接CodeMode工具调用桥接和CellActor与执行队列。前者解释嵌套工具如何往返,后者解释 actor 与 observer;本文继续追踪 wait、terminate、Turn 中断、输出适配和 analytics,不重复 V8 初始化过程。

下面的时序先给出正常等待与主动终止的共同主线。CodeModeSession 的实现可以是 V1 process/WebSocket client,也可以是 gRPC client;Core handler 只依赖统一接口。

同一个 WaitOutcome 会同时驱动三个消费者:模型可见输出、cell/dispatch 收尾和 analytics。只看返回文本,会漏掉后两类状态变化。

1. 控制工具契约 ​

公开 wait 工具是 Function 工具,只要求 cell_id。yield_time_ms 决定本次观察多久后再次返回,max_tokens 只限制本次新增输出,terminate 则把请求切换为终止操作。

源码位置:codex-rs/core/src/tools/code_mode/wait_spec.rs :: create_wait_tool

rust
let properties = BTreeMap::from([
    (
        "cell_id".to_string(),
        JsonSchema::string(Some("Identifier of the running exec cell.".to_string())),
    ),
    (
        "yield_time_ms".to_string(),
        JsonSchema::number(Some(
            "Wait before yielding more output. Defaults to 10000 ms.".to_string(),
        )),
    ),
    (
        "max_tokens".to_string(),
        JsonSchema::number(Some(
            "Output token budget for this wait call. Defaults to 10000 tokens.".to_string(),
        )),
    ),
    (
        "terminate".to_string(),
        JsonSchema::boolean(Some(
            "True stops the running exec cell; false or omitted waits for output.".to_string(),
        )),
    ),
]);

handler 侧的反序列化类型把默认等待时间固定为 DEFAULT_WAIT_YIELD_TIME_MS,当前值为 10000 毫秒。schema 的 strict 为 false,但 object schema 明确拒绝额外属性;最终仍由 serde_json::from_str 决定输入能否进入控制路径。

源码位置:codex-rs/core/src/tools/code_mode/wait_handler.rs :: ExecWaitArgs, parse_arguments

rust
#[derive(Debug, Deserialize)]
struct ExecWaitArgs {
    cell_id: String,
    #[serde(default = "default_wait_yield_time_ms")]
    yield_time_ms: u64,
    #[serde(default)]
    max_tokens: Option<usize>,
    #[serde(default)]
    terminate: bool,
}

fn parse_arguments<T>(arguments: &str) -> Result<T, FunctionCallError>
where
    T: for<'de> Deserialize<'de>,
{
    serde_json::from_str(arguments).map_err(|err| {
        FunctionCallError::RespondToModel(format!("failed to parse function arguments: {err}"))
    })
}

wait 还有一个不同于普通工具的边界:它不产生 PreToolUse 或 PostToolUse payload。等待现有 cell 是 runtime 控制动作,Hook 不能阻塞、改写或用模型反馈替换它的返回值;但 cell 内部发起的嵌套工具仍走普通工具 Hook。

源码位置:codex-rs/core/src/tools/code_mode/wait_handler.rs :: CoreToolRuntime for CodeModeWaitHandler

rust
fn pre_tool_use_payload(&self, _invocation: &ToolInvocation) -> Option<PreToolUsePayload> {
    None
}

fn post_tool_use_payload(
    &self,
    _invocation: &ToolInvocation,
    _result: &dyn ToolOutput,
) -> Option<PostToolUsePayload> {
    None
}

因此 wait 的权限与失败语义主要来自 cell、transport 和 Core Turn,而不是 Hook 决策。

2. 观察者模型 ​

runtime 内部有两种观察模式。公开 wait 使用 YieldAfter(Duration);PendingFrontier 服务于 host 的内部执行握手,用来观察脚本是否暂停在嵌套工具边界。二者不能混为一种“等待状态”。

源码位置:codex-rs/code-mode-runtime/src/session_runtime/types.rs :: ObserveMode, CellEvent

rust
pub(crate) enum ObserveMode {
    YieldAfter(Duration),
    PendingFrontier,
}

pub(crate) enum CellEvent {
    Yielded {
        content_items: Vec<OutputItem>,
    },
    Pending {
        content_items: Vec<OutputItem>,
        pending_tool_call_ids: Vec<String>,
    },
    Completed {
        content_items: Vec<OutputItem>,
        error_text: Option<String>,
    },
    Terminated {
        content_items: Vec<OutputItem>,
    },
}

公开等待从 InProcessCodeModeSession::begin_wait 进入 SessionRuntime::begin_observe。actor 返回事件时包装为 LiveCell;cell 在注册观察前已经缺失或关闭,或者观察期间变成缺失/关闭,则包装为 MissingCell,而不是把内部 MissingCell、ClosedCell 类型直接泄露给调用方。

源码位置:codex-rs/code-mode-runtime/src/service.rs :: InProcessCodeModeSession::begin_wait

rust
match self
    .runtime
    .begin_observe(
        &runtime_cell_id,
        runtime::ObserveMode::YieldAfter(self.resolve_yield_timeout(yield_time_ms)),
    )
    .await
{
    Ok(pending_event) => Box::pin(async move {
        match pending_event.event().await {
            Ok(event) => Ok(WaitOutcome::LiveCell(runtime_response(&cell_id, event)?)),
            Err(runtime::Error::MissingCell(_) | runtime::Error::ClosedCell(_)) => {
                Ok(WaitOutcome::MissingCell(missing_cell_response(cell_id)))
            }
            Err(error) => Err(error.to_string()),
        }
    }),
    Err(runtime::Error::MissingCell(_) | runtime::Error::ClosedCell(_)) => {
        missing_wait(cell_id)
    }
    Err(error) => Box::pin(async move { Err(error.to_string()) }),
}

一个 cell 同时只能有一位 active observer。actor 会先清理 receiver 已关闭的旧 observer;若旧 observer 仍有效,或者终止已经开始,新观察者收到 Busy,不会挤掉第一位观察者。

源码位置:codex-rs/code-mode-runtime/src/cell_actor/mod.rs :: CellCommand::Observe handling

rust
if observer
    .as_ref()
    .is_some_and(|observer| observer.response_tx.is_closed())
{
    observer = None;
    yield_timer = None;
}
if observer.is_some() || termination {
    let _ = response_tx.send(Err(CellError::Busy));
    continue;
}

观察者通过前,actor 还可能立即交付已经到达的 pending frontier;只有没有可立即交付事件时,才把 sender 存为 active observer,并恢复暂停的 runtime。

源码位置:codex-rs/code-mode-runtime/src/cell_actor/mod.rs :: CellCommand::Observe handling

rust
if matches!(mode, ObserveMode::PendingFrontier) && pending_frontier_ready {
    pending_frontier_ready = false;
    match send_cell_event(
        response_tx,
        CellEvent::Pending {
            content_items: std::mem::take(&mut content_items),
            pending_tool_call_ids: std::mem::take(&mut pending_tool_call_ids),
        },
    ) {
        Ok(()) => {}
        Err(CellEvent::Pending {
            content_items: undelivered_items,
            pending_tool_call_ids: undelivered_tool_call_ids,
        }) => {
            content_items = undelivered_items;
            pending_tool_call_ids = undelivered_tool_call_ids;
            pending_frontier_ready = true;
        }
        Err(event) => {
            panic!("pending delivery returned an unexpected event: {event:?}")
        }
    }
    continue;
}
observer = Some(Observer { mode, response_tx });
yield_timer = observer.as_ref().and_then(observer_timer);
if runtime_paused && matches!(mode, ObserveMode::YieldAfter(_)) {
    pending_frontier_ready = false;
    pending_tool_call_ids.clear();
}
resume_for_observation(
    mode,
    &mut runtime_paused,
    &runtime_tx,
    &runtime_control_tx,
);

timer 到期时,actor 用 mem::take 取走自上次观察后积累的内容。这就是协议所说“wait 只返回新增输出”的根源;下一次等待不会再次拿到已经成功交付的内容。如果 oneshot receiver 已经丢弃,restore_undelivered_yield 会把未交付内容放回 buffer,避免取消 wait 造成输出丢失。

源码位置:codex-rs/code-mode-runtime/src/cell_actor/mod.rs :: yield timer branch

rust
yield_timer = None;
restore_undelivered_yield(
    send_observer_event(
        observer.take(),
        CellEvent::Yielded {
            content_items: std::mem::take(&mut content_items),
        },
    ),
    &mut content_items,
);

3. 等待窗口 ​

yield_time_ms 是观察窗口,不是 cell 的总执行期限。runtime 对 10 秒及以上的请求增加 1 秒 grace,使恰好在边界附近完成的嵌套工具有机会把最终结果与本次 wait 一起返回;然后再应用 session 的 max_yield_time_ms 上限。

源码位置:codex-rs/code-mode-runtime/src/service.rs :: InProcessCodeModeSession::resolve_yield_timeout

rust
fn resolve_yield_timeout(&self, yield_time_ms: u64) -> Duration {
    let yield_time = Duration::from_millis(yield_time_ms);
    let timeout = if yield_time >= MIN_YIELD_TIME_FOR_GRACE {
        yield_time.saturating_add(YIELD_GRACE_PERIOD)
    } else {
        yield_time
    };
    self.cell_execution_limits
        .max_yield_time_ms
        .map(Duration::from_millis)
        .map_or(timeout, |limit| timeout.min(limit))
}

应用顺序是“先加 grace,再取上限”。例如请求 10000 毫秒且没有上限,观察窗口是 11000 毫秒;上限为 10500 毫秒时,最终窗口是 10500 毫秒;上限为 10000 毫秒时,grace 被完整裁掉。窗口到期只生成 Yielded,不会终止脚本。

V1 与 gRPC client 还各自增加 transport deadline。它们同样为 runtime 的一秒 grace 留出空间,随后再叠加 transport 等待预算;transport 超时意味着连接无法确认请求结果,通常会使连接失效,而不是简单返回一次普通 Yielded。

源码位置:codex-rs/code-mode/src/remote_session/connection.rs :: Connection::wait

rust
let runtime_timeout =
    Duration::from_millis(request.yield_time_ms).saturating_add(Duration::from_secs(1));
let cancellation = CallerCancellation::new();
let (response_tx, response_rx) = oneshot::channel();
let result = self
    .with_transport_deadline(runtime_timeout, "wait", async {
        self.send(DriverCommand::Wait {
            session,
            request,
            caller_cancellation: cancellation.token(),
            response_tx,
        })
        .await?;
        self.receive(response_rx).await
    })
    .await;
cancellation.disarm();
result

这里的 CallerCancellation 是 RAII guard。正常收到结果后 disarm();future 被丢弃、超时或外层取消导致 guard 提前 Drop 时,它会取消 token,通知 driver 退休这次请求。

4. 终止线性化 ​

terminate: true 不创建新的 yield observer。Core 直接调用 CodeModeSession::terminate,runtime 在 CellState::request_termination 中完成终态竞争的线性化。

源码位置:codex-rs/code-mode-runtime/src/cell_actor/types.rs :: CellState::request_termination

rust
match std::mem::replace(&mut *phase, CellPhase::Tombstone) {
    CellPhase::Running => {
        let (response_tx, response_rx) = oneshot::channel();
        *phase = CellPhase::Terminating { response_tx };
        self.cancellation_token.cancel();
        response_event(response_rx)
    }
    CellPhase::Terminating { response_tx } => {
        *phase = CellPhase::Terminating { response_tx };
        Box::pin(async { Err(CellError::AlreadyTerminating) })
    }
    CellPhase::Completed {
        pending_initial_yield_items,
        event,
    } => {
        let event = prepend_initial_yield(event, pending_initial_yield_items);
        *phase = CellPhase::CompletionClaimed(event.clone());
        self.cancellation_token.cancel();
        ready_event(event)
    }
    CellPhase::CompletionClaimed(event) => {
        *phase = CellPhase::CompletionClaimed(event);
        Box::pin(async { Err(CellError::AlreadyTerminating) })
    }
    CellPhase::Tombstone => closed_event(),
}

若 cell 仍在运行,第一次 terminate 把 phase 改成 Terminating 并取消 cell token;第二次 terminate 明确返回 AlreadyTerminating。若自然完成已经先提交,则 terminate 取得现有完成结果,而不是把已完成脚本改写成 Terminated。因此 terminate 参与的是终态所有权竞争,不是无条件覆盖。

actor 观察到取消后会停止接收新 observe 命令、通知 runtime 结束并等待 callback 清理。只有 callback task 已取消并 drain 后,finish_termination 才把 Terminated 同时交给 terminate caller 和仍在等待的 observer。

源码位置:codex-rs/code-mode-runtime/src/cell_actor/mod.rs :: cancellation branch

rust
_ = cancellation_token.cancelled(), if !termination => {
    termination = true;
    yield_timer = None;
    drop(command_rx.take());
    begin_termination(
        &runtime_tx,
        &runtime_control_tx,
        &runtime_terminate_handle,
        &cancellation_token,
    );
    if runtime_closed {
        finish_callbacks(
            &callback_cancellation_token,
            &mut notification_tasks,
            &mut tool_tasks,
            CallbackCompletion::Cancel,
            task_failure_handler.as_ref(),
        ).await;
        finish_termination(
            &cell_state,
            observer.take().map(|observer| observer.response_tx),
            CellEvent::Terminated {
                content_items: std::mem::take(&mut content_items),
            },
        );
        break;
    }
}

终止成功的含义因此很严格:不是“已发送取消信号”,而是“runtime 终止路径与 callback 清理已经到达可发布终态”。这也是重复终止在清理期间必须失败,而不能同时存在两个终止 waiter 的原因。

5. V1等待退休 ​

V1 process/WebSocket 协议用普通 request ID 关联 wait。调用方丢弃 future 时,CallerCancellation 取消 token;driver 扫描 pending request 后发送 CancelRequest。但发送取消帧不代表旧 wait 已经从 host 退休。

源码位置:codex-rs/code-mode/src/remote_session/connection.rs :: CallerCancellation

rust
impl Drop for CallerCancellation {
    fn drop(&mut self) {
        if self.armed {
            self.token.cancel();
        }
    }
}

同一 session、同一 cell 如果仍有一个已取消但尚未收到 host 响应的 wait,新 wait 会进入 deferred_waits,不会立即发往 host。

源码位置:codex-rs/code-mode/src/remote_session/connection/driver/commands.rs :: ConnectionDriver::wait

rust
if self.requests.has_cancelled_wait(&session, &request.cell_id) {
    self.requests.push_deferred_wait(DeferredWait {
        session,
        request,
        caller_cancellation,
        response_tx,
    });
    return true;
}
self.start_wait(session, request, caller_cancellation, response_tx)

每次 driver 状态推进后会尝试 flush 延迟队列。旧 cancelled wait 仍存在时继续排队;旧 response 被消费并从 pending map 删除后,下一次 wait 才真正发送。若延迟期间第二个 caller 也取消,则直接返回取消错误,不制造无主 host 请求。

源码位置:codex-rs/code-mode/src/remote_session/connection/driver/responses.rs :: ConnectionDriver::flush_deferred_waits

rust
while let Some(wait) = deferred.pop_front() {
    if wait.caller_cancellation.is_cancelled() {
        let _ = wait
            .response_tx
            .send(Err("code-mode request cancelled".to_string()));
        continue;
    }
    if self
        .requests
        .has_cancelled_wait(&wait.session, &wait.request.cell_id)
    {
        self.requests.push_deferred_wait(wait);
        continue;
    }
    if !self.start_wait(
        wait.session,
        wait.request,
        wait.caller_cancellation,
        wait.response_tx,
    ) {
        for wait in deferred {
            let _ = wait
                .response_tx
                .send(Err("code-mode host connection closed".to_string()));
        }
        return false;
    }
}
true

这不是性能优化,而是 observer 不变量的 transport 版本:host 端只允许同一 cell 一个 active observer。若第二个 wait 越过尚未退休的第一个 wait,它可能收到 Busy,甚至与取消响应发生乱序。

6. gRPC等待撤销 ​

gRPC 不复用 V1 request cancellation。client 为每个 cell 维护一个弱引用 WaitSlot,其中 active 立即拒绝本进程内的并发观察,异步 mutex 则把取消退休与下一次 wait 串行化。

源码位置:codex-rs/code-mode/src/grpc_session/operations.rs :: SessionInner::wait

rust
let slot = match slots
    .get(&request.cell_id)
    .and_then(std::sync::Weak::upgrade)
{
    Some(slot) => slot,
    None => {
        let slot = Arc::new(WaitSlot {
            lock: Arc::new(tokio::sync::Mutex::new(())),
            active: AtomicBool::new(false),
        });
        slots.insert(request.cell_id.clone(), Arc::downgrade(&slot));
        slot
    }
};
if slot.active.swap(true, Ordering::AcqRel) {
    return Err(format!(
        "exec cell {} already has an active observer",
        request.cell_id
    ));
}

取得 slot permit 后,client 创建 UUID wait_id 并发送 gRPC WaitRequest。WaitCancellation 同时拥有 wait ID、slot 与 permit;正常响应时 disarm,异常退出时由 Drop 异步发送 CancelWait。

源码位置:codex-rs/code-mode/src/grpc_session/operations.rs :: WaitCancellation::drop

rust
let Some(wait_id) = self.wait_id.take() else {
    return;
};
let permit = self.permit.take();
if self.session.stopped.is_cancelled() {
    return;
}
let session = Arc::clone(&self.session);
self.session.runtime.spawn(async move {
    let mut client = session.client();
    let result = deadline::request(
        &session,
        "wait cancellation",
        Duration::ZERO,
        client.cancel_wait(grpc::CancelWaitRequest {
            session_id: session.id.clone(),
            wait_id,
        }),
    )
    .await;
    if let Err(error) = result
        && !session.stopped.is_cancelled()
    {
        session.fail(format!(
            "failed to retire canceled gRPC code-mode wait: {error}"
        ));
    }
    drop(permit);
    drop(slot);
    session.prune_wait_slots();
});

permit 在 CancelWait RPC 完成后才释放,所以第二次 wait 即使已经把 active 改回 false,也要等旧 wait 在 host 端退休。取消退休失败会使 session fail closed;client 不能继续假设 observer 状态仍与 host 一致。

host 端用相同 wait_id 注册 WaitRegistration,并在 wait future、session close 与 registration cancellation 之间竞速。CancelWait 返回前调用 session.cancel_wait(wait_id),使对应注册的 cancellation token 生效。

源码位置:codex-rs/code-mode-host/src/grpc/mod.rs :: wait_request, cancel_wait_request

rust
let registration = WaitRegistration::new(Arc::clone(&session), request.wait_id)?;
let outcome = tokio::select! {
    biased;
    _ = registration.cancellation().cancelled() => {
        return Err(Status::cancelled("code-mode wait was cancelled"));
    }
    _ = session.closed.cancelled() => {
        return Err(Status::cancelled("code-mode session is closed"));
    }
    outcome = session.runtime.wait(request) => {
        outcome.map_err(Status::failed_precondition)?
    }
};

V1 和 gRPC 的具体消息不同,但不变量相同:旧 observer 的取消必须被远端确认并退休,下一位 observer 才能接管同一 cell。

6.1 通知收口 ​

gRPC wait/terminate 在返回前还要处理 notification。Yielded 不关闭通知;Terminated 主动取消当前 cell 的 notification;自然 Result 则等待 lease stream 上的 CellClosed 在所有已接纳通知完成后取消 execution token。

源码位置:codex-rs/code-mode/src/grpc_session/operations.rs :: SessionInner::settle_notifications

rust
match response {
    RuntimeResponse::Yielded { .. } => {}
    RuntimeResponse::Terminated { cell_id, .. } => self
        .state
        .lock()
        .unwrap_or_else(PoisonError::into_inner)
        .cancel_notifications(cell_id),
    RuntimeResponse::Result { cell_id, .. } => {
        let cancellation = self
            .state
            .lock()
            .unwrap_or_else(PoisonError::into_inner)
            .notification_cancellation(cell_id);
        if let Some(cancellation) = cancellation {
            cancellation.cancelled().await;
        }
    }
}

自然完成与主动终止采用不同策略,是因为自然完成允许已接纳的 notify() 排空;主动终止要求尽快撤销未完成通知。结果在 notification 边界尚未稳定时不能提前返回,否则模型可能先看到“脚本结束”,随后又收到属于该 cell 的迟到通知。

7. Turn中断 ​

模型显式调用 wait(..., terminate: true) 不是唯一终止入口。用户中断正在运行的 Turn 时,任务取消 token 先被取消;若 CodeModeInterrupt feature 开启,Session 再枚举 dispatch broker 中仍有 gate 的 cell,并发调用 terminate。

源码位置:codex-rs/core/src/tasks/mod.rs :: Session::handle_task_abort

rust
task.cancellation_token.cancel();
if reason == TurnAbortReason::Interrupted
    && task
        .turn_context
        .config
        .features
        .enabled(Feature::CodeModeInterrupt)
{
    self.services
        .code_mode_service
        .interrupt_active_cells()
        .await;
}

源码位置:codex-rs/core/src/tools/code_mode/mod.rs :: CodeModeService::interrupt_active_cells

rust
let Some(session) = self.session.get() else {
    return;
};
join_all(
    self.dispatch_broker
        .active_cell_ids()
        .into_iter()
        .map(|cell_id| async move {
            if let Err(error) = session.terminate(cell_id.clone()).await {
                tracing::warn!(%cell_id, %error, "failed to terminate interrupted code-mode cell");
            }
        }),
)
.await;

这里没有为了中断而初始化一个从未使用过的 Code Mode session;session.get() 为空就直接返回。多个活跃 cell 用 join_all 并行终止,单个 cell 的错误只记录 warning,不阻止其他 cell 收尾。

feature 关闭时,Turn task 仍会取消,但 Core 不主动调用 Code Mode session 的 terminate。此时 host/session 自身的取消与生命周期规则仍可能关闭 cell,不能把 feature 开关理解为“关闭所有取消”。

8. Core终态收尾 ​

Core handler 根据 terminate 选择 session 操作,并把 runtime/transport 错误转换成模型可见工具错误。只有 LiveCell 才会写入本次 wait 的 cell_id、注册 executed_tool_calls,并判断是否需要关闭 cell bookkeeping。

源码位置:codex-rs/core/src/tools/code_mode/wait_handler.rs :: CodeModeWaitHandler::handle_call

rust
let wait_response = if args.terminate {
    exec.session
        .services
        .code_mode_service
        .terminate(cell_id)
        .await
} else {
    exec.session
        .services
        .code_mode_service
        .wait(codex_code_mode::WaitRequest {
            cell_id,
            yield_time_ms: args.yield_time_ms,
        })
        .await
}
.map_err(|error| {
    telemetry.finish(/*success*/ false);
    FunctionCallError::RespondToModel(error)
})?;

Yielded 保留 dispatch gate,让同一 cell 后续继续发嵌套工具;Result 与 Terminated 则依次结束 rollout code-cell trace、移除 dispatch gate 并记录 CellClosed analytics fact。

源码位置:codex-rs/core/src/tools/code_mode/wait_handler.rs :: terminal LiveCell handling

rust
if !matches!(response, codex_code_mode::RuntimeResponse::Yielded { .. }) {
    exec.session
        .services
        .rollout_thread_trace
        .code_cell_trace_context(
            exec.turn.sub_id.as_str(),
            runtime_cell_id.as_str(),
        )
        .record_ended(response);
    exec.session
        .services
        .code_mode_service
        .finish_cell_dispatch(runtime_cell_id);
    exec.session
        .services
        .analytics_events_client
        .track_code_mode_tool_call(
            codex_analytics::CodeModeToolCallFact::CellClosed {
                thread_id: exec.session.thread_id.to_string(),
                turn_id: exec.turn.sub_id.clone(),
                cell_id: runtime_cell_id.to_string(),
            },
        );
}

MissingCell 不走这段 live-cell 收尾。runtime 会返回一个空内容、带 exec cell ... not found 错误文本的 RuntimeResponse::Result,随后统一输出适配把它显示为 Script failed。这使“不存在的 cell”既保留 WaitOutcome 层的结构差异,又能使用普通 Code Mode 结果格式呈现给模型。

源码位置:codex-rs/code-mode-runtime/src/service.rs :: missing_cell_response

rust
fn missing_cell_response(cell_id: CellId) -> RuntimeResponse {
    RuntimeResponse::Result {
        error_text: Some(format!("exec cell {cell_id} not found")),
        cell_id,
        content_items: Vec::new(),
    }
}

在结果适配前,handler 还等待当前 Session 的 elicitations 清空。若 cell 内部工具触发了仍未结束的用户交互,wait 不会先把模型推进到下一轮,再留下悬空 elicitation。

9. 输出与耗时 ​

WaitOutcome 最终转换为 RuntimeResponse,再进入与初次 exec 共用的 handle_runtime_response。Yielded、Terminated 与成功 Result 的状态文字不同;带 error_text 的 Result 会额外追加 Script error: 内容并把工具输出标记为失败。

源码位置:codex-rs/core/src/tools/code_mode/mod.rs :: handle_runtime_response

rust
match response {
    RuntimeResponse::Yielded { content_items, .. } => {
        let mut content_items = into_function_call_output_content_items(content_items);
        sanitize_runtime_image_detail(exec.turn.as_ref(), &mut content_items);
        content_items = truncate_code_mode_result(content_items, max_output_tokens);
        prepend_script_status(&mut content_items, &script_status, started_at.elapsed());
        Ok(FunctionToolOutput::from_content(content_items, Some(true)))
    }
    RuntimeResponse::Terminated { content_items, .. } => {
        let mut content_items = into_function_call_output_content_items(content_items);
        sanitize_runtime_image_detail(exec.turn.as_ref(), &mut content_items);
        content_items = truncate_code_mode_result(content_items, max_output_tokens);
        prepend_script_status(&mut content_items, &script_status, started_at.elapsed());
        Ok(FunctionToolOutput::from_content(content_items, Some(true)))
    }
    RuntimeResponse::Result {
        content_items,
        error_text,
        ..
    } => {
        let mut content_items = into_function_call_output_content_items(content_items);
        sanitize_runtime_image_detail(exec.turn.as_ref(), &mut content_items);
        let success = error_text.is_none();
        if let Some(error_text) = error_text {
            content_items.push(FunctionCallOutputContentItem::InputText {
                text: format!("Script error:\n{error_text}"),
            });
        }
        content_items = truncate_code_mode_result(content_items, max_output_tokens);
        prepend_script_status(&mut content_items, &script_status, started_at.elapsed());
        Ok(FunctionToolOutput::from_content(content_items, Some(success)))
    }
}

主动终止在工具输出层被标记为 success,因为 control request 已按要求完成;它不等同于脚本自然成功。相反,MissingCell 被投影成带错误的 Result,所以 success 为 false。

状态前缀中的 wall time 从本次 wait handler 完成参数解析后开始计时,包含 transport wait、notification/elicitation 收口与结果适配前的等待,不是 cell 从初次 exec 启动以来的累计时间。

输出预算同样属于本次 wait。纯文本使用带原始 token/行数提示的格式化截断;含图片或音频的混合内容使用 function-output 策略,并为音频单独估算 token。

源码位置:codex-rs/core/src/tools/code_mode/mod.rs :: truncate_code_mode_result

rust
let max_output_tokens = resolve_max_tokens(max_output_tokens);
let policy = TruncationPolicy::Tokens(max_output_tokens);
if items
    .iter()
    .all(|item| matches!(item, FunctionCallOutputContentItem::InputText { .. }))
{
    let (truncated_items, _) =
        formatted_truncate_text_content_items_with_policy(&items, policy);
    return truncated_items;
}

truncate_function_output_items_with_policy(&items, policy, estimate_audio_token_count)

图片 detail 还会按当前 Turn 模型能力清理。因此 runtime 产生的内容不是原样穿透到模型,必须经过类型转换、图片策略、预算和状态前缀四层投影。

10. 遥测关联 ​

每次 exec 或 wait 都创建 CodeModeToolCallGuard。它保存的是本次工具调用的 thread、Turn、call ID、可选 cell ID、开始时间与状态;rust-v0.150.0 还把 TurnAnalyticsMetadata 快照传入 guard,使工具即使在 Turn 主事件之后完成,也能保留正确的 root Turn 与 workload 归属。

源码位置:codex-rs/core/src/tools/code_mode/telemetry.rs :: CodeModeToolCallGuard

rust
pub(super) struct CodeModeToolCallGuard {
    analytics: AnalyticsEventsClient,
    thread_id: String,
    turn_id: String,
    turn_metadata: Arc<dyn TurnAnalyticsMetadata>,
    call_id: String,
    pub(super) cell_id: Option<String>,
    tool_name: &'static str,
    started_at_ms: u64,
    status: CodeModeToolCallStatus,
}

pub(super) fn finish(&mut self, success: bool) {
    self.status = if success {
        CodeModeToolCallStatus::Completed
    } else {
        CodeModeToolCallStatus::Failed
    };
}

初始状态是 Interrupted。参数解析错误和 session 错误会在提前返回前显式设为 Failed;正常拿到 FunctionToolOutput 后依据 success_for_logging() 设为 Completed 或 Failed;future 被 Turn 中断并在显式 finish 前 Drop,则保留 Interrupted。

源码位置:codex-rs/core/src/tools/code_mode/telemetry.rs :: Drop for CodeModeToolCallGuard

rust
impl Drop for CodeModeToolCallGuard {
    fn drop(&mut self) {
        self.analytics
            .track_code_mode_tool_call(CodeModeToolCallFact::Completed {
                thread_id: self.thread_id.clone(),
                turn_id: self.turn_id.clone(),
                turn_metadata: self.turn_metadata.clone(),
                call_id: self.call_id.clone(),
                cell_id: self.cell_id.clone(),
                tool_name: self.tool_name.to_string(),
                started_at_ms: self.started_at_ms,
                completed_at_ms: codex_analytics::now_unix_millis(),
                status: self.status,
            });
    }
}

Drop 只发送一次 Completed fact,fact 内的 status 再区分 Completed、Failed、Interrupted。除此之外,Code Mode 还会发送 cell 与采样层事实:

Fact产生时机关联作用
CellStarted初次 exec 获得 cell ID绑定外层 exec call 与 cell
ChildStartedcell 发起嵌套工具绑定 Core child call 与 cell
CellClosedterminal wait/exec 收尾标记 cell 在哪个 Turn 关闭
SamplingResponseCompleted模型响应完成绑定 response 与其中的工具 call
Completedguard Drop记录单次 exec/wait 的耗时与终态

reducer 只有在 thread metadata 与 connection context 都存在时才消费这些事实,并把 guard status 映射为动态工具事件的 terminal status。Interrupted 不会被伪装成普通失败,也不会设置 ToolError failure kind。

11. 测试反推边界 ​

V1 client/provider 测试不需要启动 V8,可以直接验证取消退休、transport timeout、generation 和 connection cleanup:

text
cd codex-rs
cargo test -p codex-code-mode --lib -- --nocapture --test-threads=1
cargo test -p codex-core --lib tools::code_mode::wait_spec::tests::create_wait_tool_matches_expected_spec -- --nocapture --test-threads=1
cargo test -p codex-core --lib tools::code_mode::tests:: -- --nocapture --test-threads=1
cargo test -p codex-analytics completed_background_tool_item_emits_after_turn_event -- --nocapture --test-threads=1

第一条命令的 70 项测试覆盖 V1/gRPC client 状态、deadline、generation 与连接清理;wait schema 测试 1 项、Core Code Mode 测试 4 项、analytics 目标测试 1 项也都通过。Core 的 4 项中,与本文直接相关的是文本截断和超预算音频省略,另外两项检查嵌套工具 payload。

cancelled_wait_is_retired_before_next_wait_is_sent 先发出一个 60 秒 wait,取消其 caller,再提交第二个 wait。关键断言是 outgoing queue 先只出现 cancel frame,第二个 wait frame 必须等 host 返回旧 wait 的取消响应后才出现。

源码位置:codex-rs/code-mode/src/remote_session/connection/driver_tests.rs :: cancelled_wait_is_retired_before_next_wait_is_sent

rust
first_cancellation.cancel();
drop(first_rx);

harness
    .command_tx
    .send(DriverCommand::Wait {
        session,
        request: WaitRequest {
            cell_id: CellId::new("1".to_string()),
            yield_time_ms: 1,
        },
        caller_cancellation: CancellationToken::new(),
        response_tx: second_tx,
    })
    .await
    .expect("second wait command");
harness
    .outgoing_rx
    .recv()
    .await
    .expect("cancel request frame");
assert!(matches!(
    harness.outgoing_rx.try_recv(),
    Err(mpsc::error::TryRecvError::Empty)
));

这项测试证明 V1 client 不会把第二位 observer 提前发送;它不覆盖 gRPC CancelWait,也不验证 V8 runtime 是否已经响应第一位 observer 的取消。

runtime contract 测试 termination_cancels_pending_callbacks_before_responding 的输入是先启动一个阻塞 notification,再终止无限等待的 cell。断言顺序是 NotificationCancelled、terminate 返回 Terminated、最后 CellClosed,从而限定“终止响应必须晚于 callback 清理”。

源码位置:codex-rs/code-mode-runtime/src/service_contract_tests.rs :: termination_cancels_pending_callbacks_before_responding

rust
assert_eq!(
    service.terminate(cell_id("1")).await.unwrap(),
    WaitOutcome::LiveCell(RuntimeResponse::Terminated {
        cell_id: cell_id("1"),
        content_items: Vec::new(),
    })
);
assert!(delegate.notification_finished.load(Ordering::Acquire));
assert_eq!(
    next_event(&mut events_rx).await,
    DelegateEvent::NotificationCancelled
);
assert_eq!(
    next_event(&mut events_rx).await,
    DelegateEvent::CellClosed(cell_id("1"))
);

这项测试需要构建 V8 runtime。若目标平台没有对应 rusty_v8 预构建库且未启用源码构建,它会在 build script 阶段停止;在这种情况下,transport 测试不能替代 callback 顺序的动态断言。

analytics 测试 completed_background_tool_item_emits_after_turn_event 在 Turn 已完成后再输入带 Interrupted 状态和 metadata 的 Code Mode fact,断言最终动态工具事件仍携带原 Turn 与 root Turn。这验证 metadata 快照服务于迟到工具事实,但不覆盖真实 wait handler 的 RAII Drop 时机。

排查实际问题时可以按现象定位:第二次 wait 报 active observer,先看旧 wait 是否真正退休;wait 超时且连接失效,先看 transport deadline 而不是 V8 yield timer;terminate 已调用但迟迟未返回,检查 callback/notification drain;模型看到 Script failed 且 cell 未写入 telemetry,检查是否为 MissingCell;Turn 被中断但 cell 继续存在,检查 CodeModeInterrupt 与 dispatch gate 是否仍活跃。

至此,Code Mode 从 host、V8、cell、状态保持、嵌套工具到等待收尾的执行链已经闭合。下一篇Codex威胁模型将转向安全边界,解释模型请求、审批、sandbox、网络与执行后端分别能约束什么。