Skip to content

HeadTailBuffer与输出截断

解析 Unified Exec 的 head/tail 保留算法、中间字节省略、drain 合并和 omission marker 输出。

基于rust-v0.150.0
CodexRustExecution

HeadTailBuffer与输出截断 ​

Unified Exec 不能无限保留输出,但只留开头或结尾都会损失诊断信息。HeadTailBuffer 把容量分成稳定 head 与滚动 tail,超出的中间字节进入 omitted_bytes。本文承接Unified Exec轮询与等待和AsyncWatcher后台监控,专门讲字节保留、跨 drain 合并、模型响应截断与日志输出;它不负责 UTF-8 delta 分块。

1. 双预算 ​

容量现在是 const generic,而不是对象字段。head 取得一半容量,tail 取得余数;奇数容量因此让 tail 多一个字节。head 使用 Vec,tail 使用便于从前端淘汰的 VecDeque。生产默认容量是 1 MiB,测试可直接实例化 HeadTailBuffer<0>、HeadTailBuffer<1> 或 HeadTailBuffer<10>。

源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer

rust
#[derive(Debug, Default)]
pub(crate) struct HeadTailBuffer<const MAX_BYTES: usize = UNIFIED_EXEC_OUTPUT_MAX_BYTES> {
    head: Vec<u8>,
    tail: VecDeque<u8>,
    omitted_bytes: usize,
}

impl<const MAX_BYTES: usize> HeadTailBuffer<MAX_BYTES> {
    const HEAD_BUDGET: usize = MAX_BYTES / 2;
    const TAIL_BUDGET: usize = MAX_BYTES.saturating_sub(Self::HEAD_BUDGET);
}

2. 写入算法 ​

2.1 先填head ​

源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer::push_chunk、fill_head

rust
pub(crate) fn push_chunk(&mut self, chunk: &[u8]) {
    let chunk = self.fill_head(chunk);
    self.push_tail(chunk);
}

fn fill_head<'a>(&mut self, chunk: &'a [u8]) -> &'a [u8] {
    let remaining_head = Self::HEAD_BUDGET.saturating_sub(self.head.len());
    let (chunk_head, chunk_tail) = chunk
        .split_at_checked(remaining_head)
        .unwrap_or((chunk, &[]));
    self.head.extend_from_slice(chunk_head);
    chunk_tail
}

head 填满后不再改变。split_at_checked 让短 chunk 整体进入 head,返回空余量;零容量时 head budget 为 0,整块输入自然流向 tail 逻辑,无需单独的 max_bytes == 0 分支。

2.2 滚动tail ​

源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer::push_tail

rust
let remaining_tail = Self::TAIL_BUDGET.saturating_sub(tail.len());
let excess_tail = chunk.len().saturating_sub(remaining_tail);
*omitted_bytes = omitted_bytes.saturating_add(excess_tail);

let chunk = match excess_tail.checked_sub(tail.len()) {
    None => {
        tail.drain(..excess_tail);
        chunk
    }
    Some(skip) => {
        tail.clear();
        &chunk[skip..]
    }
};
tail.extend(chunk);

excess_tail 一次计算出“为新数据腾出空间需要丢多少字节”。若旧 tail 足够承担淘汰,只从其前端 drain;若新 chunk 自身已经大于整个 tail 预算,则清空旧 tail 并跳过新 chunk 前部。零容量时 TAIL_BUDGET == 0,最终 slice 为空,但 omitted_bytes 仍累加完整输入长度。

3. 转移与合并 ​

HeadTailBuffer 已删除自己的 drain 方法。轮询器在共享锁内直接 mem::take 整个值,依靠 Default 创建同一 const capacity 的空 buffer;锁外再调用 push_buffer 合并。容量属于类型,因此无需在转移时复制 max_bytes 配置。

源码位置:

  • codex-rs/core/src/unified_exec/process_manager.rs :: collect_output_until_deadline
  • codex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer::push_buffer
rust
let mut guard = output_buffer.lock().await;
drained_output = std::mem::take(&mut *guard);

push_buffer 不能简单地把 source head 和 tail 当成普通 chunk 依次写入,因为 source 自己已经是原始流的摘要。它先累加 source omission,再保留目标已有 head;若 source tail 已占满整个 tail budget,就直接替换目标 tail,并把目标旧 tail 与 source 未用 head 计入省略。

源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer::push_buffer

rust
pub(crate) fn push_buffer(&mut self, buffer: Self) {
    let Self {
        head,
        tail,
        omitted_bytes,
    } = buffer;

    self.omitted_bytes = self.omitted_bytes.saturating_add(omitted_bytes);

    let overflow = if self.head.is_empty() {
        self.head = head;
        &[]
    } else {
        self.fill_head(&head)
    };

    if tail.len() == Self::TAIL_BUDGET {
        self.omitted_bytes = self
            .omitted_bytes
            .saturating_add(self.tail.len())
            .saturating_add(overflow.len());
        self.tail = tail;
    } else {
        self.push_tail(overflow);
        if self.tail.is_empty() {
            self.tail = tail;
        } else {
            let (first, second) = tail.as_slices();
            self.push_tail(first);
            self.push_tail(second);
        }
    }
}

这使轮询器只在锁内转移所有权,锁外再合并;多次 producer drain 表现得像连续写入同一个有界 buffer,已经发生的省略不会归零。

4. 输出呈现 ​

4.1 marker ​

to_bytes 只连接 head 与 tail;公开输出使用 to_bytes_with_omission_marker,在二者之间插入 ... N bytes omitted ...。marker 是 Codex 添加的说明,不是命令真实输出,也不能恢复丢失内容。

源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: to_bytes_with_omission_marker

rust
if self.omitted_bytes == 0 {
    return self.to_bytes();
}
let marker = format_output_omission_marker(self.omitted_bytes);
let marker_delimiter_bytes = 2;
let mut out = Vec::with_capacity(
    self.retained_bytes()
        .saturating_add(marker.len())
        .saturating_add(marker_delimiter_bytes),
);
out.extend_from_slice(&self.head);
out.push(b'\n');
out.extend_from_slice(marker.as_bytes());
out.push(b'\n');
out.extend(self.tail.iter().copied());

4.2 原始统计 ​

响应文本只保留 head/tail,但 token 估计基于 total_bytes = retained + omitted,同时把 omission bytes 单独暴露给工具输出。

源码位置:codex-rs/core/src/unified_exec/process_manager.rs :: write_stdin

rust
let original_token_count = usize::try_from(approx_tokens_from_byte_count(
    collected_output.total_bytes(),
))
.unwrap_or(usize::MAX);
let output_omitted_bytes = NonZeroUsize::new(collected_output.omitted_bytes());
let collected = collected_output.to_bytes_with_omission_marker();

4.3 模型预算 ​

HeadTailBuffer 的 1 MiB 是采集上限,不等于模型能看到 1 MiB。ExecCommandToolOutput::model_output_policy 比较模型提供的 TruncationPolicy 和调用者请求的 max_output_tokens,选择 byte budget 更小的一方。

源码位置:codex-rs/core/src/tools/context.rs :: ExecCommandToolOutput::model_output_policy、truncated_output_with_policy

rust
fn model_output_policy(&self) -> TruncationPolicy {
    let requested_policy = TruncationPolicy::Tokens(resolve_max_tokens(self.max_output_tokens));
    if requested_policy.byte_budget() < self.truncation_policy.byte_budget() {
        requested_policy
    } else {
        self.truncation_policy
    }
}

如果采集层已经省略字节,模型层必须保留 omission marker。即使 raw_output 因其他调用路径没有带 marker,只要 output_omitted_bytes 存在,序列化也会补回;模型层再次截断时还会加入 original token count warning。

rust
let marker = format_output_omission_marker(omitted_bytes.get());
if text.len() <= policy.byte_budget() {
    return if text.contains(&marker) {
        text
    } else {
        format!("{marker}\n{text}")
    };
}

let original_token_count = self
    .original_token_count
    .unwrap_or_else(|| approx_token_count(&text));
let truncated = truncate_text(&text, policy);

4.4 元数据预算 ​

模型响应还包含 chunk ID、墙钟时间、exit code、session ID 和 original token count。response_text 给完整响应保留 1.2 * truncation_policy 的 byte budget,并从中扣除 header;若输出连同 warning/marker 仍超出,循环缩小输出 policy,避免 history 对已经截断的响应再截一次。

源码位置:codex-rs/core/src/tools/context.rs :: ExecCommandToolOutput::response_text

rust
let header = self.response_header();
let output_budget = (self.truncation_policy * 1.2)
    .byte_budget()
    .saturating_sub(header.len().saturating_add(/*rhs*/ 1));
let mut policy = self.model_output_policy();
let mut output = self.truncated_output_with_policy(policy);

while output.len() > output_budget && policy.byte_budget() > 0 {
    let excess_bytes = output.len() - output_budget;
    policy = match policy {
        TruncationPolicy::Bytes(bytes) => {
            TruncationPolicy::Bytes(bytes.saturating_sub(excess_bytes))
        }
        TruncationPolicy::Tokens(tokens) => TruncationPolicy::Tokens(
            tokens.saturating_sub(TruncationPolicy::Bytes(excess_bytes).token_budget()),
        ),
    };
    output = self.truncated_output_with_policy(policy);
}

日志走另一条路径:log_output 使用完整 raw_output 和响应 header,不继承模型请求的 max_output_tokens;若 omission metadata 存在但文本缺 marker,同样补回。这让诊断日志与模型上下文拥有不同预算。

5. 三层预算 ​

6. 四组验证 ​

6.1 容量边界 ​

HeadTailBuffer 模块的 6 项测试覆盖 10 字节 head/tail、零容量、一字节容量、超大单 chunk、多 chunk 填充和空/单字节高频输入。代表性用例写入 0123456789ab,断言 marker 为 2 bytes omitted,head 保留 01234,tail 保留 789ab。

源码位置:codex-rs/core/src/unified_exec/head_tail_buffer_tests.rs

text
cd codex-rs
cargo test -p codex-core --lib 'unified_exec::head_tail_buffer::tests::' -- --test-threads=1

6.2 跨drain合并 ​

output_collection_stays_bounded_across_repeated_drains 用 10 字节 buffer 分四轮生产并强制每轮被 collector 取空,最终断言与连续写入同一个 buffer 完全相等。output_collection_preserves_omissions_from_drained_buffer 先制造 omission,再执行 mem::take,断言 source omission 被保留。

源码位置:

  • codex-rs/core/src/unified_exec/process_manager_tests.rs :: output_collection_stays_bounded_across_repeated_drains
  • codex-rs/core/src/unified_exec/process_manager_tests.rs :: output_collection_preserves_omissions_from_drained_buffer
text
cd codex-rs
cargo test -p codex-core --lib 'unified_exec::process_manager::tests::output_collection_' -- --test-threads=1

6.3 响应序列化 ​

3 项 context 测试分别断言请求 token 上限会截断模型响应但不会截断日志、完整 response 为 header 预留空间且保留 Bytes/Tokens policy 单位、采集层 omission marker 在模型二次截断和日志缺 marker 时仍只出现一次。

源码位置:

  • codex-rs/core/src/tools/context_tests.rs :: exec_command_tool_output_formats_truncated_response
  • codex-rs/core/src/tools/context_tests.rs :: exec_command_tool_output_reserves_metadata_budget_and_preserves_policy_units
  • codex-rs/core/src/tools/context_tests.rs :: exec_command_tool_output_preserves_omission_metadata_when_truncated
text
cd codex-rs
cargo test -p codex-core --lib 'tools::context::tests::exec_command_tool_output_' -- --test-threads=1

6.4 真实大输出 ​

两个集成测试分别让 exec_command 与 write_stdin 请求 70,000 tokens,但模型 policy 只有 50,断言最终只有一个 truncation notice,并保留 original token count。unified_exec_formats_large_output_summary 产生约 1.3 MiB 输出,断言结束事件和模型响应都保留 HEAD、TAIL 与 byte omission marker,original token count 按原始总字节估算。

源码位置:

  • codex-rs/core/tests/suite/unified_exec.rs :: exec_command_clamps_model_requested_max_output_tokens_to_policy
  • codex-rs/core/tests/suite/unified_exec.rs :: write_stdin_clamps_model_requested_max_output_tokens_to_policy
  • codex-rs/core/tests/suite/unified_exec.rs :: unified_exec_formats_large_output_summary
text
cd codex-rs
RUST_MIN_STACK=8388608 cargo test -p codex-core --test all exec_command_clamps_model_requested_max_output_tokens_to_policy -- --test-threads=1
RUST_MIN_STACK=8388608 cargo test -p codex-core --test all write_stdin_clamps_model_requested_max_output_tokens_to_policy -- --test-threads=1
RUST_MIN_STACK=8388608 cargo test -p codex-core --test all unified_exec_formats_large_output_summary -- --test-threads=1

这些测试证明字节位置、容量、预算和元数据,不证明按行或按 Unicode 字符截断。HeadTailBuffer 可以在多字节字符中间切断;UTF-8 安全只属于 AsyncWatcher delta 分帧,模型文本通过 lossy 转换处理 retained raw bytes。

7. 源码排查 ​

text
rg -n "HeadTailBuffer|push_chunk|fill_head|push_tail" codex-rs/core/src/unified_exec
rg -n "to_bytes_with_omission_marker|push_buffer|total_bytes" codex-rs/core/src/unified_exec
rg -n "model_output_policy|truncated_output_with_policy|response_text|log_output" codex-rs/core/src/tools/context.rs

核心不变量是:retained_bytes <= MAX_BYTES,total_bytes = retained + omitted,head 稳定、tail 滚动;mem::take 与多轮合并不能丢失 omitted 元数据;模型预算和日志预算不会改变采集层已经记录的原始统计。