HeadTailBuffer与输出截断
Unified Exec 不能无限保留输出,但只留开头或结尾都会损失诊断信息。HeadTailBuffer 把容量分成稳定 head 与滚动 tail,超出的中间字节进入 omitted_bytes。本文承接Unified Exec轮询与等待和AsyncWatcher后台监控,专门讲字节保留、跨 drain 合并、模型响应截断与日志输出;它不负责 UTF-8 delta 分块。
1. 双预算
容量现在是 const generic,而不是对象字段。head 取得一半容量,tail 取得余数;奇数容量因此让 tail 多一个字节。head 使用 Vec,tail 使用便于从前端淘汰的 VecDeque。生产默认容量是 1 MiB,测试可直接实例化 HeadTailBuffer<0>、HeadTailBuffer<1> 或 HeadTailBuffer<10>。
源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer
#[derive(Debug, Default)]
pub(crate) struct HeadTailBuffer<const MAX_BYTES: usize = UNIFIED_EXEC_OUTPUT_MAX_BYTES> {
head: Vec<u8>,
tail: VecDeque<u8>,
omitted_bytes: usize,
}
impl<const MAX_BYTES: usize> HeadTailBuffer<MAX_BYTES> {
const HEAD_BUDGET: usize = MAX_BYTES / 2;
const TAIL_BUDGET: usize = MAX_BYTES.saturating_sub(Self::HEAD_BUDGET);
}2. 写入算法
2.1 先填head
源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer::push_chunk、fill_head
pub(crate) fn push_chunk(&mut self, chunk: &[u8]) {
let chunk = self.fill_head(chunk);
self.push_tail(chunk);
}
fn fill_head<'a>(&mut self, chunk: &'a [u8]) -> &'a [u8] {
let remaining_head = Self::HEAD_BUDGET.saturating_sub(self.head.len());
let (chunk_head, chunk_tail) = chunk
.split_at_checked(remaining_head)
.unwrap_or((chunk, &[]));
self.head.extend_from_slice(chunk_head);
chunk_tail
}head 填满后不再改变。split_at_checked 让短 chunk 整体进入 head,返回空余量;零容量时 head budget 为 0,整块输入自然流向 tail 逻辑,无需单独的 max_bytes == 0 分支。
2.2 滚动tail
源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer::push_tail
let remaining_tail = Self::TAIL_BUDGET.saturating_sub(tail.len());
let excess_tail = chunk.len().saturating_sub(remaining_tail);
*omitted_bytes = omitted_bytes.saturating_add(excess_tail);
let chunk = match excess_tail.checked_sub(tail.len()) {
None => {
tail.drain(..excess_tail);
chunk
}
Some(skip) => {
tail.clear();
&chunk[skip..]
}
};
tail.extend(chunk);excess_tail 一次计算出“为新数据腾出空间需要丢多少字节”。若旧 tail 足够承担淘汰,只从其前端 drain;若新 chunk 自身已经大于整个 tail 预算,则清空旧 tail 并跳过新 chunk 前部。零容量时 TAIL_BUDGET == 0,最终 slice 为空,但 omitted_bytes 仍累加完整输入长度。
3. 转移与合并
HeadTailBuffer 已删除自己的 drain 方法。轮询器在共享锁内直接 mem::take 整个值,依靠 Default 创建同一 const capacity 的空 buffer;锁外再调用 push_buffer 合并。容量属于类型,因此无需在转移时复制 max_bytes 配置。
源码位置:
codex-rs/core/src/unified_exec/process_manager.rs :: collect_output_until_deadlinecodex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer::push_buffer
let mut guard = output_buffer.lock().await;
drained_output = std::mem::take(&mut *guard);push_buffer 不能简单地把 source head 和 tail 当成普通 chunk 依次写入,因为 source 自己已经是原始流的摘要。它先累加 source omission,再保留目标已有 head;若 source tail 已占满整个 tail budget,就直接替换目标 tail,并把目标旧 tail 与 source 未用 head 计入省略。
源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: HeadTailBuffer::push_buffer
pub(crate) fn push_buffer(&mut self, buffer: Self) {
let Self {
head,
tail,
omitted_bytes,
} = buffer;
self.omitted_bytes = self.omitted_bytes.saturating_add(omitted_bytes);
let overflow = if self.head.is_empty() {
self.head = head;
&[]
} else {
self.fill_head(&head)
};
if tail.len() == Self::TAIL_BUDGET {
self.omitted_bytes = self
.omitted_bytes
.saturating_add(self.tail.len())
.saturating_add(overflow.len());
self.tail = tail;
} else {
self.push_tail(overflow);
if self.tail.is_empty() {
self.tail = tail;
} else {
let (first, second) = tail.as_slices();
self.push_tail(first);
self.push_tail(second);
}
}
}这使轮询器只在锁内转移所有权,锁外再合并;多次 producer drain 表现得像连续写入同一个有界 buffer,已经发生的省略不会归零。
4. 输出呈现
4.1 marker
to_bytes 只连接 head 与 tail;公开输出使用 to_bytes_with_omission_marker,在二者之间插入 ... N bytes omitted ...。marker 是 Codex 添加的说明,不是命令真实输出,也不能恢复丢失内容。
源码位置:codex-rs/core/src/unified_exec/head_tail_buffer.rs :: to_bytes_with_omission_marker
if self.omitted_bytes == 0 {
return self.to_bytes();
}
let marker = format_output_omission_marker(self.omitted_bytes);
let marker_delimiter_bytes = 2;
let mut out = Vec::with_capacity(
self.retained_bytes()
.saturating_add(marker.len())
.saturating_add(marker_delimiter_bytes),
);
out.extend_from_slice(&self.head);
out.push(b'\n');
out.extend_from_slice(marker.as_bytes());
out.push(b'\n');
out.extend(self.tail.iter().copied());4.2 原始统计
响应文本只保留 head/tail,但 token 估计基于 total_bytes = retained + omitted,同时把 omission bytes 单独暴露给工具输出。
源码位置:codex-rs/core/src/unified_exec/process_manager.rs :: write_stdin
let original_token_count = usize::try_from(approx_tokens_from_byte_count(
collected_output.total_bytes(),
))
.unwrap_or(usize::MAX);
let output_omitted_bytes = NonZeroUsize::new(collected_output.omitted_bytes());
let collected = collected_output.to_bytes_with_omission_marker();4.3 模型预算
HeadTailBuffer 的 1 MiB 是采集上限,不等于模型能看到 1 MiB。ExecCommandToolOutput::model_output_policy 比较模型提供的 TruncationPolicy 和调用者请求的 max_output_tokens,选择 byte budget 更小的一方。
源码位置:codex-rs/core/src/tools/context.rs :: ExecCommandToolOutput::model_output_policy、truncated_output_with_policy
fn model_output_policy(&self) -> TruncationPolicy {
let requested_policy = TruncationPolicy::Tokens(resolve_max_tokens(self.max_output_tokens));
if requested_policy.byte_budget() < self.truncation_policy.byte_budget() {
requested_policy
} else {
self.truncation_policy
}
}如果采集层已经省略字节,模型层必须保留 omission marker。即使 raw_output 因其他调用路径没有带 marker,只要 output_omitted_bytes 存在,序列化也会补回;模型层再次截断时还会加入 original token count warning。
let marker = format_output_omission_marker(omitted_bytes.get());
if text.len() <= policy.byte_budget() {
return if text.contains(&marker) {
text
} else {
format!("{marker}\n{text}")
};
}
let original_token_count = self
.original_token_count
.unwrap_or_else(|| approx_token_count(&text));
let truncated = truncate_text(&text, policy);4.4 元数据预算
模型响应还包含 chunk ID、墙钟时间、exit code、session ID 和 original token count。response_text 给完整响应保留 1.2 * truncation_policy 的 byte budget,并从中扣除 header;若输出连同 warning/marker 仍超出,循环缩小输出 policy,避免 history 对已经截断的响应再截一次。
源码位置:codex-rs/core/src/tools/context.rs :: ExecCommandToolOutput::response_text
let header = self.response_header();
let output_budget = (self.truncation_policy * 1.2)
.byte_budget()
.saturating_sub(header.len().saturating_add(/*rhs*/ 1));
let mut policy = self.model_output_policy();
let mut output = self.truncated_output_with_policy(policy);
while output.len() > output_budget && policy.byte_budget() > 0 {
let excess_bytes = output.len() - output_budget;
policy = match policy {
TruncationPolicy::Bytes(bytes) => {
TruncationPolicy::Bytes(bytes.saturating_sub(excess_bytes))
}
TruncationPolicy::Tokens(tokens) => TruncationPolicy::Tokens(
tokens.saturating_sub(TruncationPolicy::Bytes(excess_bytes).token_budget()),
),
};
output = self.truncated_output_with_policy(policy);
}日志走另一条路径:log_output 使用完整 raw_output 和响应 header,不继承模型请求的 max_output_tokens;若 omission metadata 存在但文本缺 marker,同样补回。这让诊断日志与模型上下文拥有不同预算。
5. 三层预算
6. 四组验证
6.1 容量边界
HeadTailBuffer 模块的 6 项测试覆盖 10 字节 head/tail、零容量、一字节容量、超大单 chunk、多 chunk 填充和空/单字节高频输入。代表性用例写入 0123456789ab,断言 marker 为 2 bytes omitted,head 保留 01234,tail 保留 789ab。
源码位置:codex-rs/core/src/unified_exec/head_tail_buffer_tests.rs
cd codex-rs
cargo test -p codex-core --lib 'unified_exec::head_tail_buffer::tests::' -- --test-threads=16.2 跨drain合并
output_collection_stays_bounded_across_repeated_drains 用 10 字节 buffer 分四轮生产并强制每轮被 collector 取空,最终断言与连续写入同一个 buffer 完全相等。output_collection_preserves_omissions_from_drained_buffer 先制造 omission,再执行 mem::take,断言 source omission 被保留。
源码位置:
codex-rs/core/src/unified_exec/process_manager_tests.rs :: output_collection_stays_bounded_across_repeated_drainscodex-rs/core/src/unified_exec/process_manager_tests.rs :: output_collection_preserves_omissions_from_drained_buffer
cd codex-rs
cargo test -p codex-core --lib 'unified_exec::process_manager::tests::output_collection_' -- --test-threads=16.3 响应序列化
3 项 context 测试分别断言请求 token 上限会截断模型响应但不会截断日志、完整 response 为 header 预留空间且保留 Bytes/Tokens policy 单位、采集层 omission marker 在模型二次截断和日志缺 marker 时仍只出现一次。
源码位置:
codex-rs/core/src/tools/context_tests.rs :: exec_command_tool_output_formats_truncated_responsecodex-rs/core/src/tools/context_tests.rs :: exec_command_tool_output_reserves_metadata_budget_and_preserves_policy_unitscodex-rs/core/src/tools/context_tests.rs :: exec_command_tool_output_preserves_omission_metadata_when_truncated
cd codex-rs
cargo test -p codex-core --lib 'tools::context::tests::exec_command_tool_output_' -- --test-threads=16.4 真实大输出
两个集成测试分别让 exec_command 与 write_stdin 请求 70,000 tokens,但模型 policy 只有 50,断言最终只有一个 truncation notice,并保留 original token count。unified_exec_formats_large_output_summary 产生约 1.3 MiB 输出,断言结束事件和模型响应都保留 HEAD、TAIL 与 byte omission marker,original token count 按原始总字节估算。
源码位置:
codex-rs/core/tests/suite/unified_exec.rs :: exec_command_clamps_model_requested_max_output_tokens_to_policycodex-rs/core/tests/suite/unified_exec.rs :: write_stdin_clamps_model_requested_max_output_tokens_to_policycodex-rs/core/tests/suite/unified_exec.rs :: unified_exec_formats_large_output_summary
cd codex-rs
RUST_MIN_STACK=8388608 cargo test -p codex-core --test all exec_command_clamps_model_requested_max_output_tokens_to_policy -- --test-threads=1
RUST_MIN_STACK=8388608 cargo test -p codex-core --test all write_stdin_clamps_model_requested_max_output_tokens_to_policy -- --test-threads=1
RUST_MIN_STACK=8388608 cargo test -p codex-core --test all unified_exec_formats_large_output_summary -- --test-threads=1这些测试证明字节位置、容量、预算和元数据,不证明按行或按 Unicode 字符截断。HeadTailBuffer 可以在多字节字符中间切断;UTF-8 安全只属于 AsyncWatcher delta 分帧,模型文本通过 lossy 转换处理 retained raw bytes。
7. 源码排查
rg -n "HeadTailBuffer|push_chunk|fill_head|push_tail" codex-rs/core/src/unified_exec
rg -n "to_bytes_with_omission_marker|push_buffer|total_bytes" codex-rs/core/src/unified_exec
rg -n "model_output_policy|truncated_output_with_policy|response_text|log_output" codex-rs/core/src/tools/context.rs核心不变量是:retained_bytes <= MAX_BYTES,total_bytes = retained + omitted,head 稳定、tail 滚动;mem::take 与多轮合并不能丢失 omitted 元数据;模型预算和日志预算不会改变采集层已经记录的原始统计。
