Context归一化算法
ContextManager 保存的是运行过程中收到的 raw history,而模型请求需要满足另一组约束:工具调用必须有对应输出,输出不能脱离调用,图片和音频必须符合模型能力。Codex 没有在每次事件写入时强行修整这些形状,而是在消费快照的 for_prompt 路径中归一化。
本文回答一个具体问题:一份包含残缺工具对和不支持媒体的 history,如何变成可发送的 prompt;哪些修复只存在于这个 prompt 快照,哪些错误会在 debug 构建直接暴露?
读者需要了解 ResponseItem、ContextManager::for_prompt 和 Rust 的 Vec 原地修改。建议先读 ContextHistory读写,它解释 raw history、写时复制和快照所有权;本文不重复 record_items 的追加与 token 截断,只追踪 normalize.rs。
1. 投影入口
1.1 调用顺序
源码位置:codex-rs/core/src/context_manager/history.rs :: for_prompt、normalize_history
/// Returns the history prepared for sending to the model.
pub(crate) fn for_prompt(mut self, input_modalities: &[InputModality]) -> Vec<ResponseItem> {
self.normalize_history(input_modalities);
Arc::unwrap_or_clone(self.items)
}
fn normalize_history(&mut self, input_modalities: &[InputModality]) {
let items = Arc::make_mut(&mut self.items);
// all function/tool calls must have a corresponding output
normalize::ensure_call_outputs_present(items);
// all outputs must have a corresponding function/tool call
normalize::remove_orphan_outputs(items);
// strip images when model does not support them
normalize::strip_images_when_unsupported(input_modalities, items);
// strip audio when model does not support it
normalize::strip_audio_when_unsupported(input_modalities, items);
}for_prompt 消费的是一个 ContextManager 快照,所以 Arc::make_mut 只会修改这个快照;Session 中的 raw history 不会因为“为某个模型准备 prompt”而被偷偷写回。四步顺序也有因果关系:先建立 call/output 对,再删除无法配对的 output,最后才按能力处理媒体。媒体替换不会改变配对索引,工具修复却可能插入或删除 item,因此不能把四步任意交换。
在当前版本,ContextManager 内部保存的是 ResponseItemEnvelope,归一化只替换 envelope 内的 ResponseItem, 不会丢失 rollout 所需的 metadata。对只需要模型输入的调用方,for_prompt 返回裸 item;需要保留 envelope 的内部调用则使用 for_prompt_annotated。因此“归一化不回写 live history”同时适用于 item 内容和 envelope metadata。
1.2 消费者边界
compact、普通模型采样、远程 compact 和 prompt debug 都会在请求前调用 for_prompt。这意味着归一化结果属于一次请求的输入投影,不等于 rollout 中的 durable item,也不等于 raw_items() 看到的容器。合成输出可能只在 prompt 中出现,下一次投影会重新计算它。
2. 配对补全
2.1 输出ID收集
源码位置:codex-rs/core/src/context_manager/normalize.rs :: ensure_call_outputs_present
pub(crate) fn ensure_call_outputs_present(items: &mut Vec<ResponseItem>) {
let mut function_output_ids = HashSet::new();
let mut tool_search_output_ids = HashSet::new();
let mut custom_tool_output_ids = HashSet::new();
for item in items.iter() {
match item {
ResponseItem::FunctionCallOutput { call_id, .. } => {
function_output_ids.insert(call_id.as_str());
}
ResponseItem::ToolSearchOutput {
call_id: Some(call_id),
..
} => {
tool_search_output_ids.insert(call_id.as_str());
}
ResponseItem::CustomToolCallOutput { call_id, .. } => {
custom_tool_output_ids.insert(call_id.as_str());
}
_ => {}
}
}
let mut missing_outputs_to_insert: Vec<(usize, ResponseItem)> = Vec::new();
for (idx, item) in items.iter().enumerate() {
match item {
ResponseItem::FunctionCall { id, call_id, .. }
if !function_output_ids.contains(call_id.as_str()) =>
{
missing_outputs_to_insert.push((
idx,
ResponseItem::FunctionCallOutput {
id: synthetic_output_id("fco", id.as_deref()),
call_id: call_id.clone(),
output: FunctionCallOutputPayload::from_text("aborted".to_string()),
internal_chat_message_metadata_passthrough: None,
},
));
}
_ => {}
}
}
for (idx, output_item) in missing_outputs_to_insert.into_iter().rev() {
items.insert(idx + 1, output_item);
}
}实现先建立三个 HashSet,再扫描 call。这样判断是按 call_id 的存在关系,而不是假设 output 必须紧邻 call。收集插入位置后从后向前 insert,避免前面的插入改变后面尚未处理的位置;输出总是放在对应 call 后面,保持模型请求中的局部顺序。
2.2 四种调用
当前代码对四类 call 做补全,但输出类型并不完全相同:
| call | 匹配 output | 缺失时的 prompt item | debug/release |
|---|---|---|---|
FunctionCall | FunctionCallOutput(call_id) | aborted 文本 | 直接补全 |
ToolSearchCall(Some(id)) | ToolSearchOutput(Some(id)) | client completed、空 tools | 直接补全 |
CustomToolCall | CustomToolCallOutput | aborted 文本 | debug panic,release 补全 |
LocalShellCall(Some(id)) | FunctionCallOutput | aborted 文本 | debug panic,release 补全 |
LocalShellCall 使用 FunctionCallOutput 是协议约定,而不是漏写了 LocalShellCallOutput。ToolSearchCall 没有 call_id 时不会进入补全条件;它的 output 也可能是 server execution,后续孤儿清理会采用不同规则。
3. 稳定合成
3.1 合成ID稳定性
源码位置:codex-rs/core/src/context_manager/normalize.rs :: synthetic_output_id
// Changing this value would change model-visible IDs and invalidate prompt caches.
const SYNTHETIC_OUTPUT_ID_NAMESPACE: Uuid =
Uuid::from_u128(0x90d38d3e_6a5b_4d52_bfe2_2f1e634bfac4);
fn synthetic_output_id(prefix: &str, item_id: Option<&str>) -> Option<ResponseItemId> {
let source_id = item_id.filter(|id| !id.is_empty())?;
let name = format!("{prefix}:{source_id}");
Some(ResponseItemId::with_suffix(
prefix,
Uuid::new_v5(&SYNTHETIC_OUTPUT_ID_NAMESPACE, name.as_bytes()),
))
}合成 output 不会写回 raw history,因此每次 retry、resume 或重复调用 for_prompt 都会重新生成。若这里使用随机 UUID,同一份 raw history 的 prompt item ID 就会变化,破坏 provider 的 prompt cache 和测试中的重复投影一致性。实现用固定 namespace 加源 call 的 item ID 做 UUID v5;源 call 没有 ID 或 ID 为空时返回 None,保留旧历史的兼容行为。
3.2 调用顺序不变
补全函数只插入 output,不重排原有 call 和后续消息。对应测试用“call + later user message”检查合成 output 位于二者之间,并重复投影两次比较完整结果。这个性质很关键:稳定 ID 解决重复投影的一致性,插入位置解决模型看到的调用顺序;两者是不同不变量。
4. 孤儿清理
4.1 类型化索引
源码位置:codex-rs/core/src/context_manager/normalize.rs :: remove_orphan_outputs
pub(crate) fn remove_orphan_outputs(items: &mut Vec<ResponseItem>) {
let function_call_ids: HashSet<String> = items
.iter()
.filter_map(|i| match i {
ResponseItem::FunctionCall { call_id, .. } => Some(call_id.clone()),
_ => None,
})
.collect();
let tool_search_call_ids: HashSet<String> = items
.iter()
.filter_map(|i| match i {
ResponseItem::ToolSearchCall {
call_id: Some(call_id),
..
} => Some(call_id.clone()),
_ => None,
})
.collect();
let local_shell_call_ids: HashSet<String> = items
.iter()
.filter_map(|i| match i {
ResponseItem::LocalShellCall {
call_id: Some(call_id),
..
} => Some(call_id.clone()),
_ => None,
})
.collect();
let custom_tool_call_ids: HashSet<String> = items
.iter()
.filter_map(|i| match i {
ResponseItem::CustomToolCall { call_id, .. } => Some(call_id.clone()),
_ => None,
})
.collect();
items.retain(|item| match item {
ResponseItem::FunctionCallOutput { call_id, .. } => {
function_call_ids.contains(call_id) || local_shell_call_ids.contains(call_id)
}
ResponseItem::CustomToolCallOutput { call_id, .. } => {
custom_tool_call_ids.contains(call_id)
}
ResponseItem::ToolSearchOutput { execution, .. } if execution == "server" => true,
ResponseItem::ToolSearchOutput {
call_id: Some(call_id),
..
} => tool_search_call_ids.contains(call_id),
ResponseItem::ToolSearchOutput { call_id: None, .. } => true,
_ => true,
});
}孤儿规则按类型区分,而不是所有 output 共用一个集合:普通 function output 可以匹配 FunctionCall 或 LocalShellCall;custom output 只能匹配 custom call;client tool-search output 必须有对应 call;server output 即使在当前 vector 中没有 call 也保留;没有 call_id 的 tool-search output 也保留。最后两条是协议语义,不是“校验不严格”。
4.2 两阶段边界
先补全再清理看似多了一次扫描,但它避免了“缺失 output 被当作孤儿删除”的矛盾。补全阶段建立缺失的配对,清理阶段再处理原本就没有 call 的 output。若把清理放在前面,孤儿规则无法区分“稍后应该合成的 output”和真正来自损坏流的 output。
5. 构建边界
5.1 error_or_panic
源码位置:codex-rs/core/src/util.rs :: error_or_panic
pub(crate) fn error_or_panic(message: impl std::string::ToString) {
if cfg!(debug_assertions) {
panic!("{}", message.to_string());
} else {
error!("{}", message.to_string());
}
}CustomToolCall、带 ID 的 LocalShellCall 和孤儿 output 会调用这个函数。debug 构建把协议不变量当成开发期立即失败;release 构建记录错误并继续执行,随后补全或 retain 删除使请求仍能形成。这不是“debug 没有归一化、release 才有归一化”:两种构建都执行同一流程,只是异常分支的反馈强度不同。
因此阅读测试时必须注意 #[cfg(debug_assertions)] 与 #[cfg(not(debug_assertions))]。同一个残缺输入,在 debug 测试中预期 should_panic,在 release 测试中预期合成 output 或删除孤儿。只运行其中一组不能宣称两种构建都已验证。
6. 媒体降级
6.1 图片替换
源码位置:codex-rs/core/src/context_manager/normalize.rs :: strip_images_when_unsupported
pub(crate) fn strip_images_when_unsupported(
input_modalities: &[InputModality],
items: &mut [ResponseItem],
) {
let supports_images = input_modalities.contains(&InputModality::Image);
if supports_images {
return;
}
for item in items.iter_mut() {
match item {
ResponseItem::Message { content, .. } => {
let mut normalized_content = Vec::with_capacity(content.len());
for content_item in content.iter() {
match content_item {
ContentItem::InputImage { .. } => {
normalized_content.push(ContentItem::InputText {
text: IMAGE_CONTENT_OMITTED_PLACEHOLDER.to_string(),
});
}
_ => normalized_content.push(content_item.clone()),
}
}
*content = normalized_content;
}
ResponseItem::FunctionCallOutput { output, .. }
| ResponseItem::CustomToolCallOutput { output, .. } => {
if let Some(content_items) = output.content_items_mut() {
let mut normalized_content_items = Vec::with_capacity(content_items.len());
for content_item in content_items.iter() {
match content_item {
FunctionCallOutputContentItem::InputImage { .. } => {
normalized_content_items.push(
FunctionCallOutputContentItem::InputText {
text: IMAGE_CONTENT_OMITTED_PLACEHOLDER.to_string(),
},
);
}
_ => normalized_content_items.push(content_item.clone()),
}
}
*content_items = normalized_content_items;
}
}
ResponseItem::ImageGenerationCall { result, .. } => {
result.clear();
}
_ => {}
}
}
}图片降级覆盖三处:普通 Message 内容、function/custom tool output 的 content items,以及 ImageGenerationCall.result。前两处把图片变成明确的文本占位符,后一处清空结果字符串;它不会删除整个 item,也不会修改 revised_prompt。有 InputModality::Image 时整个函数早退,保留原 payload。
6.2 音频替换
源码位置:codex-rs/core/src/context_manager/normalize.rs :: strip_audio_when_unsupported
pub(crate) fn strip_audio_when_unsupported(
input_modalities: &[InputModality],
items: &mut [ResponseItem],
) {
if input_modalities.contains(&InputModality::Audio) {
return;
}
for item in items.iter_mut() {
match item {
ResponseItem::Message { content, .. } => {
for content_item in content.iter_mut() {
if matches!(content_item, ContentItem::InputAudio { .. }) {
*content_item = ContentItem::InputText {
text: AUDIO_CONTENT_OMITTED_PLACEHOLDER.to_string(),
};
}
}
}
ResponseItem::FunctionCallOutput { output, .. }
| ResponseItem::CustomToolCallOutput { output, .. } => {
if let Some(content_items) = output.content_items_mut() {
for content_item in content_items.iter_mut() {
if matches!(
content_item,
FunctionCallOutputContentItem::InputAudio { .. }
) {
*content_item = FunctionCallOutputContentItem::InputText {
text: AUDIO_CONTENT_OMITTED_PLACEHOLDER.to_string(),
};
}
}
}
}
_ => {}
}
}
}图片和音频的替换范围相似,但不完全相同:音频没有 ImageGenerationCall 对应分支。文本、未支持的媒体和其他 content item 的相对顺序保持不变,所以下游仍能看到原消息中的位置关系,只是媒体 payload 变成占位文本。
7. 测试与边界
以下测试覆盖归一化算法的输入、输出和异常配对边界:
just test -p codex-core normalize_adds_missing_output_for_function_call_inserts_output
just test -p codex-core for_prompt_assigns_stable_id_to_synthetic_output_without_reordering_history
just test -p codex-core normalize_keeps_server_tool_search_output_without_matching_call
just test -p codex-core for_prompt_strips_media_when_model_does_not_support_it
just test -p codex-core for_prompt_preserves_image_generation_calls_when_images_are_supported
just test -p codex-core for_prompt_clears_image_generation_result_when_images_are_unsupported| 测试 | 输入 | 关键断言 | 覆盖范围 | 未覆盖 |
|---|---|---|---|---|
normalize_adds_missing_output_for_function_call_inserts_output | 无 output 的 function call | output 紧随 call 且文本为 aborted | 补全和顺序 | custom/local shell 的 release 分支 |
for_prompt_assigns_stable_id_to_synthetic_output_without_reordering_history | 有 ID 的 call + 后续 user | 两次投影相等,synthetic output 在中间且 ID 前缀正确 | UUID v5 稳定性与位置 | prompt cache 的网络端命中率 |
normalize_keeps_server_tool_search_output_without_matching_call | 无 call 的 server output | output 保留 | server 协议例外 | server output 的上游生成时序 |
for_prompt_strips_media_when_model_does_not_support_it | message、function/custom output 中的图片和音频 | payload 变占位文本,call/output 关系保留 | 图片音频降级范围 | 模型实际 token 变化 |
for_prompt_preserves_image_generation_calls_when_images_are_supported | image generation result + Image modality | result 不被清空 | 能力开启路径 | provider 是否真的接受该项 |
for_prompt_clears_image_generation_result_when_images_are_unsupported | image generation result + Text modality | item 保留但 result 为空 | 图片生成降级特例 | TUI 展示行为 |
debug/release 专用测试还覆盖 custom/local shell 缺失 output、function/custom/client tool-search 孤儿和混合输入。debug 期望 panic,release 期望修复或删除;这证明的是构建条件下的错误策略,不代表生产流中可以任意提交残缺历史。
8. 源码导航
遇到“模型请求与 raw history 不一样”的问题,可以按以下顺序定位:
| 现象 | 入口 | 重点检查 |
|---|---|---|
call 后出现 aborted | ensure_call_outputs_present | call 类型、call_id、synthetic ID |
| output 消失 | remove_orphan_outputs | output execution、call_id、对应 call 集合 |
| 同一请求重试的 item ID 变化 | synthetic_output_id | 固定 namespace、源 item ID 是否为空 |
| 图片变成英文占位文本 | strip_images_when_unsupported | InputModality::Image 是否存在 |
| image generation 结果为空 | 同上 | 这是预期的 result 清理,不是 item 删除 |
| debug 崩溃、release 只记录错误 | error_or_panic | cfg!(debug_assertions) 与具体测试 cfg |
只读搜索命令:
rg -n "normalize_history|ensure_call_outputs_present|remove_orphan_outputs" codex-rs/core/src
rg -n "synthetic_output_id|error_or_panic|strip_images_when_unsupported|strip_audio_when_unsupported" codex-rs/core/src
rg -n "normalize_|for_prompt_.*media|synthetic_output" codex-rs/core/src/context_manager/history_tests.rs现在可以反向复述这条主线:clone_history → for_prompt → 配对补全 → 孤儿清理 → 媒体能力投影 → Prompt。如果你要修改其中一步,先判断它是在 raw history 还是 prompt snapshot 上运行,再确认是否会改变 call/output 顺序、stable ID、debug/release 行为或媒体占位文本;这四类不变量分别由不同测试保护,不能用一项测试代替全部结论。
