ApplyPatch文件更新算法
Apply Patch 的 Update 不会边找边改文件。实现先读取原内容,把 chunk 转成 (start, old_len, new_lines) replacement,全部定位成功后才在内存中应用;随后生成 unified diff,再进入 filesystem 写入或 move。当前实现有两种重建模式:传统模式把更新文件统一为 LF,保留行尾模式则通过 SourceFile 保留未改行的 LF/CRLF/CR。多文件 patch 仍然不是事务:前面已经提交的变化会被记录在 AppliedPatchDelta,后续失败不会自动回滚。
本文承接ApplyPatchInvocation模型和ApplyPatch解析器,面向理解文本 diff、索引偏移和异步 filesystem 的读者。范围是 derive_new_contents_from_chunks、compute_replacements、seek_sequence 和 apply commit,不展开安全/原子性策略。读完后,你应能解释 context 搜索的宽松层级、replacement 为什么倒序执行,以及 move 失败时 delta 为何可能已经包含目标写入。
1. 两阶段更新
1.1 先计算
derive_new_contents_from_chunks 先按 ReadFileOptions.follow_symlinks 读取完整文本,再按更新模式分支。 NormalizeToLf 使用 split('\n') 和倒序 replacement,最后保证 trailing LF;PreserveLineEndings 用 SourceFile::parse 保存每一行原来的 terminator,再按 source order 应用 replacement。两种模式共享相同 的 context 定位和 replacement 计算。
源码位置:codex-rs/apply-patch/src/file_update.rs :: derive_new_contents_from_chunks
let original_contents = fs
.read_file_text(path, ReadFileOptions { follow_symlinks }, sandbox)
.await
.map_err(|err| {
ApplyPatchError::IoError(IoError {
context: format!("Failed to read file to update {}", path.inferred_native_path_string()),
source: err,
})
})?;
let new_contents = match update_file_mode {
ApplyPatchFileUpdateMode::NormalizeToLf => {
let mut original_lines = original_contents
.split('\n')
.map(String::from)
.collect::<Vec<_>>();
if original_lines.last().is_some_and(String::is_empty) {
original_lines.pop();
}
let replacements =
compute_replacements(&original_lines, &path_text, chunks, update_file_mode)?;
let mut new_lines = apply_replacements(original_lines, &replacements);
if !new_lines.last().is_some_and(String::is_empty) {
new_lines.push(String::new());
}
new_lines.join("\n")
}
ApplyPatchFileUpdateMode::PreserveLineEndings => {
let mut source_file = SourceFile::parse(&original_contents);
let original_lines = source_file.line_texts();
let replacements =
compute_replacements(&original_lines, &path_text, chunks, update_file_mode)?;
source_file.apply_replacements(&replacements);
source_file.into_contents()
}
};1.2 再提交
计算成功只产生内存中的 AppliedPatch;apply_hunks_to_files 才调用 write/remove。这样 context 不匹配不会写出半个文件,但跨 hunk/多文件仍可能部分成功。
2. Context定位
2.1 搜索顺序
seek_sequence 依次尝试 exact、忽略尾空白、两端 trim、Unicode 标点归一化。pattern 为空返回 start,pattern 比文件长直接 None,避免越界。PreserveLineEndings 的 EOF 搜索不会退回到 line_index 之前,维持多个 chunk 的单调顺序;传统模式保留旧的 EOF 起点行为。
源码位置:codex-rs/apply-patch/src/seek_sequence.rs :: seek_sequence
if pattern.is_empty() {
return Some(start);
}
if pattern.len() > lines.len() {
return None;
}
let search_start = if eof && lines.len() >= pattern.len() {
let eof_start = lines.len() - pattern.len();
match update_file_mode {
ApplyPatchFileUpdateMode::NormalizeToLf => eof_start,
ApplyPatchFileUpdateMode::PreserveLineEndings => eof_start.max(start),
}
} else {
start
};
for i in search_start..=lines.len().saturating_sub(pattern.len()) {
if lines[i..i + pattern.len()] == *pattern {
return Some(i);
}
}
// then rstrip, trim, and Unicode punctuation normalizationeof=true 时先从文件末尾可能匹配的位置搜索;找不到再按普通范围处理。因此 EOF marker 是定位偏好,不是“不看内容直接追加”。
3. Replacement计算
每个 chunk 先按 change_context 更新 line_index,再搜索 old_lines。纯 addition 没有 old_lines,当前实现把它安排到文件末尾。EOF 替换如果 pattern 末尾是空字符串,会去掉 trailing newline sentinel 后重试。
源码位置:codex-rs/apply-patch/src/file_update.rs :: compute_replacements
if let Some(ctx_line) = &chunk.change_context {
if let Some(idx) = seek_sequence::seek_sequence(
original_lines,
std::slice::from_ref(ctx_line),
line_index,
false,
update_file_mode,
) {
line_index = idx + 1;
} else {
return Err(ApplyPatchError::ComputeReplacements(format!(
"Failed to find context '{ctx_line}' in {path}"
)));
}
}
if chunk.old_lines.is_empty() {
let insertion_idx = match update_file_mode {
ApplyPatchFileUpdateMode::NormalizeToLf => {
if original_lines.last().is_some_and(String::is_empty) {
original_lines.len() - 1
} else {
original_lines.len()
}
}
ApplyPatchFileUpdateMode::PreserveLineEndings => original_lines.len(),
};
replacements.push((insertion_idx, 0, chunk.new_lines.clone()));
continue;
}找到 pattern 后记录 replacement,并把 line_index 移到旧区域之后;后续 chunk 因此不能回到之前位置。 在 PreserveLineEndings 模式中,算法按 context_line_indices 把一个 chunk 拆成多个 replacement,让显式 context 行留在原 SourceFile 中,只替换它们之间的真正改动区域,从而保留混合行尾。
源码位置:codex-rs/apply-patch/src/file_update.rs :: compute_replacements 的 PreserveLineEndings 分支
let mut old_start = 0;
let mut new_start = 0;
for &(old_context, new_context) in &chunk.context_line_indices {
if old_context >= pattern.len() || new_context >= new_slice.len() {
break;
}
if old_start != old_context || new_start != new_context {
replacements.push((
start_idx + old_start,
old_context - old_start,
new_slice[new_start..new_context].to_vec(),
));
}
old_start = old_context + 1;
new_start = new_context + 1;
}4. Replacement应用
传统 LF 模式先按 start index 排序,再从后向前删除旧行并插入新行。倒序保证后面的索引不受前面替换 导致的长度变化影响。保留行尾模式则由 SourceFile 按 source order 消费不重叠 replacement:未触碰的 SourceLine 保留原 ending,新插入行使用文件中第一个已存在 ending;没有 ending 的文件默认 LF,最后 仍保证每个结果行有 terminator。
源码位置:codex-rs/apply-patch/src/file_update.rs :: apply_replacements
for (start_idx, old_len, new_segment) in replacements.iter().rev() {
for _ in 0..*old_len {
if *start_idx < lines.len() {
lines.remove(*start_idx);
}
}
for (offset, new_line) in new_segment.iter().enumerate() {
lines.insert(*start_idx + offset, new_line.clone());
}
}源码位置:codex-rs/apply-patch/src/text_file.rs :: SourceFile::parse、SourceFile::apply_replacements
for (start_idx, old_len, new_segment) in replacements {
for line in source_lines.by_ref().take(*start_idx - source_index) {
new_lines.push(line);
}
for _ in source_lines.by_ref().take(*old_len) {}
new_lines.extend(new_segment.iter().map(|text| SourceLine {
text: text.clone(),
ending: Some(self.preferred_ending),
}));
source_index = start_idx + old_len;
}
new_lines.extend(source_lines);5. Move与部分提交
move update 先写 destination,再删除 source。若 destination 已存在,会记录 overwritten content;若随后删除 source 失败,destination 写入已经提交,failure delta 会保留这段确定变化。它不是 rename 原子操作。
源码位置:codex-rs/apply-patch/src/lib.rs :: apply_hunks_to_files 的 UpdateFile 分支
try_write!(write_file_with_missing_parent_retry(
fs,
&dest_uri,
new_contents.clone().into_bytes(),
sandbox,
).await);
let dest_write_change_index = delta.changes.len();
delta.changes.push(AppliedPatchChange {
path: dest_uri.clone(),
change: AppliedPatchFileChange::Add {
content: new_contents.clone(),
overwritten_content: overwritten_move_content.clone(),
},
});
// source removal happens afterwards5.1 Delta精确性
AppliedPatchDelta 按提交顺序记录变化,并带 exact 标记。如果目标内容无法读取或 remove failure 是否无副作用无法确认,exact 会变 false。失败不会把已记录 delta 清空。
源码位置:codex-rs/apply-patch/src/lib.rs :: AppliedPatchDelta、ApplyPatchFailure
pub struct AppliedPatchDelta {
changes: Vec<AppliedPatchChange>,
exact: bool,
}
pub struct ApplyPatchFailure {
error: ApplyPatchError,
delta: AppliedPatchDelta,
}6. Unified diff
unified_diff_from_chunks 复用相同的新内容计算,再用 similar::TextDiff 生成展示 diff。展示 diff 是派生结果;真正提交使用 new_contents,不是重新解析 unified diff。
源码位置:codex-rs/apply-patch/src/file_update.rs :: unified_diff_from_chunks_with_context_and_mode
let AppliedPatch { original_contents, new_contents } =
derive_new_contents_from_chunks(
path,
chunks,
update_file_mode,
fs,
/*follow_symlinks*/ true,
sandbox,
).await?;
let text_diff = TextDiff::from_lines(&original_contents, &new_contents);
let unified_diff = text_diff.unified_diff().context_radius(context).to_string();
Ok(ApplyPatchFileUpdate {
unified_diff,
original_content: original_contents,
content: new_contents,
})7. 验证
7.1 匹配与换行
seek_sequence 单测覆盖 exact、rstrip、trim、pattern 过长;file-update tests 覆盖 EOF 插入、首尾行替换和 interleaved chunks。CLI tool tests 进一步覆盖 CRLF、裸 CR、混合行尾、重复 context 行、EOF chunk 顺序 和 legacy LF 模式。
源码位置:
codex-rs/apply-patch/src/seek_sequence.rs:: testscodex-rs/apply-patch/src/file_update_tests.rs:: file-update testscodex-rs/apply-patch/tests/suite/tool.rs:: line-ending and ordering tests
cd codex-rs
cargo test -p codex-apply-patch --lib seek_sequence::tests -- --test-threads=1
cargo test -p codex-apply-patch --lib file_update::tests -- --test-threads=1
cargo test -p codex-apply-patch --test all suite::tool -- --test-threads=17.2 Move失败
test_failed_move_returns_committed_destination_delta 构造 destination 写入成功、source 删除失败,断言 failure delta 包含已提交目标变化。它证明非事务边界,不证明 filesystem 支持原子 rename。
源码位置:codex-rs/apply-patch/src/lib.rs :: test_failed_move_returns_committed_destination_delta
cd codex-rs
cargo test -p codex-apply-patch --lib test_failed_move_returns_committed_destination_delta -- --test-threads=1这些测试不证明跨文件 patch 具有事务性;相反,partial-success 与 failed-move 场景明确断言已提交变化仍然 存在。外部 filesystem 的崩溃一致性、远程断连和平台 rename 原子性不能由这些测试外推。
8. 源码排查
rg -n "derive_new_contents_from_chunks|compute_replacements|apply_replacements" codex-rs/apply-patch/src/file_update.rs
rg -n "SourceFile|preferred_ending|apply_replacements" codex-rs/apply-patch/src/text_file.rs
rg -n "seek_sequence|normalise|eof" codex-rs/apply-patch/src/seek_sequence.rs
rg -n "failed_move|insert_at_eof|interleaved|preserves_crlf|mixed_line" codex-rs/apply-patch/src codex-rs/apply-patch/tests更新算法的主线是:读取原文,按逐步宽松规则定位 context/old lines,生成 replacement,再按 LF 或保留 行尾模式重建内容,随后提交 write/remove;跨文件与 move 不构成事务,失败前提交内容由 delta 记录。 下一篇ApplyPatch流式识别将分析模型参数增量如何转成进度事件。
