Shell命令解析与安全分类
执行工具拿到的是 Vec<String>,但其中一个元素可能又包含一段 shell 程序。Codex 因此维护两种不同的解析目标:一类把“完全可证明是字面参数的简单命令”拆成 argv 段,交给 ExecPolicy 做逐段匹配;另一类只从复杂脚本中找出静态可见的命令字面量,用来发现危险操作。前者失败时必须保守地保留原始 argv,后者不能反过来证明脚本安全。
这两个目标容易被混为“shell parser”,但它们的输入、允许语法和安全结论并不相同。Bash plain lowering 只接受 word-only command sequence;heredoc、重定向、变量赋值和命令替换会让它失败。危险检测则可以递归检查 sudo、env、trap 或复杂 shell 中的字面 rm -rf,同时忽略动态命令名。Windows 还会在已 lowering 的 PowerShell words 上使用 PowerShell 专用 heuristic。
本文承接ExecPolicy语言与规则和ApprovalPolicy完整参考。阅读重点是“哪一层可以得出哪一种结论”:解析器只负责建立可信的 argv 视图,危险 heuristic 只负责提高未命中规则时的决策,最终的 approval requirement 仍由 Core 结合权限 profile、审批策略和 sandbox 状态计算。
1. 两种解析目标
先看模块边界。codex-shell-command 导出 Bash、PowerShell 和通用命令摘要工具;危险命令模块在 crate 内部复用 Bash 的 literal scanner。生产代码不把 PowerShell 子进程 parser 编译进分类路径,测试才把它作为 AST 对照器使用。
源码位置:codex-rs/shell-command/src/lib.rs :: module exports
//! Command parsing and safety utilities shared across Codex crates.
pub mod shell_detect;
pub mod shell_snapshot;
pub mod bash;
pub(crate) mod command_safety;
pub mod parse_command;
pub mod powershell;
pub use command_safety::is_dangerous_command;这里有一个重要的所有权关系:parse_command 为用户界面生成可读摘要,bash 和 powershell 为策略层提供 lowering,command_safety 只提供危险判定。摘要中的 Unknown 并不等价于安全失败;它只表示摘要器无法把任意命令压缩成某个展示类别。
源码位置:codex-rs/shell-command/src/command_safety/mod.rs :: parser modules
// Keep the PowerShell subprocess parser available as a test oracle, but do not
// compile it into production command classification.
#[cfg(test)]
#[allow(dead_code)]
mod powershell_parser;
mod powershell_tree_sitter;
pub mod is_dangerous_command;
pub(crate) use powershell_tree_sitter::try_parse_powershell_commands;因此,文章后面的“PowerShell parser”指 production tree-sitter lowering;powershell_parser.rs 只用于 fixture 对照,不能被写成运行时安全 oracle。
2. Bash的字面分段
2.1 wrapper入口
extract_bash_command 对 argv 形状有严格要求:只能是三个元素,第二个是 -lc 或 -c,第一个 executable 由 shell detection 识别为 Bash、Zsh 或 sh。额外的环境参数、多个脚本参数或未知 wrapper 不会被误判成 Bash 脚本。
源码位置:codex-rs/shell-command/src/bash.rs :: extract_bash_command, parse_shell_lc_plain_commands
pub fn extract_bash_command(command: &[String]) -> Option<(&str, &str)> {
let [shell, flag, script] = command else {
return None;
};
if !matches!(flag.as_str(), "-lc" | "-c")
|| !matches!(
detect_shell_type(PathBuf::from(shell)),
Some(ShellType::Zsh) | Some(ShellType::Bash) | Some(ShellType::Sh)
)
{
return None;
}
Some((shell, script))
}
pub fn parse_shell_lc_plain_commands(command: &[String]) -> Option<Vec<Vec<String>>> {
let (_, script) = extract_bash_command(command)?;
parse_shell_script_into_commands(script)
}这里的 Option 是安全边界而不是便利返回值:None 表示“不能证明脚本能被无损降成 argv”,调用方必须继续走原始命令路径。
2.2 语法树白名单
plain parser 先用 tree-sitter-bash 建树,再遍历所有节点。只有顶层容器、command、word、字符串和 concatenation 等节点进入白名单;变量展开、控制流、重定向、子 shell、命令替换等节点一旦出现,就返回 None。连接符本身也只允许 &&、||、; 和 |。
源码位置:codex-rs/shell-command/src/bash.rs :: try_parse_word_only_commands_sequence
pub fn try_parse_word_only_commands_sequence(tree: &Tree, src: &str) -> Option<Vec<Vec<String>>> {
if tree.root_node().has_error() {
return None;
}
const ALLOWED_KINDS: &[&str] = &[
"program",
"list",
"pipeline",
"command",
"command_name",
"word",
"string",
"string_content",
"raw_string",
"number",
"concatenation",
];
const ALLOWED_PUNCT_TOKENS: &[&str] = &["&&", "||", ";", "|", "\"", "'"];
let root = tree.root_node();
let mut cursor = root.walk();
let mut stack = vec![root];
let mut command_nodes = Vec::new();
while let Some(node) = stack.pop() {
let kind = node.kind();
if node.is_named() {
if !ALLOWED_KINDS.contains(&kind) {
return None;
}
if matches!(kind, "word" | "number") && !is_literal_word_or_number(node, src) {
return None;
}
if kind == "command" {
command_nodes.push(node);
}
} else {
if kind.chars().any(|c| "&;|".contains(c))
&& !ALLOWED_PUNCT_TOKENS.contains(&kind)
{
return None;
}
if !(ALLOWED_PUNCT_TOKENS.contains(&kind) || kind.trim().is_empty()) {
return None;
}
}
for child in node.children(&mut cursor) {
stack.push(child);
}
}
command_nodes.sort_by_key(Node::start_byte);遍历使用显式 stack 而非递归调用,最后按 start_byte 排序恢复源码顺序。这样 ls && pwd; echo hi | wc -l 会得到四个策略段,而不是把整段 shell 文本当成一个 prefix。
2.3 运行时展开
节点类型白名单还不够,因为 tree-sitter 的 word 可能包含 shell 展开。is_literal_word_or_number 拒绝 glob、brace expansion、反斜杠、tilde、变量、命令替换和 Zsh 的特殊语法;双引号则额外拒绝会改变运行时值的 escape。
源码位置:codex-rs/shell-command/src/bash.rs :: is_literal_word_or_number, parse_double_quoted_string
fn is_literal_word_or_number(node: Node<'_>, src: &str) -> bool {
if !matches!(node.kind(), "word" | "number") {
return false;
}
let mut cursor = node.walk();
node.named_children(&mut cursor).next().is_none()
&& node.utf8_text(src.as_bytes()).is_ok_and(|word| {
!word.starts_with('=')
&& !word.contains(['{', '}', '*', '?', '[', ']', '\\', '~', '^', '#', '$', '`'])
})
}
fn parse_double_quoted_string(node: Node, src: &str) -> Option<String> {
if node.kind() != "string" {
return None;
}
let mut cursor = node.walk();
for part in node.named_children(&mut cursor) {
if part.kind() != "string_content" {
return None;
}
}
let raw = node.utf8_text(src.as_bytes()).ok()?;
let stripped = raw
.strip_prefix('"')
.and_then(|text| text.strip_suffix('"'))?;
if stripped
.as_bytes()
.windows(2)
.any(|pair| pair[0] == b'\\' && matches!(pair[1], b'$' | b'`' | b'"' | b'\\' | b'\n'))
{
return None;
}
Some(stripped.to_owned())
}parse_plain_command_from_node 再把 command 的 command name、word、number、quoted string 和 concatenation 逐项转为 Vec<String>。concatenation 允许 -g"*.py" 这类字面拼接,但不会接受其中的变量或命令替换。
源码位置:codex-rs/shell-command/src/bash.rs :: parse_plain_command_from_node
fn parse_plain_command_from_node(cmd: tree_sitter::Node, src: &str) -> Option<Vec<String>> {
if cmd.kind() != "command" {
return None;
}
let mut words = Vec::new();
let mut cursor = cmd.walk();
for child in cmd.named_children(&mut cursor) {
match child.kind() {
"command_name" => {
let word_node = child.named_child(0)?;
if word_node.kind() != "word" {
return None;
}
words.push(word_node.utf8_text(src.as_bytes()).ok()?.to_owned());
}
"word" | "number" => {
words.push(child.utf8_text(src.as_bytes()).ok()?.to_owned());
}
"string" => words.push(parse_double_quoted_string(child, src)?),
"raw_string" => words.push(parse_raw_string(child, src)?),
"concatenation" => {
let mut concatenated = String::new();
let mut concat_cursor = child.walk();
for part in child.named_children(&mut concat_cursor) {
match part.kind() {
"word" | "number" => {
concatenated.push_str(part.utf8_text(src.as_bytes()).ok()?);
}
"string" => concatenated.push_str(&parse_double_quoted_string(part, src)?),
"raw_string" => concatenated.push_str(&parse_raw_string(part, src)?),
_ => return None,
}
}
if concatenated.is_empty() {
return None;
}
words.push(concatenated);
}
_ => return None,
}
}
Some(words)
}3. Bash复杂语法
3.1 两种解析用途
Bash 模块还提供 parse_shell_lc_literal_commands。它不要求整棵树属于 plain allowlist,而是遍历有效语法树里的每个 command 节点,只保留能静态读出的 command name 和后续 literal words。动态词和重定向会被省略。
源码位置:codex-rs/shell-command/src/bash.rs :: parse_shell_lc_literal_commands, parse_literal_command_from_node
pub(crate) fn parse_shell_lc_literal_commands(command: &[String]) -> Option<Vec<Vec<String>>> {
let (_, script) = extract_bash_command(command)?;
let tree = try_parse_shell(script)?;
let root = tree.root_node();
if root.has_error() {
return None;
}
let mut commands = Vec::new();
let mut stack = vec![root];
while let Some(node) = stack.pop() {
if node.kind() == "command"
&& let Some(command) = parse_literal_command_from_node(node, script)
{
commands.push(command);
}
let mut cursor = node.walk();
for child in node.named_children(&mut cursor) {
stack.push(child);
}
}
Some(commands)
}
fn parse_literal_command_from_node(cmd: Node<'_>, src: &str) -> Option<Vec<String>> {
if cmd.kind() != "command" {
return None;
}
let mut words = Vec::new();
let mut found_command_name = false;
let mut cursor = cmd.walk();
for child in cmd.named_children(&mut cursor) {
if child.kind() == "command_name" {
let command_name = parse_literal_shell_word(child.named_child(0)?, src)?;
words.push(command_name);
found_command_name = true;
} else if found_command_name && let Some(word) = parse_literal_shell_word(child, src) {
words.push(word);
}
}
found_command_name.then_some(words)
}这条路径的设计意图是“发现可能危险的字面命令”,不是“批准这些字面命令”。例如 heredoc 中的 python3 可以被看到,但 shell 仍可能把输入重定向给它;变量赋值、动态命令名或命令替换也可能改变最终执行效果。因而 plain lowering 失败时,ExecPolicy 不能仅凭 inner command 的 Allow 就设置 bypass_sandbox。
3.2 原始argv回退
commands_for_exec_policy 先尝试 Bash plain lowering;Windows 构建再尝试 PowerShell lowering;两者没有得到非空结果时返回一个只包含原始 argv 的 ExecPolicyCommands。当前实现没有旧版本的“single prefix + complex flag”中间态。
源码位置:codex-rs/core/src/exec_policy.rs :: commands_for_exec_policy
fn commands_for_exec_policy(command: &[String]) -> ExecPolicyCommands {
if let Some(commands) = parse_shell_lc_plain_commands(command)
&& !commands.is_empty()
{
return ExecPolicyCommands {
commands,
command_origin: ExecPolicyCommandOrigin::Generic,
};
}
#[cfg(windows)]
{
if let Some(commands) =
codex_shell_command::powershell::parse_powershell_command_into_plain_commands(command)
&& !commands.is_empty()
{
return ExecPolicyCommands {
commands,
command_origin: ExecPolicyCommandOrigin::PowerShell,
};
}
}
ExecPolicyCommands {
commands: vec![command.to_vec()],
command_origin: ExecPolicyCommandOrigin::Generic,
}
}空字符串和只有空白的 bash -lc 因为 parser 返回空结果,都会走这个 fallback。fallback 并不把脚本变成“无操作”;它确保规则匹配和 heuristic 仍然拥有一个输入。
源码位置:codex-rs/core/src/exec_policy_tests.rs :: commands_for_exec_policy_falls_back_for_empty_shell_script, commands_for_exec_policy_falls_back_for_whitespace_shell_script
#[test]
fn commands_for_exec_policy_falls_back_for_empty_shell_script() {
let command = vec!["bash".to_string(), "-lc".to_string(), "".to_string()];
assert_eq!(
commands_for_exec_policy(&command),
ExecPolicyCommands {
commands: vec![command],
command_origin: ExecPolicyCommandOrigin::Generic,
}
);
}
#[test]
fn commands_for_exec_policy_falls_back_for_whitespace_shell_script() {
let command = vec![
"bash".to_string(),
"-lc".to_string(),
" \n\t ".to_string(),
];
assert_eq!(
commands_for_exec_policy(&command),
ExecPolicyCommands {
commands: vec![command],
command_origin: ExecPolicyCommandOrigin::Generic,
}
);
}3.3 heredoc边界
下面的测试给出了一个容易误读的案例:规则允许 python3,但实际输入是带 heredoc 的 shell wrapper。plain parser 失败后,整条 bash -lc ... 作为一个命令参与策略,因此结果可以是 Skip,但 bypass_sandbox 必须为 false,并且 amendment 候选保存完整 wrapper,而不是只保存 python3。
源码位置:codex-rs/core/src/exec_policy_tests.rs :: heredoc_script_stays_in_sandbox_despite_inner_allow_rule
let command = vec![
"bash".to_string(),
"-lc".to_string(),
"python3 <<'PY'\nprint('hello')\nPY".to_string(),
];
assert_exec_approval_requirement_for_command(
ExecApprovalRequirementScenario {
policy_src: Some(r#"prefix_rule(pattern=["python3"], decision="allow")"#.to_string()),
command,
approval_policy: AskForApproval::OnRequest,
permission_profile: PermissionProfile::read_only(),
sandbox_permissions: SandboxPermissions::UseDefault,
prefix_rule: None,
},
ExecApprovalRequirement::Skip {
bypass_sandbox: false,
proposed_execpolicy_amendment: Some(ExecPolicyAmendment::new(vec![
"bash".to_string(),
"-lc".to_string(),
"python3 <<'PY'\nprint('hello')\nPY".to_string(),
])),
},
)
.await;这说明“规则允许了可见的解释器”与“整个 shell invocation 可以脱离 sandbox”是两个不同命题。变量赋值、重定向和控制流都应按同一原则阅读。
4. 危险命令递归
4.1 wrapper扫描
危险分类器返回 DangerousCommandMatch,目前区分 ForcedRm 与 Other。它先检查已经 tokenized 的 argv,再尝试对 Bash wrapper 做 literal command scan,最后在 Windows 构建中检查 CMD/PowerShell 特有规则。
源码位置:codex-rs/shell-command/src/command_safety/is_dangerous_command.rs :: DangerousCommandMatch, dangerous_command_match_with_depth
#[derive(Clone, Copy, Debug, Eq, PartialEq)]
pub enum DangerousCommandMatch {
ForcedRm,
Other,
}
const MAX_DANGEROUS_COMMAND_WRAPPER_DEPTH: usize = 8;
pub fn dangerous_command_match(command: &[String]) -> Option<DangerousCommandMatch> {
dangerous_command_match_with_depth(command, /*wrapper_depth*/ 0)
}
fn dangerous_command_match_with_depth(
command: &[String],
wrapper_depth: usize,
) -> Option<DangerousCommandMatch> {
if wrapper_depth > MAX_DANGEROUS_COMMAND_WRAPPER_DEPTH {
return Some(DangerousCommandMatch::Other);
}
if let Some(dangerous_match) = dangerous_command_match_for_exec(command, wrapper_depth) {
return Some(dangerous_match);
}
if let Some(dangerous_match) = parse_shell_lc_literal_commands(command).and_then(|commands| {
commands
.iter()
.find_map(|command| dangerous_command_match_with_depth(command, wrapper_depth + 1))
}) {
return Some(dangerous_match);
}
#[cfg(windows)]
{
if windows_dangerous_commands::is_dangerous_command_windows(command) {
return Some(DangerousCommandMatch::Other);
}
}
None
}递归深度上限是一个不变量:env、sudo、trap 或嵌套 sh -c 不能无限消耗调用栈,也不能因为嵌套太深而被当成安全命令。超过上限直接返回 Other,让上层按危险路径处理。
4.2 rm force判定
dangerous_command_match_for_exec 通过 executable basename 识别 rm,然后只检查 --force 或包含 f 的短选项。-- 后的参数不再被当成选项,因此 rm -- -f 不会误报。sudo 会剥掉 wrapper 继续检查,env 会跳过环境赋值,trap 则把 action 重新包装成 sh -c。
源码位置:codex-rs/shell-command/src/command_safety/is_dangerous_command.rs :: dangerous_command_match_for_exec, dangerous_command_match_for_env, rm_args_include_force_option
fn dangerous_command_match_for_exec(
command: &[String],
wrapper_depth: usize,
) -> Option<DangerousCommandMatch> {
let cmd0 = command
.first()
.and_then(|command| executable_name_lookup_key(command));
match cmd0.as_deref() {
Some("rm") if rm_args_include_force_option(&command[1..]) => {
Some(DangerousCommandMatch::ForcedRm)
}
Some("sudo") => dangerous_command_match_with_depth(&command[1..], wrapper_depth + 1),
Some("env") => dangerous_command_match_for_env(command, wrapper_depth),
Some("trap") => dangerous_command_match_for_trap(command, wrapper_depth),
_ => None,
}
}
fn dangerous_command_match_for_env(
command: &[String],
wrapper_depth: usize,
) -> Option<DangerousCommandMatch> {
let mut command_index = 1;
while let Some(argument) = command.get(command_index) {
if argument == "--" {
command_index += 1;
break;
}
if matches!(argument.as_str(), "-i" | "--ignore-environment")
|| argument
.split_once('=')
.is_some_and(|(name, _)| !name.is_empty() && !name.starts_with('-'))
{
command_index += 1;
continue;
}
break;
}
dangerous_command_match_with_depth(&command[command_index..], wrapper_depth + 1)
}
fn rm_args_include_force_option(args: &[String]) -> bool {
args.iter()
.take_while(|arg| arg.as_str() != "--")
.any(|arg| {
arg == "--force"
|| arg
.strip_prefix('-')
.is_some_and(|flags| !flags.starts_with('-') && flags.contains('f'))
})
}测试覆盖了绝对路径 /bin/rm、选项分开写、sudo rm、带环境赋值的 env rm,并验证非 force 的 rm -r、rm -- -f 和动态命令名不会得到 ForcedRm。
4.3 复杂脚本样例
literal scan 的价值在于它不会因为控制流、管道或命令替换而完全放弃危险检查。以下测试中的 if、循环、管道、重定向、嵌套 bash -c 和 trap 都能找到字面 rm -rf;但变量命令名或只出现在字符串里的文本不会被当成执行命令。
源码位置:codex-rs/shell-command/src/command_safety/is_dangerous_command.rs :: forced_rm_in_complex_shell_syntax_is_dangerous, non_forced_or_non_literal_rm_is_not_dangerous
#[test]
fn forced_rm_in_complex_shell_syntax_is_dangerous() {
for script in [
"printf x | rm -rf /tmp/example",
"if test -d /tmp/example; then rm --force /tmp/example; fi",
"rm -rf \"$TARGET\" >/dev/null",
"for target in /tmp/a /tmp/b; do rm -r -f \"$target\"; done",
"echo \"$(rm -rf /tmp/example)\"",
"bash -c 'rm -rf /tmp/example'",
"trap 'rm -rf /tmp/example' EXIT",
] {
let command = vec_str(&["bash", "-lc", script]);
assert_eq!(
dangerous_command_match(&command),
Some(DangerousCommandMatch::ForcedRm),
"{script}"
);
}
}
#[test]
fn non_forced_or_non_literal_rm_is_not_dangerous() {
for command in [
vec_str(&["rm", "-r", "/tmp/example"]),
vec_str(&["rm", "--", "-f"]),
vec_str(&["bash", "-lc", "echo 'rm -rf /tmp/example'"]),
vec_str(&["bash", "-lc", "cmd=rm; $cmd -rf /tmp/example"]),
] {
assert_eq!(dangerous_command_match(&command), None, "{command:?}");
}
}这里的“字面”是故意保守的技术术语:rm -rf "$TARGET" 的 command name 和 force option 是字面可见的,所以需要提高风险;但目标路径的实际值仍未知。相反,$cmd -rf 的 command name 动态生成,分类器没有足够证据把它归为 ForcedRm,但这不构成安全证明。
5. PowerShell边界
5.1 wrapper提取
PowerShell wrapper 允许 -NoLogo、-NoProfile 等已知 flag,然后寻找 -Command 或 -c 后面的脚本。未知 flag、缺少脚本或脚本后仍有额外 argv 都返回 None。
源码位置:codex-rs/shell-command/src/powershell.rs :: extract_powershell_command, parse_powershell_command_into_plain_commands
const POWERSHELL_FLAGS: &[&str] = &["-nologo", "-noprofile", "-command", "-c"];
pub fn extract_powershell_command(command: &[String]) -> Option<(&str, &str)> {
if command.len() < 3 {
return None;
}
let shell = &command[0];
if !matches!(
detect_shell_type(PathBuf::from(shell)),
Some(ShellType::PowerShell)
) {
return None;
}
let mut i = 1usize;
while i + 1 < command.len() {
let flag = &command[i];
if !POWERSHELL_FLAGS.contains(&flag.to_ascii_lowercase().as_str()) {
return None;
}
if flag.eq_ignore_ascii_case("-Command") || flag.eq_ignore_ascii_case("-c") {
let script = &command[i + 1];
return Some((shell, script));
}
i += 1;
}
None
}
pub fn parse_powershell_command_into_plain_commands(
command: &[String],
) -> Option<Vec<Vec<String>>> {
let (_, script) = extract_powershell_command(command)?;
try_parse_powershell_commands(script)
}5.2 AST白名单
production PowerShell lowering 使用 tree-sitter 检查根节点错误、#requires 指令和未知 named node,然后收集 command 节点并逐条做 literal argv lowering。它明确拒绝动态表达式和未覆盖的 CST 形状;因此 Some 只表示当前白名单能够解释该脚本,不表示支持完整 PowerShell 语言。
源码位置:codex-rs/shell-command/src/command_safety/powershell_tree_sitter.rs :: try_parse_powershell_commands, lower_with_tree_sitter
pub(crate) fn try_parse_powershell_commands(script: &str) -> Option<Vec<Vec<String>>> {
lower_with_tree_sitter(script).ok()
}
fn lower_with_tree_sitter(script: &str) -> Result<Vec<Vec<String>>, String> {
if script
.chars()
.any(|ch| matches!(ch, '‘' | '’' | '“' | '”' | '–' | '—' | '―'))
{
return Err("PowerShell Unicode syntax alias".to_string());
}
let mut parser = Parser::new();
parser
.set_language(&tree_sitter_powershell::LANGUAGE.into())
.map_err(|error| format!("load grammar: {error}"))?;
let tree = parser
.parse(script, None)
.ok_or_else(|| "tree-sitter returned no tree".to_string())?;
let root = tree.root_node();
if root.has_error() {
return Err("tree contains ERROR or missing nodes".to_string());
}
if has_requires_directive(root, script) {
return Err("requires directives can execute before command lowering".to_string());
}
if let Some(kind) = first_unrecognized_named_kind(root) {
return Err(format!("unrecognized named node: {kind}"));
}
let mut command_nodes = Vec::new();
collect_command_nodes(root, &mut command_nodes);
if command_nodes.is_empty() {
return Err("no literal command nodes".to_string());
}源码还会用 source_is_covered_by_commands 检查 command ranges 之间的内容,只允许空白、换行、注释、分号、管道和 &&。这阻止 parser 只提取几个 command,却悄悄丢掉其它会执行的脚本部分。
源码位置:codex-rs/shell-command/src/command_safety/powershell_tree_sitter.rs :: source_is_covered_by_commands
fn source_is_covered_by_commands(script: &str, command_ranges: &[std::ops::Range<usize>]) -> bool {
let mut index = 0;
let mut range_index = 0;
let mut can_chain = false;
let mut needs_command = false;
let mut paren_depth = 0;
while index < script.len() {
if let Some(range) = command_ranges.get(range_index)
&& index == range.start
{
index = range.end;
range_index += 1;
can_chain = true;
needs_command = false;
continue;
}
let Some(ch) = script[index..].chars().next() else {
return false;
};
let next = index + ch.len_utf8();
if ch == '\r' || ch == '\n' || ch.is_whitespace() {
index = next;
continue;
}
if ch == ';' {
if needs_command {
return false;
}
can_chain = false;
index = next;
continue;
}
if ch == '|' && can_chain {
can_chain = false;
needs_command = true;
index = if script[next..].starts_with('|') {
next + '|'.len_utf8()
} else {
next
};
continue;
}
if ch == '&' && can_chain && script[next..].starts_with('&') {
can_chain = false;
needs_command = true;
index = next + '&'.len_utf8();
continue;
}
return false;
}
range_index == command_ranges.len() && !needs_command && paren_depth == 0
}5.3 参数拒绝规则
PowerShell 的 lower_command_text 只解码静态单引号、双引号和 bare word。双引号中的变量、版本相关 escape、bare word 中的结构字符以及需要 PowerShell 特殊值转换的参数都会使 lowering 失败。
源码位置:codex-rs/shell-command/src/command_safety/powershell_tree_sitter.rs :: lower_command_text
fn lower_command_text(command_text: &str) -> Result<Vec<String>, String> {
let mut words = Vec::new();
let chars: Vec<char> = command_text.trim().chars().collect();
let mut index = 0;
while index < chars.len() {
while index < chars.len() && chars[index].is_whitespace() {
index += 1;
}
if index == chars.len() || chars[index] == '#' {
break;
}
let (word, next, is_bare) = if chars[index] == '\'' {
let (word, next) = parse_single_quoted(&chars, index)?;
(word, next, false)
} else if chars[index] == '"' {
let (word, next) = parse_double_quoted(&chars, index)?;
(word, next, false)
} else {
let (word, next) = parse_bare_word(&chars, index)?;
(word, next, true)
};
index = next;
if index < chars.len() && !chars[index].is_whitespace() && chars[index] != '#' {
return Err("adjacent/concatenated command elements".to_string());
}
if word.is_empty() {
return Err("empty word".to_string());
}
if is_bare {
reject_unsupported_bare_word(&word)?;
}
words.push(word);
}
if words.is_empty() {
return Err("command lowered to no words".to_string());
}
Ok(words)
}测试 fixture 以 Some(Vec<Vec<String>>) 表示支持的字面形式,以 None 表示需要更多 PowerShell 语义的形式;#requires -Modules Evil 明确被拒绝,因为它可能在脚本主体之前加载模块或程序集。
源码位置:codex-rs/shell-command/src/command_safety/powershell_tree_sitter_tests.rs :: lowers_compact_literal_fixture, rejects_requires_directives
#[test]
fn lowers_compact_literal_fixture() {
let cases: Vec<FixtureCase> =
serde_json::from_str(include_str!("fixtures/powershell_lowering.json"))
.expect("valid PowerShell lowering fixture");
for case in cases {
assert_eq!(
try_parse_powershell_commands(&case.script),
case.expected,
"fixture case: {}",
case.name
);
}
}
#[test]
fn rejects_requires_directives() {
assert_eq!(
try_parse_powershell_commands("#requires -Modules Evil\nGet-Location"),
None
);
}6. 平台差异
6.1 origin与词法
Core 将 lowering 后的 segments 与 origin 一起保存。Generic 使用通用 dangerous_command_match;只有顶层 PowerShell wrapper 成功 lowering 时才使用 dangerous_powershell_words_match。这避免把 PowerShell cmdlet 当成 Unix executable,也避免把 Windows 词法套到原始 argv 上。
源码位置:codex-rs/core/src/exec_policy.rs :: ExecPolicyCommandOrigin, dangerous_command_match_for_origin
#[derive(Clone, Copy, Debug, Eq, PartialEq)]
pub(crate) enum ExecPolicyCommandOrigin {
Generic,
#[cfg(windows)]
PowerShell,
}
fn dangerous_command_match_for_origin(
command: &[String],
command_origin: ExecPolicyCommandOrigin,
) -> Option<DangerousCommandMatch> {
match command_origin {
ExecPolicyCommandOrigin::Generic => dangerous_command_match(command),
#[cfg(windows)]
ExecPolicyCommandOrigin::PowerShell => {
codex_shell_command::is_dangerous_command::dangerous_powershell_words_match(command)
}
}
}6.2 Windows危险项
Windows 实现先识别 PowerShell,再识别 CMD,最后检查直接 GUI launch。PowerShell words 中的 Remove-Item -Force、带 URL 的 Start-Process、Invoke-Item、mshta、浏览器或 rundll32 url.dll,fileProtocolHandler 会被标记;CMD 则拆出 /c 后的连接符,检查 start URL、del /f 和 rd /s /q。
源码位置:codex-rs/shell-command/src/command_safety/windows_dangerous_commands.rs :: is_dangerous_command_windows, is_dangerous_powershell_words
pub fn is_dangerous_command_windows(command: &[String]) -> bool {
if is_dangerous_powershell(command) {
return true;
}
if is_dangerous_cmd(command) {
return true;
}
is_direct_gui_launch(command)
}
pub(crate) fn is_dangerous_powershell_words(words: &[String]) -> bool {
let tokens_lc: Vec<String> = words
.iter()
.map(|t| t.trim_matches('\'').trim_matches('"').to_ascii_lowercase())
.collect();
let has_url = args_have_url(words);
if has_url
&& tokens_lc.iter().any(|t| {
matches!(
t.as_str(),
"start-process" | "start" | "saps" | "invoke-item" | "ii"
) || t.contains("start-process")
|| t.contains("invoke-item")
})
{
return true;
}
has_force_delete_cmdlet(&tokens_lc)
}Windows 规则是 heuristic,不是对全部 PowerShell/CMD 语义的解释器。未知 flag、动态脚本和无法解析的 URL 会保持不命中,随后仍由审批策略和 sandbox owner 决定是否允许执行。
7. ExecPolicy消费
7.1 segment匹配
create_exec_approval_requirement_for_command 调用 commands_for_exec_policy 后,给每个 segment 使用同一个 exec_policy_fallback。Core 的 fallback 读取 command_origin,再把未命中命令交给 render_decision_for_unmatched_command。显式 policy match 和 heuristic match 最后一起进入 check_multiple_with_options。
源码位置:codex-rs/core/src/exec_policy.rs :: create_exec_approval_requirement_for_command, create_exec_approval_requirement_for_parsed_commands
pub(crate) async fn create_exec_approval_requirement_for_command(
&self,
req: ExecApprovalRequest<'_>,
) -> ExecApprovalRequirement {
let commands = commands_for_exec_policy(req.command);
self.create_exec_approval_requirement_for_parsed_commands(req, commands)
.await
}
async fn create_exec_approval_requirement_for_parsed_commands(
&self,
req: ExecApprovalRequest<'_>,
ExecPolicyCommands {
commands,
command_origin,
}: ExecPolicyCommands,
) -> ExecApprovalRequirement {
let ExecApprovalRequest {
command,
approval_policy,
permission_profile,
environment_policy,
windows_sandbox_level,
sandbox_permissions,
prefix_rule,
allow_prefix_rules,
} = req;
let exec_policy = self.current_for_environment(environment_policy, allow_prefix_rules);
let auto_amendment_allowed = allow_prefix_rules == AllowPrefixRules::Honor;
let exec_policy_fallback = |cmd: &[String]| {
render_decision_for_unmatched_command(
cmd,
UnmatchedCommandContext {
approval_policy,
permission_profile: &permission_profile,
windows_sandbox_level,
sandbox_permissions,
command_origin,
},
)
};
let match_options = MatchOptions {
resolve_host_executables: true,
};
let evaluation = exec_policy.check_multiple_with_options(
commands.iter(),
&exec_policy_fallback,
&match_options,
);7.2 Allow与sandbox
当聚合结果为 Allow 时,Core 仍会重新对每个 segment 查询显式 policy match。只有每一个 segment 都有 PrefixRuleMatch 且 decision 为 Allow,才设置 bypass_sandbox: true。heuristic 得到的 Allow 只代表当前审批决策,不足以成为脱离 sandbox 的证明。
源码位置:codex-rs/core/src/exec_policy.rs :: create_exec_approval_requirement_for_parsed_commands
Decision::Allow => ExecApprovalRequirement::Skip {
bypass_sandbox: commands.iter().all(|command| {
exec_policy
.matches_for_command_with_options(
command,
/*heuristics_fallback*/ None,
&match_options,
)
.iter()
.any(|rule_match| {
is_policy_match(rule_match) && rule_match.decision() == Decision::Allow
})
}),
proposed_execpolicy_amendment: if auto_amendment_allowed {
try_derive_execpolicy_amendment_for_allow_rules(&evaluation.matched_rules)
} else {
None
},
},因此,“每段都被检查”与“只有显式 Allow 才能 bypass”同时成立。若复合脚本的一段只得到 heuristic Allow,整个调用仍在 sandbox 内执行。
7.3 未命中决策
危险命中或 Windows managed filesystem restriction 没有可用 sandbox backend 时,未命中命令不能静默 Allow:Never 变成 Forbidden,其余审批模式先变成 Prompt。非危险命令则按 AskForApproval、filesystem sandbox kind 和是否请求 sandbox override 决定 Allow/Prompt。
源码位置:codex-rs/core/src/exec_policy.rs :: render_decision_for_unmatched_command
let dangerous_command_match =
dangerous_command_match_for_origin(command, context.command_origin);
let file_system_sandbox_policy = permission_profile.file_system_sandbox_policy();
let windows_managed_fs_restrictions_without_sandbox_backend = cfg!(windows)
&& windows_sandbox_level == WindowsSandboxLevel::Disabled
&& profile_has_managed_filesystem_restrictions(permission_profile);
if dangerous_command_match.is_some() || windows_managed_fs_restrictions_without_sandbox_backend {
return match approval_policy {
AskForApproval::Never => Decision::Forbidden,
AskForApproval::OnRequest
| AskForApproval::UnlessTrusted
| AskForApproval::Granular(_) => Decision::Prompt,
};
}
match approval_policy {
AskForApproval::Never => Decision::Allow,
AskForApproval::UnlessTrusted => Decision::Prompt,
AskForApproval::OnRequest => match file_system_sandbox_policy.kind {
FileSystemSandboxKind::Unrestricted | FileSystemSandboxKind::ExternalSandbox => {
Decision::Allow
}
FileSystemSandboxKind::Restricted => {
if sandbox_permissions.requests_sandbox_override() {
Decision::Prompt
} else {
Decision::Allow
}
}
},
AskForApproval::Granular(_) => match file_system_sandbox_policy.kind {
FileSystemSandboxKind::Unrestricted | FileSystemSandboxKind::ExternalSandbox => {
Decision::Allow
}
FileSystemSandboxKind::Restricted => {
if sandbox_permissions.requests_sandbox_override() {
Decision::Prompt
} else {
Decision::Allow
}
}
},
}因此,危险 heuristic 不是一条独立的“禁止列表”:它提高未命中命令的 decision,随后仍经过 AskForApproval 和权限 profile。显式 policy rule 命中时,rule decision 优先参与 reason 和聚合。
8. Unified Exec
解析与分类发生在命令真正交给执行 backend 之前。Unified Exec handler 先解析 environment、cwd、sandbox permissions 和 shell mode,再调用 get_command 得到 resolved command;后续 manager 才接管进程创建、sandbox enforcement 和输出。
源码位置:codex-rs/core/src/tools/handlers/unified_exec/exec_command.rs :: handle, get_command call
let process_id = manager.allocate_process_id().await;
let resolved_command = get_command(
&args,
shell,
&shell_mode,
turn_environment.config().allow_login_shell,
)
.map_err(FunctionCallError::RespondToModel)?;
let command = resolved_command.command;
let shell_type = resolved_command.shell_type;
let command_for_display = codex_shell_command::parse_command::shlex_join(&command);真正发给 manager 的 request 同时携带 command、shell_type、environment、network、sandbox permissions 和 prefix rule。这里没有把 shell parser 的结果直接当成执行命令;parser 只服务策略和展示,执行仍使用 resolved command。
源码位置:codex-rs/core/src/tools/handlers/unified_exec/exec_command.rs :: UnifiedExecProcessManager::exec_command request
manager
.exec_command(
ExecCommandRequest {
command,
shell_type,
hook_command: hook_command.clone(),
process_id,
yield_time_ms,
max_output_tokens,
cwd,
sandbox_cwd: native_environment_cwd,
turn_environment: turn_environment.clone(),
shell_mode,
network: context.step_context.turn.network.clone(),
tty,
sandbox_permissions: effective_additional_permissions.sandbox_permissions,
additional_permissions: normalized_additional_permissions,
additional_permissions_preapproved: effective_additional_permissions
.permissions_preapproved,
justification,
prefix_rule,
},
&context,
)
.await远程 environment 还可能使用目标 executor 的 native shell;handler 会校验 requested shell 是否与环境报告的 shell type 一致。因而本地 Bash parser 的成功不能外推为远程 shell 的同样语法边界,远程执行最终仍由 executor-side sandbox 和 shell owner 强制。
9. 测试边界
可以用以下几组输入复述实现的证明范围:
ls && pwd; echo 'hi there' | wc -l:Bash plain parser 返回四个 argv 段,证明允许连接符和静态引号会被逐段 lowering。echo $(pwd)、ls > out.txt、FOO=bar ls:plain parser 返回None,证明动态展开、重定向和赋值不会被伪装成安全 prefix。- 空脚本或空白脚本:Core 保留原始 argv,证明 parser 失败不会跳过 policy evaluation。
python3 <<'PY'...配合python3Allow rule:可以得到Skip,但bypass_sandbox=false,证明 inner literal 与完整 shell invocation 的安全结论分离。sudo rm -rf、env TARGET=x rm -rf、trap 'rm -rf ...' EXIT:递归 wrapper 会得到ForcedRm;超过八层 wrapper 则 fail closed 为Other。- Windows 的
powershell.exe -Command "echo blocked":PowerShell lowering 后 origin 为PowerShell,规则匹配和 heuristic 使用 inner words,而不是 wrapper executable 本身。
源码位置:
codex-rs/shell-command/src/bash.rs :: testscodex-rs/shell-command/src/command_safety/is_dangerous_command.rs :: testscodex-rs/core/src/exec_policy_windows_tests.rs :: commands_for_exec_policy_parses_powershell_shell_wrapper
#[test]
fn commands_for_exec_policy_parses_powershell_shell_wrapper() {
let command = vec![
"powershell.exe".to_string(),
"-NoProfile".to_string(),
"-Command".to_string(),
"echo blocked".to_string(),
];
assert_eq!(
commands_for_exec_policy(&command),
ExecPolicyCommands {
commands: vec![vec!["echo".to_string(), "blocked".to_string()]],
command_origin: ExecPolicyCommandOrigin::PowerShell,
}
);
}这些测试证明的是当前 allowlist、wrapper 递归和策略映射;它们不能证明任意 Bash、Zsh、PowerShell 或 CMD 语法都能被无损解释,也不能证明 heuristic Allow 可以替代操作系统 sandbox。要追踪一次具体调用,应从 get_command 的 resolved argv 开始,进入 commands_for_exec_policy,再跟到 check_multiple_with_options、render_decision_for_unmatched_command 和最终的 ExecApprovalRequirement。
10. 继续阅读
读者可以用三个问题检查自己是否顺着源码走通:
- 为什么 heredoc 不能直接复用
python3的 Allow rule 来设置bypass_sandbox?请指出 plain parser、原始 argv fallback 和 per-segment 显式 Allow 的三个位置。 - 为什么
bash -lc 'echo "$(rm -rf /tmp/x)"'能被危险扫描命中,而bash -lc 'cmd=rm; $cmd -rf /tmp/x'不返回ForcedRm?请区分 literal scanner 的证据和运行时动态值。 - 为什么 PowerShell production parser 的
Some仍不能说明完整 PowerShell 语义已覆盖?请查看 node allowlist、source_is_covered_by_commands和#requires测试。
下一篇命令规范化与审批缓存会继续追踪这些 segment 如何参与 amendment 模拟、规则落盘和审批复用。
