debugfs Binder节点
debugfs 的价值不在字段数量,而在于把 proc、thread、transaction、node/ref、buffer 和统计量映射回驱动 owner。下面的图给出从症状到节点再到源码状态的阅读路径。
本文面向已读过 Binder线程耗尽排查、Binder锁竞争 和 Binder调用者身份 的读者。debugfs 不是稳定用户态 API,而是驱动把当前内核对象图和计数器导出的只读快照。本文以 Android 17 binder.c 的 show 函数为事实来源,说明每个节点展示什么、锁如何保护读取、字段如何反推阻塞和泄漏。
1. 节点注册
源码文件:kernel/common/drivers/android/binder.c
binder_init 创建 /sys/kernel/debug/binder,注册 state、state_hashed、stats、transactions、transactions_hashed、transaction_log 和 failed_transaction_log;随后创建 proc/ 子目录。所有文件模式为 0444,读取不改变驱动状态。
2. state快照
state_show 调用 print_binder_state,先打印 dead nodes,再遍历全局 binder_procs,对每个 proc 输出线程、节点、引用、allocator、todo 和特殊 work。读取全局 proc 列表使用 binder_procs_lock;proc 内部线程/node/todo 使用 inner lock,引用树使用 outer lock,allocator 另有自己的锁。
3. proc字段
源码文件:kernel/common/drivers/android/binder.c
单 proc 输出包括 proc <pid>、context、thread looper 状态、node/ref、buffer 分配信息和 pending transaction。重点字段:
requested threads: pending+started/max:驱动请求、已注册、上限。ready threads:waiting_threads中可接收 proc work 的数量。free async space:oneway buffer 额度。pending transaction:proc todo 尚未交付的事务。
这些字段要和 BD085 的线程池判断一起使用,不能用 threads 单独推断当前执行数。
4. transactions节点
transactions_show 只打印有活动事务/队列的 proc(print_all=false),适合快速定位等待链;state 则打印完整对象图。事务条目包含 debug id、from/to proc/thread、code、flags、buffer size、nested/reply 关系。若要看调用链,优先读取 transactions,再用 state 补齐 node/ref。
5. stats节点
stats_show 输出全局及每 proc 的 BC/BR 命令计数、对象 created/deleted 计数和 pending transactions。created 长期高于 deleted 可能表示对象尚未释放或快照时刻仍在使用;不能单次读取就断言泄漏,应多次采样看趋势。
6. hashed节点
state_hashed/transactions_hashed 将指针、cookie 等敏感地址哈希后输出,保留拓扑关联但降低地址泄露风险。哈希值只能在同一启动/同一采样上下文比较,不能与另一设备或重启后的值直接对应源码对象。
7. proc专属文件
创建 proc 时,驱动在 binder/proc/<pid> 下建立只读文件;进程释放时 debugfs_remove。路径适合快速查看单 proc,但 PID 可能复用;诊断报告应同时记录启动时间、context 和采样时间,避免把旧 PID 状态当作同一进程。
8. 日志对照
transaction_log 与 failed_transaction_log 是固定环形缓冲区,日志 entry 保存 debug id、from/to、target handle、code、flags、data/offset size、context。debugfs state 是当前快照,transaction_log 是最近事件;两者必须用 debug id 和时间顺序关联。
9. 典型判读
ready threads=0、pending transaction增长:线程池耗尽或线程同步等待。free async space下降、node async 队列增长:oneway 生产过快。requested pending长期非零、started 不增长:用户态未处理BR_SPAWN_LOOPER。- failed log 出现
-ENOSPC:结合 allocator 最大连续块和并发事务判断,不等于单 Parcel 超限。 - created/deleted 差距持续扩大:检查 death/ref/transaction 生命周期。
10. 采样限制
show 函数在多个锁域间分段读取,输出不是全局原子快照;事务可能在打印过程中完成,计数可能前后不一致。应连续采样并与 tracepoint、logcat、线程栈对齐,而不是把一份 state 文本当作严格时间点。
11. 阅读检查
给定一份 proc 输出,分别定位“线程池耗尽”“oneway 堆积”“引用泄漏”“单次大事务失败”四类问题,并为每类指定应联读的 debugfs 节点、内核字段和下一条 trace/log 采样。
