Binder线程管理
用户态的一个 Linux TID 在 Binder 驱动中对应一个 binder_thread,但这个对象不是进程打开设备时一次性创建的线程表项。它通常在该 TID 第一次进入 ioctl 或 poll 时按需创建,挂入所属 binder_proc->threads 红黑树;事务、等待、退出和进程死亡又会从不同路径访问它。
本文面向已经读过 Binder线程状态模型、Binder ioctl入口 和 Binder事务发送 的读者,专门回答驱动对象层面的问题:线程如何被唯一查找,哪些字段由哪把锁保护,为什么退出后对象还可能存活,以及 thread todo、proc todo、transaction stack 的 owner 如何区分。
本文不重复解释用户态线程池创建、BR_SPAWN_LOOPER 的完整策略,也不把 binder_thread 结构体做成逐字段词典。读完后应能顺着 binder_get_thread()、binder_poll()、binder_thread_read() 和 binder_thread_release() 判断一条线程状态从创建到最终释放的证据链。
1. 对象边界
| 对象 | owner | 主要索引 | 主要状态 |
|---|---|---|---|
binder_proc | 一次设备打开 | 全局 proc 表、PID | 进程 work、线程树、临时引用 |
binder_thread | 一个当前 TID | proc->threads 红黑树 | looper、todo、wait、transaction stack |
binder_transaction | 一次事务 | thread/proc/node work | from/to、buffer、父栈 |
task_struct | Linux 内核 | task 引用 | 线程身份与退出 |
线程对象的 pid 是当前 TID,不是 binder_proc 的 group leader PID。binder_thread->proc 把它接回进程;binder_transaction->from/to_thread 把事务栈和 thread work 接起来。
2. 结构字段
源码文件:kernel/common/drivers/android/binder_internal.h
相关结构:struct binder_thread
struct binder_thread {
struct binder_proc *proc;
struct rb_node rb_node;
struct list_head waiting_thread_node;
int pid;
int looper;
bool looper_need_return;
struct binder_transaction *transaction_stack;
struct list_head todo;
bool process_todo;
struct binder_error return_error;
struct binder_error reply_error;
struct binder_extended_error ee;
wait_queue_head_t wait;
struct binder_stats stats;
atomic_t tmp_ref;
bool is_dead;
struct task_struct *task;
spinlock_t prio_lock;
struct binder_priority prio_next;
enum binder_prio_state prio_state;
};字段可以按 owner 分组:rb_node 和 waiting_thread_node 连接 proc 的索引树与等待列表;todo、process_todo 和 transaction_stack 描述当前线程的工作和同步栈;wait 是阻塞读/poll 的唤醒点;错误槽位、临时引用和优先级锁分别由不同路径消费。
3. 查找创建
3.1 外层入口
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_get_thread()
static struct binder_thread *binder_get_thread(
struct binder_proc *proc) {
struct binder_thread *thread;
struct binder_thread *new_thread;
binder_inner_proc_lock(proc);
thread = binder_get_thread_ilocked(
proc, NULL);
binder_inner_proc_unlock(proc);
if (!thread) {
new_thread = kzalloc(
sizeof(*thread), GFP_KERNEL);
if (!new_thread)
return NULL;
binder_inner_proc_lock(proc);
thread = binder_get_thread_ilocked(
proc, new_thread);
binder_inner_proc_unlock(proc);
if (thread != new_thread)
kfree(new_thread);
}
return thread;
}第一次查找只持有 proc->inner_lock,锁外执行可睡眠的 kzalloc,再重新加锁查找并插入。第二次查找处理竞态:另一个入口已经为同一 TID 创建对象时,临时对象被释放。
3.2 红黑树查找
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_get_thread_ilocked()
static struct binder_thread *
binder_get_thread_ilocked(
struct binder_proc *proc,
struct binder_thread *new_thread) {
struct binder_thread *thread = NULL;
struct rb_node *parent = NULL;
struct rb_node **p =
&proc->threads.rb_node;
while (*p) {
parent = *p;
thread = rb_entry(parent,
struct binder_thread, rb_node);
if (current->pid < thread->pid)
p = &(*p)->rb_left;
else if (current->pid > thread->pid)
p = &(*p)->rb_right;
else
return thread;
}
if (!new_thread)
return NULL;
rb_link_node(&new_thread->rb_node,
parent, p);
rb_insert_color(&new_thread->rb_node,
&proc->threads);
return new_thread;
}树键是 current->pid。它不是 handle,也不是 task 地址;同一 proc 内同一 TID 始终复用一个 binder_thread。该函数要求调用者已经持有 inner lock。
3.3 初始状态
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_get_thread_ilocked()
thread = new_thread;
binder_stats_created(BINDER_STAT_THREAD);
thread->proc = proc;
thread->pid = current->pid;
get_task_struct(current);
thread->task = current;
atomic_set(&thread->tmp_ref, 0);
init_waitqueue_head(&thread->wait);
INIT_LIST_HEAD(&thread->todo);
thread->looper_need_return = true;
thread->return_error.work.type =
BINDER_WORK_RETURN_ERROR;
thread->return_error.cmd = BR_OK;
thread->reply_error.work.type =
BINDER_WORK_RETURN_ERROR;
thread->reply_error.cmd = BR_OK;
spin_lock_init(&thread->prio_lock);
thread->prio_state = BINDER_PRIO_SET;
thread->ee.command = BR_OK;
INIT_LIST_HEAD(
&thread->waiting_thread_node);创建时持有当前 task 引用;初始 todo、等待节点和错误槽为空。looper_need_return = true 使新线程第一次驱动交互先返回用户态处理注册/状态。
4. 入口复用
4.1 Ioctl入口
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_ioctl()
thread = binder_get_thread(proc);
if (thread == NULL) {
ret = -ENOMEM;
goto err;
}
switch (cmd) {
case BINDER_WRITE_READ:
ret = binder_ioctl_write_read(
filp, arg, thread);
break;
case BINDER_THREAD_EXIT:
binder_thread_release(proc, thread);
thread = NULL;
break;
}任何 ioctl 命令都可能触发 thread 创建,包括版本查询、线程上限设置和 BINDER_WRITE_READ。BINDER_THREAD_EXIT 完成 helper 后立即把局部指针清空,统一尾部不能再写 looper_need_return。
4.2 Poll入口
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_poll()
static __poll_t binder_poll(
struct file *filp,
struct poll_table_struct *wait) {
struct binder_proc *proc =
filp->private_data;
struct binder_thread *thread =
binder_get_thread(proc);
bool wait_for_proc_work;
if (!thread)
return EPOLLERR;
binder_inner_proc_lock(proc);
thread->looper |= BINDER_LOOPER_STATE_POLL;
wait_for_proc_work =
binder_available_for_proc_work_ilocked(
thread);
binder_inner_proc_unlock(proc);
poll_wait(filp, &thread->wait, wait);
if (binder_has_work(
thread, wait_for_proc_work))
return EPOLLIN;
return 0;
}poll 和 ioctl 共享按 TID 查找入口,但 poll 额外设置 POLL 位并注册同一个 wait queue。poll 返回可读不等于 read 已经取走 work;真正消费仍发生在 BINDER_WRITE_READ。
4.3 Read入口
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_thread_read()
wait_for_proc_work =
binder_available_for_proc_work_ilocked(
thread);
thread->looper |=
BINDER_LOOPER_STATE_WAITING;
if (non_block) {
ret = binder_has_work(
thread, wait_for_proc_work)
? 0 : -EAGAIN;
} else {
ret = binder_wait_for_work(
thread, wait_for_proc_work);
}
thread->looper &=
~BINDER_LOOPER_STATE_WAITING;WAITING 由当前线程设置和清除,但其他线程可能通过 flush 写 looper_need_return 并唤醒 wait。transaction_stack 非空的线程不具备普通 proc work 消费资格,仍可消费自己的 thread todo。
5. 工作所有权
5.1 Thread todo
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_enqueue_thread_work_ilocked()、binder_thread_read()
if (!binder_worklist_empty_ilocked(
&thread->todo))
list = &thread->todo;
w = binder_dequeue_work_head_ilocked(list);
if (binder_worklist_empty_ilocked(
&thread->todo))
thread->process_todo = false;thread todo 的消费者固定为这个 binder_thread。reply、指定 target_thread 的同步事务、线程错误和部分完成 work 都可以进入这里;其他线程不会因为 proc todo 为空就替它直接消费。
5.2 Proc todo
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_select_thread_ilocked()、binder_proc_transaction()
if (!thread && !pending_async)
thread = binder_select_thread_ilocked(proc);
if (thread)
binder_enqueue_thread_work_ilocked(
thread, &t->work);
else if (!pending_async)
binder_enqueue_work_ilocked(
&t->work, &proc->todo);proc todo 是进程级 work。没有指定目标线程时,驱动先从 waiting_threads 选择一个可消费 proc work 的线程;没有可用线程才把 work 留在 proc todo。
5.3 Async todo
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_proc_transaction()
if (oneway) {
if (node->has_async_transaction)
pending_async = true;
else
node->has_async_transaction = true;
}
if (!pending_async)
binder_enqueue_work_ilocked(
&t->work, &proc->todo);
else
binder_enqueue_work_ilocked(
&t->work, &node->async_todo);同一 node 的异步事务通过 has_async_transaction 和 async_todo 保持顺序。第一笔可进入 proc work;后续 pending async 由 node 持有,当前事务释放后再转移下一项。
6. Looper状态
6.1 注册命令
源码文件:kernel/common/drivers/android/binder.c
相关命令:BC_REGISTER_LOOPER、BC_ENTER_LOOPER、BC_EXIT_LOOPER
case BC_REGISTER_LOOPER:
binder_inner_proc_lock(proc);
if (thread->looper & BINDER_LOOPER_STATE_ENTERED)
thread->looper |= BINDER_LOOPER_STATE_INVALID;
else if (proc->requested_threads == 0)
thread->looper |= BINDER_LOOPER_STATE_INVALID;
else {
proc->requested_threads--;
proc->requested_threads_started++;
}
thread->looper |= BINDER_LOOPER_STATE_REGISTERED;
binder_inner_proc_unlock(proc);
break;
case BC_ENTER_LOOPER:
if (thread->looper & BINDER_LOOPER_STATE_REGISTERED)
thread->looper |= BINDER_LOOPER_STATE_INVALID;
thread->looper |= BINDER_LOOPER_STATE_ENTERED;
break;
case BC_EXIT_LOOPER:
thread->looper |= BINDER_LOOPER_STATE_EXITED;
break;注册状态由 thread->looper 保存,但 requested_threads 和 requested_threads_started 属于 proc,并由 inner lock 保护。未请求时 REGISTER 会标为 INVALID,不会立即释放 thread。
6.2 状态与消费
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_available_for_proc_work_ilocked()
static bool binder_available_for_proc_work_ilocked(
struct binder_thread *thread) {
return !thread->transaction_stack &&
binder_worklist_empty_ilocked(
&thread->todo);
}Looper 注册状态和 work 消费资格是两件事:REGISTERED/ENTERED 允许线程作为合法 looper,transaction_stack 和 thread todo 决定它此刻能否接 proc work。
7. 线程退出
7.1 显式退出
源码文件:kernel/common/drivers/android/binder.c
相关命令:BINDER_THREAD_EXIT
case BINDER_THREAD_EXIT:
binder_debug(BINDER_DEBUG_THREADS,
"%d:%d exit\n",
proc->pid, thread->pid);
binder_thread_release(proc, thread);
thread = NULL;
break;显式 ioctl 退出和用户态发送 BC_EXIT_LOOPER 不同:前者进入 binder_thread_release 释放内核 thread 对象;后者只设置 looper exited 位,线程仍可能继续存在并由进程 release 最终清理。
7.2 释放栈
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_thread_release()
static int binder_thread_release(
struct binder_proc *proc,
struct binder_thread *thread) {
struct binder_transaction *send_reply = NULL;
int active_transactions = 0;
binder_inner_proc_lock(proc);
proc->tmp_ref++;
atomic_inc(&thread->tmp_ref);
rb_erase(&thread->rb_node,
&proc->threads);
thread->is_dead = true;
binder_inner_proc_unlock(proc);
if (thread->looper & BINDER_LOOPER_STATE_POLL)
wake_up_pollfree(&thread->wait);
if (thread->looper & BINDER_LOOPER_STATE_POLL)
synchronize_rcu();
if (send_reply)
binder_send_failed_reply(
send_reply, BR_DEAD_REPLY);
binder_release_work(proc, &thread->todo);
binder_thread_dec_tmpref(thread);
return active_transactions;
}释放先从 proc->threads 摘除并标记 is_dead。若 thread 曾进入 poll 模式,驱动先用 wake_up_pollfree() 通知 poll/epoll 移除 wait queue,再用 synchronize_rcu() 等待并发读侧结束;若栈中有待回复事务, 发送 BR_DEAD_REPLY。随后才处理 thread todo。proc tmp_ref 保证 thread 摘除后 proc 仍活着;thread tmp_ref 保证清理过程不会直接触发 kfree。
7.3 事务栈
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_thread_release()
t = thread->transaction_stack;
thread->is_dead = true;
while (t) {
if (t->to_thread == thread) {
t->to_proc = NULL;
t->to_thread = NULL;
if (t->buffer) {
t->buffer->transaction = NULL;
t->buffer = NULL;
}
t = t->to_parent;
} else if (t->from == thread) {
t->from = NULL;
t = t->from_parent;
}
}线程死亡可能位于事务栈中间。作为 to_thread 的事务断开目标和 buffer;作为 from 的事务断开发送方。驱动沿 to_parent/from_parent 继续处理,而不是只清除栈顶指针。
7.4 Todo清理
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_release_work()、binder_free_thread()
binder_release_work(proc, &thread->todo);
static void binder_free_thread(
struct binder_thread *thread) {
BUG_ON(!list_empty(&thread->todo));
binder_stats_deleted(BINDER_STAT_THREAD);
binder_proc_dec_tmpref(thread->proc);
put_task_struct(thread->task);
kfree(thread);
}thread todo 中的 transaction、error、death work 不能直接丢弃;release_work 按 work type 清理对象。只有 todo 为空且 tmp_ref 归零,free_thread 才归还 task 引用并释放 thread。
8. 引用保护
8.1 Thread引用
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_thread_dec_tmpref()
static void binder_thread_dec_tmpref(
struct binder_thread *thread) {
binder_inner_proc_lock(thread->proc);
atomic_dec(&thread->tmp_ref);
if (thread->is_dead &&
!atomic_read(&thread->tmp_ref)) {
binder_inner_proc_unlock(thread->proc);
binder_free_thread(thread);
return;
}
binder_inner_proc_unlock(thread->proc);
}is_dead 只表示 thread 已从正常使用路径退出,不能直接推出内存已释放。临时引用归零才触发 free;释放函数又会减少 proc tmp_ref,可能进一步触发 proc 最终释放。
8.2 Transaction引用
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_get_txn_from()、binder_get_txn_from_and_acq_inner()
static struct binder_thread *binder_get_txn_from(
struct binder_transaction *t) {
struct binder_thread *from;
spin_lock(&t->lock);
from = t->from;
if (from)
atomic_inc(&from->tmp_ref);
spin_unlock(&t->lock);
return from;
}事务读取 from 指针时同时增加 thread tmp_ref,使用完成后调用 binder_thread_dec_tmpref。transaction lock 保护 from/to 指针,proc inner lock 保护 thread 树和 transaction stack;不能只拿其中一把锁推断对象稳定。
8.3 Proc释放关系
源码文件:kernel/common/drivers/android/binder.c
相关函数:binder_proc_dec_tmpref()
if (proc->is_dead &&
RB_EMPTY_ROOT(&proc->threads) &&
!proc->tmp_ref) {
binder_inner_proc_unlock(proc);
binder_free_proc(proc);
return;
}thread 释放会减少 proc tmp_ref;只有 proc 已死、线程树为空且 proc tmp_ref 为零,binder_proc 才能进入 BD013 的最终释放。
9. 测试输入
9.1 驱动退出
源码文件:frameworks/native/libs/binder/tests/binderDriverInterfaceTest.cpp
相关测试:ThreadExit
TEST_F(BinderDriverInterfaceTest, ThreadExit) {
int32_t dummy = 0;
binderTestIoctl(
BINDER_THREAD_EXIT, &dummy);
static_cast<BinderDriverInterfaceTestEnv *>(
binder_env)->EnterLooper();
}测试先显式退出当前 driver thread,再发送 BC_ENTER_LOOPER。它证明 fd/proc 仍可创建新的 thread 状态并进入 looper;不证明活动 transaction stack 的完整拆栈。
9.2 线程池边界
源码文件:frameworks/native/libs/binder/tests/binderLibTest.cpp
相关测试:ThreadPoolStarted、ThreadPoolAvailableThreads
TEST_F(BinderLibTest, ThreadPoolStarted) {
Parcel data, reply;
sp<IBinder> server = addServer();
ASSERT_TRUE(server != nullptr);
EXPECT_THAT(server->transact(
BINDER_LIB_TEST_IS_THREADPOOL_STARTED,
data, &reply), NO_ERROR);
EXPECT_TRUE(reply.readBool());
}这些测试从用户态观察线程池启动和可用线程上界,间接覆盖 driver thread registration、BR_SPAWN_LOOPER 和 work 消费;不能替代内核测试,也不能证明固定线程数。
9.3 线程死亡通知
源码文件:frameworks/native/libs/binder/tests/binderLibTest.cpp
相关测试:DeathNotificationThread
该测试让目标服务退出,再由另一个客户端注册死亡通知,验证通知 work 可以投递到 proc workqueue,而不是错误地绑定到一个当前阻塞的注册线程。它覆盖 thread/proc work owner 的可观察结果,但不证明所有 thread release race。
10. 失败边界
| 阶段 | 失败或竞态 | 结果 |
|---|---|---|
| 首次查找 | kzalloc 失败 | ioctl/poll 返回分配错误 |
| 双重创建 | 另一路径先插入同 TID | 释放临时对象,复用已登记 thread |
| looper注册 | 未请求或状态冲突 | 设置 INVALID,记录用户错误 |
| read等待 | 信号中断 | -EINTR,WAITING 清除 |
| thread退出 | 活动 transaction | 拆 from/to,保留临时引用 |
| todo释放 | 未投递 transaction/death/error work | 按 work type 清理 |
| 最终释放 | tmp_ref 非零 | 延迟 free_thread 和 proc 回收 |
11. 源码复现
本文主线:
ioctl/poll
→ binder_get_thread(proc)
→ current TID查找或锁外分配
→ proc->threads红黑树
→ thread todo/wait/transaction_stack
→ looper与work消费
→ BINDER_THREAD_EXIT或进程release
→ 栈拆解、todo清理、tmp_ref归零
→ binder_free_thread
→ 可能触发binder_free_proc源码搜索:
rg -n "binder_get_thread|binder_get_thread_ilocked|struct binder_thread" \
kernel/common/drivers/android/binder.c \
kernel/common/drivers/android/binder_internal.h
rg -n "binder_poll|binder_thread_read|binder_wait_for_work|waiting_thread_node" \
kernel/common/drivers/android/binder.c
rg -n "binder_thread_release|binder_thread_dec_tmpref|binder_free_thread|BINDER_THREAD_EXIT" \
kernel/common/drivers/android/binder.c
rg -n "ThreadExit|ThreadPoolStarted|ThreadPoolAvailableThreads|DeathNotificationThread" \
frameworks/native/libs/binder/tests/binderDriverInterfaceTest.cpp \
frameworks/native/libs/binder/tests/binderLibTest.cpp如果能够说明为什么同一个 TID 必须复用一个 binder_thread、为什么 poll 和 ioctl 共享 thread 但消费 owner 不同、为什么 is_dead 不等于已释放,以及 thread tmp_ref 如何牵连 proc 最终释放,就已经掌握了 Binder 驱动线程对象的生命周期。
图中的 Dead 只是禁止新 work 进入;内核仍可能通过 tmp_ref 暂时持有对象,直到所有事务和清理路径结束。
