Skip to content

Binder线程管理

追踪 binder_thread 的 TID 查找、红黑树登记、work 与 wait owner、临时引用和退出清理。

基于android-17.0.0_r1
AndroidBinderbinder_thread驱动源码阅读

Binder线程管理 ​

用户态的一个 Linux TID 在 Binder 驱动中对应一个 binder_thread,但这个对象不是进程打开设备时一次性创建的线程表项。它通常在该 TID 第一次进入 ioctl 或 poll 时按需创建,挂入所属 binder_proc->threads 红黑树;事务、等待、退出和进程死亡又会从不同路径访问它。

本文面向已经读过 Binder线程状态模型、Binder ioctl入口 和 Binder事务发送 的读者,专门回答驱动对象层面的问题:线程如何被唯一查找,哪些字段由哪把锁保护,为什么退出后对象还可能存活,以及 thread todo、proc todo、transaction stack 的 owner 如何区分。

本文不重复解释用户态线程池创建、BR_SPAWN_LOOPER 的完整策略,也不把 binder_thread 结构体做成逐字段词典。读完后应能顺着 binder_get_thread()、binder_poll()、binder_thread_read() 和 binder_thread_release() 判断一条线程状态从创建到最终释放的证据链。

1. 对象边界 ​

对象owner主要索引主要状态
binder_proc一次设备打开全局 proc 表、PID进程 work、线程树、临时引用
binder_thread一个当前 TIDproc->threads 红黑树looper、todo、wait、transaction stack
binder_transaction一次事务thread/proc/node workfrom/to、buffer、父栈
task_structLinux 内核task 引用线程身份与退出

线程对象的 pid 是当前 TID,不是 binder_proc 的 group leader PID。binder_thread->proc 把它接回进程;binder_transaction->from/to_thread 把事务栈和 thread work 接起来。

2. 结构字段 ​

源码文件:kernel/common/drivers/android/binder_internal.h

相关结构:struct binder_thread

c
struct binder_thread {
    struct binder_proc *proc;
    struct rb_node rb_node;
    struct list_head waiting_thread_node;
    int pid;
    int looper;
    bool looper_need_return;
    struct binder_transaction *transaction_stack;
    struct list_head todo;
    bool process_todo;
    struct binder_error return_error;
    struct binder_error reply_error;
    struct binder_extended_error ee;
    wait_queue_head_t wait;
    struct binder_stats stats;
    atomic_t tmp_ref;
    bool is_dead;
    struct task_struct *task;
    spinlock_t prio_lock;
    struct binder_priority prio_next;
    enum binder_prio_state prio_state;
};

字段可以按 owner 分组:rb_node 和 waiting_thread_node 连接 proc 的索引树与等待列表;todo、process_todo 和 transaction_stack 描述当前线程的工作和同步栈;wait 是阻塞读/poll 的唤醒点;错误槽位、临时引用和优先级锁分别由不同路径消费。

3. 查找创建 ​

3.1 外层入口 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_get_thread()

c
static struct binder_thread *binder_get_thread(
        struct binder_proc *proc) {
    struct binder_thread *thread;
    struct binder_thread *new_thread;

    binder_inner_proc_lock(proc);
    thread = binder_get_thread_ilocked(
            proc, NULL);
    binder_inner_proc_unlock(proc);
    if (!thread) {
        new_thread = kzalloc(
                sizeof(*thread), GFP_KERNEL);
        if (!new_thread)
            return NULL;
        binder_inner_proc_lock(proc);
        thread = binder_get_thread_ilocked(
                proc, new_thread);
        binder_inner_proc_unlock(proc);
        if (thread != new_thread)
            kfree(new_thread);
    }
    return thread;
}

第一次查找只持有 proc->inner_lock,锁外执行可睡眠的 kzalloc,再重新加锁查找并插入。第二次查找处理竞态:另一个入口已经为同一 TID 创建对象时,临时对象被释放。

3.2 红黑树查找 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_get_thread_ilocked()

c
static struct binder_thread *
binder_get_thread_ilocked(
        struct binder_proc *proc,
        struct binder_thread *new_thread) {
    struct binder_thread *thread = NULL;
    struct rb_node *parent = NULL;
    struct rb_node **p =
            &proc->threads.rb_node;

    while (*p) {
        parent = *p;
        thread = rb_entry(parent,
                struct binder_thread, rb_node);
        if (current->pid < thread->pid)
            p = &(*p)->rb_left;
        else if (current->pid > thread->pid)
            p = &(*p)->rb_right;
        else
            return thread;
    }
    if (!new_thread)
        return NULL;
    rb_link_node(&new_thread->rb_node,
                 parent, p);
    rb_insert_color(&new_thread->rb_node,
                    &proc->threads);
    return new_thread;
}

树键是 current->pid。它不是 handle,也不是 task 地址;同一 proc 内同一 TID 始终复用一个 binder_thread。该函数要求调用者已经持有 inner lock。

3.3 初始状态 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_get_thread_ilocked()

c
thread = new_thread;
binder_stats_created(BINDER_STAT_THREAD);
thread->proc = proc;
thread->pid = current->pid;
get_task_struct(current);
thread->task = current;
atomic_set(&thread->tmp_ref, 0);
init_waitqueue_head(&thread->wait);
INIT_LIST_HEAD(&thread->todo);
thread->looper_need_return = true;
thread->return_error.work.type =
        BINDER_WORK_RETURN_ERROR;
thread->return_error.cmd = BR_OK;
thread->reply_error.work.type =
        BINDER_WORK_RETURN_ERROR;
thread->reply_error.cmd = BR_OK;
spin_lock_init(&thread->prio_lock);
thread->prio_state = BINDER_PRIO_SET;
thread->ee.command = BR_OK;
INIT_LIST_HEAD(
        &thread->waiting_thread_node);

创建时持有当前 task 引用;初始 todo、等待节点和错误槽为空。looper_need_return = true 使新线程第一次驱动交互先返回用户态处理注册/状态。

4. 入口复用 ​

4.1 Ioctl入口 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_ioctl()

c
thread = binder_get_thread(proc);
if (thread == NULL) {
    ret = -ENOMEM;
    goto err;
}
switch (cmd) {
case BINDER_WRITE_READ:
    ret = binder_ioctl_write_read(
            filp, arg, thread);
    break;
case BINDER_THREAD_EXIT:
    binder_thread_release(proc, thread);
    thread = NULL;
    break;
}

任何 ioctl 命令都可能触发 thread 创建,包括版本查询、线程上限设置和 BINDER_WRITE_READ。BINDER_THREAD_EXIT 完成 helper 后立即把局部指针清空,统一尾部不能再写 looper_need_return。

4.2 Poll入口 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_poll()

c
static __poll_t binder_poll(
        struct file *filp,
        struct poll_table_struct *wait) {
    struct binder_proc *proc =
            filp->private_data;
    struct binder_thread *thread =
            binder_get_thread(proc);
    bool wait_for_proc_work;

    if (!thread)
        return EPOLLERR;
    binder_inner_proc_lock(proc);
    thread->looper |= BINDER_LOOPER_STATE_POLL;
    wait_for_proc_work =
        binder_available_for_proc_work_ilocked(
            thread);
    binder_inner_proc_unlock(proc);
    poll_wait(filp, &thread->wait, wait);
    if (binder_has_work(
            thread, wait_for_proc_work))
        return EPOLLIN;
    return 0;
}

poll 和 ioctl 共享按 TID 查找入口,但 poll 额外设置 POLL 位并注册同一个 wait queue。poll 返回可读不等于 read 已经取走 work;真正消费仍发生在 BINDER_WRITE_READ。

4.3 Read入口 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_read()

c
wait_for_proc_work =
    binder_available_for_proc_work_ilocked(
        thread);
thread->looper |=
    BINDER_LOOPER_STATE_WAITING;
if (non_block) {
    ret = binder_has_work(
            thread, wait_for_proc_work)
        ? 0 : -EAGAIN;
} else {
    ret = binder_wait_for_work(
            thread, wait_for_proc_work);
}
thread->looper &=
    ~BINDER_LOOPER_STATE_WAITING;

WAITING 由当前线程设置和清除,但其他线程可能通过 flush 写 looper_need_return 并唤醒 wait。transaction_stack 非空的线程不具备普通 proc work 消费资格,仍可消费自己的 thread todo。

5. 工作所有权 ​

5.1 Thread todo ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_enqueue_thread_work_ilocked()、binder_thread_read()

c
if (!binder_worklist_empty_ilocked(
        &thread->todo))
    list = &thread->todo;
w = binder_dequeue_work_head_ilocked(list);
if (binder_worklist_empty_ilocked(
        &thread->todo))
    thread->process_todo = false;

thread todo 的消费者固定为这个 binder_thread。reply、指定 target_thread 的同步事务、线程错误和部分完成 work 都可以进入这里;其他线程不会因为 proc todo 为空就替它直接消费。

5.2 Proc todo ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_select_thread_ilocked()、binder_proc_transaction()

c
if (!thread && !pending_async)
    thread = binder_select_thread_ilocked(proc);
if (thread)
    binder_enqueue_thread_work_ilocked(
        thread, &t->work);
else if (!pending_async)
    binder_enqueue_work_ilocked(
        &t->work, &proc->todo);

proc todo 是进程级 work。没有指定目标线程时,驱动先从 waiting_threads 选择一个可消费 proc work 的线程;没有可用线程才把 work 留在 proc todo。

5.3 Async todo ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_proc_transaction()

c
if (oneway) {
    if (node->has_async_transaction)
        pending_async = true;
    else
        node->has_async_transaction = true;
}
if (!pending_async)
    binder_enqueue_work_ilocked(
        &t->work, &proc->todo);
else
    binder_enqueue_work_ilocked(
        &t->work, &node->async_todo);

同一 node 的异步事务通过 has_async_transaction 和 async_todo 保持顺序。第一笔可进入 proc work;后续 pending async 由 node 持有,当前事务释放后再转移下一项。

6. Looper状态 ​

6.1 注册命令 ​

源码文件:kernel/common/drivers/android/binder.c

相关命令:BC_REGISTER_LOOPER、BC_ENTER_LOOPER、BC_EXIT_LOOPER

c
case BC_REGISTER_LOOPER:
    binder_inner_proc_lock(proc);
    if (thread->looper & BINDER_LOOPER_STATE_ENTERED)
        thread->looper |= BINDER_LOOPER_STATE_INVALID;
    else if (proc->requested_threads == 0)
        thread->looper |= BINDER_LOOPER_STATE_INVALID;
    else {
        proc->requested_threads--;
        proc->requested_threads_started++;
    }
    thread->looper |= BINDER_LOOPER_STATE_REGISTERED;
    binder_inner_proc_unlock(proc);
    break;

case BC_ENTER_LOOPER:
    if (thread->looper & BINDER_LOOPER_STATE_REGISTERED)
        thread->looper |= BINDER_LOOPER_STATE_INVALID;
    thread->looper |= BINDER_LOOPER_STATE_ENTERED;
    break;

case BC_EXIT_LOOPER:
    thread->looper |= BINDER_LOOPER_STATE_EXITED;
    break;

注册状态由 thread->looper 保存,但 requested_threads 和 requested_threads_started 属于 proc,并由 inner lock 保护。未请求时 REGISTER 会标为 INVALID,不会立即释放 thread。

6.2 状态与消费 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_available_for_proc_work_ilocked()

c
static bool binder_available_for_proc_work_ilocked(
        struct binder_thread *thread) {
    return !thread->transaction_stack &&
        binder_worklist_empty_ilocked(
            &thread->todo);
}

Looper 注册状态和 work 消费资格是两件事:REGISTERED/ENTERED 允许线程作为合法 looper,transaction_stack 和 thread todo 决定它此刻能否接 proc work。

7. 线程退出 ​

7.1 显式退出 ​

源码文件:kernel/common/drivers/android/binder.c

相关命令:BINDER_THREAD_EXIT

c
case BINDER_THREAD_EXIT:
    binder_debug(BINDER_DEBUG_THREADS,
            "%d:%d exit\n",
            proc->pid, thread->pid);
    binder_thread_release(proc, thread);
    thread = NULL;
    break;

显式 ioctl 退出和用户态发送 BC_EXIT_LOOPER 不同:前者进入 binder_thread_release 释放内核 thread 对象;后者只设置 looper exited 位,线程仍可能继续存在并由进程 release 最终清理。

7.2 释放栈 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_release()

c
static int binder_thread_release(
        struct binder_proc *proc,
        struct binder_thread *thread) {
    struct binder_transaction *send_reply = NULL;
    int active_transactions = 0;
    binder_inner_proc_lock(proc);
    proc->tmp_ref++;
    atomic_inc(&thread->tmp_ref);
    rb_erase(&thread->rb_node,
             &proc->threads);
    thread->is_dead = true;
    binder_inner_proc_unlock(proc);

    if (thread->looper & BINDER_LOOPER_STATE_POLL)
        wake_up_pollfree(&thread->wait);
    if (thread->looper & BINDER_LOOPER_STATE_POLL)
        synchronize_rcu();
    if (send_reply)
        binder_send_failed_reply(
                send_reply, BR_DEAD_REPLY);
    binder_release_work(proc, &thread->todo);
    binder_thread_dec_tmpref(thread);
    return active_transactions;
}

释放先从 proc->threads 摘除并标记 is_dead。若 thread 曾进入 poll 模式,驱动先用 wake_up_pollfree() 通知 poll/epoll 移除 wait queue,再用 synchronize_rcu() 等待并发读侧结束;若栈中有待回复事务, 发送 BR_DEAD_REPLY。随后才处理 thread todo。proc tmp_ref 保证 thread 摘除后 proc 仍活着;thread tmp_ref 保证清理过程不会直接触发 kfree。

7.3 事务栈 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_release()

c
t = thread->transaction_stack;
thread->is_dead = true;
while (t) {
    if (t->to_thread == thread) {
        t->to_proc = NULL;
        t->to_thread = NULL;
        if (t->buffer) {
            t->buffer->transaction = NULL;
            t->buffer = NULL;
        }
        t = t->to_parent;
    } else if (t->from == thread) {
        t->from = NULL;
        t = t->from_parent;
    }
}

线程死亡可能位于事务栈中间。作为 to_thread 的事务断开目标和 buffer;作为 from 的事务断开发送方。驱动沿 to_parent/from_parent 继续处理,而不是只清除栈顶指针。

7.4 Todo清理 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_release_work()、binder_free_thread()

c
binder_release_work(proc, &thread->todo);

static void binder_free_thread(
        struct binder_thread *thread) {
    BUG_ON(!list_empty(&thread->todo));
    binder_stats_deleted(BINDER_STAT_THREAD);
    binder_proc_dec_tmpref(thread->proc);
    put_task_struct(thread->task);
    kfree(thread);
}

thread todo 中的 transaction、error、death work 不能直接丢弃;release_work 按 work type 清理对象。只有 todo 为空且 tmp_ref 归零,free_thread 才归还 task 引用并释放 thread。

8. 引用保护 ​

8.1 Thread引用 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_dec_tmpref()

c
static void binder_thread_dec_tmpref(
        struct binder_thread *thread) {
    binder_inner_proc_lock(thread->proc);
    atomic_dec(&thread->tmp_ref);
    if (thread->is_dead &&
            !atomic_read(&thread->tmp_ref)) {
        binder_inner_proc_unlock(thread->proc);
        binder_free_thread(thread);
        return;
    }
    binder_inner_proc_unlock(thread->proc);
}

is_dead 只表示 thread 已从正常使用路径退出,不能直接推出内存已释放。临时引用归零才触发 free;释放函数又会减少 proc tmp_ref,可能进一步触发 proc 最终释放。

8.2 Transaction引用 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_get_txn_from()、binder_get_txn_from_and_acq_inner()

c
static struct binder_thread *binder_get_txn_from(
        struct binder_transaction *t) {
    struct binder_thread *from;
    spin_lock(&t->lock);
    from = t->from;
    if (from)
        atomic_inc(&from->tmp_ref);
    spin_unlock(&t->lock);
    return from;
}

事务读取 from 指针时同时增加 thread tmp_ref,使用完成后调用 binder_thread_dec_tmpref。transaction lock 保护 from/to 指针,proc inner lock 保护 thread 树和 transaction stack;不能只拿其中一把锁推断对象稳定。

8.3 Proc释放关系 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_proc_dec_tmpref()

c
if (proc->is_dead &&
        RB_EMPTY_ROOT(&proc->threads) &&
        !proc->tmp_ref) {
    binder_inner_proc_unlock(proc);
    binder_free_proc(proc);
    return;
}

thread 释放会减少 proc tmp_ref;只有 proc 已死、线程树为空且 proc tmp_ref 为零,binder_proc 才能进入 BD013 的最终释放。

9. 测试输入 ​

9.1 驱动退出 ​

源码文件:frameworks/native/libs/binder/tests/binderDriverInterfaceTest.cpp

相关测试:ThreadExit

cpp
TEST_F(BinderDriverInterfaceTest, ThreadExit) {
    int32_t dummy = 0;
    binderTestIoctl(
        BINDER_THREAD_EXIT, &dummy);
    static_cast<BinderDriverInterfaceTestEnv *>(
        binder_env)->EnterLooper();
}

测试先显式退出当前 driver thread,再发送 BC_ENTER_LOOPER。它证明 fd/proc 仍可创建新的 thread 状态并进入 looper;不证明活动 transaction stack 的完整拆栈。

9.2 线程池边界 ​

源码文件:frameworks/native/libs/binder/tests/binderLibTest.cpp

相关测试:ThreadPoolStarted、ThreadPoolAvailableThreads

cpp
TEST_F(BinderLibTest, ThreadPoolStarted) {
    Parcel data, reply;
    sp<IBinder> server = addServer();
    ASSERT_TRUE(server != nullptr);
    EXPECT_THAT(server->transact(
        BINDER_LIB_TEST_IS_THREADPOOL_STARTED,
        data, &reply), NO_ERROR);
    EXPECT_TRUE(reply.readBool());
}

这些测试从用户态观察线程池启动和可用线程上界,间接覆盖 driver thread registration、BR_SPAWN_LOOPER 和 work 消费;不能替代内核测试,也不能证明固定线程数。

9.3 线程死亡通知 ​

源码文件:frameworks/native/libs/binder/tests/binderLibTest.cpp

相关测试:DeathNotificationThread

该测试让目标服务退出,再由另一个客户端注册死亡通知,验证通知 work 可以投递到 proc workqueue,而不是错误地绑定到一个当前阻塞的注册线程。它覆盖 thread/proc work owner 的可观察结果,但不证明所有 thread release race。

10. 失败边界 ​

阶段失败或竞态结果
首次查找kzalloc 失败ioctl/poll 返回分配错误
双重创建另一路径先插入同 TID释放临时对象,复用已登记 thread
looper注册未请求或状态冲突设置 INVALID,记录用户错误
read等待信号中断-EINTR,WAITING 清除
thread退出活动 transaction拆 from/to,保留临时引用
todo释放未投递 transaction/death/error work按 work type 清理
最终释放tmp_ref 非零延迟 free_thread 和 proc 回收

11. 源码复现 ​

本文主线:

text
ioctl/poll
→ binder_get_thread(proc)
→ current TID查找或锁外分配
→ proc->threads红黑树
→ thread todo/wait/transaction_stack
→ looper与work消费
→ BINDER_THREAD_EXIT或进程release
→ 栈拆解、todo清理、tmp_ref归零
→ binder_free_thread
→ 可能触发binder_free_proc

源码搜索:

bash
rg -n "binder_get_thread|binder_get_thread_ilocked|struct binder_thread" \
  kernel/common/drivers/android/binder.c \
  kernel/common/drivers/android/binder_internal.h

rg -n "binder_poll|binder_thread_read|binder_wait_for_work|waiting_thread_node" \
  kernel/common/drivers/android/binder.c

rg -n "binder_thread_release|binder_thread_dec_tmpref|binder_free_thread|BINDER_THREAD_EXIT" \
  kernel/common/drivers/android/binder.c

rg -n "ThreadExit|ThreadPoolStarted|ThreadPoolAvailableThreads|DeathNotificationThread" \
  frameworks/native/libs/binder/tests/binderDriverInterfaceTest.cpp \
  frameworks/native/libs/binder/tests/binderLibTest.cpp

如果能够说明为什么同一个 TID 必须复用一个 binder_thread、为什么 poll 和 ioctl 共享 thread 但消费 owner 不同、为什么 is_dead 不等于已释放,以及 thread tmp_ref 如何牵连 proc 最终释放,就已经掌握了 Binder 驱动线程对象的生命周期。

图中的 Dead 只是禁止新 work 进入;内核仍可能通过 tmp_ref 暂时持有对象,直到所有事务和清理路径结束。