Skip to content

Binder死亡通知

从远程引用注册到节点死亡、BR_DEAD_BINDER 消费和清除确认,追踪 Binder 死亡通知的两次用户态内核态握手。

基于android-17.0.0_r1
AndroidBinder死亡通知源码阅读

Binder死亡通知 ​

linkToDeath() 的作用不是让驱动直接调用一个 C++ 回调。Native 层先把 BpBinder 和 DeathRecipient 保存在用户态,再通过 BC_REQUEST_DEATH_NOTIFICATION(handle, cookie) 把一个进程内的标识交给驱动。目标节点死亡后,驱动把对应的 binder_ref_death 作为 work 投递给持有引用的进程,用户态收到 BR_DEAD_BINDER(cookie),调用 BpBinder::sendObituary(),最后再写回 BC_DEAD_BINDER_DONE(cookie)。

本文面向已经读过 Binder节点管理、Binder远程引用 和 BR事务返回 的读者。前两篇解释 binder_node、binder_ref 和引用所属进程,后一篇解释 Binder work 如何变成 BR 命令;本文只追踪死亡通知这条特殊 work 的注册、投递、确认和释放,不重复事务 payload、对象翻译或 ServiceManager 的服务替换流程。

读完后,应该能从 BpBinder::linkToDeath() 定位到驱动的 ref->death,解释为什么目标进程退出后通知进入 proc->todo 而不是固定投递到注册线程,区分 BR_DEAD_BINDER 与 BR_CLEAR_DEATH_NOTIFICATION_DONE,并判断取消发生在通知出队前、出队后还是用户态确认之后。

1. 对象边界 ​

死亡通知至少有四个不同对象:

对象所有者作用生命周期终点
binder_node目标服务进程的 Binder 驱动上下文表示目标进程拥有的本地 Binder 实体进程退出后转入 dead node 或释放
binder_ref监听方进程用 handle 找到目标 node引用和附件都清理后释放
binder_ref_death监听方进程的驱动状态保存 work 和用户 cookie清除确认完成或引用回收后释放
BpBinder / Obituary监听方用户态保存 DeathRecipient 列表并发起 BC 命令proxy 销毁、死亡回调或取消后清理

handle 只在持有引用的进程内有意义;cookie 不是驱动用来查找 node 的地址,而是用户态收到 BR 命令后找回 BpBinder 的标识。在当前 libbinder 实现中,cookie 写入的是 BpBinder*,驱动只负责原样保存和返回。

源码文件:kernel/common/drivers/android/binder_internal.h

相关类型:binder_ref_death、binder_ref、binder_proc

c
struct binder_ref_death {
    struct binder_work work;
    binder_uintptr_t cookie;
};

struct binder_ref {
    struct binder_ref_data data;
    struct rb_node rb_node_desc;
    struct rb_node rb_node_node;
    struct hlist_node node_entry;
    struct binder_proc *proc;
    struct binder_node *node;
    struct binder_ref_death *death;
    struct binder_ref_freeze *freeze;
};

ref->death 由目标 node 锁保护,work 挂入哪个进程队列则由所属 proc 的 inner lock 保护。这个分工解释了后文为什么注册时要同时取得 proc、node 和 inner 相关锁,而用户态回调本身不在这些锁里执行。

源码文件:kernel/common/include/uapi/linux/android/binder.h

相关类型:binder_handle_cookie、binder_driver_return_protocol、binder_driver_command_protocol

c
struct binder_handle_cookie {
    __u32 handle;
    binder_uintptr_t cookie;
} __packed;

BR_DEAD_BINDER = _IOR('r', 15, binder_uintptr_t),
BR_CLEAR_DEATH_NOTIFICATION_DONE = _IOR('r', 16, binder_uintptr_t),

BC_REQUEST_DEATH_NOTIFICATION = _IOW(
        'c', 14, struct binder_handle_cookie),
BC_CLEAR_DEATH_NOTIFICATION = _IOW(
        'c', 15, struct binder_handle_cookie),
BC_DEAD_BINDER_DONE = _IOW(
        'c', 16, binder_uintptr_t),

2. 注册通知 ​

2.1 用户态入口 ​

源码文件:frameworks/native/libs/binder/BpBinder.cpp

相关函数:BpBinder::linkToDeath()

cpp
status_t BpBinder::linkToDeath(
        const sp<DeathRecipient>& recipient,
        void* cookie, uint32_t flags) {
    Obituary ob;
    ob.recipient = recipient;
    ob.cookie = cookie;
    ob.flags = flags;

    LOG_ALWAYS_FATAL_IF(
            recipient == nullptr,
            "linkToDeath(): recipient must be non-NULL");

    RpcMutexUniqueLock _l(mLock);
    if (!mObitsSent) {
        if (!mObituaries) {
            mObituaries = new Vector<Obituary>;
            getWeakRefs()->incWeak(this);
            IPCThreadState* self = IPCThreadState::self();
            self->requestDeathNotification(
                    binderHandle(), this);
            self->flushCommands();
        }
        ssize_t res = mObituaries->add(ob);
        return res >= (ssize_t)NO_ERROR
                ? (status_t)NO_ERROR : res;
    }
    return DEAD_OBJECT;
}

第一次注册时,BpBinder 才向驱动发送请求;同一个 proxy 上追加第二个 DeathRecipient 只增加 mObituaries 项,不会再次创建一个驱动 death attachment。getWeakRefs()->incWeak(this) 让用户态 proxy 在驱动尚未回送死亡命令时保持可定位,后续收到 BR_CLEAR_DEATH_NOTIFICATION_DONE 才对应减少这份弱引用。

linkToDeath() 对本地 BBinder 不走这条路径。目标对象由本进程拥有时,本进程与对象一起退出,不存在另一个进程可以收到该对象的死亡通知;BBinder::linkToDeath() 直接返回 INVALID_OPERATION。

源码文件:frameworks/native/libs/binder/Binder.cpp

相关函数:BBinder::linkToDeath()、BBinder::unlinkToDeath()

cpp
status_t BBinder::linkToDeath(
        const sp<DeathRecipient>& /*recipient*/,
        void* /*cookie*/, uint32_t /*flags*/) {
    return INVALID_OPERATION;
}

status_t BBinder::unlinkToDeath(
        const wp<DeathRecipient>& /*recipient*/,
        void* /*cookie*/, uint32_t /*flags*/,
        wp<DeathRecipient>* /*outRecipient*/) {
    return INVALID_OPERATION;
}

2.2 命令构造 ​

源码文件:frameworks/native/libs/binder/IPCThreadState.cpp

相关函数:IPCThreadState::requestDeathNotification()、clearDeathNotification()

cpp
status_t IPCThreadState::requestDeathNotification(
        int32_t handle, BpBinder* proxy) {
    mOut.writeInt32(BC_REQUEST_DEATH_NOTIFICATION);
    mOut.writeInt32(handle);
    mOut.writePointer((uintptr_t)proxy);
    return NO_ERROR;
}

status_t IPCThreadState::clearDeathNotification(
        int32_t handle, BpBinder* proxy) {
    mOut.writeInt32(BC_CLEAR_DEATH_NOTIFICATION);
    mOut.writeInt32(handle);
    mOut.writePointer((uintptr_t)proxy);
    return NO_ERROR;
}

这两个函数只把 BC 命令追加到 mOut;flushCommands() 才通过 BINDER_WRITE_READ 把字节交给驱动。因而 linkToDeath() 返回成功表示命令已写入并刷新成功,不表示目标进程仍然存活,也不表示未来回调已经执行。

2.3 引用查找 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_write() 中 BC_REQUEST_DEATH_NOTIFICATION 分支

c
if (get_user(target, (uint32_t __user *)ptr))
    return -EFAULT;
ptr += sizeof(uint32_t);
if (get_user(cookie,
        (binder_uintptr_t __user *)ptr))
    return -EFAULT;
ptr += sizeof(binder_uintptr_t);

death = kzalloc(sizeof(*death), GFP_KERNEL);
if (death == NULL) {
    thread->return_error.cmd = BR_ERROR;
    binder_enqueue_thread_work(
            thread, &thread->return_error.work);
    break;
}

binder_proc_lock(proc);
ref = binder_get_ref_olocked(proc, target, false);
if (ref == NULL) {
    binder_proc_unlock(proc);
    kfree(death);
    break;
}

驱动先在锁外分配 binder_ref_death,再在 proc 锁下查找 handle。这样分配失败不会持有引用树锁;无效 handle 也不会留下半初始化附件。此处没有创建新的 binder_ref,死亡通知只能挂在已经存在的远程引用上。

2.4 附件挂载 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_write() 中 ref->death 更新逻辑

c
binder_node_lock(ref->node);
if (ref->death) {
    binder_node_unlock(ref->node);
    binder_proc_unlock(proc);
    kfree(death);
    break;
}

binder_stats_created(BINDER_STAT_DEATH);
INIT_LIST_HEAD(&death->work.entry);
death->cookie = cookie;
ref->death = death;

if (ref->node->proc == NULL) {
    death->work.type = BINDER_WORK_DEAD_BINDER;
    binder_inner_proc_lock(proc);
    binder_enqueue_work_ilocked(
            &death->work, &proc->todo);
    binder_wakeup_proc_ilocked(proc);
    binder_inner_proc_unlock(proc);
}
binder_node_unlock(ref->node);
binder_proc_unlock(proc);

ref->node->proc == NULL 是“节点所属进程已经释放”的状态。注册发生在节点死亡之后时,驱动不会等待另一个退出事件,而是把 BINDER_WORK_DEAD_BINDER 立即加入监听方的 proc->todo。这条路径是 late registration 的关键,否则已经死亡的服务会让新注册者永远收不到通知。

3. 节点死亡 ​

3.1 释放入口 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_deferred_release()、binder_node_release()

c
while ((n = rb_first(&proc->nodes))) {
    struct binder_node *node =
            rb_entry(n, struct binder_node, rb_node);

    binder_inc_node_tmpref_ilocked(node);
    rb_erase(&node->rb_node, &proc->nodes);
    binder_inner_proc_unlock(proc);
    incoming_refs = binder_node_release(
            node, incoming_refs);
    binder_inner_proc_lock(proc);
}

进程文件释放并不在 binder_release() 中同步遍历所有 node;它设置 deferred release,之后 binder_deferred_release() 标记 proc 死亡、释放线程、逐个调用 binder_node_release()。因此“服务进程退出”和“所有客户端已经收到回调”不是同一时刻。

3.2 节点脱离 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_node_release()

c
binder_node_lock(node);
binder_inner_proc_lock(proc);
binder_dequeue_work_ilocked(&node->work);

if (hlist_empty(&node->refs) && node->tmp_refs == 1) {
    binder_inner_proc_unlock(proc);
    binder_node_unlock(node);
    binder_free_node(node);
    return refs;
}

node->proc = NULL;
node->local_strong_refs = 0;
node->local_weak_refs = 0;
binder_inner_proc_unlock(proc);

spin_lock(&binder_dead_nodes_lock);
hlist_add_head(&node->dead_node,
        &binder_dead_nodes);
spin_unlock(&binder_dead_nodes_lock);

如果 node 仍被远端 binder_ref 持有,驱动不能直接释放它:先清空 node->proc,再把 node 放进 dead-node 链表。远端引用仍需要通过这个 node 找到自己的 death 附件,并在之后完成引用和附件清理。

3.3 通知入队 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_node_release()

c
hlist_for_each_entry(ref, &node->refs, node_entry) {
    binder_inner_proc_lock(ref->proc);
    if (!ref->death) {
        binder_inner_proc_unlock(ref->proc);
        continue;
    }

    BUG_ON(!list_empty(&ref->death->work.entry));
    ref->death->work.type = BINDER_WORK_DEAD_BINDER;
    binder_enqueue_work_ilocked(
            &ref->death->work, &ref->proc->todo);
    binder_wakeup_proc_ilocked(ref->proc);
    binder_inner_proc_unlock(ref->proc);
}

通知挂到 ref->proc->todo,不是注册 linkToDeath() 的那条用户线程的 thread->todo。这样任意已注册 looper 的 Binder 线程都可以消费它;这也是用户态回调不能假定运行在调用 linkToDeath() 的线程上的原因。

4. BR消费 ​

4.1 work选择 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_read() 中死亡 work 分支

c
case BINDER_WORK_DEAD_BINDER:
case BINDER_WORK_DEAD_BINDER_AND_CLEAR:
case BINDER_WORK_CLEAR_DEATH_NOTIFICATION: {
    struct binder_ref_death *death;
    uint32_t cmd;
    binder_uintptr_t cookie;

    death = container_of(w,
            struct binder_ref_death, work);
    if (w->type == BINDER_WORK_CLEAR_DEATH_NOTIFICATION)
        cmd = BR_CLEAR_DEATH_NOTIFICATION_DONE;
    else
        cmd = BR_DEAD_BINDER;
    cookie = death->cookie;

驱动把两个“死亡相关” work 类型都编码成 BR_DEAD_BINDER:BINDER_WORK_DEAD_BINDER 表示正常死亡通知,BINDER_WORK_DEAD_BINDER_AND_CLEAR 表示死亡通知已经排队时又收到取消。只有 BINDER_WORK_CLEAR_DEATH_NOTIFICATION 才编码成清除确认。

4.2 交付保管 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_read() 中 proc->delivered_death 处理

c
cookie = death->cookie;
if (w->type == BINDER_WORK_CLEAR_DEATH_NOTIFICATION) {
    binder_inner_proc_unlock(proc);
    kfree(death);
    binder_stats_deleted(BINDER_STAT_DEATH);
} else {
    binder_enqueue_work_ilocked(
            w, &proc->delivered_death);
    binder_inner_proc_unlock(proc);
}

if (put_user(cmd, (uint32_t __user *)ptr))
    return -EFAULT;
ptr += sizeof(uint32_t);
if (put_user(cookie,
        (binder_uintptr_t __user *)ptr))
    return -EFAULT;
ptr += sizeof(binder_uintptr_t);

if (cmd == BR_DEAD_BINDER)
    goto done;

BR_DEAD_BINDER 写入用户 read buffer 后,work 被移到 proc->delivered_death,而不是立即释放。驱动必须保留它,直到用户态通过 cookie 回写 BC_DEAD_BINDER_DONE。相反,清除确认没有用户态二次确认,驱动在写 BR 前就释放 binder_ref_death。

goto done 还会结束本轮 read work 选择。死亡回调可能同步调用 unlinkToDeath() 或发起新的 Binder 操作,驱动不把同一轮 read 继续扩展成不可预测的通知批次。

5. 用户态回调 ​

5.1 命令分发 ​

源码文件:frameworks/native/libs/binder/IPCThreadState.cpp

相关函数:IPCThreadState::executeCommand()

cpp
case BR_DEAD_BINDER: {
    BpBinder* proxy =
            (BpBinder*)mIn.readPointer();
    proxy->sendObituary();
    mOut.writeInt32(BC_DEAD_BINDER_DONE);
    mOut.writePointer((uintptr_t)proxy);
    break;
}

case BR_CLEAR_DEATH_NOTIFICATION_DONE: {
    BpBinder* proxy =
            (BpBinder*)mIn.readPointer();
    proxy->getWeakRefs()->decWeak(proxy);
    break;
}

executeCommand() 先调用 sendObituary(),再写 BC_DEAD_BINDER_DONE。回调执行顺序因此是:驱动交付 cookie → 用户态设置 proxy 已死亡并通知 recipients → 用户态确认 work。BC_DEAD_BINDER_DONE 不是给每个 DeathRecipient 发一次,而是针对一次驱动 death attachment 发一次。

5.2 回调列表 ​

源码文件:frameworks/native/libs/binder/BpBinder.cpp

相关函数:BpBinder::sendObituary()、BpBinder::reportOneDeath()

cpp
void BpBinder::sendObituary() {
    mAlive = 0;
    if (mObitsSent) return;

    mLock.lock();
    Vector<Obituary>* obits = mObituaries;
    if (obits != nullptr) {
        IPCThreadState* self = IPCThreadState::self();
        self->clearDeathNotification(
                binderHandle(), this);
        self->flushCommands();
        mObituaries = nullptr;
    }
    mObitsSent = 1;
    mLock.unlock();

    if (obits != nullptr) {
        for (size_t i = 0; i < obits->size(); i++)
            reportOneDeath(obits->itemAt(i));
        delete obits;
    }
}

void BpBinder::reportOneDeath(const Obituary& obit) {
    sp<DeathRecipient> recipient = obit.recipient.promote();
    if (recipient == nullptr) return;
    recipient->binderDied(
            wp<BpBinder>::fromExisting(this));
}

sendObituary() 把 mObituaries 从 proxy 上摘下并设置 mObitsSent,再在锁外逐个调用 binderDied()。回调拿到的是 wp<BpBinder>;目标已经死亡,不能把它当作可继续使用的强引用。驱动侧的 BC_CLEAR_DEATH_NOTIFICATION 会和 BC_DEAD_BINDER_DONE 一起进入后续写入批次,因此用户回调完成不等于驱动附件立即释放。

5.3 多个监听者 ​

同一 BpBinder 可以有多个 Obituary,但驱动只保存一个 binder_ref_death 和一个 cookie。驱动通知一次,用户态在 mObituaries 中遍历多个 recipient。反过来,最后一个 recipient 被 unlinkToDeath() 移除时,libbinder 才写 BC_CLEAR_DEATH_NOTIFICATION;这不是每个 listener 一次清理。

6. 确认清理 ​

6.1 查找交付项 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_write() 中 BC_DEAD_BINDER_DONE 分支

c
if (get_user(cookie,
        (binder_uintptr_t __user *)ptr))
    return -EFAULT;
ptr += sizeof(cookie);

binder_inner_proc_lock(proc);
list_for_each_entry(w, &proc->delivered_death, entry) {
    struct binder_ref_death *tmp_death =
            container_of(w,
                    struct binder_ref_death, work);
    if (tmp_death->cookie == cookie) {
        death = tmp_death;
        break;
    }
}
if (death == NULL) {
    binder_inner_proc_unlock(proc);
    break;
}
binder_dequeue_work_ilocked(&death->work);

驱动按 cookie 在线性 delivered_death 链表中查找,而不是按 handle 查找。一个错误 cookie 不会释放任意 death;驱动记录用户错误并结束这个 BC 命令。相同 cookie 的多个未确认通知在正常模型中不会由同一 attachment 产生,但调用方不能伪造确认顺序来替代真实交付。

6.2 普通完成 ​

当 death->work.type == BINDER_WORK_DEAD_BINDER 时,binder_dequeue_work_ilocked() 只是把 work 从 delivered_death 移除,当前 death 对象仍由 ref->death 指向,直到用户态稍后清除或引用释放。BC_DEAD_BINDER_DONE 的职责是确认“这次 BR 已经被消费”,不是单独取消监听。

6.3 清除完成 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_read() 中 BINDER_WORK_CLEAR_DEATH_NOTIFICATION

c
if (w->type == BINDER_WORK_CLEAR_DEATH_NOTIFICATION) {
    binder_inner_proc_unlock(proc);
    kfree(death);
    binder_stats_deleted(BINDER_STAT_DEATH);
}

cmd = BR_CLEAR_DEATH_NOTIFICATION_DONE;
put_user(cmd, (uint32_t __user *)ptr);
ptr += sizeof(uint32_t);
put_user(cookie,
        (binder_uintptr_t __user *)ptr);

清除完成命令携带同一个 cookie,但用户态处理方式不同:executeCommand() 找回 BpBinder 后只减少 getWeakRefs(),不会调用 binderDied()。因此 BR_CLEAR_DEATH_NOTIFICATION_DONE 表示 driver-side attachment 的释放确认,而不是目标服务死亡。

7. 并发取消 ​

7.1 尚未出队 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_thread_write() 中 BC_CLEAR_DEATH_NOTIFICATION

c
death = ref->death;
if (death->cookie != cookie) {
    binder_node_unlock(ref->node);
    binder_proc_unlock(proc);
    break;
}

ref->death = NULL;
binder_inner_proc_lock(proc);
if (list_empty(&death->work.entry)) {
    death->work.type = BINDER_WORK_CLEAR_DEATH_NOTIFICATION;
    if (thread->looper &
            (BINDER_LOOPER_STATE_REGISTERED |
             BINDER_LOOPER_STATE_ENTERED))
        binder_enqueue_thread_work_ilocked(
                thread, &death->work);
    else {
        binder_enqueue_work_ilocked(
                &death->work, &proc->todo);
        binder_wakeup_proc_ilocked(proc);
    }
} else {
    death->work.type =
            BINDER_WORK_DEAD_BINDER_AND_CLEAR;
}
binder_inner_proc_unlock(proc);

当 work 尚未挂在任何队列时,取消会新建一个 CLEAR_DEATH_NOTIFICATION work。已注册 looper 的当前线程可以收到 thread todo;非 looper 线程则走 proc todo,交给可用的 Binder 线程。这里的“当前线程”只影响清除确认 work 的投递,不改变正常死亡通知的 proc todo 语义。

7.2 已经排队 ​

如果 death->work.entry 非空,说明死亡 work 已在 proc->todo、某个 thread todo 或 delivered 链表中。驱动不能把同一 list entry 再插入清除队列,于是只把类型改成 BINDER_WORK_DEAD_BINDER_AND_CLEAR。

当 binder_thread_read() 取到它时仍返回 BR_DEAD_BINDER;用户态照常执行 sendObituary() 和 BC_DEAD_BINDER_DONE。确认分支随后把类型改为 BINDER_WORK_CLEAR_DEATH_NOTIFICATION,再次排队,最终才返回清除完成命令。

7.3 已经交付 ​

如果取消发生在 BR_DEAD_BINDER 已交付、但 BC_DEAD_BINDER_DONE 尚未写回之后,work 位于 proc->delivered_death,同样进入 BINDER_WORK_DEAD_BINDER_AND_CLEAR。这保证用户态先完成死亡回调确认,驱动再释放 attachment;否则 sendObituary() 仍可能使用的 proxy 弱引用会失去对应的 driver 生命周期保护。

8. Proxy清理 ​

8.1 最后强引用 ​

源码文件:frameworks/native/libs/binder/BpBinder.cpp

相关函数:BpBinder::onLastStrongRef()

cpp
mLock.lock();
Vector<Obituary>* obits = mObituaries;
if (obits != nullptr) {
    IPCThreadState* ipc = IPCThreadState::self();
    if (ipc)
        ipc->clearDeathNotification(
                binderHandle(), this);
    mObituaries = nullptr;
}
mLock.unlock();

if (obits != nullptr)
    delete obits;

即使调用方没有显式 unlinkToDeath(),proxy 最后一个强引用释放时也会发清除命令。IBinder 接口文档把“必须持有 binder 才能继续收到死亡通知”写成契约;因此仅保留 DeathRecipient 而释放 BpBinder,不会留下可用的监听。

8.2 引用清理 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_cleanup_ref_olocked()、binder_free_ref()

c
if (ref->death) {
    binder_dequeue_work(ref->proc,
            &ref->death->work);
    binder_stats_deleted(BINDER_STAT_DEATH);
}

static void binder_free_ref(struct binder_ref *ref)
{
    if (ref->node)
        binder_free_node(ref->node);
    kfree(ref->death);
    kfree(ref->freeze);
    kfree(ref);
}

引用清理会先把 death work 从所属队列摘除,再由 binder_free_ref() 释放附件。若 work 已经进入 delivered_death,proc deferred release 也会释放这条链表;驱动不会因为用户态没有再确认而永久保留一个已经死亡的 proc。

8.3 进程回收 ​

源码文件:kernel/common/drivers/android/binder.c

相关函数:binder_deferred_release()

c
binder_release_work(proc, &proc->todo);
binder_release_work(proc, &special_list);
binder_release_work(proc, &proc->delivered_death);
binder_release_work(proc, &proc->delivered_freeze);

proc->delivered_death 是“等待用户确认”的正常状态,但它不是跨进程永久存储。监听进程自身退出时,deferred release 直接清理未确认的 death work;目标服务进程退出时,驱动则保留监听方的 work,直到监听方消费或自己退出。

9. 状态关系 ​

图中 proc->todo 和 proc->delivered_death 都属于监听方 proc;它们不是目标服务 proc 的队列。mObituaries 只属于用户态 proxy,驱动不会读取其中的 DeathRecipient,也不会理解 C++ 对象布局。

10. 测试边界 ​

10.1 驱动命令 ​

源码文件:frameworks/native/libs/binder/tests/binderDriverInterfaceTest.cpp

相关测试:BinderDriverInterfaceTest.RequestDeathNotification

cpp
binder_uintptr_t cookie = 1234;
bc.cmd1 = BC_REQUEST_DEATH_NOTIFICATION;
bc.arg1.handle = 0;
bc.arg1.cookie = cookie;
bc.cmd2 = BC_CLEAR_DEATH_NOTIFICATION;
bc.arg2.handle = 0;
bc.arg2.cookie = cookie;

binderTestIoctl(BINDER_WRITE_READ, &bwr);
EXPECT_EQ(BR_NOOP, br.cmd0);
EXPECT_EQ(BR_CLEAR_DEATH_NOTIFICATION_DONE,
        br.cmd1);
EXPECT_EQ(cookie, br.arg1);

这个测试把注册和取消放在一次 write 中,关键断言是清除确认携带原 cookie。它覆盖“没有发生目标进程死亡时,clear 如何产生 BR_CLEAR_DEATH_NOTIFICATION_DONE”,不覆盖 binder_node_release() 遍历远程引用的真正死亡路径。

10.2 目标进程死亡 ​

源码文件:frameworks/native/libs/binder/tests/binderLibTest.cpp

相关测试:BinderLibTest.DeathNotificationStrongRef、DeathNotificationMultiple

cpp
sp<TestDeathRecipient> recipient =
        new TestDeathRecipient();
sp<IBinder> binder = addServer();
EXPECT_THAT(binder->linkToDeath(recipient),
        StatusEq(NO_ERROR));

EXPECT_THAT(binder->transact(
        BINDER_LIB_TEST_EXIT_TRANSACTION,
        data, &reply, TF_ONE_WAY), StatusEq(OK));
IPCThreadState::self()->flushCommands();
EXPECT_THAT(recipient->waitEvent(5),
        StatusEq(NO_ERROR));

测试输入是先建立远程 server proxy,再注册 recipient,最后让 server 退出;关键断言是 recipient 收到事件。DeathNotificationMultiple 进一步让多个 client 对同一 target 注册,并要求每个 callback 收到事件,说明一个驱动 attachment 可以由用户态 fan-out 到多个 recipient。测试没有断言 callback 的具体线程 ID,也没有覆盖 clear 与 BR 交错的每一种调度顺序。

10.3 线程消费者 ​

源码文件:frameworks/native/libs/binder/tests/binderLibTest.cpp

相关测试:BinderLibTest.DeathNotificationThread

该测试先在一个 client 上注册死亡通知并等待目标退出,再把已经死亡的引用传给另一个进程,由另一个进程调用 linkToDeath()。测试注释明确要求两件事:已死亡引用的 late registration 仍能收到通知;通知进入 proc work queue,不能假设它一定推送给执行注册操作的线程。第二个断言对应驱动在 binder_node_release() 和 late registration 中都使用 proc->todo 的实现。

11. 排查路径 ​

遇到“服务已退出但没有收到 binderDied”时,可以沿这条源码路径检查:

  1. 监听方是否仍持有 BpBinder 强引用,mObituaries 是否仍有项;
  2. linkToDeath() 是否真正写入 BC_REQUEST_DEATH_NOTIFICATION,handle 是否仍存在于监听方 refs_by_desc;
  3. 目标 node 是否已经进入 node->proc == NULL 状态;
  4. 监听方 proc->todo 是否出现 BINDER_WORK_DEAD_BINDER,以及 read 线程是否在消费 proc work;
  5. 用户态是否读到 BR_DEAD_BINDER,cookie 是否仍等于该 BpBinder*;
  6. sendObituary() 是否因 mObitsSent 已置位而跳过重复回调;
  7. BC_DEAD_BINDER_DONE 是否携带同一个 cookie,或是否在 clear 交错时等待 BR_CLEAR_DEATH_NOTIFICATION_DONE。

源码搜索:

bash
rg -n "BC_REQUEST_DEATH_NOTIFICATION|BC_CLEAR_DEATH_NOTIFICATION|BC_DEAD_BINDER_DONE" \
  kernel/common/drivers/android/binder.c \
  frameworks/native/libs/binder/IPCThreadState.cpp

rg -n "binder_node_release|BINDER_WORK_DEAD_BINDER|delivered_death" \
  kernel/common/drivers/android/binder.c \
  kernel/common/drivers/android/binder_internal.h

rg -n "linkToDeath|unlinkToDeath|sendObituary|reportOneDeath" \
  frameworks/native/libs/binder/BpBinder.cpp \
  frameworks/native/libs/binder/Binder.cpp

rg -n "DeathNotification|RequestDeathNotification|BR_DEAD_BINDER" \
  frameworks/native/libs/binder/tests/binderDriverInterfaceTest.cpp \
  frameworks/native/libs/binder/tests/binderLibTest.cpp

如果能解释“为什么一个 BpBinder 只有一个驱动 death attachment 却能通知多个 recipient”“为什么 BR_DEAD_BINDER 之后还必须发送 BC_DEAD_BINDER_DONE”“为什么 clear 可能先产生 BR_DEAD_BINDER 再产生清除确认”,就已经能顺着这条协议定位大多数死亡通知时序问题。文章没有把 Binder death notification 与 freeze notification、事务 BR_DEAD_REPLY 或 ServiceManager 的服务重注册混为一谈;这些机制拥有不同的 work 类型、消费者和结束条件。

死亡通知的消费者是客户端 Binder 线程和 recipient,不是原始事务等待者;BR_DEAD_REPLY 只表示一次调用目标死亡, 不能替代完整 death attachment 协议。