Skip to content

SIGCHLD进程死亡处理

追踪 Zygote 的 SIGCHLD、waitpid、USAP 回收和 unsolicited socket,并区分退出信息记录与 AMS 进程清理。

基于android-17.0.0_r1
AndroidZygoteAMSSIGCHLD源码阅读

SIGCHLD进程死亡处理 ​

一个应用已经从 /proc 消失,为什么退出记录还没有更新?反过来,日志已经出现 Zygote 的 “exited due to signal”,为什么 AMS 中还有与它关联的状态?这两个现象要求我们区分三件事:Linux 回收子进程、收集退出原因、清理 Framework 对象。它们有不同入口,不存在一条把所有动作同步做完的 SIGCHLD 回调。

本文面向理解 fork()、PID 和 Handler 的读者。可以先读子进程PID管理,理解 PID 如何关联到 ProcessRecord。这里从父进程收到 SIGCHLD 开始,追踪 waitpid 的状态字如何进入 ApplicationExitInfo,并解释 USAP 表项回收、消息丢失、乱序到达和 system_server 死亡。AMS 内各类服务、Provider 和窗口的完整清理不在本篇展开。

1. 三种死亡状态 ​

Zygote 是普通应用和 system_server 的父进程;通过 App Zygote 等路径创建的进程则由相应父进程负责回收。退出的进程在父进程取走退出状态以前可能保持 zombie 状态。waitpid() 消费的是内核保存的子进程退出信息;ProcessRecord 是 system_server 中描述应用的 Java 对象,两者不存在共享内存里的“自动同步删除”。

状态所有者和入口完成后能说明什么
子进程已回收父 Zygote,SigChldHandler → waitpid父进程已取得该 child 的退出状态
退出原因已合并AppExitInfoTracker,Zygote/lmkd/AMS 多个输入对外记录可提供 reason、status 等信息
Framework 已清理AMS,appDiedLocked → handleAppDiedLocked连接、LRU、窗口等按各自策略处理

下图只描述相关输入如何汇合,不表示 Binder death 与 SIGCHLD 有固定先后。

Zygote 的发送是尽力通知;AMS 的清理也不以 datagram 返回作为提交屏障。阅读后续代码时,应分别标出“内核状态已消费”和“Java 记录已更新”的时点。

2. 回收循环 ​

源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp

相关符号:SetSignalHandlers / UnsetChldSignalHandler。以下为源码节选。

cpp
// 父进程安装处理器,子进程恢复默认处理,避免沿用父进程的回收职责。
static void SetSignalHandlers() {
    struct sigaction sig_chld = {.sa_flags = SA_SIGINFO, .sa_sigaction = SigChldHandler};

    if (sigaction(SIGCHLD, &sig_chld, nullptr) < 0) {
        ALOGW("Error setting SIGCHLD handler: %s", strerror(errno));
    }

  struct sigaction sig_hup = {};
  sig_hup.sa_handler = SIG_IGN;
  if (sigaction(SIGHUP, &sig_hup, nullptr) < 0) {
    ALOGW("Error setting SIGHUP handler: %s", strerror(errno));
  }
}

// Sets the SIGCHLD handler back to default behavior in zygote children.
static void UnsetChldSignalHandler() {
  struct sigaction sa;
  memset(&sa, 0, sizeof(sa));
  sa.sa_handler = SIG_DFL;

  if (sigaction(SIGCHLD, &sa, nullptr) < 0) {
    ALOGW("Error unsetting SIGCHLD handler: %s", strerror(errno));
  }
}
// ...

sigaction 使用 SA_SIGINFO,让处理器取得 siginfo_t。fork 周围还会暂时屏蔽 SIGCHLD;ForkCommon() 在父子分支整理完成后解除屏蔽。安装、屏蔽和解除屏蔽需要一起阅读,不能认为 handler 可以在任意初始化阶段安全执行。

源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp

相关符号:SigChldHandler。以下为源码节选。

cpp
// 一次信号处理循环回收多个 child;status 来自每次 waitpid,而 uid 取自本次信号的 info。
// This signal handler is for zygote mode, since the zygote must reap its children
NO_STACK_PROTECTOR
static void SigChldHandler(int /*signal_number*/, siginfo_t* info, void* /*ucontext*/) {
    pid_t pid;
    int status;
    int64_t usaps_removed = 0;

    // It's necessary to save and restore the errno during this function.
    // Since errno is stored per thread, changing it here modifies the errno
    // on the thread on which this signal handler executes. If a signal occurs
    // between a call and an errno check, it's possible to get the errno set
    // here.
    // See b/23572286 for extra information.
    int saved_errno = errno;

    while ((pid = waitpid(-1, &status, WNOHANG)) > 0) {
        // Notify system_server that we received a SIGCHLD
        sendSigChildStatus(pid, info->si_uid, status);
        // Log process-death status that we care about.
        if (WIFEXITED(status)) {
            async_safe_format_log(ANDROID_LOG_INFO, LOG_TAG, "Process %d exited cleanly (%d)", pid,
                                  WEXITSTATUS(status));

            // Check to see if the PID is in the USAP pool and remove it if it is.
            if (RemoveUsapTableEntry(pid)) {
                ++usaps_removed;
            }
        } else if (WIFSIGNALED(status)) {
            async_safe_format_log(ANDROID_LOG_INFO, LOG_TAG,
                                  "Process %d exited due to signal %d (%s)%s", pid,
                                  WTERMSIG(status), strsignal(WTERMSIG(status)),
                                  WCOREDUMP(status) ? "; core dumped" : "");

            // If the process exited due to a signal other than SIGTERM, check to see
            // if the PID is in the USAP pool and remove it if it is.  If the process
            // was closed by the Zygote using SIGTERM then the USAP pool entry will
            // have already been removed (see nativeEmptyUsapPool()).
            if (WTERMSIG(status) != SIGTERM && RemoveUsapTableEntry(pid)) {
                ++usaps_removed;
            }
        }

        // If the just-crashed process is the system_server, bring down zygote
        // so that it is restarted by init and system server will be restarted
        // from there.
        if (pid == gSystemServerPid) {
            async_safe_format_log(ANDROID_LOG_ERROR, LOG_TAG,
                                  "Exit zygote because system server (pid %d) has terminated", pid);
            kill(getpid(), SIGKILL);
        }
    }

    // Note that we shouldn't consider ECHILD an error because
    // the secondary zygote might have no children left to wait for.
    if (pid < 0 && errno != ECHILD) {
        async_safe_format_log(ANDROID_LOG_WARN, LOG_TAG, "Zygote SIGCHLD error in waitpid: %s",
                              strerror(errno));
    }

    if (usaps_removed > 0) {
        if (TEMP_FAILURE_RETRY(write(gUsapPoolEventFD, &usaps_removed, sizeof(usaps_removed))) ==
            -1) {
            // If this write fails something went terribly wrong.  We will now kill
            // the zygote and let the system bring it back up.
            async_safe_format_log(ANDROID_LOG_ERROR, LOG_TAG,
                                  "Zygote failed to write to USAP pool event FD: %s",
                                  strerror(errno));
            kill(getpid(), SIGKILL);
        }
    }

    errno = saved_errno;
}
// ...

WNOHANG 使父进程不会为了尚未退出的 child 阻塞;返回正数才表示回收了一项。普通信号可能合并,循环因此不能改成单次 waitpid。pid == 0 表示当时没有可回收的退出状态;负数时 ECHILD 被允许,其他错误才打印警告。

status 是编码后的 wait status,不是直接的退出码或信号编号。只有 WIFEXITED(status) 成立,WEXITSTATUS(status) 才表示正常退出码;信号终止需要用 WIFSIGNALED 和 WTERMSIG。例如正常 exit(7) 与因信号 7 终止,不能都写成“退出码 7”。

还有一个很容易在复述中消失的边界:源码把每次 waitpid 返回的 PID 与同一份 info->si_uid 配对。它没有逐个查询刚回收 PID 的 UID。一次 handler 收割多个退出进程时,不能单凭这段实现声称“每条 UID 都由对应 waitpid 结果提供”。这属于通知来源的边界,不能在示意代码中擅自补出一次 UID 查询。

errno 是线程状态,handler 会打断该线程原来的代码。保存和恢复它,是为了避免原代码刚执行系统调用、尚未检查错误时,被 handler 的 waitpid 或 write 改写判断依据。NO_STACK_PROTECTOR 是函数编译属性,不能把它当成“所有被调用函数都满足异步信号安全”的证明;这里也不应引入 Java 回调或 Framework 锁。

3. 池成员回收 ​

USAP 是尚未特化成具体应用的预创建进程。回收它除了消费内核状态,还要关闭池表项关联的读管道,并且只扣减一次池计数。

源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp

相关符号:UsapTableEntry::ClearForPID / RemoveUsapTableEntry。以下为源码节选。

cpp
// 只有成功把表项从当前值交换成无效值的调用,才能关闭 fd 并触发池计数扣减。
bool ClearForPID(int32_t pid) {
  EntryStorage storage = mStorage.load();

  if (storage.pid == pid) {
    /*
     * There are three possible outcomes from this compare-and-exchange:
     *   1) It succeeds, in which case we close the FD
     *   2) It fails and the new value is INVALID_ENTRY_VALUE, in which case
     *      the entry has already been cleared.
     *   3) It fails and the new value isn't INVALID_ENTRY_VALUE, in which
     *      case the entry has already been cleared and re-used.
     *
     * In all three cases the goal of the caller has been met, but only in
     * the first case do we need to decrement the pool count.
     */
    if (mStorage.compare_exchange_strong(storage, INVALID_ENTRY_VALUE)) {
      close(storage.read_pipe_fd);
      return true;
    } else {
      return false;
    }

  } else {
    return false;
  }
}
// ...
static bool RemoveUsapTableEntry(pid_t usap_pid) {
  for (UsapTableEntry& entry : gUsapTable) {
    if (entry.ClearForPID(usap_pid)) {
      --gUsapPoolCount;
      return true;
    }
  }

  return false;
}
// ...

这里的原子比较交换保护的是“观察到的表项是否仍然有效”。另一条路径可能已经清除该项,或者清除后将其复用;CAS 失败时返回 false,调用方不再扣减。仅有 PID 比较并不能概括这段并发约束,更不能据此宣称消除了系统中所有 PID 复用问题。

SigChldHandler() 只累计本次确实移除的数量,循环结束后写一次 gUsapPoolEventFD。eventfd 通知把池成员变化带回事件循环;补池不是在信号处理器中直接调用 Java。写失败会触发 kill(getpid(), SIGKILL),因为 native 池状态变化已经发生,继续运行会让外部消费者错过变化。

对 SIGTERM 的特殊分支也不能删掉:Zygote 主动清空池的 nativeEmptyUsapPool() 会先移除表项再发送 SIGTERM,handler 对这个终止信号不再执行常规扣减。不能把所有 WIFSIGNALED 都等价理解为“池计数减一”。

4. 数据报契约 ​

源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp

相关符号:UnsolicitedZygoteMessageSigChld / initUnsolSocketToSystemServer / sendSigChildStatus。以下为源码节选。

cpp
// 通知使用非阻塞 Unix datagram,发送失败不会撤销前面的 waitpid。
struct UnsolicitedZygoteMessageSigChld {
    struct {
        UnsolicitedZygoteMessageTypes type;
    } header;
    struct {
        pid_t pid;
        uid_t uid;
        int status;
    } payload;
};

// Keep sync with services/core/java/com/android/server/am/ProcessList.java
static constexpr struct sockaddr_un kSystemServerSockAddr =
        {.sun_family = AF_LOCAL, .sun_path = "/data/system/unsolzygotesocket"};
// ...
static void initUnsolSocketToSystemServer() {
    gSystemServerSocketFd = socket(AF_LOCAL, SOCK_DGRAM | SOCK_NONBLOCK, 0);
    if (gSystemServerSocketFd >= 0) {
        ALOGV("Zygote:systemServerSocketFD = %d", gSystemServerSocketFd);
    } else {
        ALOGE("Unable to create socket file descriptor to connect to system_server");
    }
}

static void sendSigChildStatus(const pid_t pid, const uid_t uid, const int status) {
    int socketFd = gSystemServerSocketFd;
    if (socketFd >= 0) {
        // fill the message buffer
        struct UnsolicitedZygoteMessageSigChld data =
                {.header = {.type = UNSOLICITED_ZYGOTE_MESSAGE_TYPE_SIGCHLD},
                 .payload = {.pid = pid, .uid = uid, .status = status}};
        if (TEMP_FAILURE_RETRY(
                    sendto(socketFd, &data, sizeof(data), 0,
                           reinterpret_cast<const struct sockaddr*>(&kSystemServerSockAddr),
                           sizeof(kSystemServerSockAddr))) == -1) {
            async_safe_format_log(ANDROID_LOG_ERROR, LOG_TAG,
                                  "Zygote failed to write to system_server FD: %s",
                                  strerror(errno));
        }
    }
}
// ...

数据报由消息类型和三个整数构成。它不是Zygote参数传递协议中的文本启动请求,也没有请求编号、确认响应或失败重发队列。SOCK_NONBLOCK 避免因接收方跟不上而长期占住 signal handler,但也意味着退出原因通知可能丢失。

gSystemServerSocketFd < 0 时根本不发送;sendto 返回错误时只记录日志。无论哪一种,子进程已经被回收,不会再次通过 waitpid 自动补发。排查“应用已退出但缺少精确状态”时,这条失败路径比猜测 AMS 没收到 Binder death 更直接。

5. 接收与校验 ​

源码文件:frameworks/base/services/core/java/com/android/server/am/ProcessList.java

相关符号:createSystemServerSocketForZygote / handleZygoteMessages。以下为源码节选。

java
// 创建失败关闭已创建 socket;只有解析器返回三项才把数据交给 Tracker。
private LocalSocket createSystemServerSocketForZygote() {
    // The file system entity for this socket is created with 0666 perms, owned
    // by system:system. selinux restricts things so that only zygotes can
    // access it.
    final File socketFile = new File(UNSOL_ZYGOTE_MSG_SOCKET_PATH);
    if (socketFile.exists()) {
        socketFile.delete();
    }

    LocalSocket serverSocket = null;
    try {
        serverSocket = new LocalSocket(LocalSocket.SOCKET_DGRAM);
        serverSocket.bind(new LocalSocketAddress(
                UNSOL_ZYGOTE_MSG_SOCKET_PATH, LocalSocketAddress.Namespace.FILESYSTEM));
        Os.chmod(UNSOL_ZYGOTE_MSG_SOCKET_PATH, 0666);
    } catch (Exception e) {
        if (serverSocket != null) {
            try {
                serverSocket.close();
            } catch (IOException ex) {
            }
            serverSocket = null;
        }
    }
    return serverSocket;
}

/**
 * Handle the unsolicited message from zygote.
 */
private int handleZygoteMessages(FileDescriptor fd, int events) {
    final int eventFd = fd.getInt$();
    if ((events & EVENT_INPUT) != 0) {
        // An incoming message from zygote
        try {
            final int len = Os.read(fd, mZygoteUnsolicitedMessage, 0,
                    mZygoteUnsolicitedMessage.length);
            if (len > 0 && mZygoteSigChldMessage.length == Zygote.nativeParseSigChld(
                    mZygoteUnsolicitedMessage, len, mZygoteSigChldMessage)) {
                mAppExitInfoTracker.handleZygoteSigChld(
                        mZygoteSigChldMessage[0] /* pid */,
                        mZygoteSigChldMessage[1] /* uid */,
                        mZygoteSigChldMessage[2] /* status */);
            }
        } catch (Exception e) {
            Slog.w(TAG, "Exception in reading unsolicited zygote message: " + e);
        }
    }
    return EVENT_INPUT;
}
// ...

创建成功后,ProcessList 在 sKillHandler 的 MessageQueue 上注册 EVENT_INPUT fd listener;返回 EVENT_INPUT 表示继续监听。创建失败返回 null,注册步骤也就被跳过。这是与“发送端正常,但 system_server 未监听”对应的真实失败分支。

socket 节点的 0666 权限不是来源认证。这里还依赖 SELinux 访问控制;native parser 的职责则是验证消息结构,不能用长度检查代替调用者身份校验。

源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp

相关符号:com_android_internal_os_Zygote_nativeParseSigChld。以下为源码节选。

cpp
// 未知类型最终返回 -1;正确类型也必须经过输出数组长度检查。
static jint com_android_internal_os_Zygote_nativeParseSigChld(JNIEnv* env, jclass, jbyteArray in,
                                                              jint length, jintArray out) {
    if (length != sizeof(struct UnsolicitedZygoteMessageSigChld)) {
        // Apparently it's not the message we are expecting.
        return -1;
    }
    if (in == nullptr || out == nullptr) {
        // Invalid parameter
        jniThrowException(env, "java/lang/IllegalArgumentException", nullptr);
        return -1;
    }
    ScopedByteArrayRO source(env, in);
    if (source.size() < static_cast<size_t>(length)) {
        // Invalid parameter
        jniThrowException(env, "java/lang/IllegalArgumentException", nullptr);
        return -1;
    }
    const struct UnsolicitedZygoteMessageSigChld* msg =
            reinterpret_cast<const struct UnsolicitedZygoteMessageSigChld*>(source.get());

    switch (msg->header.type) {
        case UNSOLICITED_ZYGOTE_MESSAGE_TYPE_SIGCHLD: {
            ScopedIntArrayRW buf(env, out);
            if (buf.size() != 3) {
                jniThrowException(env, "java/lang/IllegalArgumentException", nullptr);
                return UNSOLICITED_ZYGOTE_MESSAGE_TYPE_RESERVED;
            }
            buf[0] = msg->payload.pid;
            buf[1] = msg->payload.uid;
            buf[2] = msg->payload.status;
            return 3;
        }
        default:
            break;
    }
    return -1;
}
// ...

这段 JNI 方法给出了可逐项反推的输入边界:

输入条件返回或异常Java 消费结果
length 不等于结构体大小-1不进入 Tracker
数组为 null,或输入容量小于 length抛 IllegalArgumentExceptionfd 回调捕获异常
未知消息类型-1不进入 Tracker
SIGCHLD 类型,但输出数组长度不是 3抛异常,并在 native 路径返回 reserved 值不构成有效三元组
大小、类型和数组均符合约束写入 PID、UID、status,返回 3提交退出状态

这些失败分支反向说明:数据报可读不等于消息有效。若要扩展协议,必须同时修改结构体、长度判断、类型分派和 Java 消费者;只向 C++ payload 加字段会让旧解析器拒收。

6. 退出原因合并 ​

源码文件:frameworks/base/services/core/java/com/android/server/am/AppExitInfoTracker.java

相关符号:scheduleChildProcDied / handleZygoteSigChld。以下为源码节选。

java
// fd 回调只投递消息;真正的合并在 KillHandler 消费路径。
private void scheduleChildProcDied(int pid, int uid, int status) {
    mKillHandler.obtainMessage(KillHandler.MSG_CHILD_PROC_DIED, pid, uid, (Integer) status)
            .sendToTarget();
}

/** Calls when zygote sends us SIGCHLD */
void handleZygoteSigChld(int pid, int uid, int status) {
    if (DEBUG_PROCESSES) {
        Slog.i(TAG, "Got SIGCHLD from zygote: pid=" + pid + ", uid=" + uid
                + ", status=" + Integer.toHexString(status));
    }
    scheduleChildProcDied(pid, uid, status);
}
// ...

MSG_CHILD_PROC_DIED 的 arg1、arg2 和 obj 分别保存 PID、UID 和装箱后的 status。KillHandler.handleMessage() 将它们交给 mAppExitInfoSourceZygote.onProcDied()。即使投递和消费使用同一个 Looper,发送消息仍然引入队列时序,不能把投递返回当成数据已合并。

源码文件:frameworks/base/services/core/java/com/android/server/am/AppExitInfoTracker.java

相关符号:AppExitInfoExternalSource.onProcDied。以下为源码节选。

java
// 已有退出记录则更新;记录尚未建立则暂存外部状态,等待另一条输入到达。
void onProcDied(final int pid, final int uid, final Integer status, final Long rssKb) {
    if (DEBUG_PROCESSES) {
        Slog.i(TAG, mTag + ": proc died: pid=" + pid + " uid=" + uid
                + ", status=" + status);
    }

    if (mService == null) {
        return;
    }

    // Unlikely but possible: the record has been created
    // Let's update it if we could find a ApplicationExitInfo record
    synchronized (mLock) {
        if (!updateExitInfoIfNecessaryLocked(pid, uid, status, mPresetReason, rssKb)) {
            if (rssKb != null) {
                addLocked(pid, uid, rssKb);     // lmkd
            } else {
                addLocked(pid, uid, status);    // zygote
            }
        }

        // Notify any interesed party regarding the lmkd kills
        final BiConsumer<Integer, Integer> listener = mProcDiedListener;
        if (listener != null) {
            mService.mHandler.post(()-> listener.accept(pid, uid));
        }
    }
}
// ...

mLock 保护退出信息和外部缓存。这不是 AMS 的全局进程锁,持有它并不代表已获得修改所有 ProcessRecord 的权限。缓存分支正是应对乱序:Zygote 状态先到时,系统还不一定拥有带包名等信息的完整退出记录。

源码文件:frameworks/base/services/core/java/com/android/server/am/AppExitInfoTracker.java

相关符号:handleNoteProcessDiedLocked。以下为源码节选。

java
// 消费暂存状态时同时 remove;lmkd 信息在此分支优先于 Zygote 状态。
@GuardedBy("mLock")
void handleNoteProcessDiedLocked(final ApplicationExitInfo raw) {
    if (raw != null) {
        if (DEBUG_PROCESSES) {
            Slog.i(TAG, "Update process exit info for " + raw.getPackageName()
                    + "(" + raw.getPid() + "/u" + raw.getRealUid() + ")");
        }

        ApplicationExitInfo info = getExitInfoLocked(raw.getPackageName(),
                raw.getPackageUid(), raw.getPid());

        // query zygote and lmkd to get the exit info, and clear the saved info
        Pair<Long, Object> zygote = mAppExitInfoSourceZygote.remove(
                raw.getPid(), raw.getRealUid());
        Pair<Long, Object> lmkd = mAppExitInfoSourceLmkd.remove(
                raw.getPid(), raw.getRealUid());

        if (info == null) {
            info = addExitInfoLocked(raw);
        }

        mIsolatedUidRecords.removeIsolatedUidLocked(raw.getRealUid());

        if (lmkd != null) {
            updateExistingExitInfoRecordLocked(info, null,
                    ApplicationExitInfo.REASON_LOW_MEMORY, (Long) lmkd.second);
        } else if (zygote != null) {
            updateExistingExitInfoRecordLocked(info, (Integer) zygote.second, null, null);
        } else {
            scheduleLogToStatsdLocked(info, false);
        }
    }
}
// ...

反方向也成立:AMS 记录先到,后来 onProcDied() 可以更新现有记录;Zygote 先到,这里取出缓存并清除。两个方向最终都进入退出原因更新逻辑。lmkd 和 Zygote 同时有信息时,这段消费者优先写低内存原因,避免把“系统为回收内存杀进程”降格为只有信号号的描述。

源码文件:frameworks/base/services/core/java/com/android/server/am/AppExitInfoTracker.java

相关符号:updateExistingExitInfoRecordLocked。以下为源码节选。

java
// 过旧记录不更新;信号终止不覆盖所有已有的更具体 reason。
@GuardedBy("mLock")
private void updateExistingExitInfoRecordLocked(ApplicationExitInfo info,
        Integer status, Integer reason, Long rssKb) {
    if (info == null || !isFresh(info.getTimestamp())) {
        // if the record is way outdated, don't update it then (because of potential pid reuse)
        return;
    }
    boolean immediateLog = false;
    if (status != null) {
        if (OsConstants.WIFEXITED(status)) {
            info.setReason(ApplicationExitInfo.REASON_EXIT_SELF);
            info.setStatus(OsConstants.WEXITSTATUS(status));
            immediateLog = true;
        } else if (OsConstants.WIFSIGNALED(status)) {
            if (info.getReason() == ApplicationExitInfo.REASON_UNKNOWN) {
                info.setReason(ApplicationExitInfo.REASON_SIGNALED);
                info.setStatus(OsConstants.WTERMSIG(status));
            } else if (info.getReason() == ApplicationExitInfo.REASON_CRASH_NATIVE) {
                info.setStatus(OsConstants.WTERMSIG(status));
                immediateLog = true;
            }
        }
    }
    if (reason != null) {
        info.setReason(reason);
        if (reason == ApplicationExitInfo.REASON_LOW_MEMORY) {
            immediateLog = true;
        }
    }
    if (rssKb != null) {
        info.setRss(rssKb.longValue());
    }
    scheduleLogToStatsdLocked(info, immediateLog);
}
// ...

isFresh() 的守卫限制旧记录被迟来的 PID 状态污染。正常退出写 REASON_EXIT_SELF 与退出码;信号终止只在 reason 未知时填 REASON_SIGNALED,已有 native crash 则补信号编号。退出原因不是“最后到达的消息无条件覆盖前一个消息”。

可以用两个场景检验理解:若先建立未知原因记录,再收到信号终止状态,应进入未知原因补全分支;若已有 native crash,收到同一终止信号,不应把 reason 改成普通的 signaled。它们是源码分支推演,真实设备上还要结合记录新鲜度、UID 映射和输入到达顺序验证。

7. 清理与重启 ​

源码文件:frameworks/base/services/core/java/com/android/server/am/ActivityManagerService.java

相关符号:AppDeathRecipient.binderDied / handleAppDiedLocked。以下为源码节选。

java
// Binder death 使用保存的 app、pid、thread;组件清理另有 owner 和锁。
@Override
public void binderDied() {
    if (DEBUG_ALL) Slog.v(
        TAG, "Death received in " + this
        + " for thread " + mAppThread.asBinder());
    synchronized (mGlobalLock) {
        appDiedLocked(mApp, mPid, mAppThread, true, null);
    }
}
// ...
@GuardedBy("this")
final void handleAppDiedLocked(ProcessRecord app, int pid,
        boolean restarting, boolean allowRestart, boolean fromBinderDied) {
    boolean kept = cleanUpApplicationRecordLocked(app, pid, restarting, allowRestart, -1,
            false /*replacingPid*/, fromBinderDied);
    if (!kept && !restarting) {
        removeLruProcessLocked(app);
        if (pid > 0) {
            ProcessList.remove(pid);
        }
    }

    mAppProfiler.onAppDiedLocked(app);

    mAtmInternal.handleAppDied(app.getWindowProcessController(), restarting, () -> {
        Slog.w(TAG, "Crash of app " + app.processName
                + " running instrumentation " + app.getActiveInstrumentation().mClass);
        Bundle info = new Bundle();
        info.putString("shortMsg", "Process crashed.");
        finishInstrumentationLocked(app, Activity.RESULT_CANCELED, info);
    });
}
// ...

appDiedLocked() 再进入 handleAppDiedLocked(),后者调用 cleanUpApplicationRecordLocked,并根据 kept、restarting 等状态决定是否移除 LRU 项。对窗口侧的通知又委托给 ATMS。这些动作说明 Framework 清理不是把 PID 从一个表里删掉那么简单。

system_server 自身死亡则走不同恢复边界。处理器发现 pid == gSystemServerPid 后直接杀死父 Zygote,把恢复交给 init 的服务管理。创建 system_server 的父分支还会在发布 PID 后重新检查一次:

源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp

相关符号:com_android_internal_os_Zygote_nativeForkSystemServer。以下为源码节选。

cpp
// 发布 system_server PID 后再次 waitpid,覆盖创建阶段的提前死亡窗口。
// The zygote process checks whether the child process has died or not.
ALOGI("System server process %d has been created", pid);
gSystemServerPid = pid;
// There is a slight window that the system server process has crashed
// but it went unnoticed because we haven't published its pid yet. So
// we recheck here just to make sure that all is well.
int status;
if (waitpid(pid, &status, WNOHANG) == pid) {
    ALOGE("System server process %d has died. Restarting Zygote!", pid);
    RuntimeAbort(env, __LINE__, "System server process has died. Restarting Zygote!");
}
// ...

这段检查和信号处理器中的 PID 比较共同服务启动恢复,但不能据此推导设备一定执行完整冷启动。init 对 Zygote 服务及其依赖的重启行为属于服务配置层;这里证明的是 native 端决定退出,而不是恢复耗时或用户可见效果。

8. 故障时间线 ​

从一段退出日志开始,先记下 PID、用户身份和时间,再把证据放回不同的 owner。单独看到一条 SIGCHLD 日志只能证明父进程处理到了退出状态。

现象优先追踪不能直接下的结论
子进程留下 zombie父进程、handler 安装、waitpid 返回值AMS 对象一定未清理
Zygote 有退出日志,缺少精确退出码sendto 错误、socket 创建、parser、外部缓存Binder death 一定丢失
退出原因是 low memory,status 不像原始信号lmkd 优先分支及 reason 更新数据报解析一定出错
PID 仍出现在 Framework 输出AMS death、重启、旧新 ProcessRecord 对照该 Linux 进程仍活着
USAP 计数异常CAS 成功与否、SIGTERM 清池、eventfd每条退出日志都应减一

以下搜索在 AOSP 源码根目录执行,用于重建本篇链路。不要直接向系统 socket 写伪造数据报来测试解析器;边界测试应在隔离的测试代码中构造输入。

bash
rg -n 'SigChldHandler|sendSigChildStatus|nativeParseSigChld|ClearForPID'   frameworks/base/core/jni/com_android_internal_os_Zygote.cpp
rg -n 'handleZygoteMessages|MSG_CHILD_PROC_DIED|onProcDied|handleNoteProcessDiedLocked'   frameworks/base/services/core/java/com/android/server/am/{ProcessList,AppExitInfoTracker}.java
rg -n 'binderDied|appDiedLocked|handleAppDiedLocked'   frameworks/base/services/core/java/com/android/server/am/ActivityManagerService.java

找到这些入口后,可以分别演算“Zygote 消息先到”“AMS 记录先到”和“数据报丢失”三条时间线。每一步写出当前记录在哪个对象中、是否持有 mLock、是否已消费缓存,便能判断缺少的是退出原因,还是进程清理本身。