SIGCHLD进程死亡处理
一个应用已经从 /proc 消失,为什么退出记录还没有更新?反过来,日志已经出现 Zygote 的 “exited due to signal”,为什么 AMS 中还有与它关联的状态?这两个现象要求我们区分三件事:Linux 回收子进程、收集退出原因、清理 Framework 对象。它们有不同入口,不存在一条把所有动作同步做完的 SIGCHLD 回调。
本文面向理解 fork()、PID 和 Handler 的读者。可以先读子进程PID管理,理解 PID 如何关联到 ProcessRecord。这里从父进程收到 SIGCHLD 开始,追踪 waitpid 的状态字如何进入 ApplicationExitInfo,并解释 USAP 表项回收、消息丢失、乱序到达和 system_server 死亡。AMS 内各类服务、Provider 和窗口的完整清理不在本篇展开。
1. 三种死亡状态
Zygote 是普通应用和 system_server 的父进程;通过 App Zygote 等路径创建的进程则由相应父进程负责回收。退出的进程在父进程取走退出状态以前可能保持 zombie 状态。waitpid() 消费的是内核保存的子进程退出信息;ProcessRecord 是 system_server 中描述应用的 Java 对象,两者不存在共享内存里的“自动同步删除”。
| 状态 | 所有者和入口 | 完成后能说明什么 |
|---|---|---|
| 子进程已回收 | 父 Zygote,SigChldHandler → waitpid | 父进程已取得该 child 的退出状态 |
| 退出原因已合并 | AppExitInfoTracker,Zygote/lmkd/AMS 多个输入 | 对外记录可提供 reason、status 等信息 |
| Framework 已清理 | AMS,appDiedLocked → handleAppDiedLocked | 连接、LRU、窗口等按各自策略处理 |
下图只描述相关输入如何汇合,不表示 Binder death 与 SIGCHLD 有固定先后。
Zygote 的发送是尽力通知;AMS 的清理也不以 datagram 返回作为提交屏障。阅读后续代码时,应分别标出“内核状态已消费”和“Java 记录已更新”的时点。
2. 回收循环
源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp
相关符号:SetSignalHandlers / UnsetChldSignalHandler。以下为源码节选。
// 父进程安装处理器,子进程恢复默认处理,避免沿用父进程的回收职责。
static void SetSignalHandlers() {
struct sigaction sig_chld = {.sa_flags = SA_SIGINFO, .sa_sigaction = SigChldHandler};
if (sigaction(SIGCHLD, &sig_chld, nullptr) < 0) {
ALOGW("Error setting SIGCHLD handler: %s", strerror(errno));
}
struct sigaction sig_hup = {};
sig_hup.sa_handler = SIG_IGN;
if (sigaction(SIGHUP, &sig_hup, nullptr) < 0) {
ALOGW("Error setting SIGHUP handler: %s", strerror(errno));
}
}
// Sets the SIGCHLD handler back to default behavior in zygote children.
static void UnsetChldSignalHandler() {
struct sigaction sa;
memset(&sa, 0, sizeof(sa));
sa.sa_handler = SIG_DFL;
if (sigaction(SIGCHLD, &sa, nullptr) < 0) {
ALOGW("Error unsetting SIGCHLD handler: %s", strerror(errno));
}
}
// ...sigaction 使用 SA_SIGINFO,让处理器取得 siginfo_t。fork 周围还会暂时屏蔽 SIGCHLD;ForkCommon() 在父子分支整理完成后解除屏蔽。安装、屏蔽和解除屏蔽需要一起阅读,不能认为 handler 可以在任意初始化阶段安全执行。
源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp
相关符号:SigChldHandler。以下为源码节选。
// 一次信号处理循环回收多个 child;status 来自每次 waitpid,而 uid 取自本次信号的 info。
// This signal handler is for zygote mode, since the zygote must reap its children
NO_STACK_PROTECTOR
static void SigChldHandler(int /*signal_number*/, siginfo_t* info, void* /*ucontext*/) {
pid_t pid;
int status;
int64_t usaps_removed = 0;
// It's necessary to save and restore the errno during this function.
// Since errno is stored per thread, changing it here modifies the errno
// on the thread on which this signal handler executes. If a signal occurs
// between a call and an errno check, it's possible to get the errno set
// here.
// See b/23572286 for extra information.
int saved_errno = errno;
while ((pid = waitpid(-1, &status, WNOHANG)) > 0) {
// Notify system_server that we received a SIGCHLD
sendSigChildStatus(pid, info->si_uid, status);
// Log process-death status that we care about.
if (WIFEXITED(status)) {
async_safe_format_log(ANDROID_LOG_INFO, LOG_TAG, "Process %d exited cleanly (%d)", pid,
WEXITSTATUS(status));
// Check to see if the PID is in the USAP pool and remove it if it is.
if (RemoveUsapTableEntry(pid)) {
++usaps_removed;
}
} else if (WIFSIGNALED(status)) {
async_safe_format_log(ANDROID_LOG_INFO, LOG_TAG,
"Process %d exited due to signal %d (%s)%s", pid,
WTERMSIG(status), strsignal(WTERMSIG(status)),
WCOREDUMP(status) ? "; core dumped" : "");
// If the process exited due to a signal other than SIGTERM, check to see
// if the PID is in the USAP pool and remove it if it is. If the process
// was closed by the Zygote using SIGTERM then the USAP pool entry will
// have already been removed (see nativeEmptyUsapPool()).
if (WTERMSIG(status) != SIGTERM && RemoveUsapTableEntry(pid)) {
++usaps_removed;
}
}
// If the just-crashed process is the system_server, bring down zygote
// so that it is restarted by init and system server will be restarted
// from there.
if (pid == gSystemServerPid) {
async_safe_format_log(ANDROID_LOG_ERROR, LOG_TAG,
"Exit zygote because system server (pid %d) has terminated", pid);
kill(getpid(), SIGKILL);
}
}
// Note that we shouldn't consider ECHILD an error because
// the secondary zygote might have no children left to wait for.
if (pid < 0 && errno != ECHILD) {
async_safe_format_log(ANDROID_LOG_WARN, LOG_TAG, "Zygote SIGCHLD error in waitpid: %s",
strerror(errno));
}
if (usaps_removed > 0) {
if (TEMP_FAILURE_RETRY(write(gUsapPoolEventFD, &usaps_removed, sizeof(usaps_removed))) ==
-1) {
// If this write fails something went terribly wrong. We will now kill
// the zygote and let the system bring it back up.
async_safe_format_log(ANDROID_LOG_ERROR, LOG_TAG,
"Zygote failed to write to USAP pool event FD: %s",
strerror(errno));
kill(getpid(), SIGKILL);
}
}
errno = saved_errno;
}
// ...WNOHANG 使父进程不会为了尚未退出的 child 阻塞;返回正数才表示回收了一项。普通信号可能合并,循环因此不能改成单次 waitpid。pid == 0 表示当时没有可回收的退出状态;负数时 ECHILD 被允许,其他错误才打印警告。
status 是编码后的 wait status,不是直接的退出码或信号编号。只有 WIFEXITED(status) 成立,WEXITSTATUS(status) 才表示正常退出码;信号终止需要用 WIFSIGNALED 和 WTERMSIG。例如正常 exit(7) 与因信号 7 终止,不能都写成“退出码 7”。
还有一个很容易在复述中消失的边界:源码把每次 waitpid 返回的 PID 与同一份 info->si_uid 配对。它没有逐个查询刚回收 PID 的 UID。一次 handler 收割多个退出进程时,不能单凭这段实现声称“每条 UID 都由对应 waitpid 结果提供”。这属于通知来源的边界,不能在示意代码中擅自补出一次 UID 查询。
errno 是线程状态,handler 会打断该线程原来的代码。保存和恢复它,是为了避免原代码刚执行系统调用、尚未检查错误时,被 handler 的 waitpid 或 write 改写判断依据。NO_STACK_PROTECTOR 是函数编译属性,不能把它当成“所有被调用函数都满足异步信号安全”的证明;这里也不应引入 Java 回调或 Framework 锁。
3. 池成员回收
USAP 是尚未特化成具体应用的预创建进程。回收它除了消费内核状态,还要关闭池表项关联的读管道,并且只扣减一次池计数。
源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp
相关符号:UsapTableEntry::ClearForPID / RemoveUsapTableEntry。以下为源码节选。
// 只有成功把表项从当前值交换成无效值的调用,才能关闭 fd 并触发池计数扣减。
bool ClearForPID(int32_t pid) {
EntryStorage storage = mStorage.load();
if (storage.pid == pid) {
/*
* There are three possible outcomes from this compare-and-exchange:
* 1) It succeeds, in which case we close the FD
* 2) It fails and the new value is INVALID_ENTRY_VALUE, in which case
* the entry has already been cleared.
* 3) It fails and the new value isn't INVALID_ENTRY_VALUE, in which
* case the entry has already been cleared and re-used.
*
* In all three cases the goal of the caller has been met, but only in
* the first case do we need to decrement the pool count.
*/
if (mStorage.compare_exchange_strong(storage, INVALID_ENTRY_VALUE)) {
close(storage.read_pipe_fd);
return true;
} else {
return false;
}
} else {
return false;
}
}
// ...
static bool RemoveUsapTableEntry(pid_t usap_pid) {
for (UsapTableEntry& entry : gUsapTable) {
if (entry.ClearForPID(usap_pid)) {
--gUsapPoolCount;
return true;
}
}
return false;
}
// ...这里的原子比较交换保护的是“观察到的表项是否仍然有效”。另一条路径可能已经清除该项,或者清除后将其复用;CAS 失败时返回 false,调用方不再扣减。仅有 PID 比较并不能概括这段并发约束,更不能据此宣称消除了系统中所有 PID 复用问题。
SigChldHandler() 只累计本次确实移除的数量,循环结束后写一次 gUsapPoolEventFD。eventfd 通知把池成员变化带回事件循环;补池不是在信号处理器中直接调用 Java。写失败会触发 kill(getpid(), SIGKILL),因为 native 池状态变化已经发生,继续运行会让外部消费者错过变化。
对 SIGTERM 的特殊分支也不能删掉:Zygote 主动清空池的 nativeEmptyUsapPool() 会先移除表项再发送 SIGTERM,handler 对这个终止信号不再执行常规扣减。不能把所有 WIFSIGNALED 都等价理解为“池计数减一”。
4. 数据报契约
源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp
相关符号:UnsolicitedZygoteMessageSigChld / initUnsolSocketToSystemServer / sendSigChildStatus。以下为源码节选。
// 通知使用非阻塞 Unix datagram,发送失败不会撤销前面的 waitpid。
struct UnsolicitedZygoteMessageSigChld {
struct {
UnsolicitedZygoteMessageTypes type;
} header;
struct {
pid_t pid;
uid_t uid;
int status;
} payload;
};
// Keep sync with services/core/java/com/android/server/am/ProcessList.java
static constexpr struct sockaddr_un kSystemServerSockAddr =
{.sun_family = AF_LOCAL, .sun_path = "/data/system/unsolzygotesocket"};
// ...
static void initUnsolSocketToSystemServer() {
gSystemServerSocketFd = socket(AF_LOCAL, SOCK_DGRAM | SOCK_NONBLOCK, 0);
if (gSystemServerSocketFd >= 0) {
ALOGV("Zygote:systemServerSocketFD = %d", gSystemServerSocketFd);
} else {
ALOGE("Unable to create socket file descriptor to connect to system_server");
}
}
static void sendSigChildStatus(const pid_t pid, const uid_t uid, const int status) {
int socketFd = gSystemServerSocketFd;
if (socketFd >= 0) {
// fill the message buffer
struct UnsolicitedZygoteMessageSigChld data =
{.header = {.type = UNSOLICITED_ZYGOTE_MESSAGE_TYPE_SIGCHLD},
.payload = {.pid = pid, .uid = uid, .status = status}};
if (TEMP_FAILURE_RETRY(
sendto(socketFd, &data, sizeof(data), 0,
reinterpret_cast<const struct sockaddr*>(&kSystemServerSockAddr),
sizeof(kSystemServerSockAddr))) == -1) {
async_safe_format_log(ANDROID_LOG_ERROR, LOG_TAG,
"Zygote failed to write to system_server FD: %s",
strerror(errno));
}
}
}
// ...数据报由消息类型和三个整数构成。它不是Zygote参数传递协议中的文本启动请求,也没有请求编号、确认响应或失败重发队列。SOCK_NONBLOCK 避免因接收方跟不上而长期占住 signal handler,但也意味着退出原因通知可能丢失。
gSystemServerSocketFd < 0 时根本不发送;sendto 返回错误时只记录日志。无论哪一种,子进程已经被回收,不会再次通过 waitpid 自动补发。排查“应用已退出但缺少精确状态”时,这条失败路径比猜测 AMS 没收到 Binder death 更直接。
5. 接收与校验
源码文件:frameworks/base/services/core/java/com/android/server/am/ProcessList.java
相关符号:createSystemServerSocketForZygote / handleZygoteMessages。以下为源码节选。
// 创建失败关闭已创建 socket;只有解析器返回三项才把数据交给 Tracker。
private LocalSocket createSystemServerSocketForZygote() {
// The file system entity for this socket is created with 0666 perms, owned
// by system:system. selinux restricts things so that only zygotes can
// access it.
final File socketFile = new File(UNSOL_ZYGOTE_MSG_SOCKET_PATH);
if (socketFile.exists()) {
socketFile.delete();
}
LocalSocket serverSocket = null;
try {
serverSocket = new LocalSocket(LocalSocket.SOCKET_DGRAM);
serverSocket.bind(new LocalSocketAddress(
UNSOL_ZYGOTE_MSG_SOCKET_PATH, LocalSocketAddress.Namespace.FILESYSTEM));
Os.chmod(UNSOL_ZYGOTE_MSG_SOCKET_PATH, 0666);
} catch (Exception e) {
if (serverSocket != null) {
try {
serverSocket.close();
} catch (IOException ex) {
}
serverSocket = null;
}
}
return serverSocket;
}
/**
* Handle the unsolicited message from zygote.
*/
private int handleZygoteMessages(FileDescriptor fd, int events) {
final int eventFd = fd.getInt$();
if ((events & EVENT_INPUT) != 0) {
// An incoming message from zygote
try {
final int len = Os.read(fd, mZygoteUnsolicitedMessage, 0,
mZygoteUnsolicitedMessage.length);
if (len > 0 && mZygoteSigChldMessage.length == Zygote.nativeParseSigChld(
mZygoteUnsolicitedMessage, len, mZygoteSigChldMessage)) {
mAppExitInfoTracker.handleZygoteSigChld(
mZygoteSigChldMessage[0] /* pid */,
mZygoteSigChldMessage[1] /* uid */,
mZygoteSigChldMessage[2] /* status */);
}
} catch (Exception e) {
Slog.w(TAG, "Exception in reading unsolicited zygote message: " + e);
}
}
return EVENT_INPUT;
}
// ...创建成功后,ProcessList 在 sKillHandler 的 MessageQueue 上注册 EVENT_INPUT fd listener;返回 EVENT_INPUT 表示继续监听。创建失败返回 null,注册步骤也就被跳过。这是与“发送端正常,但 system_server 未监听”对应的真实失败分支。
socket 节点的 0666 权限不是来源认证。这里还依赖 SELinux 访问控制;native parser 的职责则是验证消息结构,不能用长度检查代替调用者身份校验。
源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp
相关符号:com_android_internal_os_Zygote_nativeParseSigChld。以下为源码节选。
// 未知类型最终返回 -1;正确类型也必须经过输出数组长度检查。
static jint com_android_internal_os_Zygote_nativeParseSigChld(JNIEnv* env, jclass, jbyteArray in,
jint length, jintArray out) {
if (length != sizeof(struct UnsolicitedZygoteMessageSigChld)) {
// Apparently it's not the message we are expecting.
return -1;
}
if (in == nullptr || out == nullptr) {
// Invalid parameter
jniThrowException(env, "java/lang/IllegalArgumentException", nullptr);
return -1;
}
ScopedByteArrayRO source(env, in);
if (source.size() < static_cast<size_t>(length)) {
// Invalid parameter
jniThrowException(env, "java/lang/IllegalArgumentException", nullptr);
return -1;
}
const struct UnsolicitedZygoteMessageSigChld* msg =
reinterpret_cast<const struct UnsolicitedZygoteMessageSigChld*>(source.get());
switch (msg->header.type) {
case UNSOLICITED_ZYGOTE_MESSAGE_TYPE_SIGCHLD: {
ScopedIntArrayRW buf(env, out);
if (buf.size() != 3) {
jniThrowException(env, "java/lang/IllegalArgumentException", nullptr);
return UNSOLICITED_ZYGOTE_MESSAGE_TYPE_RESERVED;
}
buf[0] = msg->payload.pid;
buf[1] = msg->payload.uid;
buf[2] = msg->payload.status;
return 3;
}
default:
break;
}
return -1;
}
// ...这段 JNI 方法给出了可逐项反推的输入边界:
| 输入条件 | 返回或异常 | Java 消费结果 |
|---|---|---|
length 不等于结构体大小 | -1 | 不进入 Tracker |
| 数组为 null,或输入容量小于 length | 抛 IllegalArgumentException | fd 回调捕获异常 |
| 未知消息类型 | -1 | 不进入 Tracker |
| SIGCHLD 类型,但输出数组长度不是 3 | 抛异常,并在 native 路径返回 reserved 值 | 不构成有效三元组 |
| 大小、类型和数组均符合约束 | 写入 PID、UID、status,返回 3 | 提交退出状态 |
这些失败分支反向说明:数据报可读不等于消息有效。若要扩展协议,必须同时修改结构体、长度判断、类型分派和 Java 消费者;只向 C++ payload 加字段会让旧解析器拒收。
6. 退出原因合并
源码文件:frameworks/base/services/core/java/com/android/server/am/AppExitInfoTracker.java
相关符号:scheduleChildProcDied / handleZygoteSigChld。以下为源码节选。
// fd 回调只投递消息;真正的合并在 KillHandler 消费路径。
private void scheduleChildProcDied(int pid, int uid, int status) {
mKillHandler.obtainMessage(KillHandler.MSG_CHILD_PROC_DIED, pid, uid, (Integer) status)
.sendToTarget();
}
/** Calls when zygote sends us SIGCHLD */
void handleZygoteSigChld(int pid, int uid, int status) {
if (DEBUG_PROCESSES) {
Slog.i(TAG, "Got SIGCHLD from zygote: pid=" + pid + ", uid=" + uid
+ ", status=" + Integer.toHexString(status));
}
scheduleChildProcDied(pid, uid, status);
}
// ...MSG_CHILD_PROC_DIED 的 arg1、arg2 和 obj 分别保存 PID、UID 和装箱后的 status。KillHandler.handleMessage() 将它们交给 mAppExitInfoSourceZygote.onProcDied()。即使投递和消费使用同一个 Looper,发送消息仍然引入队列时序,不能把投递返回当成数据已合并。
源码文件:frameworks/base/services/core/java/com/android/server/am/AppExitInfoTracker.java
相关符号:AppExitInfoExternalSource.onProcDied。以下为源码节选。
// 已有退出记录则更新;记录尚未建立则暂存外部状态,等待另一条输入到达。
void onProcDied(final int pid, final int uid, final Integer status, final Long rssKb) {
if (DEBUG_PROCESSES) {
Slog.i(TAG, mTag + ": proc died: pid=" + pid + " uid=" + uid
+ ", status=" + status);
}
if (mService == null) {
return;
}
// Unlikely but possible: the record has been created
// Let's update it if we could find a ApplicationExitInfo record
synchronized (mLock) {
if (!updateExitInfoIfNecessaryLocked(pid, uid, status, mPresetReason, rssKb)) {
if (rssKb != null) {
addLocked(pid, uid, rssKb); // lmkd
} else {
addLocked(pid, uid, status); // zygote
}
}
// Notify any interesed party regarding the lmkd kills
final BiConsumer<Integer, Integer> listener = mProcDiedListener;
if (listener != null) {
mService.mHandler.post(()-> listener.accept(pid, uid));
}
}
}
// ...mLock 保护退出信息和外部缓存。这不是 AMS 的全局进程锁,持有它并不代表已获得修改所有 ProcessRecord 的权限。缓存分支正是应对乱序:Zygote 状态先到时,系统还不一定拥有带包名等信息的完整退出记录。
源码文件:frameworks/base/services/core/java/com/android/server/am/AppExitInfoTracker.java
相关符号:handleNoteProcessDiedLocked。以下为源码节选。
// 消费暂存状态时同时 remove;lmkd 信息在此分支优先于 Zygote 状态。
@GuardedBy("mLock")
void handleNoteProcessDiedLocked(final ApplicationExitInfo raw) {
if (raw != null) {
if (DEBUG_PROCESSES) {
Slog.i(TAG, "Update process exit info for " + raw.getPackageName()
+ "(" + raw.getPid() + "/u" + raw.getRealUid() + ")");
}
ApplicationExitInfo info = getExitInfoLocked(raw.getPackageName(),
raw.getPackageUid(), raw.getPid());
// query zygote and lmkd to get the exit info, and clear the saved info
Pair<Long, Object> zygote = mAppExitInfoSourceZygote.remove(
raw.getPid(), raw.getRealUid());
Pair<Long, Object> lmkd = mAppExitInfoSourceLmkd.remove(
raw.getPid(), raw.getRealUid());
if (info == null) {
info = addExitInfoLocked(raw);
}
mIsolatedUidRecords.removeIsolatedUidLocked(raw.getRealUid());
if (lmkd != null) {
updateExistingExitInfoRecordLocked(info, null,
ApplicationExitInfo.REASON_LOW_MEMORY, (Long) lmkd.second);
} else if (zygote != null) {
updateExistingExitInfoRecordLocked(info, (Integer) zygote.second, null, null);
} else {
scheduleLogToStatsdLocked(info, false);
}
}
}
// ...反方向也成立:AMS 记录先到,后来 onProcDied() 可以更新现有记录;Zygote 先到,这里取出缓存并清除。两个方向最终都进入退出原因更新逻辑。lmkd 和 Zygote 同时有信息时,这段消费者优先写低内存原因,避免把“系统为回收内存杀进程”降格为只有信号号的描述。
源码文件:frameworks/base/services/core/java/com/android/server/am/AppExitInfoTracker.java
相关符号:updateExistingExitInfoRecordLocked。以下为源码节选。
// 过旧记录不更新;信号终止不覆盖所有已有的更具体 reason。
@GuardedBy("mLock")
private void updateExistingExitInfoRecordLocked(ApplicationExitInfo info,
Integer status, Integer reason, Long rssKb) {
if (info == null || !isFresh(info.getTimestamp())) {
// if the record is way outdated, don't update it then (because of potential pid reuse)
return;
}
boolean immediateLog = false;
if (status != null) {
if (OsConstants.WIFEXITED(status)) {
info.setReason(ApplicationExitInfo.REASON_EXIT_SELF);
info.setStatus(OsConstants.WEXITSTATUS(status));
immediateLog = true;
} else if (OsConstants.WIFSIGNALED(status)) {
if (info.getReason() == ApplicationExitInfo.REASON_UNKNOWN) {
info.setReason(ApplicationExitInfo.REASON_SIGNALED);
info.setStatus(OsConstants.WTERMSIG(status));
} else if (info.getReason() == ApplicationExitInfo.REASON_CRASH_NATIVE) {
info.setStatus(OsConstants.WTERMSIG(status));
immediateLog = true;
}
}
}
if (reason != null) {
info.setReason(reason);
if (reason == ApplicationExitInfo.REASON_LOW_MEMORY) {
immediateLog = true;
}
}
if (rssKb != null) {
info.setRss(rssKb.longValue());
}
scheduleLogToStatsdLocked(info, immediateLog);
}
// ...isFresh() 的守卫限制旧记录被迟来的 PID 状态污染。正常退出写 REASON_EXIT_SELF 与退出码;信号终止只在 reason 未知时填 REASON_SIGNALED,已有 native crash 则补信号编号。退出原因不是“最后到达的消息无条件覆盖前一个消息”。
可以用两个场景检验理解:若先建立未知原因记录,再收到信号终止状态,应进入未知原因补全分支;若已有 native crash,收到同一终止信号,不应把 reason 改成普通的 signaled。它们是源码分支推演,真实设备上还要结合记录新鲜度、UID 映射和输入到达顺序验证。
7. 清理与重启
源码文件:frameworks/base/services/core/java/com/android/server/am/ActivityManagerService.java
相关符号:AppDeathRecipient.binderDied / handleAppDiedLocked。以下为源码节选。
// Binder death 使用保存的 app、pid、thread;组件清理另有 owner 和锁。
@Override
public void binderDied() {
if (DEBUG_ALL) Slog.v(
TAG, "Death received in " + this
+ " for thread " + mAppThread.asBinder());
synchronized (mGlobalLock) {
appDiedLocked(mApp, mPid, mAppThread, true, null);
}
}
// ...
@GuardedBy("this")
final void handleAppDiedLocked(ProcessRecord app, int pid,
boolean restarting, boolean allowRestart, boolean fromBinderDied) {
boolean kept = cleanUpApplicationRecordLocked(app, pid, restarting, allowRestart, -1,
false /*replacingPid*/, fromBinderDied);
if (!kept && !restarting) {
removeLruProcessLocked(app);
if (pid > 0) {
ProcessList.remove(pid);
}
}
mAppProfiler.onAppDiedLocked(app);
mAtmInternal.handleAppDied(app.getWindowProcessController(), restarting, () -> {
Slog.w(TAG, "Crash of app " + app.processName
+ " running instrumentation " + app.getActiveInstrumentation().mClass);
Bundle info = new Bundle();
info.putString("shortMsg", "Process crashed.");
finishInstrumentationLocked(app, Activity.RESULT_CANCELED, info);
});
}
// ...appDiedLocked() 再进入 handleAppDiedLocked(),后者调用 cleanUpApplicationRecordLocked,并根据 kept、restarting 等状态决定是否移除 LRU 项。对窗口侧的通知又委托给 ATMS。这些动作说明 Framework 清理不是把 PID 从一个表里删掉那么简单。
system_server 自身死亡则走不同恢复边界。处理器发现 pid == gSystemServerPid 后直接杀死父 Zygote,把恢复交给 init 的服务管理。创建 system_server 的父分支还会在发布 PID 后重新检查一次:
源码文件:frameworks/base/core/jni/com_android_internal_os_Zygote.cpp
相关符号:com_android_internal_os_Zygote_nativeForkSystemServer。以下为源码节选。
// 发布 system_server PID 后再次 waitpid,覆盖创建阶段的提前死亡窗口。
// The zygote process checks whether the child process has died or not.
ALOGI("System server process %d has been created", pid);
gSystemServerPid = pid;
// There is a slight window that the system server process has crashed
// but it went unnoticed because we haven't published its pid yet. So
// we recheck here just to make sure that all is well.
int status;
if (waitpid(pid, &status, WNOHANG) == pid) {
ALOGE("System server process %d has died. Restarting Zygote!", pid);
RuntimeAbort(env, __LINE__, "System server process has died. Restarting Zygote!");
}
// ...这段检查和信号处理器中的 PID 比较共同服务启动恢复,但不能据此推导设备一定执行完整冷启动。init 对 Zygote 服务及其依赖的重启行为属于服务配置层;这里证明的是 native 端决定退出,而不是恢复耗时或用户可见效果。
8. 故障时间线
从一段退出日志开始,先记下 PID、用户身份和时间,再把证据放回不同的 owner。单独看到一条 SIGCHLD 日志只能证明父进程处理到了退出状态。
| 现象 | 优先追踪 | 不能直接下的结论 |
|---|---|---|
| 子进程留下 zombie | 父进程、handler 安装、waitpid 返回值 | AMS 对象一定未清理 |
| Zygote 有退出日志,缺少精确退出码 | sendto 错误、socket 创建、parser、外部缓存 | Binder death 一定丢失 |
| 退出原因是 low memory,status 不像原始信号 | lmkd 优先分支及 reason 更新 | 数据报解析一定出错 |
| PID 仍出现在 Framework 输出 | AMS death、重启、旧新 ProcessRecord 对照 | 该 Linux 进程仍活着 |
| USAP 计数异常 | CAS 成功与否、SIGTERM 清池、eventfd | 每条退出日志都应减一 |
以下搜索在 AOSP 源码根目录执行,用于重建本篇链路。不要直接向系统 socket 写伪造数据报来测试解析器;边界测试应在隔离的测试代码中构造输入。
rg -n 'SigChldHandler|sendSigChildStatus|nativeParseSigChld|ClearForPID' frameworks/base/core/jni/com_android_internal_os_Zygote.cpp
rg -n 'handleZygoteMessages|MSG_CHILD_PROC_DIED|onProcDied|handleNoteProcessDiedLocked' frameworks/base/services/core/java/com/android/server/am/{ProcessList,AppExitInfoTracker}.java
rg -n 'binderDied|appDiedLocked|handleAppDiedLocked' frameworks/base/services/core/java/com/android/server/am/ActivityManagerService.java找到这些入口后,可以分别演算“Zygote 消息先到”“AMS 记录先到”和“数据报丢失”三条时间线。每一步写出当前记录在哪个对象中、是否持有 mLock、是否已消费缓存,便能判断缺少的是退出原因,还是进程清理本身。
