Process lifecycle — birth, execution, death

This document describes exactly how a user task is created, scheduled, and destroyed in P1Start today, including known races that were fixed and residual risks.


1. Task model (TaskControlBlock)

Each schedulable unit is a row in SCHEDULER.tasks (src/kernel/task.rs):

Field (conceptual)

Role

id

PID / TID — stable identifier for IPC and kill

ppid

Parent; used by kill_recursively to include children

cr3

Root of user page tables; threads share one cr3

state

READY / RUNNING / BLOCKED / TERMINATED

reaped

Soft “zombie struct cleaned” flag — must not imply address space freed by itself

fd_table

Per-task FD map into global VFS handles

rsp, kernel_stack_*

Context switch state

Invariant: tasks[0] is the kernel/idle anchor; tasks[0].cr3 is treated as kernel CR3 in several cleanup comparisons.


2. Birth

2.1 First user process

Kernel path loads \\bin\\init.elf into a fresh address space and schedules it (see kernel shell / boot docs).

2.2 sys_execve (59)

P1Start’s syscall 59 is “load program and spawn a new task”: the caller keeps running; rax returns the child PID on success (spawn_task with a fresh user PML4). It is not a Linux-style exec that replaces the calling process image.

  • Pre-spawn stray zombie sweep (safety net before loading the next binary): any task already TERMINATED and not yet reaped is cleaned with the same ordering concerns as sys_kill.

    • Critical lock rule: teardown_sender locks SCHEDULER internally — it must not be invoked while the caller holds SCHEDULER. The sweep collects zombie PIDs + distinct user cr3 values under one scheduler lock, drops the lock, runs teardown_sender once per PID, re-acquires the lock for reap_task_fields / fd-table clear for those PIDs only, then runs cleanup_process_memory for the collected cr3 list. Holding SCHEDULER across teardown_sender deadlocks every CPU (historical bug: UI freeze when opening apps after a short-lived child exited).

  • Loads ELF from VFS (ramdisk cache fast path when available).

  • Allocates a new user PML4 for the loaded image; spawn_task wires stack, args, env, auxv, optional FD inheritance from the parent (inherit_fds_from).

2.3 Threads — sys_clone (documented in kernel)

spawn-style threads share cr3 with the parent; they are separate TaskControlBlock rows with distinct id and ppid linkage.


3. Life — scheduling and preemption

  • Per-core current_task_per_core[core_id] selects the runnable task index.

  • Timer IRQ drives schedule_next_task (task.rs): quota, priority classes, CR3 switch, FPU save/restore.

SMP fact: two different id values with the same cr3 may be RUNNING on different cores in the same wall-clock window until each core observes TERMINATED and switches away.


4. Death — three mechanisms

4.1 Voluntary: sys_exit (60)

  1. Notify parent (message type 99).

  2. teardown_sender(my_pid) — present maps + synthetic PEER_DIED path for WM clients.

  3. purge_target_dead(my_pid) on IPC broker.

  4. If no other live task shares this CR3, switch to kernel CR3 and cleanup_process_memory.

  5. If threads still share CR3, skip freeing page tables (last survivor must free).

  6. reap_task_fields + terminate_current.

4.2 Forced: sys_kill (62) — SIGKILL path (sig == 9)

Ordered phases today (syscalls.rs):

  1. Under SCHEDULER lock (interrupts disabled on this CPU):

    • kill_recursively(pid, sig, &mut killed_pids) — builds kill set (parent + ppid descendants), marks matching tasks TERMINATED for SIGKILL.

    • Collect only tasks whose id is in killed_pids, still TERMINATED, not yet reaped, with non-zero user cr3.

    • Record (pid_reaped[], cr3_to_clean[]) — do not call reap_task_fields yet (see §5).

  2. Drop scheduler lock — other cores may now run and must drain out of the doomed address space.

  3. teardown_sender once per reaped PID — unmaps syscall-50 spans in receivers and enqueues WIRE_WM_PEER_DIED.

  4. Dedupe by CR3 — threads share one root; cleanup_process_memory must run once per distinct user CR3.

  5. Before each cleanup_process_memory: quiesce_user_cr3_before_free spins until no current_task_per_core[*] has tasks[idx].cr3 == target_cr3 (bounded spin + timeout warn).

  6. cleanup_process_memory — walks user page tables and returns physical frames.

  7. Re-acquire SCHEDULER lock: reap_task_fields for each killed PID (zeros cr3 in TCB, clears stacks/strings).

  8. purge_target_dead per killed PID on the IPC broker.

4.3 Legacy / partial: reap_zombies (timer housekeeping)

On BSP, every N ticks (task.rs), reap_zombies clears soft zombie fields for TERMINATED tasks.

Critical rule (fixed): do not set reaped = true while the task still holds a non-kernel cr3. Otherwise a later sys_kill pass can skip full teardown (reap_task_fields / cleanup_process_memory / teardown_sender) and leave stale present maps and leaked address spaces.

4.4 Parent wait: sys_waitpid (61)

A task can block on a specific child PID until that child has been reaped. When the wait completes successfully, the syscall return value rax is the child’s exit code — the value the child passed in rdi to sys_exit (60) (stored on the child’s TaskControlBlock before teardown). Userland treats this as a u64 carrying a sign-extended i32 where applicable. See syscalls-reference.md and p1scode-ide-sdk-toolchain.md (IDE p1sc exit status).


5. Why reap_task_fields moved after free (SMP)

Older ordering called reap_task_fields (which zeros cr3 in the TCB) before quiescence and cleanup_process_memory. On multi-core (smp > 1), another CPU could still be the current task for that cr3 while the struct already lied about ownership — quiescence checks became unreliable.

Current ordering: quiesce using live cr3 in TCBs → free pages → then zero TCB cr3.


6. Edge cases (explicit)

Case

Behavior

Threads share CR3

kill_recursively includes ppid descendants; cleanup_process_memory once per unique cr3; teardown_sender per PID (present map keys are per sender task id).

kill_recursively cap (32)

Deep trees may truncate kill list — operational limit.

Quiescence timeout

Kernel logs a warning and proceeds — last-resort; can still corrupt if a CPU was stuck.

sys_execve stray zombie sweep

Same building blocks as sys_kill (teardown_sender → reap_task_fields → cleanup_process_memory), but only for tasks already TERMINATED+!reaped. Must never nest SCHEDULER + teardown_sender.


7. See also