Threading model (kernel + userland)

This page is the stable contract for threads on P1Start: what the scheduler guarantees, how TLS works, how kernel stacks are assigned, and lock ordering so subsystems stay composable. Pair it with User virtual address layout (x86_64) for address-space rules.

Note

Canonical threading/scheduler/process docs: Scheduler and SMP, Process lifecycle — birth, execution, death, Memory and address spaces (sources in newdocs/). Desktop boot today runs single-core scheduler bring-up (APs not started in start_desktop); the invariants below still hold, but only the BSP runs user tasks until SMP is re-enabled (see scheduler page § boot profile).

Threads vs processes

  • A process here is a task with its own CR3 (page tables), brk_current / mmap_current, fd_table, and VMM metadata.

  • A thread is a separate task id that shares CR3 with siblings (same user address space). Created via spawn_thread (kernel) / the user clone-style syscall path. Threads share the mmap cursor and file descriptor table semantics documented in the VA layout page.

Until stress tests pass, treat multi-threaded user code as beta: use cstd::pthread (below) or raw sys_clone.

Current status (what you can use today)

API

Status

sys_clone (56)

Stable kernel ABI; cstd::sys::sys_clone matches it.

sys_waitpid (61)

Blocks until target tid exits; kernel yields after marking BLOCKED.

cstd::pthread

pthread_create / join / mutex_* implemented: per-thread mmap stack + ARCH_SET_FS to thread-local slab; mutex = spin (AtomicU32). __errno_location uses %fs:0 on P1Start (target_os = "none"); main thread calls errno::init_main_thread_tls_errno from sys_main_trampoline. Legacy errno remains for linkers / fallback when FSBASE == 0.

sys_thread_create

Rust fn(u64) helper; still valid; uses the same sys_clone.

Scheduler invariants

Invariant

Meaning

One current task per logical CPU

current_task_per_core[core] is the only task whose user context is “live” on that core between IRQ/timer boundaries.

Preempt path

Timer IRQ (hardware_timer_tick_and_schedule) and software yield (int 0x81) call schedule_next_task, which may switch tasks.

Saved GPR / RSP

Outgoing task’s RSP (kernel stack position) is stored in its TCB before another task runs.

FPU

fxsave on outgoing, fxrstor on incoming — XMM state is per-task.

CR3

Switched when the next task’s CR3 differs from the current (process switch). Threads sharing an mm keep the same CR3.

TLS (FS base)

MSR_FS_BASE is per-CPU, not per-task. The kernel reads FS base into the outgoing TCB and writes the incoming TCB’s saved value on every real context switch. Without this, ARCH_SET_FS would corrupt sibling threads.

Kernel stack (RSP0)

Each runnable task has kernel_stack_top; set_kernel_stack runs before resuming so #DF / IST paths see the right stack. Idle may use 0 where documented.

Scheduler lock vs allocation

Do not hold SCHEDULER while doing VMM / page-table work that might recurse through the same lock on timer preempt (see syscall mmap paths).

Zombie reaping

Only BSP (core 0) runs periodic reap_zombies inside the scheduler tick to avoid cross-core races on the global task table.

TLS (ARCH_SET_FS)

  • Syscall 158 (sys_arch_prctl): rdi = 0x1002 (ARCH_SET_FS), rsi = user virtual base for %fs-relative TLS (musl / typical Linux user ABI).

  • The kernel:

    1. Validates the pointer with user_fsbase_allowed_for_tls (canonical low user VA, including P1’s mmap band — same upper bound as MAX_USER_CANONICAL_LOW_VADDR in task.rs).

    2. Stores it in the current task’s TaskControlBlock::tls_fsbase.

    3. Writes MSR_FS_BASE immediately so the running thread sees the update.

  • New tasks and new kernel threads start with tls_fsbase = 0 until user code sets TLS.

  • Context switch (schedule_next_task): outgoing task’s live FsBase::read() is copied into tls_fsbase; incoming task’s tls_fsbase is written with FsBase::write. If a TCB ever contained a bad value, the switch path forces 0 (fail-safe).

Kernel stack per thread

  • spawn_task (new process) and spawn_thread each allocate a dedicated 64 KiB kernel stack buffer and set kernel_stack_top to the high end for that task’s ring-0 work.

  • Do not reuse another task’s kernel stack for a concurrently runnable thread.

Lock ordering (do not invert)

These rules prevent deadlocks with timer-driven preempt and syscall paths:

  1. SCHEDULER — only for short TCB updates and run-queue decisions. Never take VFS or allocate VMM pages while holding it unless the callee is proven non-blocking.

  2. VFS — may be taken from syscalls after dropping the scheduler lock. Never acquire SCHEDULER while holding VFS.

  3. PTY / session — expect a single global VFS lock today; session foreground and io_wait wakeups must stay consistent with task id (tid), not only “process” id, when multiple threads share an address space.

  4. Compositor / WM IPC — treat per-window state as needing either one logical UI thread in userland or explicit kernel-side serialization; concurrent present from two threads without protocol is undefined until a queue or mutex is specified.

Mnemonic: scheduler first and alone for scheduling; VFS and page tables afterward.

Enforcement (debug builds)

In cfg(debug_assertions), vmm_map_page, vmm_unmap_user_pages, and vmm_unmap_user_pages_free_phys (src/mm/memory.rs) call debug_assert_scheduler_unlocked_for_vmm: the scheduler mutex must be free before page-table mutation, matching the rule above. Release builds skip the check for speed.

Userland guidance

  • Prefer cstd::pthread over ad-hoc syscall numbers for new code.

  • Use __errno_location() (or libc wrappers that call it) for errno on threaded code; the plain errno symbol can lag behind TLS on the main thread if mixed with direct static access.

  • Prefer one thread talking to the WM v1 client for a window unless you document shared access.

Implementation map

Piece

Code (approx.)

TCB, preempt, FS save/restore

src/kernel/task.rs — TaskControlBlock::tls_fsbase, schedule_next_task, user_fsbase_allowed_for_tls

ARCH_SET_FS

src/kernel/syscalls.rs — syscall 158

VA / mmap sharing

src/kernel/task.rs — spawn_thread, user-va-layout.md

VMM vs scheduler (debug)

src/mm/memory.rs — debug_assert_scheduler_unlocked_for_vmm

Pthread shim

SDK/lib/cstd/src/pthread.rs, sys::sys_clone, sys::sys_waitpid, sys::sys_arch_prctl_set_fs

IRQ → userspace (Phase 1)

sys_irq_subscribe (226) — src/kernel/irq_subscribe.rs, PS/2 drain src/drivers/input/ps2.rs, apps/ps2d

Known gaps (today)

  • VFS / PTY / IPC / compositor: lock order is documented here but not fully enforced with mutex layering or static checks beyond the VMM entry assert above.

  • errno / %fs: implemented for P1Start userland; cstd syscall error paths write through __errno_location. Direct reads of the legacy errno static are discouraged in threaded programs.

  • Stress harness: apps/stress_threads — multi-threaded VFS + PTY + IPC + mutex load (see its main.rs).