Threading model (kernel + userland)¶
This page is the stable contract for threads on P1Start: what the scheduler guarantees, how TLS works, how kernel stacks are assigned, and lock ordering so subsystems stay composable. Pair it with User virtual address layout (x86_64) for address-space rules.
Note
Canonical threading/scheduler/process docs: Scheduler and SMP, Process lifecycle — birth, execution, death, Memory and address spaces (sources in newdocs/).
Desktop boot today runs single-core scheduler bring-up (APs not started in start_desktop); the invariants below still hold, but only the BSP runs user tasks until SMP is re-enabled (see scheduler page § boot profile).
Threads vs processes¶
A process here is a task with its own CR3 (page tables),
brk_current/mmap_current,fd_table, and VMM metadata.A thread is a separate task id that shares CR3 with siblings (same user address space). Created via
spawn_thread(kernel) / the user clone-style syscall path. Threads share the mmap cursor and file descriptor table semantics documented in the VA layout page.
Until stress tests pass, treat multi-threaded user code as beta: use cstd::pthread (below) or raw sys_clone.
Current status (what you can use today)¶
API |
Status |
|---|---|
|
Stable kernel ABI; |
|
Blocks until target tid exits; kernel yields after marking |
|
|
|
Rust |
Scheduler invariants¶
Invariant |
Meaning |
|---|---|
One current task per logical CPU |
|
Preempt path |
Timer IRQ ( |
Saved GPR / RSP |
Outgoing task’s RSP (kernel stack position) is stored in its TCB before another task runs. |
FPU |
|
CR3 |
Switched when the next task’s CR3 differs from the current (process switch). Threads sharing an mm keep the same CR3. |
TLS (FS base) |
|
Kernel stack (RSP0) |
Each runnable task has |
Scheduler lock vs allocation |
Do not hold |
Zombie reaping |
Only BSP (core 0) runs periodic |
TLS (ARCH_SET_FS)¶
Syscall 158 (
sys_arch_prctl):rdi = 0x1002(ARCH_SET_FS),rsi =user virtual base for%fs-relative TLS (musl / typical Linux user ABI).The kernel:
Validates the pointer with
user_fsbase_allowed_for_tls(canonical low user VA, including P1’s mmap band — same upper bound asMAX_USER_CANONICAL_LOW_VADDRintask.rs).Stores it in the current task’s
TaskControlBlock::tls_fsbase.Writes
MSR_FS_BASEimmediately so the running thread sees the update.
New tasks and new kernel threads start with
tls_fsbase = 0until user code sets TLS.Context switch (
schedule_next_task): outgoing task’s liveFsBase::read()is copied intotls_fsbase; incoming task’stls_fsbaseis written withFsBase::write. If a TCB ever contained a bad value, the switch path forces 0 (fail-safe).
Kernel stack per thread¶
spawn_task(new process) andspawn_threadeach allocate a dedicated 64 KiB kernel stack buffer and setkernel_stack_topto the high end for that task’s ring-0 work.Do not reuse another task’s kernel stack for a concurrently runnable thread.
Lock ordering (do not invert)¶
These rules prevent deadlocks with timer-driven preempt and syscall paths:
SCHEDULER— only for short TCB updates and run-queue decisions. Never takeVFSor allocate VMM pages while holding it unless the callee is proven non-blocking.VFS— may be taken from syscalls after dropping the scheduler lock. Never acquireSCHEDULERwhile holdingVFS.PTY / session — expect a single global
VFSlock today; session foreground andio_waitwakeups must stay consistent with task id (tid), not only “process” id, when multiple threads share an address space.Compositor / WM IPC — treat per-window state as needing either one logical UI thread in userland or explicit kernel-side serialization; concurrent
presentfrom two threads without protocol is undefined until a queue or mutex is specified.
Mnemonic: scheduler first and alone for scheduling; VFS and page tables afterward.
Enforcement (debug builds)¶
In cfg(debug_assertions), vmm_map_page, vmm_unmap_user_pages, and vmm_unmap_user_pages_free_phys (src/mm/memory.rs) call debug_assert_scheduler_unlocked_for_vmm: the scheduler mutex must be free before page-table mutation, matching the rule above. Release builds skip the check for speed.
Userland guidance¶
Prefer
cstd::pthreadover ad-hoc syscall numbers for new code.Use
__errno_location()(or libc wrappers that call it) for errno on threaded code; the plainerrnosymbol can lag behind TLS on the main thread if mixed with direct static access.Prefer one thread talking to the WM v1 client for a window unless you document shared access.
Implementation map¶
Piece |
Code (approx.) |
|---|---|
TCB, preempt, FS save/restore |
|
|
|
VA / mmap sharing |
|
VMM vs scheduler (debug) |
|
Pthread shim |
|
IRQ → userspace (Phase 1) |
|
Known gaps (today)¶
VFS / PTY / IPC / compositor: lock order is documented here but not fully enforced with mutex layering or static checks beyond the VMM entry assert above.
errno/%fs: implemented for P1Start userland;cstdsyscall error paths write through__errno_location. Direct reads of the legacyerrnostatic are discouraged in threaded programs.Stress harness:
apps/stress_threads— multi-threaded VFS + PTY + IPC + mutex load (see itsmain.rs).