PTY v1 — design & implementation plan¶
Status (Apr 2026): kernel A.1 + A.2, Phase B, and Phase D (userspace I/O cleanup) are implemented for the shipping terminal path. A.3 (per-PTY foreground_pgid / sys_signal_foreground scoped to the right PTY) and Phase C (WM-driven TIOCSWINSZ + SIGWINCH end-to-end) remain planned / partial (winsize ioctl exists; full resize signal story is not closed in this doc’s “done” sense).
Kernel (done): refcount/teardown/hup; PtyTable::read_waiters (BTreeMap<(pty_id, PtyReadEdge), VecDeque<task_id>>); blocking sys_read + IoWait::PtyRead on TCB; VfsReadResult::PtyWouldBlock; PTY write to hung-up peer → -EPIPE in RAX with cstd errno mapping; compositor PTY_PAIRS_ALIVE= telemetry.
Read flags (done, v1): sys_read fourth argument r10: bit 0 = READ_FLAG_NONBLOCK_PTY — PTY master/slave reads on an empty queue return 0 immediately instead of blocking. Exposed in SDK/lib/cstd/src/sys.rs as sys_read_with_flags. Intended for gterm’s IPC-driven main loop (poll master without stalling the WM client).
Userspace (done): dps uses blocking sys_read on stdin (PTY slave) for line input; gterm drains the PTY master with sys_read_with_flags(..., READ_FLAG_NONBLOCK_PTY) in a tight loop. gterm applies fold_pty_bytes_for_display to every completed scroll line so CSI / OSC / SS3 / stray ESC never reach the glyph rasterizer (see §9). cstd documents SYS_OPENPTY (220), SYS_PTY_REBIND_STD (221), SYS_IOCTL (16), SYS_SIGNAL_FOREGROUND (223), SYS_IPC_UNMAP_PRESENT_SLOT (224), SYS_CPU_CORE_LOAD (225) — last is compositor HUD, not PTY, but shares the same syscall table.
Audience: kernel + apps/termd (gterm) + apps/dps + compositor (deskgui).
Related: terminal-stack.md (byte spine + ownership + display sanitizer), system-architecture.md §7 (shell gaps), wm-v1-platform-roadmap.md (WM present stride, task manager, HUD).
0. Why this doc exists¶
You already have kernel byte queues (src/fs/pty.rs), pair allocation (VFS::alloc_pty), sys_openpty / sys_pty_rebind_std, ioctl winsize (TIOCGWINSZ / TIOCSWINSZ), and sys_signal_foreground. That is not “no PTY” — it is a minimal half-PTY.
The pain in daily use comes from everything around that minimum:
Gap |
Symptom |
|---|---|
Non-blocking semantics everywhere |
|
No wake / wait queue |
Writer and reader are not paired in the scheduler; CPU burn or lost latency guarantees. |
Winsize not wired to geometry |
|
No |
POSIX shells and curses-ish things expect resize signal, not only ioctl query. |
|
Ctrl+C and job control cannot be made POSIX-correct without session / pgroup binding to a specific PTY slave. |
PTY lifetime on |
|
Stdio rebind is a hammer |
|
PTY v1 = keep the good parts of the current design, then close the semantic gaps so terminals behave like Unix, not like a demo wired with busy-loops.
Non-goals for v1: full termios(3), canonical mode in kernel, hardware console mux, /dev/pts/ namespaced device nodes (can be v2), Linux-compatible job control everywhere.
1. Target architecture (what “done” looks like)¶
1.1 Objects¶
PtyPair(kernel) — one per logical terminal session.Byte paths (unchanged meaning):
in_q: master write → slave read;out_q: slave write → master read.New:
refcount— number of global fds (master + slave opens + dups) pointing at this pair. Whenrefcount == 0, drop queues and id.New:
session_leader(optional v1.0) — PID of the controlling session leader (usually dps); used forSIGWINCHand later job control.New:
foreground_pgid— process group receiving SIGINT from line discipline /sys_signal_foregroundscoped to this PTY (replaces or narrows the globalFOREGROUND_PIDfor PTY-backed shells).Existing:
ws_*— authoritative winsize;TIOCSWINSZupdates it;TIOCGWINSZreads it.
Fds (VFS) — unchanged shape: each open of master or slave is a
FileDescriptorwith(pty_id, pty_is_master). Dup increments refcount on the samepty_id.Userspace roles (unchanged philosophy from terminal-stack):
gterm — owns master fd; never parses ANSI for display policy; pushes keys to master; reads master for output bytes; owns WM surface.
dps — holds slave as stdin/stdout after exec; is the session leader candidate for signals driven from the PTY.
1.2 syscall / libc surface (v1)¶
Mechanism |
Purpose |
|---|---|
|
Unchanged contract: allocate pair, return two local fds. |
|
Keep for exec of child shell: bind 0,1,2 to slave in that task only. |
|
Blocking on PTY (and pipes) when the queue is empty unless |
|
Extend beyond winsize when needed ( |
|
On write to a hung-up peer (slave gone for master→ |
|
Change semantics: either (A) take implicit “current task’s controlling PTY” from slave fd used for stdin, or (B) new arg |
New optional: |
Until you have a unified poll, a minimal “block current task until fd readable” syscall is acceptable for v1 if it is scoped and bounded. Partial substitute: non-blocking PTY master read in gterm + blocking slave read in dps. |
2. Phased delivery (no cope)¶
Phase A — Correctness & lifetime (kernel)¶
Implementation note (Apr 2026): refcount lives on each VFS FileDescriptor row (ref_count), not on PtyPair. sys_pty_rebind_std bumps the slave row; spawn_task / spawn_thread bump each inherited global fd. PtyPair carries master_hup / slave_hup; the pair map entry is removed when both ends have fully closed.
A.1 Refcount + teardown
[x] Implemented: per–global-fd
ref_countonopen_filesrows; lastcloseremoves row and runs PTY endpoint teardown; pair removed when both ends gone.On
alloc_pty: each new master/slave row starts withref_count = 1(dup/fork paths bump as needed).On dup (if you implement dup): increment.
On
closeof any fd that referencespty_id: decrement; when zero:Remove
PtyPairfromPtyTable.Optionally wake any tasks blocked on that pair with EOF semantics (return
0permanently, or-1with a distinct errno story — pick one and match libc expectations later).
A.2 Close semantics (POSIX-ish)
[x] Implemented:
slave_hup/master_hup;write_pairreturns failure when writing to a hung-up peer; master sees persistent0on read after slave gone andout_qdrained.When slave last closes: master reads should see EOF (return
0after drainingout_q, then persistent0or aslave_goneflag).When master last closes: slave writes can fail or count as SIGPIPE later — v1 can do “return error from write” without full POSIX signal on write if that is too heavy.
A.3 Global FOREGROUND_PID → per-PTY foreground
Store
foreground_pgidonPtyPair(orforeground_pidif you do not have pgroups yet).sys_signal_foreground: resolve caller’s stdin → slave global fd →pty_id→ deliver signal toforeground_pgidonly.gterm / dps contract: on each exec of a new foreground command, kernel or parent sets foreground (document in Phase D / shell spawn path once A.3 fields exist).
Phase B — Blocking & wake (kernel + scheduler)¶
Implementation note (Apr 2026): PtyTable holds read_waiters: BTreeMap<(u32, PtyReadEdge), VecDeque<usize>> (FIFO task ids, no fixed cap). read_pair / write_pair / on_pty_endpoint_closed own enqueue/drain; wake_tasks_ready runs after VFS unlock (scheduler first, then optional prune under VFS). TaskControlBlock::io_wait = IoWait::PtyRead { pty_id, edge }; timer wake skips io_wait tasks. sys_read loops on VfsReadResult::PtyWouldBlock. Hung-up PTY sys_write → RAX = (-32i64) as u64; SDK/lib/cstd write sets errno from -RAX when RAX, interpreted as i64, is less than -1.
B.1 Wait queues
[x] Implemented:
BTreeMapregistry +VecDequeper(pty_id, edge)under theVFSlock.readon empty queue: mark task blocked, switch — do not return0unless true EOF orr10 & READ_FLAG_NONBLOCK_PTYon a PTY fd (returns0immediately). This is not fullO_NONBLOCKon all file types; it is a deliberate v1 escape hatch for the compositor client loop.writeon full queue: v1 can keep “drop oldest” or block writer; pick one and document. Blocking writer + bounded queue is closer to Linux.
B.2 Wake rules
[x] Implemented: after
write_pairadds bytes, pop one FIFO waiter on the peer read edge;close/ partial hang-up drains all relevant waiters.
B.3 Integration point
[x]
VirtualFileSystem::read→VfsReadResult;sys_readretry loop;write→VfsWriteResult+wake_tasks_readyon success.
Phase C — Winsize + SIGWINCH (kernel + compositor + gterm)¶
C.1 TIOCSWINSZ side effects
After successful
TIOCSWINSZon slave (or master — pick one canonical fd for ioctl; Linux uses slave for controlling tty): deliverSIGWINCHto:session leader if set, else foreground_pgid, else all tasks sharing that slave stdio (worst fallback — avoid if possible).
C.2 WM → geometry
Implement
WIRE_WM_EVENT_CLIENT_BOUNDS(37) (already in roadmap): compositor sends client-area width/height in cells (gterm derives cols/rows from font metrics) or in pixels + gterm converts.gterm on receive:
ioctl(slave_or_master, TIOCSWINSZ, &ws)withws_xpixel/ws_ypixelfilled from framebuffer if useful.
C.3 dps
On startup and after each
SIGWINCH, callTIOCGWINSZand reflow prompt /readlinewidth if you have horizontal line editing.
Phase D — Userspace cleanup¶
D.1 gterm (apps/termd/src/main.rs) — [x] Apr 2026
[x] PTY master read:
poll_dps_outputloopssys_read_with_flags(master, buf, len, READ_FLAG_NONBLOCK_PTY)so drainingout_qnever blocks the WM/IPC thread; when the queue is empty,readreturns0and the loop exits.[x] Idle behavior: after draining PTY output and handling IPC, gterm yields / short sleeps where needed so the system stays responsive (exact policy lives in
main.rs).[x] Scroll lines:
append_pty_master_bytesbuffers until\n, strips trailing\r, thenfold_pty_bytes_for_displayon every pushed line so escape sequences and C0 controls never hitdraw_charas raw bytes (prevents column drift and “rainbow” garbage from mis-parsed CSI).[x] WM present metadata: client buffer may be padded to a capability size (
FB_CAP_W×FB_CAP_H).present_wm_v1sendsstride_px = FB_CAP_Wandbuffer_bytes = stride * cap_h * 4so the compositor’sdraw_external_frameuses the correct source row stride (seewm-v1-platform-roadmap.md§6b).
D.2 dps (apps/dps/src/main.rs) — [x] Apr 2026
[x] Stdin: blocking
sys_read(0, …)into a chunk buffer; bytes are queued intostdin_pendingfor the readline state machine.[x] EOF vs idle:
read == 0: ifTIOCGWINSZon fd 0 succeeds, treat as PTY EOF andSYS_EXIT; otherwise treat as non-PTY empty andsys_sleep(1)(legacy pipe / bootstrap — avoids a tight spin).[ ]
SIGWINCHhandler (optional proof): still a Phase C / polish item unless already wired.
D.3 libc (cstd) — [x] documented / partial wrappers
[x]
sys_read_with_flags,READ_FLAG_NONBLOCK_PTY, stableSYS_*constants inSDK/lib/cstd/src/sys.rs.[x]
sys_openpty/sys_pty_rebind_std/sys_ioctlas used by gterm + dps.[ ] Full
poll/select/sigactionsurface — still out of scope for v1 unless added deliberately.
3. Data structures (concrete)¶
3.1 PtyPair extensions (src/fs/pty.rs)¶
PtyPair {
in_q, out_q, // existing
ws_row, ws_col, ws_xpixel, ws_ypixel,
refcount: u32, // NEW
session_leader: u64, // NEW; 0 = unset
foreground_pgid: u64, // NEW; 0 = unset (or use pid until pgroups exist)
slave_closed: bool, // NEW — after last slave fd closes
master_closed: bool, // NEW
// wait lists: either intrusive linked list of task ids or fixed-cap arrays
wait_master_read: [...], // NEW — tasks blocked reading master (want out_q)
wait_slave_read: [...], // NEW — tasks blocked reading slave (want in_q)
...
}
Keep lock ordering documented: VFS lock vs scheduler lock — avoid deadlock when waking from write_pair (typically: take scheduler lock only to mark runnable after dropping VFS lock, or use a “pending wake” queue).
3.2 FileDescriptor (src/fs/vfs.rs)¶
Already carries
pty_id,pty_is_master. Addpty_open_cookieif you need generation counters to detect stale wakeups after free — optional.
4. ioctl roadmap (after winsize)¶
Request |
When |
|---|---|
|
Fields + ioctl path already exist — finish call-site wiring (gterm/dps) anytime; |
|
Nice with blocking reads; small, usually after B. |
|
v2 unless you need raw mode for a specific app. |
|
v2 with |
5. Testing & “done” criteria¶
Two dock terminals: independent stdin; typing in A never appears in B; closing A does not EOF B’s slave unless you explicitly share (you should not).
Resize: shrink / grow WM window → gterm updates winsize → dps sees new cols (log or prompt wrap).
SIGWINCH: handler in test binary prints once per resize (prove delivery).
Blocking read: dps blocks on
read(0)with zero CPU spin in kernel idle metrics (serial marker or QEMU icount if you use it).Close graph: close master first / slave first / both orders — no leak of
PtyPairin a debug counter; no UAF in wait queues.Ctrl+C: with two terminals, signal only foreground of the PTY whose slave is stdin for the focused gterm’s child — not the other terminal’s shell.
Multi-terminal load (ship bar): open four dock terminals and type in all of them at once — there must be no noticeable input latency and no sustained CPU churn in the kernel idle path (no spin-wait “fake blocking”; idle stays idle under simultaneous keystrokes). This is the bar for “terminals that don’t feel like a science project.”
Phase B dual-terminal soak (manual): open two dock terminals; type heavily in both at the same time for ~30s. Pass: top-bar
PTY_PAIRS_ALIVEstays equal to the number of open PTY sessions (returns to baseline after closing both); host/QEMU idle CPU does not ramp to a sustained spin (contrast with pre–Phase-B busy-waitread).
6. Risks & explicit trade-offs¶
Blocking read without a timeout or poll can make shutdown harder — ensure
closefrom another thread/task wakes blocked readers (kernel should abort wait with EOF).SIGWINCH(Phase C) before per-PTY foreground (A.3) is solid = risk of spurious or wrong-target delivery — ship A.3 before aggressiveSIGWINCHfanout; implement delivery to one well-identified PID first, then generalize.Per-PTY foreground interacts with
sys_execveand orphaning — document what happens toforeground_pgidwhen child exits (kernel sets to 0 or parent shell pid).
7. Suggested implementation order (execution)¶
Same as §2.0 — repeated here for skimmers:
A.1 + A.2 — refcount, teardown, EOF — bulletproof before anything else.
Phase B — blocking read/write + wake — terminal feel.
A.3 — per-PTY foreground +
sys_signal_foreground— right Ctrl+C.Phase C — WM bounds,
TIOCSWINSZplumbing,SIGWINCH— polish vs 1–3, still required for TUIs.Phase D — gterm/dps/
cstd— drop spin-loops once the kernel lies truthfully.
8. Doc ownership¶
When code lands, update in the same PR:
This file (checkboxes / status).
terminal-stack.md§1a “reality check” + §2/§3 (ownership vs display folding — must matchfold_pty_bytes_for_display).wm-v1-platform-roadmap.mdwhen present stride, IPC opcodes, or compositor HUD syscalls change.
9. gterm display path (holistic, not VT emulation)¶
This section exists so nobody confuses “no full ANSI/VT emulator in gterm” with “raw bytes go straight to the framebuffer.”
What gterm is: a line-oriented viewer: scroll buffer of completed lines + one tail line (incomplete PTY read up to the next \n), WM v1 present, PS/2-style keyboard → bytes on the PTY master.
What gterm is not: it does not interpret cursor motion, SGR colors in the bitmap, alternate screen, mouse reporting, etc. There is no scroll region / DECTCEM stack.
What we still do in software (fold_pty_bytes_for_display):
CSI
ESC [ … final— scan with a tentative index: only bytes0x20..=0x3fthen one0x40..=0x7e. If the sequence is malformed or interrupted by a C0 (e.g.\nslipped inside), onlyESCis skipped and[+ remainder are left for normal handling so newlines are not swallowed inside a bogus CSI.SS3
ESC O+ one final0x40..=0x7e(cursor/function keys from gterm’s own key mapping).OSC
ESC ] … BELorESC ] … ESC \\.Lone
ESC: skip one byte (avoid leakingESCas a glyph).BS / DEL: pop last display byte from the folded output buffer.
TAB → space;
\r/\n: dropped in the fold pass (line assembly already split on\n; trailing\rstripped before fold).
Rendering: draw_string iterates chars(), skips control characters, maps non-ASCII to ?, advances one column per displayed character. wrapped_row_ranges uses the same column model so wrapping and blitting stay aligned (UTF-8 in prompts/paths no longer desyncs cx from glyph count).
Color heuristics: line_color_for_draw still keys off substring markers like \x1B[32m in stored Strings — after folding, those substrings are usually absent from scroll lines. Treat row coloring as best-effort until a dedicated style model exists.
10. “Done” criteria — addendum for Phase D¶
Extend §5 with:
gterm IPC loop: with output flowing,
poll_dps_outputmust not wedge waiting on the PTY master; verified by code path usingREAD_FLAG_NONBLOCK_PTY.dps: shell blocks on stdin with near-zero spin when idle;
read(0)==0on a PTY exits the process.Visual: no stray multicolor columns or pre-title-bar noise after
ls, plain Enter, or burst output — regression guard for fold + stride + blit alignment.
If this doc disagrees with src/fs/pty.rs, src/fs/vfs.rs, apps/termd/src/main.rs, or SDK/lib/cstd/src/sys.rs, the repo wins — update this file in one pass.