PTY v1 — design & implementation plan

Status (Apr 2026): kernel A.1 + A.2, Phase B, and Phase D (userspace I/O cleanup) are implemented for the shipping terminal path. A.3 (per-PTY foreground_pgid / sys_signal_foreground scoped to the right PTY) and Phase C (WM-driven TIOCSWINSZ + SIGWINCH end-to-end) remain planned / partial (winsize ioctl exists; full resize signal story is not closed in this doc’s “done” sense).

Kernel (done): refcount/teardown/hup; PtyTable::read_waiters (BTreeMap<(pty_id, PtyReadEdge), VecDeque<task_id>>); blocking sys_read + IoWait::PtyRead on TCB; VfsReadResult::PtyWouldBlock; PTY write to hung-up peer → -EPIPE in RAX with cstd errno mapping; compositor PTY_PAIRS_ALIVE= telemetry.

Read flags (done, v1): sys_read fourth argument r10: bit 0 = READ_FLAG_NONBLOCK_PTY — PTY master/slave reads on an empty queue return 0 immediately instead of blocking. Exposed in SDK/lib/cstd/src/sys.rs as sys_read_with_flags. Intended for gterm’s IPC-driven main loop (poll master without stalling the WM client).

Userspace (done): dps uses blocking sys_read on stdin (PTY slave) for line input; gterm drains the PTY master with sys_read_with_flags(..., READ_FLAG_NONBLOCK_PTY) in a tight loop. gterm applies fold_pty_bytes_for_display to every completed scroll line so CSI / OSC / SS3 / stray ESC never reach the glyph rasterizer (see §9). cstd documents SYS_OPENPTY (220), SYS_PTY_REBIND_STD (221), SYS_IOCTL (16), SYS_SIGNAL_FOREGROUND (223), SYS_IPC_UNMAP_PRESENT_SLOT (224), SYS_CPU_CORE_LOAD (225) — last is compositor HUD, not PTY, but shares the same syscall table.

Audience: kernel + apps/termd (gterm) + apps/dps + compositor (deskgui).
Related: terminal-stack.md (byte spine + ownership + display sanitizer), system-architecture.md §7 (shell gaps), wm-v1-platform-roadmap.md (WM present stride, task manager, HUD).

0. Why this doc exists

You already have kernel byte queues (src/fs/pty.rs), pair allocation (VFS::alloc_pty), sys_openpty / sys_pty_rebind_std, ioctl winsize (TIOCGWINSZ / TIOCSWINSZ), and sys_signal_foreground. That is not “no PTY” — it is a minimal half-PTY.

The pain in daily use comes from everything around that minimum:

Gap

Symptom

Non-blocking semantics everywhere

read returns 0 on empty; userland spins or sleeps; “fake blocking” in loops.

No wake / wait queue

Writer and reader are not paired in the scheduler; CPU burn or lost latency guarantees.

Winsize not wired to geometry

PtyPair has ws_row / ws_col, but resize does not reliably flow WM client rect → gterm → TIOCSWINSZ → dps.

No SIGWINCH

POSIX shells and curses-ish things expect resize signal, not only ioctl query.

FOREGROUND_PID is global, not per-session

Ctrl+C and job control cannot be made POSIX-correct without session / pgroup binding to a specific PTY slave.

PTY lifetime on close

VFS::close drops FileDescriptor rows but does not obviously tear down PtyPair or refcount master/slave; risk of leaks or stale pty_id.

Stdio rebind is a hammer

sys_pty_rebind_std maps 0,1,2 to the same global slave fd — correct bootstrap, wrong for “multiple sessions / dup2 stories” long term.

PTY v1 = keep the good parts of the current design, then close the semantic gaps so terminals behave like Unix, not like a demo wired with busy-loops.

Non-goals for v1: full termios(3), canonical mode in kernel, hardware console mux, /dev/pts/ namespaced device nodes (can be v2), Linux-compatible job control everywhere.

1. Target architecture (what “done” looks like)

1.1 Objects

  1. PtyPair (kernel) — one per logical terminal session.

    • Byte paths (unchanged meaning): in_q: master write → slave read; out_q: slave write → master read.

    • New: refcount — number of global fds (master + slave opens + dups) pointing at this pair. When refcount == 0, drop queues and id.

    • New: session_leader (optional v1.0) — PID of the controlling session leader (usually dps); used for SIGWINCH and later job control.

    • New: foreground_pgid — process group receiving SIGINT from line discipline / sys_signal_foreground scoped to this PTY (replaces or narrows the global FOREGROUND_PID for PTY-backed shells).

    • Existing: ws_* — authoritative winsize; TIOCSWINSZ updates it; TIOCGWINSZ reads it.

  2. Fds (VFS) — unchanged shape: each open of master or slave is a FileDescriptor with (pty_id, pty_is_master). Dup increments refcount on the same pty_id.

  3. Userspace roles (unchanged philosophy from terminal-stack):

    • gterm — owns master fd; never parses ANSI for display policy; pushes keys to master; reads master for output bytes; owns WM surface.

    • dps — holds slave as stdin/stdout after exec; is the session leader candidate for signals driven from the PTY.

1.2 syscall / libc surface (v1)

Mechanism

Purpose

sys_openpty

Unchanged contract: allocate pair, return two local fds.

sys_pty_rebind_std

Keep for exec of child shell: bind 0,1,2 to slave in that task only.

sys_read

Blocking on PTY (and pipes) when the queue is empty unless r10 & READ_FLAG_NONBLOCK_PTY — then return 0 on empty PTY read (same “no data” shape as non-blocking POSIX, but PTY-scoped for v1). cstd: sys_read_with_flags(fd, buf, len, flags).

sys_ioctl

Extend beyond winsize when needed (FIONREAD optional; no giant termios in v1 unless you bite it off deliberately). TIOCGWINSZ / TIOCSWINSZ used from dps / gterm paths where wired.

sys_write (PTY)

On write to a hung-up peer (slave gone for master→in_q, or master gone for slave→out_q), return -32 (-EPIPE) in RAX (not !0); cstd write sets errno = EPIPE. Generic write errors still use RAX = !0, errno = EIO.

sys_signal_foreground

Change semantics: either (A) take implicit “current task’s controlling PTY” from slave fd used for stdin, or (B) new arg pty_id / slave_fd — must not be one global PID for multi-terminal correctness. Today: still global foreground PID until A.3 lands.

New optional: sys_wait_fd or integrate into sys_poll_ui_events

Until you have a unified poll, a minimal “block current task until fd readable” syscall is acceptable for v1 if it is scoped and bounded. Partial substitute: non-blocking PTY master read in gterm + blocking slave read in dps.

2. Phased delivery (no cope)

2.0 Execution order (authoritative)

Do not reorder this lightly. Earlier phases are foundations for later ones; skipping A.1 + A.2 to chase “feel” or resize polish will cost weeks in UAFs and leaked PtyPairs.

Step

What

Why

1

A.1 + A.2 — refcount, teardown, EOF on close

Foundation. Wrong lifetime poisons every wait queue and every wake. Bulletproof this before blocking reads or signals touch the pair. Done (Apr 2026).

2

Phase B — wait queues, blocking read/write, wake rules

What makes the terminal actually pleasant (no spin-loops). Depends on clean close/teardown so blocked tasks are not dangling. Done (Apr 2026).

3

A.3 — per-PTY foreground_pgid / session binding + sys_signal_foreground fixed

Ctrl+C killing the wrong shell is unacceptable; this must not sit on a global PID once A+B are real.

4

Phase C — WM bounds → TIOCSWINSZ, SIGWINCH delivery

Polish vs B + A.3 — still required for real TUI apps, but secondary to “read blocks” and “signal hits the right process”.

5

Phase D — gterm/dps/cstd cleanup

Delete fake blocking in the hot path after kernel blocking exists. Done (Apr 2026) for gterm master poll + dps slave read; see §2 Phase D.

Phase A — Correctness & lifetime (kernel)

Implementation note (Apr 2026): refcount lives on each VFS FileDescriptor row (ref_count), not on PtyPair. sys_pty_rebind_std bumps the slave row; spawn_task / spawn_thread bump each inherited global fd. PtyPair carries master_hup / slave_hup; the pair map entry is removed when both ends have fully closed.

A.1 Refcount + teardown

  • [x] Implemented: per–global-fd ref_count on open_files rows; last close removes row and runs PTY endpoint teardown; pair removed when both ends gone.

  • On alloc_pty: each new master/slave row starts with ref_count = 1 (dup/fork paths bump as needed).

  • On dup (if you implement dup): increment.

  • On close of any fd that references pty_id: decrement; when zero:

    • Remove PtyPair from PtyTable.

    • Optionally wake any tasks blocked on that pair with EOF semantics (return 0 permanently, or -1 with a distinct errno story — pick one and match libc expectations later).

A.2 Close semantics (POSIX-ish)

  • [x] Implemented: slave_hup / master_hup; write_pair returns failure when writing to a hung-up peer; master sees persistent 0 on read after slave gone and out_q drained.

  • When slave last closes: master reads should see EOF (return 0 after draining out_q, then persistent 0 or a slave_gone flag).

  • When master last closes: slave writes can fail or count as SIGPIPE later — v1 can do “return error from write” without full POSIX signal on write if that is too heavy.

A.3 Global FOREGROUND_PID → per-PTY foreground

  • Store foreground_pgid on PtyPair (or foreground_pid if you do not have pgroups yet).

  • sys_signal_foreground: resolve caller’s stdin → slave global fd → pty_id → deliver signal to foreground_pgid only.

  • gterm / dps contract: on each exec of a new foreground command, kernel or parent sets foreground (document in Phase D / shell spawn path once A.3 fields exist).

Phase B — Blocking & wake (kernel + scheduler)

Implementation note (Apr 2026): PtyTable holds read_waiters: BTreeMap<(u32, PtyReadEdge), VecDeque<usize>> (FIFO task ids, no fixed cap). read_pair / write_pair / on_pty_endpoint_closed own enqueue/drain; wake_tasks_ready runs after VFS unlock (scheduler first, then optional prune under VFS). TaskControlBlock::io_wait = IoWait::PtyRead { pty_id, edge }; timer wake skips io_wait tasks. sys_read loops on VfsReadResult::PtyWouldBlock. Hung-up PTY sys_write → RAX = (-32i64) as u64; SDK/lib/cstd write sets errno from -RAX when RAX, interpreted as i64, is less than -1.

B.1 Wait queues

  • [x] Implemented: BTreeMap registry + VecDeque per (pty_id, edge) under the VFS lock.

  • read on empty queue: mark task blocked, switch — do not return 0 unless true EOF or r10 & READ_FLAG_NONBLOCK_PTY on a PTY fd (returns 0 immediately). This is not full O_NONBLOCK on all file types; it is a deliberate v1 escape hatch for the compositor client loop.

  • write on full queue: v1 can keep “drop oldest” or block writer; pick one and document. Blocking writer + bounded queue is closer to Linux.

B.2 Wake rules

  • [x] Implemented: after write_pair adds bytes, pop one FIFO waiter on the peer read edge; close / partial hang-up drains all relevant waiters.

B.3 Integration point

  • [x] VirtualFileSystem::read → VfsReadResult; sys_read retry loop; write → VfsWriteResult + wake_tasks_ready on success.

Phase C — Winsize + SIGWINCH (kernel + compositor + gterm)

C.1 TIOCSWINSZ side effects

  • After successful TIOCSWINSZ on slave (or master — pick one canonical fd for ioctl; Linux uses slave for controlling tty): deliver SIGWINCH to:

    • session leader if set, else foreground_pgid, else all tasks sharing that slave stdio (worst fallback — avoid if possible).

C.2 WM → geometry

  • Implement WIRE_WM_EVENT_CLIENT_BOUNDS (37) (already in roadmap): compositor sends client-area width/height in cells (gterm derives cols/rows from font metrics) or in pixels + gterm converts.

  • gterm on receive: ioctl(slave_or_master, TIOCSWINSZ, &ws) with ws_xpixel / ws_ypixel filled from framebuffer if useful.

C.3 dps

  • On startup and after each SIGWINCH, call TIOCGWINSZ and reflow prompt / readline width if you have horizontal line editing.

Phase D — Userspace cleanup

D.1 gterm (apps/termd/src/main.rs) — [x] Apr 2026

  • [x] PTY master read: poll_dps_output loops sys_read_with_flags(master, buf, len, READ_FLAG_NONBLOCK_PTY) so draining out_q never blocks the WM/IPC thread; when the queue is empty, read returns 0 and the loop exits.

  • [x] Idle behavior: after draining PTY output and handling IPC, gterm yields / short sleeps where needed so the system stays responsive (exact policy lives in main.rs).

  • [x] Scroll lines: append_pty_master_bytes buffers until \n, strips trailing \r, then fold_pty_bytes_for_display on every pushed line so escape sequences and C0 controls never hit draw_char as raw bytes (prevents column drift and “rainbow” garbage from mis-parsed CSI).

  • [x] WM present metadata: client buffer may be padded to a capability size (FB_CAP_W × FB_CAP_H). present_wm_v1 sends stride_px = FB_CAP_W and buffer_bytes = stride * cap_h * 4 so the compositor’s draw_external_frame uses the correct source row stride (see wm-v1-platform-roadmap.md §6b).

D.2 dps (apps/dps/src/main.rs) — [x] Apr 2026

  • [x] Stdin: blocking sys_read(0, …) into a chunk buffer; bytes are queued into stdin_pending for the readline state machine.

  • [x] EOF vs idle: read == 0: if TIOCGWINSZ on fd 0 succeeds, treat as PTY EOF and SYS_EXIT; otherwise treat as non-PTY empty and sys_sleep(1) (legacy pipe / bootstrap — avoids a tight spin).

  • [ ] SIGWINCH handler (optional proof): still a Phase C / polish item unless already wired.

D.3 libc (cstd) — [x] documented / partial wrappers

  • [x] sys_read_with_flags, READ_FLAG_NONBLOCK_PTY, stable SYS_* constants in SDK/lib/cstd/src/sys.rs.

  • [x] sys_openpty / sys_pty_rebind_std / sys_ioctl as used by gterm + dps.

  • [ ] Full poll / select / sigaction surface — still out of scope for v1 unless added deliberately.

3. Data structures (concrete)

3.1 PtyPair extensions (src/fs/pty.rs)

PtyPair {
    in_q, out_q,                    // existing
    ws_row, ws_col, ws_xpixel, ws_ypixel,
    refcount: u32,                  // NEW
    session_leader: u64,            // NEW; 0 = unset
    foreground_pgid: u64,           // NEW; 0 = unset (or use pid until pgroups exist)
    slave_closed: bool,             // NEW — after last slave fd closes
    master_closed: bool,            // NEW
    // wait lists: either intrusive linked list of task ids or fixed-cap arrays
    wait_master_read: [...],       // NEW — tasks blocked reading master (want out_q)
    wait_slave_read: [...],        // NEW — tasks blocked reading slave (want in_q)
    ...
}

Keep lock ordering documented: VFS lock vs scheduler lock — avoid deadlock when waking from write_pair (typically: take scheduler lock only to mark runnable after dropping VFS lock, or use a “pending wake” queue).

3.2 FileDescriptor (src/fs/vfs.rs)

  • Already carries pty_id, pty_is_master. Add pty_open_cookie if you need generation counters to detect stale wakeups after free — optional.

4. ioctl roadmap (after winsize)

Request

When

TIOCGWINSZ / TIOCSWINSZ

Fields + ioctl path already exist — finish call-site wiring (gterm/dps) anytime; SIGWINCH on set and WM-driven resize stay in Phase C (step 4), after A.1+A.2, B, and A.3.

FIONREAD / TIOCOUTQ

Nice with blocking reads; small, usually after B.

TCGETS / TCSETS

v2 unless you need raw mode for a specific app.

TIOCSCTTY / TIOCNOTTY

v2 with /dev/tty story.

5. Testing & “done” criteria

  1. Two dock terminals: independent stdin; typing in A never appears in B; closing A does not EOF B’s slave unless you explicitly share (you should not).

  2. Resize: shrink / grow WM window → gterm updates winsize → dps sees new cols (log or prompt wrap).

  3. SIGWINCH: handler in test binary prints once per resize (prove delivery).

  4. Blocking read: dps blocks on read(0) with zero CPU spin in kernel idle metrics (serial marker or QEMU icount if you use it).

  5. Close graph: close master first / slave first / both orders — no leak of PtyPair in a debug counter; no UAF in wait queues.

  6. Ctrl+C: with two terminals, signal only foreground of the PTY whose slave is stdin for the focused gterm’s child — not the other terminal’s shell.

  7. Multi-terminal load (ship bar): open four dock terminals and type in all of them at once — there must be no noticeable input latency and no sustained CPU churn in the kernel idle path (no spin-wait “fake blocking”; idle stays idle under simultaneous keystrokes). This is the bar for “terminals that don’t feel like a science project.”

  8. Phase B dual-terminal soak (manual): open two dock terminals; type heavily in both at the same time for ~30s. Pass: top-bar PTY_PAIRS_ALIVE stays equal to the number of open PTY sessions (returns to baseline after closing both); host/QEMU idle CPU does not ramp to a sustained spin (contrast with pre–Phase-B busy-wait read).

6. Risks & explicit trade-offs

  • Blocking read without a timeout or poll can make shutdown harder — ensure close from another thread/task wakes blocked readers (kernel should abort wait with EOF).

  • SIGWINCH (Phase C) before per-PTY foreground (A.3) is solid = risk of spurious or wrong-target delivery — ship A.3 before aggressive SIGWINCH fanout; implement delivery to one well-identified PID first, then generalize.

  • Per-PTY foreground interacts with sys_execve and orphaning — document what happens to foreground_pgid when child exits (kernel sets to 0 or parent shell pid).

7. Suggested implementation order (execution)

Same as §2.0 — repeated here for skimmers:

  1. A.1 + A.2 — refcount, teardown, EOF — bulletproof before anything else.

  2. Phase B — blocking read/write + wake — terminal feel.

  3. A.3 — per-PTY foreground + sys_signal_foreground — right Ctrl+C.

  4. Phase C — WM bounds, TIOCSWINSZ plumbing, SIGWINCH — polish vs 1–3, still required for TUIs.

  5. Phase D — gterm/dps/cstd — drop spin-loops once the kernel lies truthfully.

8. Doc ownership

When code lands, update in the same PR:

  • This file (checkboxes / status).

  • terminal-stack.md §1a “reality check” + §2/§3 (ownership vs display folding — must match fold_pty_bytes_for_display).

  • wm-v1-platform-roadmap.md when present stride, IPC opcodes, or compositor HUD syscalls change.

9. gterm display path (holistic, not VT emulation)

This section exists so nobody confuses “no full ANSI/VT emulator in gterm” with “raw bytes go straight to the framebuffer.”

What gterm is: a line-oriented viewer: scroll buffer of completed lines + one tail line (incomplete PTY read up to the next \n), WM v1 present, PS/2-style keyboard → bytes on the PTY master.

What gterm is not: it does not interpret cursor motion, SGR colors in the bitmap, alternate screen, mouse reporting, etc. There is no scroll region / DECTCEM stack.

What we still do in software (fold_pty_bytes_for_display):

  • CSI ESC [ … final — scan with a tentative index: only bytes 0x20..=0x3f then one 0x40..=0x7e. If the sequence is malformed or interrupted by a C0 (e.g. \n slipped inside), only ESC is skipped and [ + remainder are left for normal handling so newlines are not swallowed inside a bogus CSI.

  • SS3 ESC O + one final 0x40..=0x7e (cursor/function keys from gterm’s own key mapping).

  • OSC ESC ] … BEL or ESC ] … ESC \\.

  • Lone ESC: skip one byte (avoid leaking ESC as a glyph).

  • BS / DEL: pop last display byte from the folded output buffer.

  • TAB → space; \r / \n: dropped in the fold pass (line assembly already split on \n; trailing \r stripped before fold).

Rendering: draw_string iterates chars(), skips control characters, maps non-ASCII to ?, advances one column per displayed character. wrapped_row_ranges uses the same column model so wrapping and blitting stay aligned (UTF-8 in prompts/paths no longer desyncs cx from glyph count).

Color heuristics: line_color_for_draw still keys off substring markers like \x1B[32m in stored Strings — after folding, those substrings are usually absent from scroll lines. Treat row coloring as best-effort until a dedicated style model exists.

10. “Done” criteria — addendum for Phase D

Extend §5 with:

  1. gterm IPC loop: with output flowing, poll_dps_output must not wedge waiting on the PTY master; verified by code path using READ_FLAG_NONBLOCK_PTY.

  2. dps: shell blocks on stdin with near-zero spin when idle; read(0)==0 on a PTY exits the process.

  3. Visual: no stray multicolor columns or pre-title-bar noise after ls, plain Enter, or burst output — regression guard for fold + stride + blit alignment.

If this doc disagrees with src/fs/pty.rs, src/fs/vfs.rs, apps/termd/src/main.rs, or SDK/lib/cstd/src/sys.rs, the repo wins — update this file in one pass.