Skip to content

Local "This device" backend drops under host memory pressure; needs WS keepalive + tolerant reconnect + backend heap headroom #21

Description

@eddy-curly

Summary

On a memory-constrained host, the desktop app's local ("This device") backend connection drops intermittently and is slow to recover, surfacing as <machine> is not connected. The backend process stays alive — it is the renderer↔backend WebSocket transport that drops and reconnection that lags. A device-side workaround (Electron background-throttling flags) reduces but does not fix it. The durable fix is app-side.

Environment (where observed)

  • OS: Windows 11 Pro
  • App: T3 Code (Alpha) 0.0.28, local build of this fork @ 4ab11db4, x64
  • Backend: local ("This device")
  • Host is memory-constrained: 28.5 GB total, ~4 GB free at times; node + claude + VS Code + Edge/WebView2 dominate. T3 Code itself is only ~1.4 GB — the pressure comes from other tooling.
  • Trigger: correlates with high host memory pressure ("when other apps take a lot of memory").

Observed behavior

  • "This device" shows disconnected / "reconnecting…"; recovery takes many seconds to minutes under load.
  • The backend does not crash: observed stable 2+ days, RSS ~400–575 MB, no OOM signature. The connection drops while the backend stays up.

Root cause (code, at 4ab11db4)

  1. No stall tolerance / liveness signal
    • apps/server/src/ws.ts — the WS server sets no idleTimeout / keep-alive / ping-pong.
    • packages/client-runtime/src/connection/ — no client heartbeat.
    • Result: a stalled socket is only detected on an actual transport close, never on silence.
  2. Slow recovery under load
    • packages/client-runtime/src/connection/supervisor.ts — each reconnect attempt has a 15s establishment + 15s probe budget (CONNECTION_ESTABLISHMENT_TIMEOUT / CONNECTION_PROBE_TIMEOUT = "15 seconds") with backoff. While the host thrashes, the backend answers too slowly and attempts keep timing out before they can succeed.
  3. No backend headroom
    • apps/desktop/src/backend/DesktopBackendManager.ts — the backend child is spawned with the default V8 heap (no --max-old-space-size / NODE_OPTIONS). The readiness check is non-fatal and the backend only restarts on an actual child exit.

Current device-side workaround (insufficient)

Added --disable-renderer-backgrounding --disable-background-timer-throttling --disable-backgrounding-occluded-windows to the app's launch shortcuts. This only stops Windows from throttling the backgrounded renderer; it adds no liveness detection, no faster recovery, and no backend headroom — and it is reset by every reinstall.

Proposed durable fixes

  • Add WS keep-alive / ping-pong + an idleTimeout on the server (ws.ts) and a client heartbeat in connection/, so a stalled socket is detected quickly and proactively.
  • Make reconnect more tolerant under load: probe fast to detect a truly dead socket, but retry sooner / use adaptive budgets so a briefly slow backend isn't repeatedly abandoned. Distinguish "slow" from "dead".
  • Give the backend heap headroom (--max-old-space-size) and/or a health-based restart so it isn't starved under host memory pressure.

Related upstream issues (pingdotgg/t3code)

Symptom-adjacent; likely share robustness gaps — worth cross-referencing / de-duping:

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions