Skip to content

FAB-380: container lifecycle ops UI (instance + cluster), safe mode, resend invite#1429

Merged
DavidCockerill merged 8 commits into
stagefrom
FAB-380
Jul 16, 2026
Merged

FAB-380: container lifecycle ops UI (instance + cluster), safe mode, resend invite#1429
DavidCockerill merged 8 commits into
stagefrom
FAB-380

Conversation

@DavidCockerill

@DavidCockerill DavidCockerill commented Jul 8, 2026

Copy link
Copy Markdown
Member

Summary

Studio UI for the FAB-380 container lifecycle operations (stop / start / restart, with safe mode), plus a small resend-invite addition for pending org users. This is the front-end for the central-manager / host-manager work in HarperFast/central-manager#404 and HarperFast/host-manager#126.

Container ops are a distinct class from the existing proxied Harper "restart" — they're handled host-side via the CM, are async (the endpoint returns a transitional status and the resting state lands as the poll catches up), and support safeMode (boot Harper core without loading user apps/components).

What's included

Per-instance container ops (useInstanceMenuItems)

  • A Container section in the instance ⋯ menu — Start / Start in safe mode (when STOPPED), Restart / Restart in safe mode / Stop (when RUNNING). Hidden mid-transition.
  • New useInstanceContainerOps hook + instanceContainerOperation client (POST /HDBInstance/{id}/container/{action}).

Cluster-wide container ops (ClusterCard)

  • A Container section in the cluster ⋯ menu, gated by cluster state (RUNNING → Restart/Restart-safe/Stop, STOPPED → Start/Start-safe, PARTIAL → Start/Stop/Restart).
  • Safe-mode ops force strategy: parallel (the CM rejects safeMode + rolling). Stop shows a confirm dialog; non-safe Restart opens a rolling/parallel picker.
  • New useClusterContainerOps hook + clusterContainerOperation client (POST /Cluster/{id}/container/{action}), and ClusterContainerOpModals.

Cluster overview — discoverability (ClusterStateMenu)

  • A labeled "Cluster actions" dropdown (AWS EC2 "Instance state" style) on the cluster overview that surfaces every op — Start / Start in safe mode / Restart / Restart in safe mode / Stop — plus Terminate, in one place. Ops that don't apply to the current status are shown but disabled, so the feature set stays discoverable (feedback was that safe mode in particular was hard to find). Hidden for self-managed clusters.
  • Safe-mode explain-and-confirm dialog (SafeModeConfirmDialog): "safe mode" is jargon and a recovery action, so rather than a hover-only tooltip we explain what it does at the moment of use. Shared by the start-in-safe-mode and restart-in-safe-mode entries on both the overview and the card.
  • Transient op status on the clusters-list card: while an op is in flight the card shows a Stopping / Starting / Restarting badge (with spinner), and a Stopped label when the cluster is stopped — so users can see what's happening without opening the cluster. Clears on its own as the 10s poll catches the resting state.
  • Plain (non-safe) Start/Restart send safeMode: false so a cluster that's in safe mode returns to normal mode. (Previously the UI omitted safeMode, and the CM preserves the current mode when it's omitted — so a plain restart of a safe-mode cluster stayed in safe mode.)

Safe mode indicators

  • Per-instance amber "Safe mode" badge on the Instances page.
  • A "Safe mode" pill next to the status pill on the cluster overview when every instance is in safe mode.

Stopped-state handling

  • isStoppedOrTransitioning helper; stopped/transitioning instances stop get_status polling (the ops API is unreachable), show a muted dot instead of a stale "Available"/endless spinner, and a "—" in the sign-in cell.
  • STOPPED badge is now destructive (red); cluster status pill colors STOPPED red / PARTIAL yellow. Stopped cards open the instances page; partial cards open the overview.

Resend invite

  • A "Resend invite" action in the org users table, shown only for PENDING users. Reuses the existing invite mutation (POST /UserInvite with { email, roleId }); the CM treats a repeat invite for a pending user as a resend.

Screenshots

Cluster overview — Running + Safe mode pill

image

Instances — Safe mode badges

image

Per-instance ⋯ menu — Container section

image

Cluster ⋯ menu — Container section

image

Restart strategy picker (non-safe cluster restart)

image

Resend invite (pending users only)

image

Backend dependencies / notes

  • Depends on the CM/HM container endpoints (HarperFast/central-manager#404, HarperFast/host-manager#126). The container clients are hand-typed past the generated SDK — the /container/{action} paths and safeMode/strategy body aren't in the CM OpenAPI yet (same pattern as terminateCluster). Follow-up: regenerate the SDK once the CM OpenAPI exposes them and drop the casts.
  • Instance.safeMode and Cluster.instances (now the patched Instance[]) are added in api.patch.d.ts for the same reason.
  • Resend invite needs the CM invite-endpoint change deployed. On the dev CM a repeat invite for a pending user still returns 409 "User already has a pending invite"; the UI is correct and will work once that build is out.

Out of scope / follow-ups

  • Error toasts: op/terminate failures in this PR route through the global MutationCache.onError (which shows the server's message) with no local error toast, so there's no double-toasting. The same pre-existing local-onError double-toast pattern lives elsewhere in the app (e.g. ClusterCard terminate, org settings/roles/users, ClusterForm) — intentionally left for a separate repo-wide cleanup rather than widened here.
  • Disable-vs-hide: the "Cluster actions" dropdown disables inapplicable ops (better discoverability); the ⋯ menus hide them. Worth aligning on one convention in a follow-up.

Testing

  • tsc -p tsconfig.app.json, oxlint, dprint all pass; full unit suite green (pre-commit, 1239 tests).
  • Two gemini-code-assist review passes addressed — genuine findings fixed (double-toasts, [clusterId] invalidation guard, Cancel-button variants, focus-visible a11y ring); icon-size comments verified as no-ops (DropdownMenuItem/Badge already size their svg children).
  • Verified live against dev/stage: instance + cluster ops fire and are accepted, safe-mode restart round-trips (badges + pill render), a plain restart of a safe-mode cluster returns it to normal mode, stopped-state polling stops, transient card status renders and clears, restart picker works, and the resend button shows only for pending rows.

🤖 Generated with Claude Code

DavidCockerill and others added 3 commits July 7, 2026 16:32
…te handling

Studio UI for the FAB-380 container operations:
- Per-instance container ops in the instance ⋯ menu (Start / Start in safe mode /
  Restart / Restart in safe mode / Stop) under a "Container" section, gated by manage
  permission + instance status. New CM apiClient mutation + toast/invalidate hook.
- Safe-mode badge on running instances (hidden when stopped — a UI start clears it).
- Stopped/transitioning instances: stop polling get_status (ops API unreachable), show a
  muted availability dot and a "—" placeholder instead of a perpetual spinner.
- STOPPED status badge -> red.
- Cluster card stays reachable when not running: fully-stopped opens the instances page,
  partial opens the cluster overview.
- Cluster overview StatusPill: STOPPED -> red, PARTIAL -> yellow.
- Add safeMode to the Instance type (returned by the CM, not yet in the generated schema).

The container calls are hand-typed past the generated SDK (the /HDBInstance/{id}/container/
{action} paths and the safeMode field aren't in the CM OpenAPI yet). Follow-up: regenerate
the SDK once the CM spec exposes them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add cluster-level Start/Stop/Restart to the ClusterCard ⋯ menu, status-gated
(RUNNING → Restart/Restart-safe/Stop, STOPPED → Start/Start-safe, PARTIAL →
Start/Stop/Restart). Safe-mode ops force parallel; Stop shows a confirm dialog;
non-safe Restart opens a rolling/parallel picker. Wired via useClusterContainerOps
+ clusterContainerOperation (hand-typed past the generated SDK, same as terminate).

Also show a "Safe mode" pill next to the status pill on the cluster overview when
every instance is in safe mode, and override Cluster.instances to the patched
Instance type so safeMode is available on cluster.instances.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a "Resend invite" action to the org users table, shown only for PENDING
users. Reuses the existing invite mutation (POST /UserInvite with { email,
roleId }, targeting the user's existing role); the backend treats a repeat
invite for a pending user as a resend. Renamed tableDefinition.ts -> .tsx to
host the JSX action cell.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces container lifecycle operations (start, stop, restart, and safe mode) for both individual instances and entire clusters, alongside a new "Resend invite" button for pending organization users. The review feedback focuses on improving code robustness and UI consistency, recommending the removal of local error handlers to prevent double-toasting, guarding query invalidations against undefined parameters, adding size classes to icons, utilizing appropriate button variants, and enhancing accessibility with focus-visible ring styles on custom buttons.

Comment thread src/features/organization/users/components/ResendInviteButton.tsx
Comment thread src/hooks/useClusterContainerOps.ts
Comment thread src/hooks/useInstanceContainerOps.ts
Comment thread src/hooks/useInstanceContainerOps.ts Outdated
Comment thread src/features/cluster/Instances.tsx
Comment thread src/features/clusters/components/ClusterContainerOpModals.tsx Outdated
Comment thread src/features/clusters/components/ClusterContainerOpModals.tsx Outdated
Comment thread src/features/clusters/components/ClusterContainerOpModals.tsx
Comment thread src/features/clusters/components/ClusterContainerOpModals.tsx
@dawsontoth

Copy link
Copy Markdown
Contributor

@DavidCockerill Can you get it to rebase, and then take a look at the new nav paradigm for clusters? Surfacing things in the ... menu for clusters is great, but you can also surface them on the cluster dashboard, before you navigate into the cluster. That can aid discoverability of these features. Take a pass at the instance level features too to see if any of those can be brought out of the ... on instances, and perhaps surfaced around that instances table for better discoverability.

@dawsontoth

Copy link
Copy Markdown
Contributor

FWIW, I tend to rebase instead of merging changes in. That will better represent what will happen when we move to bring this PR into the main branch. Plus, it gives you a clearer view of your changes being additive to the changes that have landed already.

DavidCockerill and others added 2 commits July 9, 2026 13:05
…afe-mode UX

- Cluster overview: a "Cluster actions" dropdown (AWS "Instance state" style) with all
  container ops plus Terminate; ops invalid for the current status are shown but disabled
  so the feature set is discoverable.
- Safe mode: an explain-and-confirm dialog (what it does + when to use it) that both the
  overview dropdown and the clusters-list card menu route through, so safe mode always
  explains itself instead of firing silently.
- Clusters-list card: a temporary status badge for container-op states —
  Stopping/Starting/Restarting (amber + spinner) and Stopped/Partial — that clears on its
  own as the list polls; RUNNING stays clean.
- Fix: plain Start/Restart now send safeMode:false so they return a cluster to normal mode.
  Previously they omitted safeMode and the CM preserves the current mode, so restarting an
  in-safe-mode cluster stayed in safe mode.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…rId, a11y polish

- Remove local error toasts in ResendInviteButton, useClusterContainerOps,
  and useInstanceContainerOps; the global MutationCache.onError already
  surfaces failures with the server's message (avoids double-toasting).
- Guard the [clusterId] query invalidation (useParams strict:false can be
  undefined) so we don't invalidate [undefined].
- Cancel buttons -> defaultOutline (house style) across the container-op and
  safe-mode dialogs so they don't compete with the primary/destructive action.
- Add focus-visible ring to the raw restart-strategy <button>s for keyboard a11y.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@DavidCockerill

Copy link
Copy Markdown
Member Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces cluster-wide and instance-level container lifecycle operations (stop, start, restart, and safe mode) along with confirmation dialogs, status badges, and polling suppression for stopped or transitioning instances. It also adds a "Resend invite" button for pending organization users. The review feedback suggests removing a local onError handler in ClusterStateMenu to prevent duplicate error notifications, and adding size and margin classes to several icons inside dropdown items and badges to ensure proper layout and alignment.

Comment thread src/features/clusters/components/ClusterStateMenu.tsx
Comment thread src/features/clusters/components/ClusterStateMenu.tsx
Comment thread src/features/cluster/Instances.tsx
Comment thread src/features/clusters/components/ClusterCard.tsx
Gemini re-review: the terminate mutation's local onError toast duplicates
the global MutationCache.onError toast. Keep only the modal close.
(The 3 icon-size comments are no-ops: DropdownMenuItem already applies
[&_svg]:size-4 + gap-2, and Badge applies [&>svg]:size-3.)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@DavidCockerill
DavidCockerill marked this pull request as ready for review July 9, 2026 19:18
@DavidCockerill
DavidCockerill requested a review from a team as a code owner July 9, 2026 19:18
@DavidCockerill
DavidCockerill requested a review from dawsontoth July 9, 2026 19:18

@dawsontoth dawsontoth left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like it! Let's leave it in draft until the other PRs we depend are land, though, so no one's bot gets feisty and merges it.

@dawsontoth
dawsontoth marked this pull request as draft July 10, 2026 14:20

@kriszyp kriszyp left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice feature — the container lifecycle ops (start/stop/restart across instance and cluster scope) are cleanly organized and reuse the existing polling/invalidation/toast conventions correctly, and the polling-suppression for stopped/transitioning instances (badgeStatus.tsx:22) is a genuinely nice perf cleanup on top of the feature work, not just incidental. Good shape overall.

Three things worth fixing before this ships, though — flagging as a plain comment rather than a blocking review since it's already approved and held in draft pending dependent PRs, but I'd want these addressed before/shortly after merge:

1. Safe-mode confirm dialog is cluster-only — per-instance ops skip it entirely. useInstanceMenuItems.tsx:335 fires runContainerOp('start'/'restart', { safeMode: true }) directly with no dialog, while the identical action in ClusterCard.tsx:462 routes through SafeModeConfirmDialog. This is the one place the asymmetry really matters: the dialog exists specifically because "safe mode" is jargon for a recovery action that drops all user apps/components, and that's exactly as true for a single instance as for a cluster. Clicking "Start/Restart in safe mode" from the Instances page ⋯ menu today does it with zero warning.

2. Resend invite has no client-side permission gate. tableDefinition.tsx:44 renders ResendInviteButton unconditionally for any PENDING row, but every other write affordance on this same page (OrgConfigUsersIndex) is gated on useOrganizationRolePermissions(organizationId).update. ResendInviteButton (ResendInviteButton.tsx:1) fires a real POST /UserInvite and is clickable by view-only users today — the backend presumably rejects it, but that's a defense-in-depth gap and an inconsistent affordance compared to its siblings on the same page. One-line fix: thread update into the column def or check it inside the button via the same hook.

Calling these out specifically because "safe mode" and permission gating are exactly the invariants this PR is introducing to prevent accidental/unauthorized destructive actions — so the gaps are a bit ironic relative to the feature's own goal, not a knock on the overall design.

3. Smaller one, lower urgency: the container-op hooks (useClusterContainerOps.ts:32) flip isPending back to false as soon as the "accepted" response lands, before the follow-up invalidateQueries (fired with void, unawaited) resolves — a narrow window where a fast double-click could dispatch a second conflicting op before the UI reflects the first. Same shape as the pre-existing useRestartInstanceClick pattern, so not new, but now applied to two more destructive actions (Stop, safe-mode Restart). Worth a quick confirm that central-manager/host-manager reject or queue overlapping ops server-side.

None of this is blocking given the approval already in place and the draft hold for dependent PRs — just want it in front of the team before merge. Happy to help wire up #1/#2 if useful.

— Claude (Sonnet 5), reviewed via review-queue

@dawsontoth

Copy link
Copy Markdown
Contributor

@DavidCockerill will we need version gates for these features?

@DavidCockerill

Copy link
Copy Markdown
Member Author

@dawsontoth Good instinct, but I don't think we need version gates here.

  • Start / stop / restart are docker-level (host-manager compose), so they don't depend on the Harper version running inside the container at all.
  • Safe mode is the only part with a version dependency — it needs Harper to honor HARPER_SAFE_MODE (boot without loading user apps/components). That landed in 4.7.14 on the 4.x line (commit ac66632a1, released 2026-01-06) and has been in every 5.x since 5.0.0. Every Fabric instance runs 4.7.x-latest (well past 4.7.14) or 5.x, so they all support it.

The only thing a gate would guard against is an instance older than 4.7.14 silently ignoring the env var and booting normally while the UI says "safe mode" — and we don't run anything that old.

— DAIvid (Opus 4.8)

@DavidCockerill
DavidCockerill marked this pull request as ready for review July 16, 2026 17:06
… gate resend invite

- Per-instance Start/Restart in safe mode now route through SafeModeConfirmDialog
  (generalized with a `scope` prop for cluster vs instance) instead of firing
  safeMode:true with no warning — the explanation matters as much for one instance
  as for a cluster.
- ResendInviteButton is hidden unless the user has org-role `update` permission,
  matching the page's other write affordances (it was clickable by view-only users).

useInstanceMenuItems now returns { items, dialog } and both menu consumers render the
dialog. Kris's 3rd item (isPending flips before the unawaited invalidateQueries — a
narrow double-click window) is covered server-side: the CM admission gate rejects
overlapping ops with 409 CONFLICT.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

Coverage Report

Status Category Percentage Covered / Total
🔵 Lines 50.63% 5336 / 10538
🔵 Statements 51.21% 5715 / 11159
🔵 Functions 42.74% 1305 / 3053
🔵 Branches 44.03% 3583 / 8137
File Coverage
File Stmts Branches Functions Lines Uncovered Lines
Changed Files
src/components/ui/utils/badgeStatus.tsx 0% 0% 0% 0% 22-108
src/features/cluster/ClusterHome.tsx 0% 0% 0% 0% 37-479
src/features/cluster/InstanceActionsMenu.tsx 0% 0% 0% 0% 15-35
src/features/cluster/InstanceLogInCell.tsx 0% 0% 0% 0% 18-85
src/features/cluster/InstanceRowContextMenu.tsx 0% 0% 0% 0% 20-34
src/features/cluster/InstanceStatusCell.tsx 0% 0% 0% 0% 18-116
src/features/cluster/Instances.tsx 0% 0% 0% 0% 28-174
src/features/cluster/useInstanceMenuItems.tsx 2.17% 0% 0% 2.27% 49-220
src/features/clusters/components/ClusterCard.tsx 0% 0% 0% 0% 53-436
src/features/clusters/components/ClusterContainerOpModals.tsx 0% 100% 0% 0% 37-92
src/features/clusters/components/ClusterStateMenu.tsx 0% 0% 0% 0% 32-152
src/features/clusters/components/SafeModeConfirmDialog.tsx 0% 0% 0% 0% 35-66
src/features/organization/users/components/ResendInviteButton.tsx 0% 0% 0% 0% 19-57
src/hooks/useClusterContainerOps.ts 6.25% 0% 0% 6.25% 27-56
src/hooks/useInstanceContainerOps.ts 5.26% 0% 0% 5.26% 28-62
src/integrations/api/cluster/containerOperation.ts 0% 0% 0% 0% 39-47
src/integrations/api/instance/containerOperation.ts 0% 0% 0% 0% 38-47
Generated in workflow #1529 for commit 07e8525 by the Vitest Coverage Report Action

@DavidCockerill

Copy link
Copy Markdown
Member Author

Thanks Kris — all three addressed in 07e8525d:

  1. Safe-mode confirm on instance ops — generalized SafeModeConfirmDialog with a scope prop and routed the per-instance Start/Restart-in-safe-mode through it, so the instance ⋯ and right-click menus now get the same explain-and-confirm as the cluster path (the copy drops the cluster-only "all instances at once" line for a single instance).
  2. Resend invite — now hidden unless the viewer has org-role update, matching Add-user/edit on the same page.
  3. isPending double-click window — confirmed it's closed server-side: CM's container-op admission gate rejects an op on an already-transitioning instance with a 409 CONFLICT (documented as accepted-by-design in hdbInstance/containerOperation.js). Left the client pattern as-is since it matches the existing useRestartInstanceClick shape.

— DAIvid (Opus 4.8)

@DavidCockerill
DavidCockerill added this pull request to the merge queue Jul 16, 2026
Merged via the queue into stage with commit c867f3a Jul 16, 2026
2 checks passed
@DavidCockerill
DavidCockerill deleted the FAB-380 branch July 16, 2026 18:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants