Skip to content

docs(self-host): make the cloud catalog diagnostic actionable on-prem - #3132

Merged
benjaminshafii merged 1 commit into
devfrom
docs/onprem-diagnostics-trusted-origins
Jul 25, 2026
Merged

docs(self-host): make the cloud catalog diagnostic actionable on-prem#3132
benjaminshafii merged 1 commit into
devfrom
docs/onprem-diagnostics-trusted-origins

Conversation

@benjaminshafii

Copy link
Copy Markdown
Member

The problem

Settings → Debug runs a cloud catalog check: a bounded initializetools/list against the managed openwork-cloud MCP entry, asserting Cloud exposes exactly search_capabilities + execute_capability. It is the single most useful "is Cloud MCP actually working" signal we have.

It carries an organization MCP bearer token, so agent-context-cloud-probe.ts refuses to send that credential to an origin outside a trust list. DEFAULT_TRUSTED_ORIGINS contains only:

https://app.openworklabs.com
https://api.openworklabs.com

So on every self-hosted deployment the endpoint origin is the customer's own Den, prepare() returns untrusted_endpoint, and the probe never runs. The diagnostic is dead for on-prem.

There is an escape hatch — OPENWORK_AGENT_DIAGNOSTICS_TRUSTED_ORIGINS — which appeared in exactly two places in the entire repo: its own source file and its own test. No documentation anywhere. And the guidance the operator actually saw was:

Use the hosted OpenWork Cloud endpoint or have the server administrator explicitly trust the development origin.

That calls a production on-prem Den a development origin and never names the variable. An operator cannot act on it.

What this changes

Two string literals and a docs section. That is the entire code diff.

Before After
message The managed OpenWork Cloud endpoint is outside the server's diagnostics trust policy. The server did not send the credentialed cloud catalog request because the Cloud endpoint origin is not in the diagnostics trust list.
action Use the hosted OpenWork Cloud endpoint or have the server administrator explicitly trust the development origin. Set OPENWORK_AGENT_DIAGNOSTICS_TRUSTED_ORIGINS on the OpenWork desktop/server process to include the Cloud endpoint origin, then rerun diagnostics.

Plus a ### Trust self-hosted Cloud catalog diagnostics section in packages/docs/start-here/self-host.mdx documenting the real format rules enforced by configuredTrustedOrigins(): comma-separated bare origins; scheme + host + optional port only, no path/query/fragment/credentials; https: required except loopback; invalid entries are ignored rather than failing open; entries are added to the hosted defaults, not replacing them.

The docs state explicitly that this variable enables the diagnostic only and that Cloud MCP itself works without it — so operators do not read a skipped probe as a broken deployment.

Compatibility

Cannot break an existing self-hosted deployment. apps/server/src/agent-context-cloud-probe.ts is untouched — verified with git diff --stat, empty. Which origins are trusted is therefore provably identical before and after. DEFAULT_TRUSTED_ORIGINS, safeCatalogEndpoint, the loopback rule, and configuredTrustedOrigins() parsing are all unchanged. Nothing was added to PERSISTABLE_INTERNAL_KEYS.

Why I did not "fix" the trust model

The obvious-looking fix — trust the configured Den origin — is circular and unsafe. The probe exists to validate the runtime MCP config; deriving trust from that same config would let a tampered config exfiltrate the org MCP bearer token to an attacker origin. The file makes this philosophy explicit elsewhere: "Absence of a deny in passively inspected config is never authority to send a credentialed request."

Doing it properly needs an independent source of truth for the operator-provisioned Den origin. apps/server has none today — the only candidate is the desktop bootstrap config, which lives in the renderer and would require a diagnostics request-schema change across client and server. That is a security-model decision for a maintainer, not something to slip into a backward-compatibility-sensitive PR. Flagged as a follow-up below.

Tests

```
pnpm --filter openwork-server test # 521 pass, 10 skip, 0 fail, 2827 expect() calls, 531 tests / 75 files
pnpm typecheck # pass
```

Coverage added — note both tests pin the no-credentialed-request property, not just the string:

  • agent-context-cloud-probe.test.ts: an unlisted self-hosted origin returns untrusted_endpoint and blockedCalls === 0, proving no request leaves the process. (Acceptance when the variable is set was already covered; I extended that test rather than duplicating it.)
  • agent-context-diagnostics.test.ts: the report's action must keep naming OPENWORK_AGENT_DIAGNOSTICS_TRUSTED_ORIGINS, and fetchCalls is empty. This stops the guidance regressing to something unactionable.

Report-safety respected: both strings are static, well under the bounded-text limit, and interpolate no endpoint, credential, org id, or host.

Follow-ups (reported, deliberately not fixed)

  1. The variable cannot be set from inside the app. apps/desktop/electron/runtime.mjs:470-485 strips all OPENWORK_* keys from the Settings → Environment store when building the sidecar child env, and apps/server/src/env-file.ts:22-28,205-206 rejects non-persistable OPENWORK_* keys. It must be set in the process environment before OpenWork launches (runtime.mjs:764-772 overlays real process.env, so pre-set values survive). Fixing this means touching a reserved-key policy that is mirrored across three files under a documented "must agree byte-for-byte" contract — out of scope here.

  2. Should a provisioned Den origin be trusted automatically? Needs a maintainer decision on the threat model plus a diagnostics request-schema change. Worth doing; not safe to decide unilaterally.

Fraimz

Skipped, stated explicitly per AGENTS.md. No runtime path changes — the observable surface is two operator-facing strings in a diagnostics report, which are asserted directly by the new unit test above.

Settings > Debug runs a bounded initialize -> tools/list check against the
managed openwork-cloud MCP entry. It carries an organization MCP bearer token,
so the probe refuses to send that credential to an origin outside its trust
list. DEFAULT_TRUSTED_ORIGINS holds only the hosted origins, so on every
self-hosted deployment the check returns untrusted_endpoint and never runs.

The escape hatch, OPENWORK_AGENT_DIAGNOSTICS_TRUSTED_ORIGINS, appeared in
exactly two places in the repo: its own source file and its own test. The
guidance an operator actually saw was "have the server administrator explicitly
trust the development origin" -- which calls a production on-prem Den a
development origin and never names the variable, so nobody could act on it.

Name the variable in the operator action, and document it in the self-hosting
guide with the real format rules enforced by configuredTrustedOrigins: bare
origins, comma-separated, https except loopback, invalid entries ignored rather
than failing open, additive to the hosted defaults. The docs state explicitly
that this enables the diagnostic only and that Cloud MCP itself works without
it, so operators do not read a skipped probe as a broken deployment.

The trust model is deliberately unchanged: agent-context-cloud-probe.ts is
untouched, so which origins are trusted cannot differ from before. Trusting the
endpoint's own origin would be circular, since the probe exists to validate the
very config that names it. Tests pin that an unlisted self-hosted origin still
performs no fetch, and that the action keeps naming the variable.
@vercel

vercel Bot commented Jul 25, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
openwork-app Ready Ready Preview, Comment Jul 25, 2026 5:29pm
openwork-den Ready Ready Preview, Comment Jul 25, 2026 5:29pm
openwork-den-worker-proxy Ready Ready Preview, Comment Jul 25, 2026 5:29pm
openwork-diagnostics Ready Ready Preview, Comment Jul 25, 2026 5:29pm
openwork-landing Ready Ready Preview, Comment, Open in v0 Jul 25, 2026 5:29pm

@mintlify

mintlify Bot commented Jul 25, 2026

Copy link
Copy Markdown
Contributor

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated (UTC)
differentai 🟢 Ready View Preview Jul 25, 2026, 5:28 PM

💡 Tip: Enable Workflows to automatically generate PRs for you.

@mintlify

mintlify Bot commented Jul 25, 2026

Copy link
Copy Markdown
Contributor

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated (UTC)
differentai 🟡 Building Jul 25, 2026, 5:28 PM

💡 Tip: Enable Workflows to automatically generate PRs for you.

@benjaminshafii
benjaminshafii merged commit bf32618 into dev Jul 25, 2026
12 checks passed
benjaminshafii added a commit that referenced this pull request Jul 25, 2026
…ce (#3139)

* docs(enterprise): document outbound network access and guard it in CI

Enterprise networks block outbound connections by default and IT must approve
each destination, but nothing in this repo listed what OpenWork needs to reach.
Two destinations were named nowhere at all:

- objects.githubusercontent.com, the redirect target of every release-asset
  download. Approving only github.com lets a download start and then fail
  partway through.
- registry.npmjs.org, used by `npx -y openwork-ui-mcp` in packaged builds.
  Blocked, the UI-control MCP is silently unavailable while the app still opens.

Add docs/enterprise/outbound-access.md for customer IT, in plain language, with
a copy-pasteable minimum allowlist and a what-breaks-when-blocked column. Desktop
client access is documented separately from self-hosted Den access because
different teams approve them on different networks.

Back it with docs/enterprise/outbound-access.json and
scripts/check-outbound-access.mjs, wired into ci-tests.yml, so a feature that
introduces a new destination must document its enterprise impact. The check also
reports manifest entries that no longer exist in code, so the list cannot rot.
Hosts reached indirectly (redirect targets, subprocesses, links, schema strings)
are classified by kind and exempt from the staleness rule.

Docs, one check script, one CI step, one package script. No product behavior,
configuration, API, or chart surface changes.

* docs(self-host): make the cloud catalog diagnostic actionable on-prem

Settings > Debug runs a bounded initialize -> tools/list check against the
managed openwork-cloud MCP entry. It carries an organization MCP bearer token,
so the probe refuses to send that credential to an origin outside its trust
list. DEFAULT_TRUSTED_ORIGINS holds only the hosted origins, so on every
self-hosted deployment the check returns untrusted_endpoint and never runs.

The escape hatch, OPENWORK_AGENT_DIAGNOSTICS_TRUSTED_ORIGINS, appeared in
exactly two places in the repo: its own source file and its own test. The
guidance an operator actually saw was "have the server administrator explicitly
trust the development origin" -- which calls a production on-prem Den a
development origin and never names the variable, so nobody could act on it.

Name the variable in the operator action, and document it in the self-hosting
guide with the real format rules enforced by configuredTrustedOrigins: bare
origins, comma-separated, https except loopback, invalid entries ignored rather
than failing open, additive to the hosted defaults. The docs state explicitly
that this enables the diagnostic only and that Cloud MCP itself works without
it, so operators do not read a skipped probe as a broken deployment.

The trust model is deliberately unchanged: agent-context-cloud-probe.ts is
untouched, so which origins are trusted cannot differ from before. Trusting the
endpoint's own origin would be circular, since the probe exists to validate the
very config that names it. Tests pin that an unlisted self-hosted origin still
performs no fetch, and that the action keeps naming the variable.

* docs(helm): expose the private-network MCP opt-out to operators

Den's SSRF guard refuses to fetch private/reserved addresses on behalf of
users, which is correct on hosted multi-tenant deployments. But a common
enterprise topology puts Den inside the customer's private network with their
MCP servers on private addresses, and there the guard blocks every internal
server. The operator sees MCP_URL_BLOCKED and is told to "use a public HTTPS
MCP URL or change the deployment's private-network policy through security
review", with no indication that a supported opt-out exists.

env.ts has provided DEN_ALLOW_PRIVATE_MCP_URLS=1 all along, but it appeared
only in source, unit tests, and one eval flow: not in .env.example, not in the
chart, not in any doc. Surface it in all three places so the fix is reachable
without reading the source.

Scoped to den-api only. den-api, den-web, and inference all envFrom the same
ConfigMap, so adding the key there would inject an SSRF-bypass flag into two
workloads that never read it; the value is emitted directly on the den-api
container instead. Only den-api consumes allowPrivateMcpUrls (mcp-connections,
plugin-system/store, enterprise-mcp-client-adapter, external-mcp-client).

Nothing renders unless an operator opts in, so helm template with default
values is byte-identical to the previous chart: no ConfigMap checksum change
and no pod restart on upgrade. Blank, 0, and false render nothing; only "1"
enables it, matching the trimmed string comparison in env.ts; any other value
fails the render rather than guessing. url-guard.ts and env.ts are untouched.

* docs(enterprise): name the current GitHub release-asset host

Verified against a live OpenWork release asset: github.com now redirects
release downloads to release-assets.githubusercontent.com, not
objects.githubusercontent.com. An allowlist naming only the old host would
still let downloads start and then fail partway through, which is the exact
symptom this page exists to prevent.

List both. release-assets.githubusercontent.com is what GitHub serves today;
objects.githubusercontent.com stays documented for older clients and any
staged rollback.

* docs(self-host): document the private-network (semi air-gapped) deployment

Most enterprise installs are not fully air-gapped. The common shape is laptops
with internet plus VPN, and Den web and Den API inside the customer's private
network. That split was undocumented, so every customer rediscovered its
gotchas by hitting failures.

Add a published page covering the two network paths separately, because
different teams approve them: what the desktop must reach, what Den must reach
(only MySQL and ghcr.io at pull time are hard requirements), and the three
things that surprise operators, each with symptom, cause, and fix:

- Internal MCP servers are refused by the SSRF guard until
  DEN_ALLOW_PRIVATE_MCP_URLS=1, with the security tradeoff stated.
- The Settings > Debug cloud catalog probe is skipped because only hosted
  origins are trusted. It states explicitly that this affects the diagnostic
  only and Cloud MCP still works, so a skipped probe is not read as an outage,
  and that the variable must be set before launch because the in-app
  environment store strips OPENWORK_* keys.
- Den's outbound self-check targets a public host by default and can be
  self-hosted with DEN_DIAGNOSTICS_ORIGIN.

Also resolve a contradiction that decided whether IT opens one hostname or
two. docs/org-install-links.md claimed "the desktop must be able to reach both
origins"; resolveDenBaseUrls ignores an explicit API origin when a base URL is
present and derives <baseUrl>/api/den, proven by the existing den-mcp-url test.
Steady-state desktop traffic needs one origin. A separate API origin is only
reachable-critical for the one-time install-link exchange, which builds its
endpoint from the apiBaseUrl carried in the link, and for external MCP clients.
Both docs now state their context instead of disagreeing.

* docs(self-host): consolidate network, air-gap, and certificate guidance

Promote Self-host from a small Start here group to a dedicated docs tab while
preserving every existing start-here URL. Consolidate scattered network
material into five focused published pages: air-gapped deployment, outbound
network access, certificate trust and proxies, installer delivery, and network
diagnostics. Keep the private-network topology page as the high-level entry.

Make boundaries explicit instead of conflating them:

- mounted installers solve installer delivery, not full product isolation;
- desktop Chromium, spawned sidecars, Den containers, and MySQL use different
  trust surfaces;
- Linux Electron may require NSS/user trust even when /etc/ssl is correct;
- sslmode=require encrypts but is not certificate identity verification;
- a skipped Cloud catalog diagnostic does not mean Cloud MCP is broken.

Promote the machine-guarded outbound manifest into customer-facing docs,
retain repository operator pages and their existing anchors, and slim the
self-host/private-network overviews so there is one canonical owner per detail.
The Helm and diagnostics READMEs remain authoritative for low-level mechanics.

This branch is stacked on PRs #3131, #3132, #3133, and #3135. It changes docs
and navigation only; no product, Helm template, environment, skill, or eval
behavior changes.
@benjaminshafii
benjaminshafii deleted the docs/onprem-diagnostics-trusted-origins branch July 25, 2026 20:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant