Describe the bug
Several smaller issues found while diagnosing agents that had gone silent. Grouped into one issue since none warrants its own — happy to split if you'd prefer.
1. %PATH% in agent env_vars is not expanded (Windows)
Setting an agent env var to C:/Users/Amy/bin;%PATH% passes the literal string %PATH% to the child process — cmd-style expansion doesn't happen when the parent sets the child environment directly. The agent ends up with one valid PATH entry and one garbage entry, losing the entire system PATH.
Symptom: the agent silently cannot find the binaries it needs — in my case buzz itself, which is exactly what it needs in order to reply.
Suggested fix: expand %VAR% / $VAR when composing the child environment, or validate on save and warn that %PATH% won't work, or offer an "append to inherited PATH" checkbox instead of a raw string field.
2. Deleting a persona record bricks its agent, with a misleading error
managed-agents.json stores two record types in one array with no type discriminator:
|
pubkey |
slug |
system_prompt / display_name |
role |
| Persona |
absent |
present |
present |
the agent's definition |
| Agent |
present |
absent |
absent |
the runnable instance |
Agents reference their persona via persona_id → persona slug, with no referential-integrity check. Removing a persona leaves its agent unstartable with:
This agent's configuration is missing — it may still be syncing or was deleted on another device.
That message is misleading — nothing was syncing and there was no other device. The record was missing locally, which the app can determine with certainty.
Compounding it: builtin:* personas are silently re-seeded on startup while custom personas are not. So three of my four agents self-healed and only the one with a custom persona stayed broken — which made it look like an agent-specific fault rather than a config-integrity problem.
Suggested fix: validate persona_id references on load and report the specific missing slug; reword the error to distinguish "referenced persona not found locally" from a genuine sync issue; consider separate files or an explicit "type" field.
3. Default parallelism: 10 is high for local single-user installs
Every agent defaulted to parallelism: 10, so starting one agent spawned 10 worker trees. On the Hermes runtime that's 10 × (hermes-acp.exe + 2 python.exe) ≈ 30 processes per agent, all opening the same SQLite state file:
state.db … is already in WAL mode — leaving WAL in place (no live downgrade under concurrent openers)
On a hard kill these are also what get left behind, which reads as a process leak. (On a graceful shutdown Buzz reaps them correctly — that part works well.)
parallelism: 1 seems more appropriate for a single-user local install.
4. Some UI config edits don't persist
A parallelism change made through the UI did not survive. This may be the same family as the known ModelPicker.tsx bug where envVars was serialised as null on model change. Might be worth auditing the agent-update payload path for other fields being dropped or nulled.
5. Stale PID files
agents/agent-pids/*.json are left behind after a clean shutdown, so they imply a running agent that isn't. Minor, but it costs time when triaging.
6. start_on_app_launch defaults to false
After any restart every agent is down until manually started — which presents identically to "the agent has gone mute". Given how many restarts a debugging session involves, this produced several false alarms. A first-run prompt or a more discoverable default would help.
7. minio-setup destroys the media bucket on every docker compose up
docker-compose.yml runs this as an init container:
minio-setup:
entrypoint: >
/bin/sh -c "
/usr/bin/mc alias set myminio http://minio:9000 buzzdev buzzdevpass;
/usr/bin/mc rb --force myminio/buzz-media || true;
/usr/bin/mc mb myminio/buzz-media;
exit 0;
"
mc rb --force removes the bucket and everything in it, unconditionally. The relay then depends on it completing:
relay:
depends_on:
minio-setup:
condition: service_completed_successfully
So any docker compose up that brings up the relay also re-runs the wipe. This is fine on first setup and destructive on every run after — and it isn't obvious from the outside, because the destructive step is in a service you never invoke directly. It is reached by the ordinary command for starting the stack.
I nearly lost my media to this while recovering from a power outage: the intuitive recovery is docker compose up -d, and that is exactly the command that erases the bucket. docker start <container> is safe; docker compose up is not. Nothing warns you.
Suggested fix: make bucket creation idempotent rather than destructive — mc mb --ignore-existing myminio/buzz-media alone does the job. If a reset path is genuinely wanted, gate it behind an explicit opt-in (BUZZ_RESET_MEDIA=1) or move it to a separate profile that isn't in the relay's dependency chain.
8. minio is the only long-running service with no restart policy
In docker-compose.yml, postgres, redis and relay all declare restart: unless-stopped. minio does not. (minio-setup correctly has restart: "no" — it's a one-shot.)
The consequence after an unclean shutdown — in my case a power outage — is that MinIO stays down while everything else comes back. The relay then boots normally through Postgres, Redis, migrations and search, and dies on the S3 conformance probe:
Error: git conformance probe failed: s3 backend error: reqwest: error sending request
for url (http://minio:9000/buzz-media/packs/...)
Because the relay does have restart: unless-stopped, it crash-loops indefinitely, and nothing ever listens on :3000. The stack looks alive — three of four containers healthy — while being completely unusable. The relay's log is also misleading at a glance, since Media storage connected is printed shortly before the probe fails.
Downstream effect worth noting: a Hermes agent gateway pointed at this relay retried every 5 minutes for over an hour, shelling out to the buzz CLI each time. On Windows every retry spawned a visible console window, which is how I noticed anything was wrong at all.
Suggested fix: add restart: unless-stopped to the minio service, for consistency with the other three. One line:
minio:
image: minio/minio
restart: unless-stopped
command: server /data
Related: #2723 / #2724 cover the probe hanging when a backend connection stalls. This is the adjacent case — the connection is refused outright, the probe fails fast, and the relay exits into a restart loop. A clearer fatal message naming the unreachable S3 endpoint as the cause would have saved a lot of time here.
Version and platform
- Buzz version: local dev build,
block/buzz @ 318fbf896, with local uncommitted changes
- claude-agent-acp: 0.64.2 (node v24.18.0)
- OS: Windows 10
- Stack: Postgres 16, Redis, MinIO via Docker Desktop
Describe the bug
Several smaller issues found while diagnosing agents that had gone silent. Grouped into one issue since none warrants its own — happy to split if you'd prefer.
1.
%PATH%in agentenv_varsis not expanded (Windows)Setting an agent env var to
C:/Users/Amy/bin;%PATH%passes the literal string%PATH%to the child process — cmd-style expansion doesn't happen when the parent sets the child environment directly. The agent ends up with one valid PATH entry and one garbage entry, losing the entire system PATH.Symptom: the agent silently cannot find the binaries it needs — in my case
buzzitself, which is exactly what it needs in order to reply.Suggested fix: expand
%VAR%/$VARwhen composing the child environment, or validate on save and warn that%PATH%won't work, or offer an "append to inherited PATH" checkbox instead of a raw string field.2. Deleting a persona record bricks its agent, with a misleading error
managed-agents.jsonstores two record types in one array with no type discriminator:pubkeyslugsystem_prompt/display_nameAgents reference their persona via
persona_id→ personaslug, with no referential-integrity check. Removing a persona leaves its agent unstartable with:That message is misleading — nothing was syncing and there was no other device. The record was missing locally, which the app can determine with certainty.
Compounding it:
builtin:*personas are silently re-seeded on startup while custom personas are not. So three of my four agents self-healed and only the one with a custom persona stayed broken — which made it look like an agent-specific fault rather than a config-integrity problem.Suggested fix: validate
persona_idreferences on load and report the specific missingslug; reword the error to distinguish "referenced persona not found locally" from a genuine sync issue; consider separate files or an explicit"type"field.3. Default
parallelism: 10is high for local single-user installsEvery agent defaulted to
parallelism: 10, so starting one agent spawned 10 worker trees. On the Hermes runtime that's 10 × (hermes-acp.exe+ 2python.exe) ≈ 30 processes per agent, all opening the same SQLite state file:On a hard kill these are also what get left behind, which reads as a process leak. (On a graceful shutdown Buzz reaps them correctly — that part works well.)
parallelism: 1seems more appropriate for a single-user local install.4. Some UI config edits don't persist
A
parallelismchange made through the UI did not survive. This may be the same family as the knownModelPicker.tsxbug whereenvVarswas serialised asnullon model change. Might be worth auditing the agent-update payload path for other fields being dropped or nulled.5. Stale PID files
agents/agent-pids/*.jsonare left behind after a clean shutdown, so they imply a running agent that isn't. Minor, but it costs time when triaging.6.
start_on_app_launchdefaults tofalseAfter any restart every agent is down until manually started — which presents identically to "the agent has gone mute". Given how many restarts a debugging session involves, this produced several false alarms. A first-run prompt or a more discoverable default would help.
7.
minio-setupdestroys the media bucket on everydocker compose updocker-compose.ymlruns this as an init container:mc rb --forceremoves the bucket and everything in it, unconditionally. The relay then depends on it completing:So any
docker compose upthat brings up the relay also re-runs the wipe. This is fine on first setup and destructive on every run after — and it isn't obvious from the outside, because the destructive step is in a service you never invoke directly. It is reached by the ordinary command for starting the stack.I nearly lost my media to this while recovering from a power outage: the intuitive recovery is
docker compose up -d, and that is exactly the command that erases the bucket.docker start <container>is safe;docker compose upis not. Nothing warns you.Suggested fix: make bucket creation idempotent rather than destructive —
mc mb --ignore-existing myminio/buzz-mediaalone does the job. If a reset path is genuinely wanted, gate it behind an explicit opt-in (BUZZ_RESET_MEDIA=1) or move it to a separate profile that isn't in the relay's dependency chain.8.
miniois the only long-running service with no restart policyIn
docker-compose.yml,postgres,redisandrelayall declarerestart: unless-stopped.miniodoes not. (minio-setupcorrectly hasrestart: "no"— it's a one-shot.)The consequence after an unclean shutdown — in my case a power outage — is that MinIO stays down while everything else comes back. The relay then boots normally through Postgres, Redis, migrations and search, and dies on the S3 conformance probe:
Because the relay does have
restart: unless-stopped, it crash-loops indefinitely, and nothing ever listens on:3000. The stack looks alive — three of four containers healthy — while being completely unusable. The relay's log is also misleading at a glance, sinceMedia storage connectedis printed shortly before the probe fails.Downstream effect worth noting: a Hermes agent gateway pointed at this relay retried every 5 minutes for over an hour, shelling out to the
buzzCLI each time. On Windows every retry spawned a visible console window, which is how I noticed anything was wrong at all.Suggested fix: add
restart: unless-stoppedto theminioservice, for consistency with the other three. One line:Related: #2723 / #2724 cover the probe hanging when a backend connection stalls. This is the adjacent case — the connection is refused outright, the probe fails fast, and the relay exits into a restart loop. A clearer fatal message naming the unreachable S3 endpoint as the cause would have saved a lot of time here.
Version and platform
block/buzz@318fbf896, with local uncommitted changes