10 tools. 6 task types. Failure diagnosis. Multi-agent handoff. Physics-verified.
An open framework for LLM agent orchestration of robotic tasks — the agent
chooses tools, orders steps, diagnoses failures, and hands off when stuck.
A physics-simulation gate validates every plan before the robot moves.
LLM selects tools → critic prunes → physics validates → failure analyst repairs → deterministic nav executes
T2: Locked-door delivery (Isaac Sim 3D) — the Go2 picks up the key (gold), navigates to the locked door (brown panel), unlocks and opens it (door disappears), then fetches the box and delivers it to the shelf. 4 different tools used.
T3: Blocked-path clearance (Isaac Sim 3D) — the Go2 pushes the blocking crate (brown) aside via physics collision, then fetches the box and delivers it to the shelf past the obstacle (red pillar).
T1: Baseline fetch-and-place (Isaac Sim 3D) — a trained Go2 locomotion policy walks an A*-planned route, picks the fallen box, carries it around the obstacle, and places it on the shelf.
Warning
Research software — NOT a certified safety system. physgate is a research prototype. It does not implement or replace any safety function under ISO 10218-1:2025, ISO 13849-1, IEC 61508, or ISO 3691-4, and it runs on a general-purpose Linux + GPU stack that cannot meet PL/SIL requirements. A simulation PASS is not a safety proof. See Safety & Scope before any deployment.
LLM agents that drive robots execute hazardous instructions ~95% of the time even when capable of refusing (SafeAgentBench). And LLM-generated plans fail in characteristic ways: wrong tool selection, wrong step ordering, no recovery after a failed action, and confidently attempting impossible tasks.
physgate is a dual-system architecture with a physics-verification gate between the LLM agent and the robot:
- HIGH level — the LLM agent (Claude) selects from 10 scene-conditional tools (navigate, pick, place, open/unlock doors, call elevators, push obstacles, inspect objects, request assistance) and orders them into multi-step plans across 6 task types. It does NOT do geometry: no coordinates, no waypoints, no obstacle avoidance.
- A safety-critic agent adversarially prunes plans that violate safety contracts.
- The Sim-Gate validates surviving plans in parallel NVIDIA Isaac Lab physics simulation — tool selection, step ordering, preconditions, and physical outcome.
- A failure analyst diagnoses execution failures with structured root-cause analysis and suggests targeted repairs (keep working prefix, fix only the broken part) — replacing blind regeneration.
- LOW level — deterministic navigation (A* occupancy-grid planner behind a Nav2-compatible interface) executes the winning plan. Obstacle avoidance is GUARANTEED here, never guessed by the LLM.
- Cooperative handoff — when the agent recognizes a task is infeasible or
partially complete, it calls
request_assistanceand a second agent can continue from the handoff context.
The core thesis: physics simulation is the highest-quality verifier of agent
orchestration — it catches wrong tool selection, unmet preconditions, and
infeasible requests more reliably than a neural verifier or LLM self-critique.
What physics does NOT need to do is rescue bad route geometry: with navigation in
the right layer, any well-formed plan is executable (measured feasibility ~100%,
see benchmarks/rebuild/).
- Not a certified safety component. It is an application-layer soft interlock, not a safety-rated function. See Safety & Scope.
- Not a real-time controller. The planner operates at seconds-per-decision; it never enters the motor control loop.
- Not a claim that simulation guarantees real-world safety. The sim-to-real gap is real; a sim PASS reduces risk, it does not prove safety.
flowchart TD
task["Natural-language task<br/><i>6 types: fetch, locked-door, elevator,<br/>blocked-path, sequential, infeasible</i>"]
subgraph orch ["Orchestrator (LangGraph state machine)"]
direction LR
orch_detail["phase routing · replan budget · approval gate · handoff detection"]
end
subgraph high ["HIGH level — semantic planning (10 tools)"]
direction LR
planner["Planner (Claude Opus 4.8)<br/>→ N candidates from 10 tools"]
critic["Safety Critic (SAFER)<br/>→ prune unsafe"]
planner --> critic
end
subgraph gate ["Sim-Gate (soft interlock — NOT a safety fn)"]
direction TB
l1["L1 kinematic limits (URDF) · <1 ms"]
l3["L3 scene-graph preconditions · <1 ms"]
l2["L2 parallel physics (Isaac Lab, N envs)"]
select["→ select a verified plan"]
l1 --> l3 --> l2 --> select
end
subgraph low ["LOW level — deterministic execution"]
direction LR
nav["A* / Nav2-ready navigation"]
exec["Executor (Isaac sim / walking policy)"]
ros["(planned: ROS 2 → real Go2)"]
nav --> exec
exec -.-> ros
end
subgraph recovery ["Failure recovery"]
direction LR
analyst["Failure Analyst<br/><i>root cause · affected steps · suggested fix</i>"]
strategy{"targeted<br/>repair?"}
analyst --> strategy
strategy -->|"attempt 1: keep prefix"| targeted["Targeted repair"]
strategy -->|"attempt 2: full regen"| regen["Full regeneration"]
strategy -->|"attempt 3"| escalate["Escalate / handoff"]
end
handoff["Cooperative Handoff<br/><i>partial_success → Agent 2 continues</i>"]
audit["Audit: three-stream records + Merkle checkpoint"]
task --> orch
orch -->|"N candidate plans"| high
high -->|"surviving candidates"| gate
gate -->|"verified plan (JSON)"| low
low -->|"success"| audit
low -->|"failure"| recovery
recovery -->|"repaired plan"| high
recovery -->|"handoff"| handoff
style orch fill:#f0f4ff,stroke:#4e79a7,stroke-width:2px
style high fill:#fff8f0,stroke:#f28e2b,stroke-width:2px
style gate fill:#fff0f0,stroke:#e15759,stroke-width:2px
style low fill:#f0fff0,stroke:#59a14f,stroke-width:2px
style recovery fill:#fff5f5,stroke:#e15759,stroke-width:2px,stroke-dasharray: 5 5
style handoff fill:#f0f8ff,stroke:#4e79a7,stroke-width:1px
Tool inventory (10 scene-conditional tools)
| Tool | Args | Precondition | Layer |
|---|---|---|---|
query_scene |
— | — | HIGH (perception) |
move_to_pose |
target, standoff, speed | target in scene | LOW (A* nav) |
execute_skill |
skill (pick/place), target | nearness, gripper state | LOW (manipulation) |
open_door |
door_id | near door, door unlocked | HIGH (agent decision) |
unlock_door |
door_id, key_id | near door, holding key | HIGH (agent decision) |
press_button |
button_id | near button | HIGH (agent decision) |
call_elevator |
elevator_id, target_floor | near elevator, same floor | HIGH (agent decision) |
push_object |
object_id, direction | near object, object pushable | HIGH (agent decision) |
inspect_object |
object_id | object in scene | HIGH (perception) |
request_assistance |
message | — | HIGH (meta-cognition) |
Scene-conditional: query_scene returns available_tools based on what objects are present.
A scene with no elevator will not list call_elevator.
Full design: docs/design/AGENT_UPGRADE.md,
docs/design/architecture.md, and
docs/design/REBUILD.md.
| Capability | Before (v1) | Current (v2) |
|---|---|---|
| Tools | 3 fixed | 10 scene-conditional |
| Task types | 1 (fetch & place) | 6 (locked door, elevator, blocked path, sequential, infeasible) |
| Replanning | Blind regeneration | Structured failure diagnosis → targeted repair |
| Multi-agent | None | Cooperative handoff (request_assistance → Agent 2) |
| Eval scenarios | 10 | 18 (target pass rate 60-80%) |
| Tests | 269 | 352 |
Design: docs/design/AGENT_UPGRADE.md |
Decisions: DECISIONS.md |
Results: benchmarks/
| physgate | SayCan (Google) | Inner Monologue (Google) | Reflexion | |
|---|---|---|---|---|
| Orchestration | LangGraph state machine, typed state, conditional edges | Affordance scoring loop | Text feedback loop | Verbal memory loop |
| Control flow | Explicit graph with DI-injected components | Implicit while loop | Implicit while loop | Implicit while loop |
| Verification | 4-layer gate: critic → L1 kinematic → L3 scene → L2 physics | Flat affordance check | Text self-check | Verbal self-critique |
| Failure recovery | Structured FailureReport → targeted repair (keep prefix) → full regen → escalate |
None | Text replan | Verbal reflection |
| Multi-agent | Cooperative handoff (request_assistance → Agent 2) |
None | None | None |
| Testability | All components injected; Mock/LLM/Isaac swappable per call site | Tied to hardware | Tied to hardware | Tied to LLM |
| Tool selection | LLM selects from 10 scene-conditional tools (no template) | Fixed skill set | Fixed skill set | Fixed action set |
| Hardware evidence | Isaac Sim 3D (Go2 walking policy) | Real robot | Real robot | Simulation only |
Key design decisions:
- Plan-then-execute, not ReAct. Robotics requires a complete plan for physics validation before any motor command — step-by-step ReAct cannot be gate-validated.
- State machine, not ad-hoc loop. The replan budget × failure analyst × handoff × approval gate flow is complex enough to justify LangGraph over a while loop.
- Dependency injection.
planner_fn/critic_fn/gate_fn/executor_fn/failure_analyst_fnare all injected callables — the orchestrator graph never touches an LLM or GPU directly. Tests swap in mocks; production swaps in Claude + Isaac Sim.
The core experiment: the same contaminated plan pool (clean + defect-injected plans) and the same generated task instances, run through four pipeline configurations. Without validation, defective plans execute and fail; without physics/navigation-level validation, impossible tasks are executed instead of rejected.
| Pipeline | Success rate (feasible tasks, 95% CI) | Rejection rate (infeasible tasks) |
|---|---|---|
| A0 — No validation | 0.43 [0.43, 0.43] | 0.00 |
| A1 — + LLM critic | 0.50 [0.50, 0.50] | 0.00 |
| A2 — + Symbolic gate | 1.00 [0.93, 1.00] | 0.00 |
| A3 — + Nav-aware gate (symbolic + A*) | 1.00 [0.93, 1.00] | 1.00 |
| A3_physics — + Isaac Sim GPU gate† | 1.00 [0.57, 1.00] | — |
†A3_physics: real Isaac Sim rigid-body physics on GPU (n = 5+2, RTX PRO 6000 Blackwell). Success matches the symbolic gate — the physics simulation validates the same plans the symbolic gate accepts. This confirms the symbolic gate is a faithful surrogate for the full physics simulation at much lower cost.
n = 50 feasible + 20 infeasible
procedurally generated layouts (A*-certified labels); plan pool of
21 plans. CIs are instance-level
(n = number of independent generated layouts, not plan × instance):
Wilson CI for binary per-instance outcomes (A2/A3), normal-approximation
CI for fractional per-instance rates (A0/A1). Methodology:
docs/design/EVALUATION_METHODOLOGY.md.
Every validation layer scored as a binary classifier over a labeled corpus of clean and systematically corrupted plans (defect taxonomy from the plan-verification literature). This substantiates the layered-defense claim: each layer catches strictly more defect classes, with zero false positives.
| Validation layer | Precision | Recall | F1 | False-positive rate | Wall time |
|---|---|---|---|---|---|
| LLM critic | 1.00 | 0.20 | 0.33 | 0.00 | 0.0 s |
| + L1/L3 deterministic checks | 1.00 | 0.60 | 0.75 | 0.00 | 0.0 s |
| + Symbolic L2 (full gate) | 1.00 | 0.80 | 0.89 | 0.00 | 0.0 s |
| + Nav-aware gate (symbolic + A*) | 1.00 | 1.00 | 1.00 | 0.00 | 0.7 s |
| Isaac physics L2 | 1.00 | 1.00 | 1.00 | 0.00 | 76.4 s |
The symbolic gate catches all plan-level defects (D1–D4) but misses world-level defects (D5 unreachable goal: 0% rejection). Only the nav-aware gate (A* path compilation) and physics gate catch D5 — this is the quantified, unique value of world-level validation.
Before the rebuild, "feasibility" measured whether the LLM happened to name a hand-placed rescue waypoint (obstacle avoidance in the wrong layer). After the rebuild (deterministic A* navigation), every well-formed decomposition is physically feasible:
| Planner | Pre-rebuild (artifact) | Post-rebuild | GPU wall |
|---|---|---|---|
| Mock planner | 2/8 (25%) | 6/6 (100%) | 13.2 s |
| Real Claude | 1/7 (14%) | 8/8 (100%) | 14.1 s |
Repeat statistics across 3 independent runs (fresh LLM generations each time): mock 18/18 plans; Claude 24/24 plans — 100% ± 0%.
18 scenarios across 6 task types (T1-T6) with 10 tools, run through the LangGraph orchestrator with fault injection, failure analysis, and cooperative handoff. Tested with real Claude Opus 4.8 (not mock).
| Before (demo) | After (upgraded) | |
|---|---|---|
| Tools | 3 fixed | 10 scene-conditional |
| Task types | 1 | 6 (multi-step dependencies) |
| Replanning | Blind regeneration | Failure diagnosis → targeted repair |
| Multi-agent | None | Cooperative handoff |
| Planner | MockPlanner (template) | Claude Opus 4.8 (real LLM) |
| Mock baseline | Claude (blind) | Claude (+ failure analyst) | |
|---|---|---|---|
| Scenarios correct | 13/18 (72%) | 17/18 (94%) | 18/18 (100%) |
| End-to-end success | 0.583 | 0.917 | 1.00 |
| Replan efficiency | 0.0 | 0.571 | 0.857 |
| Tool selection accuracy | 0.714 | 1.00 | 1.00 |
Infeasible recognition, recovery, and handoff accuracy are 1.00 across all three conditions — the current scenarios are not hard enough on those axes to differentiate planners. The meaningful signal is in end-to-end success and tool selection, where real LLM reasoning dominates the template baseline.
Per-category breakdown (Claude Opus 4.8):
| Task type | Scenarios | Blind | + Analyst | What the LLM must figure out |
|---|---|---|---|---|
| T1 fetch & place | 5 | 4/5 | 5/5 | Step ordering, preconditions, recovery |
| T2 locked door | 2 | 2/2 | 2/2 | key → unlock_door → open_door chain |
| T3 blocked path | 2 | 2/2 | 2/2 | inspect → push_object or A* detour |
| T4 sequential | 1 | 1/1 | 1/1 | Multi-object ordering dependency |
| T5 elevator | 1 | 1/1 | 1/1 | call_elevator for cross-floor delivery |
| T6 infeasible | 4 | 4/4 | 4/4 | inspect → request_assistance |
| All | 18 | 17/18 | 18/18 |
| Condition | Correct | The difference |
|---|---|---|
| Blind replan | 17/18 | Scene 3 (occupied gripper): LLM replans from scratch, repeats same mistake → escalated |
| + Failure analyst | 18/18 | Analyst: "gripper occupied at step 4, insert place step" → LLM fixes only that → done |
Per-scenario results (Claude Opus 4.8, full confirmed run)
| # | Scenario | Category | Blind | Targeted |
|---|---|---|---|---|
| 1 | fetch_and_place_basic | ordering | done ✓ (91s) | done ✓ (167s) |
| 2 | gate_catches_place_before_pick | ordering | done ✓ (13s) | done ✓ (14s) |
| 3 | precondition_occupied_gripper | preconditions | escalated ✗ (595s) | done ✓ (145s) |
| 4 | recovery_transient_pick_failure | recovery | done ✓ (176s) | done ✓ (171s) |
| 5 | recovery_persistent_failure_escalates | recovery | escalated ✓ (315s) | escalated ✓ (254s) |
| 6 | multi_step_two_boxes | multi_step | done ✓ (238s) | done ✓ (201s) |
| 7 | infeasible_ungraspable_object | infeasible | partial_success ✓ (64s) | partial_success ✓ (68s) |
| 8 | ambiguous_multi_shelf | multi_step | done ✓ (82s) | done ✓ (66s) |
| 9 | partial_infeasible_two_tasks | infeasible | partial_success ✓ (98s) | partial_success ✓ (86s) |
| 10 | recovery_place_failure_replan | recovery | done ✓ (331s) | done ✓ (240s) |
| 11 | t2_locked_door_delivery | locked_door | done ✓ (156s) | done ✓ (311s) |
| 12 | t2_locked_door_no_key_escalate | locked_door | partial_success ✓ (191s) | partial_success ✓ (161s) |
| 13 | t3_blocked_path_push | blocked_path | done ✓ (136s) | done ✓ (137s) |
| 14 | t3_blocked_path_immovable | blocked_path | done ✓ (445s) | done ✓ (221s) |
| 15 | t4_sequential_delivery | sequential | done ✓ (167s) | done ✓ (151s) |
| 16 | t5_elevator_delivery | elevator | done ✓ (99s) | done ✓ (94s) |
| 17 | t6_too_heavy | assistance | partial_success ✓ (68s) | partial_success ✓ (67s) |
| 18 | t6_sealed_room | assistance | partial_success ✓ (70s) | partial_success ✓ (59s) |
Scene 3 is the A5 differentiator: blind escalated in 595s, targeted repair succeeded in 145s. Total wall time: blind 56 min, targeted 44 min (structured repair avoids wasted retries).
| Parallel envs | env-steps/s | Scaling efficiency |
|---|---|---|
| 1 | 281 | 1.00 |
| 4 | 1,118 | 0.99 |
| 16 | 4,360 | 0.97 |
| 64 | 16,237 | 0.90 |
| 256 | 64,249 | 0.89 |
| 1024 | 250,227 | 0.87 |
Efficiency declines monotonically; the saturation knee was not reached in the measured range — 1024 envs is the largest measured point, not a limit.
This section and its charts are generated by benchmarks/render_readme.py from the JSONs in benchmarks/*/results/ — regenerate after re-running any benchmark; do not edit by hand.
# Pure-logic pipeline (no GPU, no API key needed)
pip install -e ".[dev]"
pytest # 352 tests, all pure-logic
python examples/fetch_and_place.py # T1: offline fetch-and-place demo
python examples/cooperative_delivery.py # T3→T1: two-agent cooperative handoff
python benchmarks/orchestration/run_eval.py --planner mock # 18-scenario eval
# With real LLM planning — either credential works:
ANTHROPIC_API_KEY=sk-ant-api03-... python examples/fetch_and_place.py # API key
CLAUDE_CODE_OAUTH_TOKEN=sk-ant-oat01-... python examples/fetch_and_place.py
# ^ Claude subscription token (from `claude setup-token`); routed through
# Claude Code headless mode automatically — see DECISIONS.md D-015
# With Isaac Sim physics validation + execution (requires the env_isaaclab venv,
# see scripts/install_sim_stack.sh and DECISIONS.md D-005)
source ~/env_isaaclab/bin/activate
pytest tests/test_isaac_sim_gate.py # Isaac integration tests
python examples/fetch_and_place.py --isaacVerified known-good stack on NVIDIA RTX Pro 6000 Blackwell (sm_120), June 2026:
| Layer | Version |
|---|---|
| OS | Ubuntu 22.04.5 LTS |
| NVIDIA driver | 580.65.06 (not 595.x) |
| CUDA / PyTorch | 12.8 / 2.7.0+cu128 |
| Isaac Sim / Isaac Lab | 5.1.0 / 2.3.2 |
| ROS 2 | Humble |
| Planner LLM | Claude Opus 4.8 (cloud API — not local) |
Note
The RTX Pro 6000 has a known sustained-compute chip-reset issue (drivers 570–595,
unresolved by NVIDIA as of mid-2026). Phase 0 benchmark #1 stress-tests for this;
mitigations are power-capping (nvidia-smi -pl 400) and keeping LLM inference on a
cloud API rather than the local GPU.
physgate is an application-layer (Layer 4) soft interlock. It sits above — and does not replace — the hardware safety layers that a certified deployment requires:
Layer 4 physgate (LLM planner + Sim-Gate + RL fallback) ← soft interlock, NOT certified
Layer 3 Motion control / ROS 2 / unitree_sdk2 ← not safety-rated
Layer 2 Safety PLC / FSoE (PLd/SIL2) ← does NOT exist on Go2
Layer 1 Hardware E-stop / SRSF ← does NOT exist on Go2
Layer 0 Mechanical limits / motor current cutoff (firmware)
- Not a safety function under ISO 10218-1:2025 / ISO 13849-1 / IEC 61508 / ISO 3691-4.
- The RL fallback is not formally verified — no barrier certificate; Black-Box Simplex (arXiv:2102.12981) / ASTM F3269-21 guarantees do not hold. It is a software watchdog with an RL fallback.
- The Unitree Go2 is a research platform with no certified hardware safety layer (FCC/CE radio compliance only).
- EU Machinery Regulation 2023/1230 (mandatory 2027-01): AI safety components with self-evolving behavior require Notified Body assessment. physgate does not satisfy this and must not be represented as safety validation in EU deployments.
What physgate CAN provide: a meaningful reduction in the probability of kinematically infeasible, contextually inappropriate, or unreviewed commands reaching the robot, compared to unmediated LLM-to-robot execution — an engineering safeguard for research settings, not a certified safety function.
PROVIDED "AS IS" WITHOUT WARRANTY. Do not deploy near persons without hardware safety infrastructure independent of this software.
- Symbolic-only for new tools. T2-T6 task types (doors, elevators, pushing) run on the MockWorldBackend. Isaac Sim physics integration for these tools is planned but not yet implemented — T1 (fetch-and-place) is the only task validated in full rigid-body physics.
- Single robot. Validated only on the Unitree Go2 quadruped with the rsl_rl flat locomotion policy. Other morphologies (arms, wheels, humanoids) are out of scope.
- Cooperative handoff, not negotiation. Multi-agent coordination is a two-orchestrator handoff pattern. Full multi-agent negotiation (bidding, conflict resolution, shared world state) is not implemented.
- No real-robot deployment. All validation is in simulation (Isaac Lab PhysX). ROS 2 executor interface is defined but NOT implemented.
- Sim-to-real gap. A simulation PASS reduces risk but does not prove real-world safety. Contact dynamics, sensor noise, and communication latency are not modeled.
- Sample size. Pipeline ablation uses n=50+20 procedurally generated instances; orchestration eval uses 18 hand-designed scenarios. Both are below cited standards (SafeAgentBench 750, PlanBench 600).
MIT © 2026 Wei Cheng (Wayne) Chiu






