Skip to content

Latest commit

 

History

History
306 lines (227 loc) · 16.4 KB

File metadata and controls

306 lines (227 loc) · 16.4 KB

nullstate: Autonomous Purple-Team IaC Sandbox on AMD MI300X

1. Executive summary

I built nullstate during the AMD x lablab.ai hackathon to test whether infrastructure security validation could move beyond static IaC findings into a repeatable red-team/blue-team loop. The CLI reads Terraform projects, starts a local sandbox target, detects dangerous cloud storage exposure, asks a red-team model to reason about the attack path, executes a constrained generated attack script against the local target, asks a blue-team model to explain remediation, applies a deterministic Terraform patch, reruns validation, and writes judge-ready evidence artifacts. The final demo used LocalStack sandboxes, Terraform automation commands, a Python Typer/Rich CLI, OpenAI-compatible vLLM endpoints, ROCm, and an AMD Instinct MI300X GPU droplet. The key engineering decision was to make the security verdict deterministic while using the model for adversarial reasoning, explanation, and report quality. The result was a working hackathon prototype with final AWS and Azure runs showing vulnerable infrastructure, exploit evidence, patched Terraform, and blocked attack paths.

2. Problem

IaC scanners are useful, but they often stop at "this configuration looks risky." A platform or security engineer still has to answer harder operational questions:

  • Is the finding actually exploitable in this context?
  • What evidence proves the exposure?
  • What exact Terraform change remediates it?
  • Did the attack path fail after the patch?
  • Can this be demonstrated without touching production cloud accounts?

nullstate targets that gap. The user should be able to point the CLI at an IaC project and get a reproducible evidence trail: finding, attack reasoning, remediation, validation, and a report that a security, cloud, or DevSecOps reviewer can inspect.

3. Context and constraints

  • Hackathon build window: roughly 48 hours.
  • Final submission pressure: working evidence needed before the deadline, not a broad but unfinished platform.
  • Cloud target: local sandboxes first, no real Azure or AWS credentials by default.
  • Model target: self-hosted endpoint on AMD MI300X through vLLM/ROCm.
  • Runtime target: LocalStack for AWS and Azure-style scenarios.
  • UX target: CLI-first instead of a full web application.
  • Security constraint: do not expose vLLM, LocalStack, cloud tokens, Terraform state, or .env values publicly.
  • Reliability constraint: model output must not be the source of truth for the pass/fail verdict.

4. Requirements

Functional requirements

  • Analyze Terraform IaC.
  • Infer or accept a scenario target.
  • Start and inspect local sandbox backends.
  • Detect high-risk public cloud storage exposure.
  • Generate red-team attack reasoning with a self-hosted model endpoint.
  • Execute a constrained generated attack script before and after remediation.
  • Log attack command, stdout, stderr, return code, target URL, stage, start time, end time, and duration into events.jsonl.
  • Generate blue-team remediation explanation with a self-hosted model endpoint.
  • Apply deterministic Terraform remediation.
  • Re-run validation after remediation.
  • Produce report.md, findings.json, events.jsonl, metrics.json, attack.py, remediation.patch, and remediation.json.
  • Support offline/mock mode so the demo can still run without GPU or sandbox access.

Non-functional requirements

  • No real cloud execution by default.
  • No arbitrary red-agent shell execution in V1; red command execution is limited to generated run-directory attack.py scripts.
  • Commands should be understandable under hackathon demo pressure.
  • Run artifacts should be suitable for later case-study evidence.
  • Secrets and provider state must be kept out of Git.
  • CI/CD should use PR checks, dependency review, CodeQL, and branch protection.

5. Architecture

flowchart LR
    A[Terraform IaC input] --> B[Terraform init / plan / show JSON]
    B --> C[Deterministic detector]
    C --> D[Sandbox adapter]
    D --> E[Red-team model reasoning]
    E --> F[Attack evidence]
    F --> G[Blue-team model remediation]
    G --> H[Deterministic Terraform patch]
    H --> I[Re-plan and apply]
    I --> J[Validation attack]
    J --> K[Report and metrics artifacts]
Loading

The CLI separates reliable security decisions from model-generated explanation:

  • Terraform automation runs terraform init -input=false, terraform plan -out=tfplan -input=false, terraform show -json tfplan, and terraform apply -auto-approve -input=false tfplan for live runs.
  • The deterministic detector identifies supported exposures from Terraform configuration and plan JSON.
  • The red-team model receives the finding evidence and writes attack reasoning.
  • The blue-team model receives the finding evidence and writes remediation guidance.
  • The deterministic remediation engine applies the Terraform patch.
  • The validator checks whether findings remain and records the post-remediation attack result.

6. What the red agent actually does today

This is important to describe accurately.

The red model itself does not receive shell access. It receives an internal system prompt and scenario findings, then returns attack reasoning such as anonymous S3 reads or Azure Blob curl requests.

After the reasoning step, nullstate executes only the generated attack.py script inside the run directory through a constrained runner. The runner invokes the current Python interpreter directly, passes only --target-url and --stage, and records command, stdout, stderr, return code, target URL, stage, start time, end time, and duration into events.jsonl.

This is intentionally narrower than a free-form agent shell. It gives the demo real command evidence against local sandbox endpoints while preserving a reliable and auditable boundary.

7. Security model

Risk Control Evidence
Accidental real cloud attack LocalStack and plan-only targets by default sandbox adapters and run commands
Model makes unsafe recommendation deterministic detector and versioned remediation remain source of truth findings.json, remediation.patch, remediation.json
Arbitrary exploit execution generated attack.py only, no arbitrary shell red-tool events, src/nullstate/attack_runner.py
Secret leakage .env and Terraform state ignored; screenshots must be redacted repo hygiene and submission checklist
Public model endpoint exposure vLLM bound through SSH tunnel, not public ingress droplet setup and tunnel workflow
Unreviewed code changes PR workflow, branch protection, CI checks GitHub PRs and checks

8. Deployment and DevSecOps workflow

The project was built with a branch and PR workflow instead of direct main edits:

feature branch
-> PR
-> unit tests, lint, type checks, dependency/security checks
-> review/merge
-> release tag / release notes

Repository practices added during the build:

  • Python package metadata and console entrypoint.
  • Structured docs: architecture, runbook, threat model, security model, CI/CD, cost notes, and failure modes.
  • GitHub Actions for quality checks.
  • CodeQL and dependency review.
  • Dependabot-style dependency hygiene.
  • Case-study evidence checklist kept out of Git.
  • Release-oriented workflow with tagged submission artifacts.

9. Final demo results

AWS public S3 scenario

Run artifact:

runs/final-aws-gemma26b/20260510-170931

Finding:

HIGH AWS_S3_PUBLIC_ACCESS_BLOCK_DISABLED
Resource: aws_s3_bucket_public_access_block.public_logs
Evidence: block_public_acls, block_public_policy, ignore_public_acls, and restrict_public_buckets were disabled.

Remediation patch:

-  block_public_acls       = false
-  block_public_policy     = false
-  ignore_public_acls      = false
-  restrict_public_buckets = false
+  block_public_acls       = true
+  block_public_policy     = true
+  ignore_public_acls      = true
+  restrict_public_buckets = true

Result:

Red before: success
Red after: blocked
Verdict: Exploit blocked after remediation

Model metrics from metrics.json:

Role Model Completion tokens Latency Output speed
Red nullstate-gemma4-26b-a4b 942 9.018s 104.452 tok/s
Blue nullstate-gemma4-26b-a4b 640 3.977s 160.937 tok/s

Azure public Blob scenario

Run artifact:

runs/final-azure-gemma26b/20260510-174858

Finding:

HIGH AZURE_STORAGE_PUBLIC_BLOB
Resource: azurerm_storage_container.secrets
Evidence: container_access_type is "container"; storage account allows nested items to be public.

Remediation patch:

-  allow_nested_items_to_be_public  = true
+  allow_nested_items_to_be_public  = false

-  container_access_type = "container"
+  container_access_type = "private"

Result:

Red before: success
Red after: blocked
Verdict: Exploit blocked after remediation

Model metrics from metrics.json:

Role Model Completion tokens Latency Output speed
Red nullstate-gemma4-26b-a4b 709 4.112s 172.424 tok/s
Blue nullstate-gemma4-26b-a4b 794 4.548s 174.586 tok/s

10. Evidence artifacts

Each run creates:

  • events.jsonl: timeline of Terraform, analysis, red-tool command execution, red-team, blue-team, and validation events.
  • findings.json: structured finding data.
  • attack.py: generated scenario attack artifact.
  • remediation.patch: deterministic Terraform diff.
  • remediation.json: deterministic remediation ruleset version, changed files, and applied rule IDs.
  • metrics.json: model call token counts, latency, throughput, and endpoint metrics.
  • report.md: human-readable case-study report.
  • run-bundle.json: schema-addressed portable evidence contract for local dashboards, support bundles, CI upload, and future cloud ingestion.
  • dashboard.html: free local single-run dashboard for non-terminal review.
  • workspace/: copied Terraform workspace for reproducibility.

Screenshots to add before publishing:

  • nullstate root command showing logo and next commands.
  • nullstate run summary for AWS or Azure.
  • report.md showing Exploit blocked after remediation.
  • DigitalOcean MI300X droplet GPU page.
  • vLLM /v1/models response for nullstate-gemma4-26b-a4b.
  • vLLM /metrics or metrics.json token throughput evidence.
  • GitHub PR/checks/release page.

11. Cost analysis

The main cost driver was the AMD MI300X GPU droplet used for model serving. LocalStack and Terraform ran locally, and the CLI did not require paid cloud resources for AWS or Azure infrastructure. The cost-control decision was to keep real cloud execution out of scope and use the GPU only long enough to collect model-serving evidence and final run metrics.

Final portfolio version should include:

  • Total DigitalOcean credits used.
  • Droplet runtime hours.
  • Whether the droplet was destroyed after evidence collection.
  • Any failed boot/model attempts that consumed GPU time.

12. Tradeoffs

Decision Alternative Why chosen Downside
CLI-first product Next.js dashboard Faster to build, better for terminal evidence Less visual for non-technical judges
Deterministic security core Pure LLM scanner Reliable pass/fail and reproducible patching Narrower scenario coverage in V1
LocalStack sandbox Real AWS/Azure Safer and no production credentials Emulator differences and setup friction
One large model for red and blue Separate red/blue models Avoided dual-container VRAM/runtime failures Less role specialization
OpenAI-compatible endpoint with provider presets Provider-specific SDK per model vendor Works with vLLM, SGLang, managed Gemini, Claude compatibility, or custom proxies Native provider features need dedicated adapters later
Constrained red runner Free-form shell agent Real command evidence without arbitrary tool access Scripts are narrow and scenario-specific

13. Failure modes and lessons learned

  • DigitalOcean account access blocked early setup, so GPU work started later than planned.
  • LocalStack containers can leave port 4566 reserved; the CLI now gives a specific recovery hint.
  • Editable Python installs break when the repo path changes; production installs should use the package entrypoint.
  • Running two model containers on one MI300X caused VRAM pressure and cache allocation failures.
  • ROCm AITER fused-MoE paths failed for one Gemma startup attempt, so the final path used a single larger model endpoint.
  • SGLang/Qwen3.5 experimentation was deprioritized after runtime import issues.
  • Model metrics from a local Windows run cannot capture amd-smi directly unless collected on the droplet; endpoint metrics and droplet screenshots fill that evidence gap.

14. What I would improve next

  • Expand allowlisted attack.py scripts into richer scenario-specific exploit probes.
  • Add per-scenario command allowlists beyond Python scripts when needed.
  • Add nullstate run --auto-sandbox for known local targets.
  • Add custom LocalStack port support so AWS and Azure sandboxes can run side by side.
  • Expand artifact redaction coverage before publishing reports.
  • Expand release SBOM quality, package signing, and provenance verification.
  • Add more Azure, AWS, Kubernetes, Docker Compose, and on-prem digital-twin scenarios.
  • Build a portfolio demo page with embedded screenshots, architecture diagram, and video.

15. Enterprise next iteration

The hackathon build proves the architecture: deterministic IaC detection, LocalStack sandbox workflow, model-assisted red/blue reasoning, constrained red command execution, deterministic remediation, and evidence artifacts.

The next enterprise-grade step is making the command evidence deeper. Today, the safe runner boundary exists and every red command is logged, but the AWS/Azure attack scripts are still narrow probes. The product should evolve each scenario from “configuration finding plus constrained probe” into “observed runtime exploit evidence plus deterministic validation.”

For the AWS S3 scenario, that means creating an intentionally exposed evidence object in LocalStack, attempting anonymous object reads before remediation, blocking the path after remediation, and recording both command outputs in the report. For Azure Blob, the same pattern applies where LocalStack Azure supports the required blob APIs; otherwise the report should clearly label emulator limitations instead of overclaiming runtime exploitation.

The business-grade principle is simple: if a report says “exploited,” the run should contain observed command evidence. If it is offline or emulator-limited, the report should label the result as deterministic simulation or inconclusive runtime evidence. That distinction is what makes the tool credible for security teams.

16. Repository and demo links

  • GitHub: https://github.com/Ker102/nullstate-cli
  • Suggested portfolio slug: /case-studies/nullstate-autonomous-purple-team-iac-sandbox
  • Demo video: pending
  • Final AWS report: runs/final-aws-gemma26b/20260510-170931/report.md
  • Final Azure report: runs/final-azure-gemma26b/20260510-174858/report.md

17. Interview explanation

I built nullstate because IaC security tools often identify misconfigurations without proving exploitability or remediation effectiveness. The architecture uses Terraform plan analysis, LocalStack sandboxes, deterministic detection, constrained attack script execution, and self-hosted model agents to create a repeatable purple-team loop. The most important security decision was to keep the pass/fail verdict deterministic and limit red command execution to generated attack.py scripts instead of arbitrary shell access. The hardest tradeoff was choosing reliability over full autonomy under hackathon time pressure. If I rebuilt it, I would expand the constrained runner into richer per-scenario allowlists with deeper network probes against the local sandbox.

18. Resume bullets

  • Built nullstate, a Python DevSecOps CLI that validates Terraform IaC through a local red-team/blue-team loop with deterministic remediation and evidence artifacts.
  • Served a large model on AMD MI300X with vLLM/ROCm and captured token, latency, throughput, and endpoint metrics for final AWS and Azure sandbox runs.
  • Implemented Terraform analysis, LocalStack sandbox workflows, structured run artifacts, report generation, CI checks, threat model, runbook, and release-oriented PR workflow.
  • Produced final demo evidence showing public S3 and Azure Blob exposures remediated and revalidated as blocked attack paths.