diff --git a/.github/aw/create-agentic-workflow.md b/.github/aw/create-agentic-workflow.md index 77b14174722..2bac7195792 100644 --- a/.github/aw/create-agentic-workflow.md +++ b/.github/aw/create-agentic-workflow.md @@ -86,6 +86,7 @@ Use this mode for exploratory testing, persona walkthroughs, and "what workflow - Do not create or edit workflow files. - Return a compact recommendation covering trigger, any scoped `paths:` filters for file-event triggers, read tools, safe outputs, permissions, and explicit `noop` criteria. - For recurring reports or digests, always include the report window, grouping dimensions, and deduplication key. See [triggers.md](triggers.md) for key-format examples. +- If custom-agent invocation is unavailable in the environment, use fallback evaluation mode: provide the recommendation directly from local gh-aw guidance, record scoring/response quality as unavailable, and continue with partial-results reporting. - Exit ad hoc evaluation mode only when the user explicitly asks to create, implement, or write the workflow file. - End by offering to turn the recommendation into `.github/workflows/.md` if the user wants to proceed. diff --git a/.github/aw/github-agentic-workflows.md b/.github/aw/github-agentic-workflows.md index 7fb20e60a82..01f5012b62d 100644 --- a/.github/aw/github-agentic-workflows.md +++ b/.github/aw/github-agentic-workflows.md @@ -101,6 +101,7 @@ See also: [workflow-constraints.md](workflow-constraints.md) Installed gh-aw agents should support scenario evaluation requests that do not create workflow files. - Treat prompts such as `agentic-workflows evaluate this scenario without creating files` as ad hoc evaluation mode. +- For explicit research/evaluation requests, invoke with wording such as `agentic-workflows evaluate this scenario without creating files` or `agentic-workflows research this workflow pattern and return recommendations only`. - Return a compact design recommendation covering trigger, scope, tools, permissions, safe outputs, `noop` behavior, and any report window / grouping / deduplication requirements. - Offer to turn the recommendation into `.github/workflows/.md` only if the user asks to proceed. diff --git a/.github/workflows/agent-persona-explorer.lock.yml b/.github/workflows/agent-persona-explorer.lock.yml index 70fd22b937b..d6a79b7fc54 100644 --- a/.github/workflows/agent-persona-explorer.lock.yml +++ b/.github/workflows/agent-persona-explorer.lock.yml @@ -1,4 +1,4 @@ -# gh-aw-metadata: {"schema_version":"v4","frontmatter_hash":"e04f94d2d06b8e3de4359e5ba7c7ffea32ab78ec4c970c054500b770e1e320ef","body_hash":"8bf2094bc1d73614458f20e18d8042f9b3b3783dbb7b57642e7827658d05c975","strict":true,"agent_id":"pi","agent_model":"copilot/gpt-5.4","detection_agent_id":"copilot","engine_versions":{"copilot":"1.0.75","pi":"0.82.1"}} +# gh-aw-metadata: {"schema_version":"v4","frontmatter_hash":"e04f94d2d06b8e3de4359e5ba7c7ffea32ab78ec4c970c054500b770e1e320ef","body_hash":"3982db2a61c21eb39fc1255b79d5d4eb63843e0ed02f3dd7e5d2b52c52e3c78d","strict":true,"agent_id":"pi","agent_model":"copilot/gpt-5.4","detection_agent_id":"copilot","engine_versions":{"copilot":"1.0.75","pi":"0.82.1"}} # gh-aw-manifest: {"version":1,"secrets":["COPILOT_GITHUB_TOKEN","GH_AW_GITHUB_MCP_SERVER_TOKEN","GH_AW_GITHUB_TOKEN","GH_AW_OTEL_GRAFANA_AUTHORIZATION","GH_AW_OTEL_GRAFANA_ENDPOINT","GH_AW_OTEL_SENTRY_AUTHORIZATION","GH_AW_OTEL_SENTRY_ENDPOINT","GITHUB_TOKEN"],"actions":[{"repo":"actions/cache/restore","sha":"55cc8345863c7cc4c66a329aec7e433d2d1c52a9","version":"v6.1.0"},{"repo":"actions/cache/save","sha":"55cc8345863c7cc4c66a329aec7e433d2d1c52a9","version":"v6.1.0"},{"repo":"actions/checkout","sha":"3d3c42e5aac5ba805825da76410c181273ba90b1","version":"v7.0.1"},{"repo":"actions/download-artifact","sha":"3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c","version":"v8.0.1"},{"repo":"actions/github-script","sha":"3a2844b7e9c422d3c10d287c895573f7108da1b3","version":"v9.0.0"},{"repo":"actions/setup-go","sha":"b7ad1dad31e06c5925ef5d2fc7ad053ef454303e","version":"v7.0.0"},{"repo":"actions/setup-node","sha":"820762786026740c76f36085b0efc47a31fe5020","version":"v7.0.0"},{"repo":"actions/upload-artifact","sha":"043fb46d1a93c77aae656e7c1c64a875d1fc6a0a","version":"v7.0.1"},{"repo":"docker/build-push-action","sha":"53b7df96c91f9c12dcc8a07bcb9ccacbed38856a","version":"v7.3.0"},{"repo":"docker/setup-buildx-action","sha":"bb05f3f5519dd87d3ba754cc423b652a5edd6d2c","version":"v4.2.0"}],"containers":[{"image":"ghcr.io/github/gh-aw-firewall/agent:0.27.43","digest":"sha256:04e2d1987a565000a8f114b89d806ae7a3864dd4f944be65275b28c93d8690e6","pinned_image":"ghcr.io/github/gh-aw-firewall/agent:0.27.43@sha256:04e2d1987a565000a8f114b89d806ae7a3864dd4f944be65275b28c93d8690e6"},{"image":"ghcr.io/github/gh-aw-firewall/api-proxy:0.27.43","digest":"sha256:d85f57975af5ea23af4996e41ed73fbc8f5b4a47402472bfe82e508f352cb0c1","pinned_image":"ghcr.io/github/gh-aw-firewall/api-proxy:0.27.43@sha256:d85f57975af5ea23af4996e41ed73fbc8f5b4a47402472bfe82e508f352cb0c1"},{"image":"ghcr.io/github/gh-aw-firewall/cli-proxy:0.27.43","digest":"sha256:65c45ea2967984d0024f3df61bc71335658a77ede96c8d9665da7a5f33a795ab","pinned_image":"ghcr.io/github/gh-aw-firewall/cli-proxy:0.27.43@sha256:65c45ea2967984d0024f3df61bc71335658a77ede96c8d9665da7a5f33a795ab"},{"image":"ghcr.io/github/gh-aw-firewall/squid:0.27.43","digest":"sha256:26be5e0b8c8f4c41c8a59126b29bb5d80b07253597472ded2a16bdd75abcbf9d","pinned_image":"ghcr.io/github/gh-aw-firewall/squid:0.27.43@sha256:26be5e0b8c8f4c41c8a59126b29bb5d80b07253597472ded2a16bdd75abcbf9d"},{"image":"ghcr.io/github/gh-aw-mcpg:v0.4.7","digest":"sha256:7545220a9aca134b71e51193ee0eaf4c50756ebf8fbd25a63ae7556e62815c00","pinned_image":"ghcr.io/github/gh-aw-mcpg:v0.4.7@sha256:7545220a9aca134b71e51193ee0eaf4c50756ebf8fbd25a63ae7556e62815c00"},{"image":"ghcr.io/github/gh-aw-node","digest":"sha256:a8082161d7dceda14b68f32eb39d0eaa96b825d07f5895b096afab9d9e0c7748","pinned_image":"ghcr.io/github/gh-aw-node@sha256:a8082161d7dceda14b68f32eb39d0eaa96b825d07f5895b096afab9d9e0c7748"},{"image":"ghcr.io/github/github-mcp-server:v1.8.0","digest":"sha256:d5a18c04b92714c309eb46a2305087e91a4dbd80420f6e462656699f95093520","pinned_image":"ghcr.io/github/github-mcp-server:v1.8.0@sha256:d5a18c04b92714c309eb46a2305087e91a4dbd80420f6e462656699f95093520"}]} # This file was automatically generated by gh-aw. DO NOT EDIT. To debug this workflow, load the skill at https://github.com/github/gh-aw/blob/main/debug.md # diff --git a/.github/workflows/agent-persona-explorer.md b/.github/workflows/agent-persona-explorer.md index 2689feac314..ea9cf066e16 100644 --- a/.github/workflows/agent-persona-explorer.md +++ b/.github/workflows/agent-persona-explorer.md @@ -148,6 +148,7 @@ Store all scenarios in cache memory. - Prompt clarity - Completeness - Notable patterns or issues (be concise) +- If invocation fails: mark the scenario as `invocation_unavailable`, set quality scoring to `N/A`, and continue. **Assessment questions:** Does the suggestion include appropriate triggers (`on:`)? Correct tools (github, web-fetch, playwright, etc.)? Proper safe-outputs? Security best practices (minimal permissions, network restrictions)? A clear, actionable prompt? @@ -161,7 +162,7 @@ For each selected scenario, invoke the "agentic-workflows" custom agent tool, pr - You are ONLY testing the agent's responses, NOT creating actual workflows - **Keep responses focused and concise** - summarize findings instead of verbose descriptions - Aim for quality over quantity - fewer well-analyzed scenarios are better than many shallow ones -- **If any tool call fails, record the error briefly and move on to the next scenario** - do NOT retry or get stuck +- **If any tool call fails, record the error briefly, mark scoring as unavailable for that scenario, and move on to the next scenario** - do NOT retry or get stuck ## Phase 4: Analyze Results (4 minutes) @@ -175,6 +176,7 @@ Review all captured responses and identify: ### Quality Insights (summarize briefly) - Which scenarios received the best responses (average score > 4)? - Which scenarios received weak responses (average score < 3)? +- If scenario invocation failed, note that scoring is unavailable for affected scenarios and exclude them from the numeric average. ### Potential Issues (only list critical issues) - Does the agent ever suggest insecure configurations? @@ -189,7 +191,7 @@ Review all captured responses and identify: **MANDATORY OUTPUT**: Regardless of how many phases completed successfully, you MUST call either the `create issue` or the `noop` safe-output tool before finishing. Failing to call a safe-output tool is the most common cause of workflow failures. -Create a GitHub issue with a **concise** summary report. Use the `create issue` safe-output to publish your findings. Even if only 1-2 scenarios were tested, create the issue with partial results. +Create a GitHub issue with a **concise** summary report. Use the `create issue` safe-output to publish your findings. Even if only 1-2 scenarios were tested, create the issue with partial results. Treat invocation failures as a standard partial-results outcome and explicitly mark scoring as unavailable where applicable. **Issue title**: "Agent Persona Exploration - [DATE]" (e.g., "Agent Persona Exploration - 2024-01-16") @@ -220,7 +222,7 @@ Example: - **Agent**: [name] - **Personas This Run**: [3 persona names] - **Scenarios Tested**: [count - should be 3-4, selected from the 6 generated in Phase 2 (2 per persona × 3 personas)] -- **Average Quality Score**: [X.X/5.0] +- **Average Quality Score**: [X.X/5.0 or N/A when invocation/scoring is unavailable] ### Key Findings (3-5 bullet points max) [High-level insights - keep concise]