| [aw] Failure Investigator (6h) |
1 |
100.0% |
0.0% |
"Were fix sub-issues created for unresolved failures, or were resolved tracking issues closed?" (0.0% YES) |
| AI Moderator |
1 |
100.0% |
0.0% |
"Did the agent apply at least one label (spam, ai-generated, link-spam, or ai-inspected) or call noop?" (0.0% YES) |
| Auto-Triage Issues |
1 |
100.0% |
100.0% |
"Did the agent apply at least one label to an unlabeled issue, or correctly call noop when no unlabeled issues were found?" (100.0% YES) |
| CLI Version Checker |
1 |
100.0% |
0.0% |
"Did the agent check for new versions and digest changes of Docker images in pkg/cli/docker_images.go (actionlint, syft, grype, grant, zizmor, poutine, runner-guard, yamllint)?" (0.0% YES) |
| Code Scanning Fixer |
1 |
100.0% |
0.0% |
"Was a pull request created with a remediation for a code scanning alert, or was noop used when no fixable alerts existed?" (0.0% YES) |
| Contribution Check |
1 |
100.0% |
0.0% |
"Did the agent dispatch PRs to the contribution-checker subagent and produce evaluation results for each PR?" (0.0% YES) |
| Copilot Session Insights |
1 |
100.0% |
0.0% |
"Was a report produced with usage patterns, success rates, and performance metrics?" (0.0% YES) |
| Daily AgentRx Trace Optimizer |
1 |
100.0% |
0.0% |
"Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed?" (0.0% YES) |
| Daily Assign Issue To User |
1 |
100.0% |
0.0% |
"Did the agent post a comment explaining the assignment decision?" (0.0% YES) |
| Daily Cli Tools Tester |
2 |
100.0% |
100.0% |
"Did the agent run exploratory tests on the audit, logs, and compile CLI tools?" (100.0% YES) |
| Daily Container Image Security Scan |
1 |
100.0% |
100.0% |
"Did the agent analyze container images for vulnerabilities, updates, and rejected licenses?" (100.0% YES) |
| Daily Safe Outputs Conformance Checker |
1 |
100.0% |
0.0% |
"Were agentic tasks created for Copilot to address critical/high/medium/low issues, or was noop used when implementation was compliant?" (0.0% YES) |
| Daily VulnHunter Scan |
1 |
100.0% |
0.0% |
"Was a security issue created for verified exploitable findings, or was noop used when VulnHunter found nothing actionable?" (0.0% YES) |
| Daily Windows Terminal Integration Builder |
1 |
100.0% |
100.0% |
"Did the agent create an issue for an actionable integration failure, or use noop when no action was required?" (100.0% YES) |
| Daily Workflow Updater |
1 |
100.0% |
0.0% |
"Did the agent create a pull request for required updates, or report that no changes were needed?" (0.0% YES) |
| Dependabot Burner |
1 |
100.0% |
0.0% |
"Did the agent analyze the selected grouped Dependabot remediation batch?" (0.0% YES) |
| Design Decision Gate 🏗️ |
2 |
100.0% |
0.0% |
"Did the agent add a PR comment, push a draft ADR, or call noop?" (0.0% YES) |
| Discussion Task Miner - Code Quality Improvement Agent |
1 |
100.0% |
0.0% |
"Does the agent output confirm that the created issues include the expected labels (code-quality, automation, task-mining)?" (0.0% YES) |
| ESLint Refiner |
1 |
100.0% |
0.0% |
"Did the agent report actionable ESLint rule refinements or explain why no refinement was needed?" (0.0% YES) |
| Issue Arborist |
1 |
100.0% |
0.0% |
"Did the agent analyze recent issues and identify related issue relationships?" (0.0% YES) |
| Issue Monster |
2 |
100.0% |
0.0% |
"Does the agent output show that at most one issue was assigned to Copilot per run?" (0.0% YES) |
| Multi-Device Docs Tester |
1 |
100.0% |
100.0% |
"Did the agent test the documentation site across the requested device form factors?" (100.0% YES) |
| PR Code Quality Reviewer |
2 |
100.0% |
0.0% |
"Does the agent output show that the review findings are limited to changes in the pull request diff rather than unrelated code?" (0.0% YES) |
| PR Sous Chef |
2 |
100.0% |
100.0% |
"Did the agent add a comment to at least one pull request?" (100.0% YES) |
| PR Triage Agent |
1 |
100.0% |
0.0% |
"Does the agent output include a triage report summarizing the PRs processed?" (0.0% YES) |
| Sub-Issue Closer |
1 |
100.0% |
100.0% |
"Did the agent check parent issues for the completion status of all their sub-issues?" (100.0% YES) |
| Test Quality Sentinel |
2 |
100.0% |
0.0% |
"Does the agent output show that the objective for experiment model_size was successfully completed?" (0.0% YES) |
| Tidy |
1 |
100.0% |
100.0% |
"Was a pull request created with formatting and tidying fixes, or was noop used when no changes were needed?" (100.0% YES) |
Executive Summary
Across 28 workflows and 34 runs, the evals feature is
DEGRADED: the eval job itself is healthy, but the overall YES rate is only 48.6%. That low YES rate is driven by 38 non-YES answers, including 32UNKNOWNlabels and 6 explicitNOs.Note
Status: DEGRADED - 100.0% eval-job success and 48.6% overall YES rate.
Key Metrics
Per-Workflow Pass Rates
Per-Question Breakdown per Workflow
[aw] Failure Investigator (6h)
AI Moderator
Auto-Triage Issues
CLI Version Checker
Code Scanning Fixer
Contribution Check
Copilot Session Insights
Daily AgentRx Trace Optimizer
Daily Assign Issue To User
Daily Cli Tools Tester
Daily Container Image Security Scan
Daily Safe Outputs Conformance Checker
Daily VulnHunter Scan
Daily Windows Terminal Integration Builder
Daily Workflow Updater
Dependabot Burner
Design Decision Gate 🏗️
Discussion Task Miner - Code Quality Improvement Agent
ESLint Refiner
Issue Arborist
Issue Monster
Multi-Device Docs Tester
PR Code Quality Reviewer
PR Sous Chef
PR Triage Agent
Sub-Issue Closer
Test Quality Sentinel
Tidy
Quality Signals
UNKNOWNlabels, so the judge is often unable to emit a binary outcome.action-taken: Did the agent add a PR comment, push a draft ADR, or call noop? - 0.0% YES across 2 workflow(s): AI Moderator, Design Decision Gate 🏗️adr-check-performed: Does the agent output confirm that it checked for existing ADRs before deciding on an action? - 0.0% YES across 1 workflow(s): Design Decision Gate 🏗️comment-posted: Did the agent post a comment explaining the assignment decision? - 0.0% YES across 1 workflow(s): Daily Assign Issue To UserRecommendations
Design Decision Gate 🏗️,AI Moderator, andDaily Assign Issue To User.UNKNOWNoutcomes by making the eval criteria more explicit or by adding clearer evidence requirements for the judge.References