docs(evaluator): add agent (task-driven) evaluation guide - #710
Conversation
|
🌿 Preview your docs: https://nvidia-preview-agent-eval-docs-schapman.docs.buildwithfern.com/nemo-platform |
|
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (11)
🚧 Files skipped from review as they are similar to previous changes (9)
📝 WalkthroughWalkthroughAdds Agent Evaluation documentation covering concepts, quickstart usage, HTTP and Harbor targets, metric authoring, trajectory scoring, result artifacts, and evaluator navigation. ChangesEvaluation foundation
Metric contract and local quickstart
Targets and deployed HTTP agents
Component scoring
Harbor evaluation
Possibly related PRs
Suggested labels: Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/evaluator/agent-eval/targets-and-runners.mdx`:
- Around line 54-72: The agent field documentation is incomplete for both
GenericAgent and NemoAgentToolkitAgent. Update the GenericAgent table to include
inherited name and format fields, and add a field table for
NemoAgentToolkitAgent covering format and nat, using the existing
field-description conventions.
In `@docs/fern/versions/latest.yml`:
- Around line 325-336: Remove the four agent-eval page entries—Evaluate a
Deployed Agent over HTTP, Evaluate a Harbor Task Suite, Score by Component, and
Targets and Runners—from the latest-version navigation, while retaining them in
the gated navigation configuration.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 9695657a-8c9f-4006-a78b-067e35434cbf
📒 Files selected for processing (11)
docs/evaluator/agent-eval/evaluate-deployed-agent.mdxdocs/evaluator/agent-eval/harbor-runner.mdxdocs/evaluator/agent-eval/index.mdxdocs/evaluator/agent-eval/quickstart.mdxdocs/evaluator/agent-eval/reading-results.mdxdocs/evaluator/agent-eval/score-by-component.mdxdocs/evaluator/agent-eval/targets-and-runners.mdxdocs/evaluator/agent-eval/writing-metrics.mdxdocs/evaluator/evaluation-approaches.mdxdocs/evaluator/index.mdxdocs/fern/versions/latest.yml
6904e01 to
046e565
Compare
|
@SandyChapman this ties into what I asked about in standup, so apologies that this is a bit repetitive, but I think the docs are a little confusing that they're tied to the sdk layer. For example, I don't follow how to run evals on k8s deployments from this as the quickstart is a local python script. |
046e565 to
f6372b7
Compare
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/evaluator/agent-eval/score-by-component.mdx`:
- Around line 121-124: Align the guessing example with its emitted trace by
consistently describing the response as using the “guess” tool, or remove that
tool call from the trace. Apply the correction in
docs/evaluator/agent-eval/score-by-component.mdx at lines 121-124, 87-89,
188-191, and 295-301, updating the inline comment, narrative, result
interpretation, and full script consistently.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 64b90bef-7ce8-4a28-ac85-d3d8e4c500aa
📒 Files selected for processing (11)
docs/evaluator/agent-eval/evaluate-deployed-agent.mdxdocs/evaluator/agent-eval/harbor-runner.mdxdocs/evaluator/agent-eval/index.mdxdocs/evaluator/agent-eval/quickstart.mdxdocs/evaluator/agent-eval/reading-results.mdxdocs/evaluator/agent-eval/score-by-component.mdxdocs/evaluator/agent-eval/targets-and-runners.mdxdocs/evaluator/agent-eval/writing-metrics.mdxdocs/evaluator/evaluation-approaches.mdxdocs/evaluator/index.mdxdocs/fern/versions/latest.yml
🚧 Files skipped from review as they are similar to previous changes (8)
- docs/evaluator/agent-eval/reading-results.mdx
- docs/evaluator/agent-eval/writing-metrics.mdx
- docs/evaluator/evaluation-approaches.mdx
- docs/evaluator/agent-eval/evaluate-deployed-agent.mdx
- docs/fern/versions/latest.yml
- docs/evaluator/agent-eval/index.mdx
- docs/evaluator/agent-eval/harbor-runner.mdx
- docs/evaluator/agent-eval/quickstart.mdx
f6372b7 to
767dc02
Compare
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/evaluator/agent-eval/quickstart.mdx`:
- Around line 58-62: Update the AgentEvalTask documentation to clarify that
in-process callable agents receiving the full task object can access the
grader-only reference; either explicitly state that such callbacks are trusted
or describe the sanitized task object they receive. Keep the existing
explanation of what external agents see unchanged.
In `@docs/evaluator/agent-eval/score-by-component.mdx`:
- Around line 85-89: Update the documentation paragraph around the deployed
agent and Harbor references to remove the claim that real runners emit traces
automatically. Clarify that the endpoint must return trajectory evidence, such
as by setting trajectory_path for HTTP targets, while Harbor is reward-based
rather than trace-based; preserve the local callable and TrialDraft explanation.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 771b6c80-9866-4abc-a84b-836140127c15
📒 Files selected for processing (11)
docs/evaluator/agent-eval/evaluate-deployed-agent.mdxdocs/evaluator/agent-eval/harbor-runner.mdxdocs/evaluator/agent-eval/index.mdxdocs/evaluator/agent-eval/quickstart.mdxdocs/evaluator/agent-eval/reading-results.mdxdocs/evaluator/agent-eval/score-by-component.mdxdocs/evaluator/agent-eval/targets-and-runners.mdxdocs/evaluator/agent-eval/writing-metrics.mdxdocs/evaluator/evaluation-approaches.mdxdocs/evaluator/index.mdxdocs/fern/versions/latest.yml
🚧 Files skipped from review as they are similar to previous changes (9)
- docs/evaluator/agent-eval/reading-results.mdx
- docs/evaluator/agent-eval/writing-metrics.mdx
- docs/evaluator/evaluation-approaches.mdx
- docs/evaluator/agent-eval/evaluate-deployed-agent.mdx
- docs/evaluator/agent-eval/index.mdx
- docs/evaluator/agent-eval/harbor-runner.mdx
- docs/fern/versions/latest.yml
- docs/evaluator/index.mdx
- docs/evaluator/agent-eval/targets-and-runners.mdx
Add an "Agent Evaluation" section under Evaluate Models & Agents, plus a "Dataset-Driven vs Task-Driven Evaluation" overview, and wire both into the Fern nav: - evaluation-approaches: dataset-driven vs task-driven, and how to choose - agent-eval/index: the task -> runner -> trial -> metrics -> result model - agent-eval/quickstart: runnable, zero-dependency local example - agent-eval/evaluate-deployed-agent: a GenericAgent over HTTP - agent-eval/harbor-runner: Harbor task suites in Docker - agent-eval/score-by-component: a trajectory metric plus views - agent-eval/targets-and-runners, writing-metrics, reading-results: reference Also refresh evaluator/index to drop the removed industry-benchmark framing and point at the two evaluation shapes. Validated end to end: the quickstart, deployed-agent, and score-by-component full scripts each reproduce their documented output against the SDK. HTTP targets authenticate via api_key_secret (the api_key_env field does not exist); Model format values are documented as ModelFormat members; and ATIF is expanded on first use with a link to the trajectory-format RFC. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Signed-off-by: Sandy Chapman <schapman@nvidia.com>
767dc02 to
1daefb3
Compare
What
Adds documentation for agent (task-driven) evaluation in NeMo Evaluator, plus a conceptual
overview that frames it against the existing dataset-driven (metrics) path, and wires everything into
the Fern nav under a renamed Evaluate Models & Agents section.
New pages:
evaluation-approachesagent-eval/indexagent-eval/quickstartagent-eval/evaluate-deployed-agentGenericAgentover HTTPagent-eval/harbor-runneragent-eval/score-by-componentagent-eval/targets-and-runnersagent-eval/writing-metricsMetricprotocolagent-eval/reading-resultsAlso refreshes
evaluator/indexto drop the removed industry-benchmark framing and point at the twoevaluation shapes. The
evaluate-modelssection slug is pinned so existing/documentation/evaluate-models/...links keep resolving despite the title change.Validation
The three offline, zero-dependency full scripts were extracted verbatim from the pages and executed
against the SDK; each reproduces its documented output:
keyword_match.score: 1.0keyword_match.score: 1.0+capital-france: 'Paris'/capital-japan: 'Tokyo'keyword_match.score: 1.0/used_expected_tool.tool_use: 0.5/view.quality: 0.75Review fixes folded in
api_key_secret(a real, settable field) —api_key_envis a derivedread-only property and would raise
ValidationErrorunderextra="forbid".Modelformat: documented asModelFormatmembers (serializednim/openai/llama_stack),not the invalid
nvidia_nim/open_aistrings.run_verifier(command)(wascmd); metric examples read the reference defensively via.get(...);reading-resultsnotes thewrite_dashboardflag.Notes
AGENT-EVAL-DOCS-UX-FINDINGS.mdtriage log is intentionally not included.🤖 Generated with Claude Code
Summary by CodeRabbit