feat: Add Agent Behavior scorer for Python and TypeScript - #208
Open
Abhijeet Prasad (AbhiPrasad) wants to merge 3 commits into
Open
feat: Add Agent Behavior scorer for Python and TypeScript#208Abhijeet Prasad (AbhiPrasad) wants to merge 3 commits into
Abhijeet Prasad (AbhiPrasad) wants to merge 3 commits into
Conversation
Braintrust eval report
|
Show Behavior directly in the scores array, matching the common Autoevals pattern. Keep explicit behavior selection as the multi-spec example.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Agent Behavior scorer
Adds
Behavior, an LLM judge for evaluating an agent against an Agent Behavior spec.Add one behavior spec at
.agents/behaviors/<name>/BEHAVIOR.md, then use the scorer like any other Autoevals scorer.TypeScript
Python
Braintrust passes the dataset case as
input, the task return value asoutput, and the instrumented trace thread automatically. The task can return a final answer or a structured trajectory.When a project contains multiple behaviors, select one explicitly:
Callers can also select a behavior with a
BEHAVIOR.mdpath, complete file content, or loaded behavior object. Scores are1for compliance,0for non-compliance, andnull/Nonewhen the behavior is not applicable or cannot be judged.