Skip to content

Python: Fix FoundryAgent telemetry gaps for name and model (Fixes #5088) - #5094

Closed
Charanvardhan wants to merge 2 commits into
microsoft:mainfrom
Charanvardhan:fix-foundry-telemetry-5088
Closed

Python: Fix FoundryAgent telemetry gaps for name and model (Fixes #5088)#5094
Charanvardhan wants to merge 2 commits into
microsoft:mainfrom
Charanvardhan:fix-foundry-telemetry-5088

Conversation

@Charanvardhan

@Charanvardhan Charanvardhan commented Apr 4, 2026

Copy link
Copy Markdown

Motivation and Context

This PR addresses two telemetry gaps present in the FoundryAgent implementation where essential OpenTelemetry instrumentation fields were either rendering erroneously as a UUID or not populating at all.

Fixes #5088

Description

This change ensures that gen_ai.agent.name and gen_ai.request.model output appropriately for telemetry and Application Insights tracing.

  • Agent Name fallback: In _agent.py, RawFoundryAgent.__init__ now correctly falls back to getattr(client, 'agent_name') if name is omitted, rather than allowing the base classes to arbitrarily generate a random UUID.
  • Request Model Lazy-Loading: Because Managed Foundry Agents dictate their models entirely server-side, RawFoundryAgentChatClient now implements an asynchronous lazy invocation to project_client.agents.get_agent inside _prepare_options. It securely caches the return response locally, directly updates the opentelemetry current span dynamically before tracing is flushed, and exposes a .model property for all subsequent calls.
  • Unit Tests: Additionally includes the test_foundry_agent_telemetry_defaults test using AsyncMock to rigorously assert these edge configurations.

Contribution Checklist

  • The code builds clean without any errors or warnings
  • The PR follows the Contribution Guidelines
  • All unit tests pass, and I have added new tests where possible
  • Is this a breaking change? If yes, add "[BREAKING]" prefix to the title of the PR.

…rosoft#5088)

Resolves gen_ai.agent.name rendering as UUID and gen_ai.request.model rendering as 'unknown' by falling back to the client agent definition and lazily resolving the Azure AI Projects SDK request model on first fetch.
Copilot AI review requested due to automatic review settings April 4, 2026 06:32
@markwallace-microsoft markwallace-microsoft added the python Usage: [Issues, PRs], Target: Python label Apr 4, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR aims to close telemetry gaps in the Python Foundry agent integration by ensuring gen_ai.agent.name uses a stable, human-readable fallback (agent name) and by resolving gen_ai.request.model from the Azure AI Projects agent definition (rather than reporting "unknown").

Changes:

  • Add a fallback for FoundryAgent.name to use the underlying client’s agent_name when name=None.
  • Add lazy model resolution via project_client.agents.get_agent(...) and attempt to patch the current OTel span with the resolved model.
  • Add a unit test validating the name fallback and lazy model fetch behavior.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 6 comments.

File Description
python/packages/foundry/agent_framework_foundry/_agent.py Adds name fallback behavior and introduces lazy model resolution + span mutation logic.
python/packages/foundry/tests/foundry/test_foundry_agent.py Adds a test covering the new name/model telemetry defaults behavior.

Comment thread python/packages/foundry/agent_framework_foundry/_agent.py
Comment thread python/packages/foundry/agent_framework_foundry/_agent.py Outdated
Comment thread python/packages/foundry/tests/foundry/test_foundry_agent.py Outdated
Comment thread python/packages/foundry/tests/foundry/test_foundry_agent.py Outdated
Comment thread python/packages/foundry/tests/foundry/test_foundry_agent.py Outdated
Comment thread python/packages/foundry/tests/foundry/test_foundry_agent.py Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 2 out of 2 changed files in this pull request and generated 5 comments.


@model.setter
def model(self, value: str) -> None:
pass

Copilot AI Apr 5, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

model is defined as a property, but the setter is a no-op. Because RawFoundryAgentChatClient subclasses RawOpenAIChatClient, RawOpenAIChatClient.__init__ assigns self.model = ... (openai/_chat_client.py:446); with the current no-op setter, that assignment is silently discarded, and any later client.model = ... will also do nothing. Implement the setter to persist the value (e.g., set the backing field used by the getter) or remove the property and use a normal attribute/backing field so assignments behave correctly.

Suggested change
pass
self._fetched_model = value

Copilot uses AI. Check for mistakes.
Comment on lines +281 to +288
# Lazily fetch the model name if not already cached
if not hasattr(self, "_fetched_model"):
try:
agent = await self.project_client.agents.get_agent(self.agent_name)
self._fetched_model = getattr(agent, "model", "unknown")
except Exception:
self._fetched_model = "unknown"

Copilot AI Apr 5, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The lazy model fetch uses if not hasattr(self, "_fetched_model") and then awaits project_client.agents.get_agent(...). If two calls enter _prepare_options concurrently before _fetched_model is set, both will perform the network call. Consider guarding the lazy initialization with an asyncio.Lock or an in-flight Task/Future so only one fetch occurs and others await it.

Copilot uses AI. Check for mistakes.
Comment on lines +283 to +302
try:
agent = await self.project_client.agents.get_agent(self.agent_name)
self._fetched_model = getattr(agent, "model", "unknown")
except Exception:
self._fetched_model = "unknown"

# Try to update the current OpenTelemetry span directly
if hasattr(self, "_fetched_model") and self._fetched_model != "unknown":
try:
from opentelemetry import trace
from agent_framework.observability import OtelAttr

current_span = trace.get_current_span()
if current_span and current_span.is_recording():
current_span.set_attribute(OtelAttr.REQUEST_MODEL, self._fetched_model)
span_name_parts = current_span.name.split(" ", 1)
if len(span_name_parts) > 0:
current_span.update_name(f"{span_name_parts[0]} {self._fetched_model}")
except Exception:
pass

Copilot AI Apr 5, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The except Exception blocks here swallow all failures (including auth/config/service errors) and then suppress them again when updating OTEL, with no logging. This can make real regressions very hard to diagnose and permanently cache "unknown" after a transient error. Prefer catching expected exception types (e.g., ImportError separately, and the specific Azure SDK error types for get_agent) and logging at least a debug/warning message once when resolution fails.

Copilot uses AI. Check for mistakes.
Comment on lines +289 to +301
# Try to update the current OpenTelemetry span directly
if hasattr(self, "_fetched_model") and self._fetched_model != "unknown":
try:
from opentelemetry import trace
from agent_framework.observability import OtelAttr

current_span = trace.get_current_span()
if current_span and current_span.is_recording():
current_span.set_attribute(OtelAttr.REQUEST_MODEL, self._fetched_model)
span_name_parts = current_span.name.split(" ", 1)
if len(span_name_parts) > 0:
current_span.update_name(f"{span_name_parts[0]} {self._fetched_model}")
except Exception:

Copilot AI Apr 5, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updating the OpenTelemetry span via trace.get_current_span() won’t work for streaming requests: ChatTelemetryLayer intentionally creates streaming spans without context attachment (core/agent_framework/observability.py:1293-1299), so there may be no “current span” to update. As a result, gen_ai.request.model can still remain unknown for stream=True. Consider a mechanism that ensures the span used for the request is the one being updated (e.g., providing the model before span creation, or passing/propagating the span explicitly).

Copilot uses AI. Check for mistakes.
agent = FoundryAgent(
project_client=mock_project,
agent_name="my-telemetry-agent",
name=None # Explicitly None to test fallback

Copilot AI Apr 5, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This multiline FoundryAgent(...) call is missing a trailing comma after the last argument. In this repo most multiline call sites include trailing commas (and formatters like Black will typically add them), so adding it will avoid churn / formatting diffs.

Suggested change
name=None # Explicitly None to test fallback
name=None, # Explicitly None to test fallback

Copilot uses AI. Check for mistakes.
@Charanvardhan

Copy link
Copy Markdown
Author

@copilot apply changes based on the comments in this thread

@TaoChenOSU

Copy link
Copy Markdown
Contributor

Thank you for your contribution! But we have decided to close this PR in favor of #6160

@TaoChenOSU TaoChenOSU closed this May 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

python Usage: [Issues, PRs], Target: Python

Projects

None yet

Development

Successfully merging this pull request may close these issues.

FoundryAgent telemetry: gen_ai.agent.name defaults to UUID, gen_ai.request.model is "unknown"

4 participants