這很重要
本文中標示為預覽的項目目前仍在預覽中。 此預覽版未簽訂服務等級協議,Microsoft 不建議用於生產工作負載。 某些功能可能不被支援或功能受限。 欲了解更多資訊,請參閱 Microsoft Azure 預覽版補充使用條款。
從已部署的 Agent 和模型評估儲存的回覆或 OpenTelemetry 追蹤,無需重新執行原始要求。
先決條件
- 完成 雲端評估的前提 條件及 客戶端設定。
- 儲存回應 ID 用於回應評估,或是連結至 Foundry 專案的 Application Insights 資源以進行追蹤評估。
- OpenTelemetry 跨度符合您評估追蹤時的資料需求。
範例中使用了在 「設定 SDK 客戶端」中設定的 SDK 客戶端。
以回應識別來評估互動
利用資料來源類型,透過回應 ID azure_ai_responses 檢索並評估 Foundry 代理的回應。 利用此情境評估特定代理人互動發生後的情況。
Tip
在開始之前,先完成 客戶端設定。
回應 ID 是每次 Foundry 代理產生回應時回傳的唯一識別碼。 你可以透過 Responses API 或應用程式的追蹤日誌,從代理互動中收集回應 ID。 將識別碼直接做為檔案內容提供。
這很重要
代理回應評估(azure_ai_responses)僅支援 file_content 用於提供回應 ID。 不支援 file_id 來源類型,並傳回 400 Bad Request 錯誤。
收集回應識別碼
每次呼叫回應 API 都會回傳一個具有唯一 id 欄位的回應物件。 從應用程式的互動中收集這些識別碼,或直接產生:
# Generate response IDs by calling a model through the Responses API
response = openai_client.responses.create(
model=model_deployment_name,
input="What is machine learning?",
)
print(response.id) # Example: resp_abc123
你也可以從應用程式的追蹤日誌或監控管線中,從代理互動中收集回應 ID。 每個回應 ID 唯一識別一個儲存的回應,評估服務可檢索。
建立評估並執行
from azure.ai.projects.models import TestingCriterionAzureAIEvaluator
data_source_config = {"type": "azure_ai_source", "scenario": "responses"}
testing_criteria = [
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="coherence",
evaluator_name="builtin.coherence",
initialization_parameters={"model": model_deployment_name},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="violence",
evaluator_name="builtin.violence",
),
]
eval_object = openai_client.evals.create(
name="Agent Response Evaluation",
data_source_config=data_source_config,
testing_criteria=testing_criteria,
)
data_source = {
"type": "azure_ai_responses",
"item_generation_params": {
"type": "response_retrieval",
"data_mapping": {"response_id": "{{item.resp_id}}"},
"source": {
"type": "file_content",
"content": [
{"item": {"resp_id": "resp_abc123"}},
{"item": {"resp_id": "resp_def456"}},
]
},
},
}
eval_run = openai_client.evals.runs.create(
eval_id=eval_object.id,
name="agent-response-evaluation",
data_source=data_source,
)
完整可執行範例請參見GitHub上的 sample_agent_response_evaluation.py。 若要投票完成並解讀結果,請參閱 「取得雲端評估結果」。
完整的.NET反應評估範例,請參見GitHub上的 Sample7_EvaluationsAgent.md。
評估追蹤 (預覽版)
評估 Application Insights 已經捕捉到的客服互動。 使用資料來源類型 azure_ai_traces 。 此情境對於部署後實際生產流量的評估非常有用。 您可以從監視管線選取追蹤,並針對這些追蹤執行評估工具,而不重播任何要求。
這很重要
追蹤評估是評估並非使用 Microsoft Foundry Agent Service 建置的代理程式(包括 LangChain 和自訂框架)的建議方法。 只要你的代理將 遵循 GenAI 語意慣例的 OpenTelemetry span 傳送至 Application Insights,追蹤評估就能使用 Foundry 代理可用的相同評估器來評估其互動。
追蹤評估支援兩種模式:
-
依據追蹤識別碼 - 提供來自 Application Insights 的
operation_Id值,以評估特定代理程式的互動。 - 透過代理篩選 器 - 自動發現並評估特定代理的近期追蹤,無需手動收集追蹤 ID。
Tip
在開始之前,先完成 客戶端設定。 此情境也需要與 你的 Foundry 專案連結的 Application Insights 資源。
智慧取樣
痕跡評估支援智慧取樣,選擇具代表性的痕跡子集進行評估,而非評估每一條捕獲的痕跡。 設定追蹤評估執行作業時,請在 Foundry 入口網站中開啟 智慧取樣 切換按鈕。 智慧取樣降低評估成本,同時保留痕跡多樣性——確保邊緣案例、錯誤路徑及多樣的對話模式被納入評估集合中。
智慧取樣的運作原理
該抽樣演算法採用 MinHash 最遠優先多樣性方法,分多個階段執行:
- 精確去重 - 從集區中移除重複的追蹤記錄。
- 硬性篩選器 - 移除不適合評估的損壞工作階段、截斷追蹤和格式錯誤的工具呼叫。
- 聚合 - 將痕跡層級訊號合併成統一表示法。
- MinHash 最遠優先選擇 - 計算使用者文本的區域敏感雜湊值(MinHash 簽名)以估算痕跡間的相似度,然後從剩餘的痕跡池中迭代選擇最不相似的痕跡。 每次後續選取都會使其與先前所有已選軌跡的距離達到最大。
此方法相較於隨機抽樣,能帶來顯著更高的詞彙多樣性與更廣泛的詞彙覆蓋,意味著評估的集合能更完整地呈現代理間的互動範圍——包括隨機抽樣容易忽略的罕見、困難及新穎案例。
智慧抽樣對以下情況特別有效:
- 評估與基準測試 - 最大化輸入分布的覆蓋範圍,使評估分數反映真實世界的多樣性。
- 評分規準生成 ——透過揭示多元對話模式,產生更具聚焦性且可執行的評分標準。
- 微調資料集整理 - 選擇能幫助模型更有效率學習的軌跡資料。
該演算法完全以本地運算運行,沒有額外的 API 呼叫,因此除了評估本身外,不會產生額外的模型推論成本。
智慧取樣範例
# Eval group for trace-based evaluations
data_source_config = {
"type": "azure_ai_source",
"scenario": "traces",
}
print("Creating trace-based evaluation group")
eval_object = client.evals.create(
name="Trace Evaluation (Agent Smart Filter)",
data_source_config=data_source_config, # type: ignore
testing_criteria=testing_criteria,
)
print(f"Evaluation created (id: {eval_object.id})")
# Compute time window in unix seconds
# Pad end_time by +600s (10 min) to avoid ingestion-delay edge exclusion
now_unix = int(time.time())
end_time = now_unix + 600
start_time = now_unix - (args.lookback_hours * 3600)
# Build trace_source based on mode
trace_source: dict = {
"type": "agent_filter",
"start_time": start_time,
"end_time": end_time,
"max_traces": args.max_traces,
"filter_strategy": "smart_filtering"
}
# Add agent name/version or agent id
trace_source["agent_name"] = agent_name
trace_source["agent_version"] = agent_version
## trace_source["agent_id"] = args.agent_id
data_source = {
"type": "azure_ai_trace_data_source_preview",
"trace_source": trace_source,
}
eval_run = client.evals.runs.create(
eval_id=eval_object.id,
name="trace-evaluation-agent-smart-filter-run",
data_source=data_source, # type: ignore
)
追蹤資料需求
追蹤評估要求你的代理程式發出符合 OpenTelemetry 生成式 AI 語意慣例的跨度。 具體來說,評估服務會invoke_agent讀取 Application Insights 的範圍,並從其屬性中擷取對話資料。
以下跨度屬性被使用:
| Attribute | Required | Description |
|---|---|---|
gen_ai.operation.name |
Yes | 必須等於 "invoke_agent"。 服務會忽略所有其他的跨度。 |
gen_ai.agent.id |
適用於代理程式篩選模式 | 唯一代理識別碼(格式: agent-name:version)。 |
gen_ai.agent.name |
適用於代理程式篩選模式 | 人類可讀的代理名稱。 |
gen_ai.input.messages |
適用於評估器查詢輸入 | 遵循 生成式人工智慧語意慣例訊息格式的 JSON 輸入訊息陣列。 帶有角色 user 或 system 的訊息會對應到 query。 帶有角色 assistant 或 tool 的訊息會對應到 response。 |
gen_ai.output.messages |
適用於評估器查詢輸入 | 模型產生的輸出訊息 JSON 陣列。 所有輸出訊息對應到 response。 若輸出也包含 type: tool_call 或 type: tool_result,則映射為 tool_calls。 |
gen_ai.tool.definitions |
選用 | 代理程式可用的 JSON 工具結構陣列。 若缺少,服務會嘗試從工具呼叫訊息推斷工具定義,但推斷出的結構可能不完整。 |
gen_ai.conversation.id |
選用 | 對話識別碼,會傳送至評估結果比對關聯性。 |
Note
若 gen_ai.input.messages 和 gen_ai.output.messages 為空或缺失,則品質評估器(一致性、流暢性、相關性、意圖解析)返回 score=None。 安全評估者(暴力、自傷、性、仇恨/不公平)仍能產出部分數據的分數,但可能不會產生有意義的結果。
對於使用 Azure AI Agent Server SDK 建置的 Python agent,請額外加裝 [tracing] 以啟用自動區間發射:
pip install "azure-ai-agentserver-core[tracing]"
痕跡評估的前提條件
除了一般 的先決條件外,痕跡評估還要求:
- 一個與你的 Foundry 專案相關的 應用洞察資源 。 請參見 Microsoft Foundry 中的追蹤設定。
- 專案的管理身份必須在 Application Insights 資源中擁有 Reader 角色 。 如果儲存追蹤資料的資料表為 受保護(其保護層級設為 受保護),也請在該資源上指派 Privileged Monitoring Data Reader 角色。 如果連結的 Log Analytics 工作區設定為「需要工作區權限」,也要在工作區上設定 Log Analytics Reader。 關於所有追蹤評估角色需求及其他工作空間設定,請參閱 「設定評估工作流程權限」。
-
azure-monitor-queryPython 套件(僅在手動收集追蹤標識符時需要)。
pip install "azure-ai-projects>=2.2.0" azure-monitor-query
設定以下環境變數:
-
APPINSIGHTS_RESOURCE_ID— Application Insights 資源 ID(例如/subscriptions/<subscription_id>/resourceGroups/<rg_name>/providers/Microsoft.Insights/components/<resource_name>)。 -
AGENT_ID— 由追蹤整合gen_ai.agent.id屬性所發出的代理識別碼,用於過濾追蹤記錄。 格式:agent-name:version。 -
TRACE_LOOKBACK_HOURS— (可選)查詢痕跡時可回溯的時數。 預設為1。
選項A:使用代理人過濾器進行評估
最簡單的方法是讓服務自動發現並評估特定代理人的最新追蹤紀錄。 你不需要手動收集追蹤 ID。
import os
agent_id = os.environ["AGENT_ID"] # e.g., "my-weather-agent:1"
trace_lookback_hours = int(os.environ.get("TRACE_LOOKBACK_HOURS", "1"))
# Create the evaluation
data_source_config = {
"type": "azure_ai_source",
"scenario": "traces",
}
eval_object = openai_client.evals.create(
name="Agent Trace Evaluation (by agent)",
data_source_config=data_source_config,
testing_criteria=testing_criteria, # See "Set up evaluators" below
)
# Create a run — the service queries App Insights for matching traces
data_source = {
"type": "azure_ai_traces",
"agent_id": agent_id,
"max_traces": 50, # Maximum number of traces to evaluate
"lookback_hours": trace_lookback_hours,
}
eval_run = openai_client.evals.runs.create(
eval_id=eval_object.id,
name="agent-trace-eval-run",
data_source=data_source,
)
print(f"Evaluation run started: {eval_run.id}")
該服務會以invoke_agent Span,並依據gen_ai.agent.id屬性進行篩選,取樣最多max_traces個不同複的追蹤識別碼,然後評估來自這些追蹤的所有 Span。
選項B:依追蹤識別碼評估
為了更精確的控制,請從 Application Insights 收集特定的追蹤 ID,並進行評估。 當您想要評估精選互動集時,例如由警示標記或為品質審查取樣的追蹤,此方法很有用。
從 Application Insights 收集追蹤識別碼
查詢 Application Insights,以取得來自您代理程式追蹤的operation_Id值。 每個 operation_Id 代表一個完整的代理互動:
import os
from datetime import datetime, timedelta, timezone
from azure.identity import DefaultAzureCredential
from azure.monitor.query import LogsQueryClient, LogsQueryStatus
appinsights_resource_id = os.environ["APPINSIGHTS_RESOURCE_ID"]
agent_id = os.environ["AGENT_ID"]
trace_query_hours = int(os.environ.get("TRACE_LOOKBACK_HOURS", "1"))
end_time = datetime.now(timezone.utc)
start_time = end_time - timedelta(hours=trace_query_hours)
query = f"""dependencies
| where timestamp between (datetime({start_time.isoformat()}) .. datetime({end_time.isoformat()}))
| extend agent_id = tostring(customDimensions["gen_ai.agent.id"])
| where agent_id == "{agent_id}"
| distinct operation_Id"""
credential = DefaultAzureCredential()
logs_client = LogsQueryClient(credential)
response = logs_client.query_resource(
appinsights_resource_id,
query=query,
timespan=None, # Time range is specified in the query itself
)
trace_ids = []
if response.status == LogsQueryStatus.SUCCESS:
for table in response.tables:
for row in table.rows:
trace_ids.append(row[0])
print(f"Found {len(trace_ids)} trace IDs")
建立評估並使用追蹤 ID 執行
# Create the evaluation
data_source_config = {
"type": "azure_ai_source",
"scenario": "traces",
}
eval_object = openai_client.evals.create(
name="Agent Trace Evaluation (by trace IDs)",
data_source_config=data_source_config,
testing_criteria=testing_criteria, # See "Set up evaluators" below
)
# Create a run using the collected trace IDs
data_source = {
"type": "azure_ai_traces",
"trace_ids": trace_ids,
"lookback_hours": trace_query_hours,
}
eval_run = openai_client.evals.runs.create(
eval_id=eval_object.id,
name="agent-trace-eval-run",
metadata={
"agent_id": agent_id,
"start_time": start_time.isoformat(),
"end_time": end_time.isoformat(),
},
data_source=data_source,
)
print(f"Evaluation run started: {eval_run.id}")
設置評估器與資料映射
當你評估追蹤時,服務會自動從 OpenTelemetry 的 span 屬性中擷取對話資料。 直接使用這些欄位名稱 data_mapping (在其他情境中不使用 item. 或 sample. 前綴):
| 變數 | 來源屬性 | Description |
|---|---|---|
{{item.query}} |
gen_ai.input.messages (使用者/系統角色) |
從追蹤中提取的使用者查詢。 |
{{item.response}} |
gen_ai.input.messages(助理/工具角色)+ gen_ai.output.messages |
從追蹤中提取的代理程式回應。 |
{{item.tool_definitions}} |
gen_ai.tool.definitions |
可供代理使用的工具架構。 僅工具相關評估工具需要。 |
{{item.tool_calls}} |
從助理訊息中擷取的資料 gen_ai.input.messages / gen_ai.output.messages |
代理人在互動過程中進行的工具呼叫。 供工具評估者使用。 僅工具相關評估工具需要。 |
from azure.ai.projects.models import TestingCriterionAzureAIEvaluator
testing_criteria = [
# Quality evaluators — require query and response from trace data
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="intent_resolution",
evaluator_name="builtin.intent_resolution",
data_mapping={
"query": "{{item.query}}",
"response": "{{item.response}}",
"tool_definitions": "{{item.tool_definitions}}",
},
initialization_parameters={"model": model_deployment_name},
),
# Tool evaluators — assess tool usage quality
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="tool_call_accuracy",
evaluator_name="builtin.tool_call_accuracy",
data_mapping={
"query": "{{item.query}}",
"response": "{{item.response}}",
"tool_calls": "{{item.tool_calls}}",
"tool_definitions": "{{item.tool_definitions}}",
},
initialization_parameters={"model": model_deployment_name},
),
]
用 .NET 執行追蹤評估
輸入你想評估的 Application Insights 追蹤 ID。 該服務會從連接的 Application Insights 資源中擷取對應的區間。
string[] traceIds = ["trace-id-1", "trace-id-2"];
object[] testingCriteria =
[
new
{
type = "azure_ai_evaluator",
name = "intent_resolution",
evaluator_name = "builtin.intent_resolution",
initialization_parameters = new { model = modelDeploymentName },
data_mapping = new
{
query = "{{item.query}}",
response = "{{item.response}}",
tool_definitions = "{{item.tool_definitions}}"
}
},
new
{
type = "azure_ai_evaluator",
name = "violence",
evaluator_name = "builtin.violence",
data_mapping = new
{
query = "{{item.query}}",
response = "{{item.response}}"
},
initialization_parameters = new { threshold = 4 }
}
];
BinaryData evaluationData = BinaryData.FromObjectAsJson(new
{
name = "Agent Trace Evaluation",
data_source_config = new
{
type = "azure_ai_source",
scenario = "traces"
},
testing_criteria = testingCriteria
});
using BinaryContent evaluationContent = BinaryContent.Create(evaluationData);
ClientResult evaluation = await evaluationClient.CreateEvaluationAsync(
evaluationContent);
string evaluationId = GetString(evaluation, "id");
BinaryData runData = BinaryData.FromObjectAsJson(new
{
name = "agent-trace-evaluation",
data_source = new
{
type = "azure_ai_traces",
trace_ids = traceIds,
lookback_hours = 1
}
});
using BinaryContent runContent = BinaryContent.Create(runData);
ClientResult evaluationRun = await evaluationClient.CreateEvaluationRunAsync(
evaluationId: evaluationId,
content: runContent);
Console.WriteLine($"Evaluation run created: {GetString(evaluationRun, "id")}");
參考資料: EvaluationClient 協定方法
下一步
- 若要投票完成並解讀結果,請參閱 「取得雲端評估結果」。
- 完整可執行範例請參見GitHub上的 sample_evaluations_builtin_with_traces.py。