使用 Microsoft Foundry SDK 評估部署模型與代理的個別互動

這很重要

本文中標示為預覽的項目目前仍在預覽中。 此預覽版未簽訂服務等級協議,Microsoft 不建議用於生產工作負載。 某些功能可能不被支援或功能受限。 欲了解更多資訊,請參閱 Microsoft Azure 預覽版補充使用條款。

從已部署的 Agent 和模型評估儲存的回覆或 OpenTelemetry 追蹤,無需重新執行原始要求。

先決條件

  • 完成 雲端評估的前提 條件及 客戶端設定。
  • 儲存回應 ID 用於回應評估,或是連結至 Foundry 專案的 Application Insights 資源以進行追蹤評估。
  • OpenTelemetry 跨度符合您評估追蹤時的資料需求。

範例中使用了在 「設定 SDK 客戶端」中設定的 SDK 客戶端。

以回應識別來評估互動

利用資料來源類型,透過回應 ID azure_ai_responses 檢索並評估 Foundry 代理的回應。 利用此情境評估特定代理人互動發生後的情況。

Tip

在開始之前,先完成 客戶端設定。

回應 ID 是每次 Foundry 代理產生回應時回傳的唯一識別碼。 你可以透過 Responses API 或應用程式的追蹤日誌,從代理互動中收集回應 ID。 將識別碼直接做為檔案內容提供。

這很重要

代理回應評估(azure_ai_responses)僅支援 file_content 用於提供回應 ID。 不支援 file_id 來源類型,並傳回 400 Bad Request 錯誤。

收集回應識別碼

每次呼叫回應 API 都會回傳一個具有唯一 id 欄位的回應物件。 從應用程式的互動中收集這些識別碼,或直接產生:

# Generate response IDs by calling a model through the Responses API
response = openai_client.responses.create(
    model=model_deployment_name,
    input="What is machine learning?",
)
print(response.id)  # Example: resp_abc123

你也可以從應用程式的追蹤日誌或監控管線中,從代理互動中收集回應 ID。 每個回應 ID 唯一識別一個儲存的回應,評估服務可檢索。

建立評估並執行

from azure.ai.projects.models import TestingCriterionAzureAIEvaluator

data_source_config = {"type": "azure_ai_source", "scenario": "responses"}

testing_criteria = [
    TestingCriterionAzureAIEvaluator(
        type="azure_ai_evaluator",
        name="coherence",
        evaluator_name="builtin.coherence",
        initialization_parameters={"model": model_deployment_name},
    ),
    TestingCriterionAzureAIEvaluator(
        type="azure_ai_evaluator",
        name="violence",
        evaluator_name="builtin.violence",
    ),
]

eval_object = openai_client.evals.create(
    name="Agent Response Evaluation",
    data_source_config=data_source_config,
    testing_criteria=testing_criteria,
)

data_source = {
    "type": "azure_ai_responses",
    "item_generation_params": {
        "type": "response_retrieval",
        "data_mapping": {"response_id": "{{item.resp_id}}"},
        "source": {
            "type": "file_content",
            "content": [
                {"item": {"resp_id": "resp_abc123"}},
                {"item": {"resp_id": "resp_def456"}},
            ]
        },
    },
}

eval_run = openai_client.evals.runs.create(
    eval_id=eval_object.id,
    name="agent-response-evaluation",
    data_source=data_source,
)

完整可執行範例請參見GitHub上的 sample_agent_response_evaluation.py。 若要投票完成並解讀結果,請參閱 「取得雲端評估結果」。

完整的.NET反應評估範例,請參見GitHub上的 Sample7_EvaluationsAgent.md。

評估追蹤 (預覽版)

評估 Application Insights 已經捕捉到的客服互動。 使用資料來源類型 azure_ai_traces 。 此情境對於部署後實際生產流量的評估非常有用。 您可以從監視管線選取追蹤,並針對這些追蹤執行評估工具,而不重播任何要求。

這很重要

追蹤評估是評估並非使用 Microsoft Foundry Agent Service 建置的代理程式(包括 LangChain 和自訂框架)的建議方法。 只要你的代理將 遵循 GenAI 語意慣例的 OpenTelemetry span 傳送至 Application Insights,追蹤評估就能使用 Foundry 代理可用的相同評估器來評估其互動。

追蹤評估支援兩種模式:

  • 依據追蹤識別碼 - 提供來自 Application Insights 的 operation_Id 值,以評估特定代理程式的互動。
  • 透過代理篩選 器 - 自動發現並評估特定代理的近期追蹤,無需手動收集追蹤 ID。

Tip

在開始之前,先完成 客戶端設定。 此情境也需要與 你的 Foundry 專案連結的 Application Insights 資源。

智慧取樣

痕跡評估支援智慧取樣,選擇具代表性的痕跡子集進行評估,而非評估每一條捕獲的痕跡。 設定追蹤評估執行作業時,請在 Foundry 入口網站中開啟 智慧取樣 切換按鈕。 智慧取樣降低評估成本,同時保留痕跡多樣性——確保邊緣案例、錯誤路徑及多樣的對話模式被納入評估集合中。

智慧取樣的運作原理

該抽樣演算法採用 MinHash 最遠優先多樣性方法,分多個階段執行:

  1. 精確去重 - 從集區中移除重複的追蹤記錄。
  2. 硬性篩選器 - 移除不適合評估的損壞工作階段、截斷追蹤和格式錯誤的工具呼叫。
  3. 聚合 - 將痕跡層級訊號合併成統一表示法。
  4. MinHash 最遠優先選擇 - 計算使用者文本的區域敏感雜湊值(MinHash 簽名)以估算痕跡間的相似度,然後從剩餘的痕跡池中迭代選擇最不相似的痕跡。 每次後續選取都會使其與先前所有已選軌跡的距離達到最大。

此方法相較於隨機抽樣,能帶來顯著更高的詞彙多樣性與更廣泛的詞彙覆蓋,意味著評估的集合能更完整地呈現代理間的互動範圍——包括隨機抽樣容易忽略的罕見、困難及新穎案例。

智慧抽樣對以下情況特別有效:

  • 評估與基準測試 - 最大化輸入分布的覆蓋範圍,使評估分數反映真實世界的多樣性。
  • 評分規準生成 ——透過揭示多元對話模式,產生更具聚焦性且可執行的評分標準。
  • 微調資料集整理 - 選擇能幫助模型更有效率學習的軌跡資料。

該演算法完全以本地運算運行,沒有額外的 API 呼叫,因此除了評估本身外,不會產生額外的模型推論成本。

智慧取樣範例

# Eval group for trace-based evaluations
data_source_config = {
    "type": "azure_ai_source",
    "scenario": "traces",
}

print("Creating trace-based evaluation group")
eval_object = client.evals.create(
    name="Trace Evaluation (Agent Smart Filter)",
    data_source_config=data_source_config,  # type: ignore
    testing_criteria=testing_criteria,
)
print(f"Evaluation created (id: {eval_object.id})")

# Compute time window in unix seconds
# Pad end_time by +600s (10 min) to avoid ingestion-delay edge exclusion
now_unix = int(time.time())
end_time = now_unix + 600
start_time = now_unix - (args.lookback_hours * 3600)

# Build trace_source based on mode
trace_source: dict = {
    "type": "agent_filter",
    "start_time": start_time,
    "end_time": end_time,
    "max_traces": args.max_traces,
    "filter_strategy": "smart_filtering"
}

# Add agent name/version or agent id
trace_source["agent_name"] = agent_name
trace_source["agent_version"] = agent_version
## trace_source["agent_id"] = args.agent_id

data_source = {
    "type": "azure_ai_trace_data_source_preview",
    "trace_source": trace_source,
}

eval_run = client.evals.runs.create(
    eval_id=eval_object.id,
    name="trace-evaluation-agent-smart-filter-run",
    data_source=data_source,  # type: ignore
)

追蹤資料需求

追蹤評估要求你的代理程式發出符合 OpenTelemetry 生成式 AI 語意慣例的跨度。 具體來說,評估服務會invoke_agent讀取 Application Insights 的範圍,並從其屬性中擷取對話資料。

以下跨度屬性被使用:

Attribute Required Description
gen_ai.operation.name Yes 必須等於 "invoke_agent"。 服務會忽略所有其他的跨度。
gen_ai.agent.id 適用於代理程式篩選模式 唯一代理識別碼(格式: agent-name:version)。
gen_ai.agent.name 適用於代理程式篩選模式 人類可讀的代理名稱。
gen_ai.input.messages 適用於評估器查詢輸入 遵循 生成式人工智慧語意慣例訊息格式的 JSON 輸入訊息陣列。 帶有角色 user 或 system 的訊息會對應到 query。 帶有角色 assistant 或 tool 的訊息會對應到 response。
gen_ai.output.messages 適用於評估器查詢輸入 模型產生的輸出訊息 JSON 陣列。 所有輸出訊息對應到 response。 若輸出也包含 type: tool_call 或 type: tool_result,則映射為 tool_calls。
gen_ai.tool.definitions 選用 代理程式可用的 JSON 工具結構陣列。 若缺少,服務會嘗試從工具呼叫訊息推斷工具定義,但推斷出的結構可能不完整。
gen_ai.conversation.id 選用 對話識別碼,會傳送至評估結果比對關聯性。

Note

若 gen_ai.input.messages 和 gen_ai.output.messages 為空或缺失,則品質評估器(一致性、流暢性、相關性、意圖解析)返回 score=None。 安全評估者(暴力、自傷、性、仇恨/不公平)仍能產出部分數據的分數,但可能不會產生有意義的結果。

對於使用 Azure AI Agent Server SDK 建置的 Python agent,請額外加裝 [tracing] 以啟用自動區間發射:

pip install "azure-ai-agentserver-core[tracing]"

痕跡評估的前提條件

除了一般 的先決條件外,痕跡評估還要求:

pip install "azure-ai-projects>=2.2.0" azure-monitor-query

設定以下環境變數:

  • APPINSIGHTS_RESOURCE_ID — Application Insights 資源 ID(例如 /subscriptions/<subscription_id>/resourceGroups/<rg_name>/providers/Microsoft.Insights/components/<resource_name>)。
  • AGENT_ID — 由追蹤整合gen_ai.agent.id屬性所發出的代理識別碼,用於過濾追蹤記錄。 格式: agent-name:version。
  • TRACE_LOOKBACK_HOURS — (可選)查詢痕跡時可回溯的時數。 預設為 1。

選項A:使用代理人過濾器進行評估

最簡單的方法是讓服務自動發現並評估特定代理人的最新追蹤紀錄。 你不需要手動收集追蹤 ID。

import os

agent_id = os.environ["AGENT_ID"]  # e.g., "my-weather-agent:1"
trace_lookback_hours = int(os.environ.get("TRACE_LOOKBACK_HOURS", "1"))

# Create the evaluation
data_source_config = {
    "type": "azure_ai_source",
    "scenario": "traces",
}

eval_object = openai_client.evals.create(
    name="Agent Trace Evaluation (by agent)",
    data_source_config=data_source_config,
    testing_criteria=testing_criteria,  # See "Set up evaluators" below
)

# Create a run — the service queries App Insights for matching traces
data_source = {
    "type": "azure_ai_traces",
    "agent_id": agent_id,
    "max_traces": 50,           # Maximum number of traces to evaluate
    "lookback_hours": trace_lookback_hours,
}

eval_run = openai_client.evals.runs.create(
    eval_id=eval_object.id,
    name="agent-trace-eval-run",
    data_source=data_source,
)

print(f"Evaluation run started: {eval_run.id}")

該服務會以invoke_agent Span,並依據gen_ai.agent.id屬性進行篩選,取樣最多max_traces個不同複的追蹤識別碼,然後評估來自這些追蹤的所有 Span。

選項B:依追蹤識別碼評估

為了更精確的控制,請從 Application Insights 收集特定的追蹤 ID,並進行評估。 當您想要評估精選互動集時,例如由警示標記或為品質審查取樣的追蹤,此方法很有用。

從 Application Insights 收集追蹤識別碼

查詢 Application Insights,以取得來自您代理程式追蹤的operation_Id值。 每個 operation_Id 代表一個完整的代理互動:

import os
from datetime import datetime, timedelta, timezone
from azure.identity import DefaultAzureCredential
from azure.monitor.query import LogsQueryClient, LogsQueryStatus

appinsights_resource_id = os.environ["APPINSIGHTS_RESOURCE_ID"]
agent_id = os.environ["AGENT_ID"]
trace_query_hours = int(os.environ.get("TRACE_LOOKBACK_HOURS", "1"))

end_time = datetime.now(timezone.utc)
start_time = end_time - timedelta(hours=trace_query_hours)

query = f"""dependencies
| where timestamp between (datetime({start_time.isoformat()}) .. datetime({end_time.isoformat()}))
| extend agent_id = tostring(customDimensions["gen_ai.agent.id"])
| where agent_id == "{agent_id}"
| distinct operation_Id"""

credential = DefaultAzureCredential()
logs_client = LogsQueryClient(credential)
response = logs_client.query_resource(
    appinsights_resource_id,
    query=query,
    timespan=None,  # Time range is specified in the query itself
)

trace_ids = []
if response.status == LogsQueryStatus.SUCCESS:
    for table in response.tables:
        for row in table.rows:
            trace_ids.append(row[0])

print(f"Found {len(trace_ids)} trace IDs")

建立評估並使用追蹤 ID 執行

# Create the evaluation
data_source_config = {
    "type": "azure_ai_source",
    "scenario": "traces",
}

eval_object = openai_client.evals.create(
    name="Agent Trace Evaluation (by trace IDs)",
    data_source_config=data_source_config,
    testing_criteria=testing_criteria,  # See "Set up evaluators" below
)

# Create a run using the collected trace IDs
data_source = {
    "type": "azure_ai_traces",
    "trace_ids": trace_ids,
    "lookback_hours": trace_query_hours,
}

eval_run = openai_client.evals.runs.create(
    eval_id=eval_object.id,
    name="agent-trace-eval-run",
    metadata={
        "agent_id": agent_id,
        "start_time": start_time.isoformat(),
        "end_time": end_time.isoformat(),
    },
    data_source=data_source,
)

print(f"Evaluation run started: {eval_run.id}")

設置評估器與資料映射

當你評估追蹤時,服務會自動從 OpenTelemetry 的 span 屬性中擷取對話資料。 直接使用這些欄位名稱 data_mapping (在其他情境中不使用 item. 或 sample. 前綴):

變數 來源屬性 Description
{{item.query}} gen_ai.input.messages (使用者/系統角色) 從追蹤中提取的使用者查詢。
{{item.response}} gen_ai.input.messages(助理/工具角色)+ gen_ai.output.messages 從追蹤中提取的代理程式回應。
{{item.tool_definitions}} gen_ai.tool.definitions 可供代理使用的工具架構。 僅工具相關評估工具需要。
{{item.tool_calls}} 從助理訊息中擷取的資料 gen_ai.input.messages / gen_ai.output.messages 代理人在互動過程中進行的工具呼叫。 供工具評估者使用。 僅工具相關評估工具需要。
from azure.ai.projects.models import TestingCriterionAzureAIEvaluator

testing_criteria = [
    # Quality evaluators — require query and response from trace data
    TestingCriterionAzureAIEvaluator(
        type="azure_ai_evaluator",
        name="intent_resolution",
        evaluator_name="builtin.intent_resolution",
        data_mapping={
            "query": "{{item.query}}",
            "response": "{{item.response}}",
            "tool_definitions": "{{item.tool_definitions}}",
        },
        initialization_parameters={"model": model_deployment_name},
    ),
    # Tool evaluators — assess tool usage quality
    TestingCriterionAzureAIEvaluator(
        type="azure_ai_evaluator",
        name="tool_call_accuracy",
        evaluator_name="builtin.tool_call_accuracy",
        data_mapping={
            "query": "{{item.query}}",
            "response": "{{item.response}}",
            "tool_calls": "{{item.tool_calls}}",
            "tool_definitions": "{{item.tool_definitions}}",
        },
        initialization_parameters={"model": model_deployment_name},
    ),
]

用 .NET 執行追蹤評估

輸入你想評估的 Application Insights 追蹤 ID。 該服務會從連接的 Application Insights 資源中擷取對應的區間。

string[] traceIds = ["trace-id-1", "trace-id-2"];
object[] testingCriteria =
[
    new
    {
        type = "azure_ai_evaluator",
        name = "intent_resolution",
        evaluator_name = "builtin.intent_resolution",
        initialization_parameters = new { model = modelDeploymentName },
        data_mapping = new
        {
            query = "{{item.query}}",
            response = "{{item.response}}",
            tool_definitions = "{{item.tool_definitions}}"
        }
    },
    new
    {
        type = "azure_ai_evaluator",
        name = "violence",
        evaluator_name = "builtin.violence",
        data_mapping = new
        {
            query = "{{item.query}}",
            response = "{{item.response}}"
        },
        initialization_parameters = new { threshold = 4 }
    }
];
BinaryData evaluationData = BinaryData.FromObjectAsJson(new
{
    name = "Agent Trace Evaluation",
    data_source_config = new
    {
        type = "azure_ai_source",
        scenario = "traces"
    },
    testing_criteria = testingCriteria
});
using BinaryContent evaluationContent = BinaryContent.Create(evaluationData);
ClientResult evaluation = await evaluationClient.CreateEvaluationAsync(
    evaluationContent);
string evaluationId = GetString(evaluation, "id");

BinaryData runData = BinaryData.FromObjectAsJson(new
{
    name = "agent-trace-evaluation",
    data_source = new
    {
        type = "azure_ai_traces",
        trace_ids = traceIds,
        lookback_hours = 1
    }
});
using BinaryContent runContent = BinaryContent.Create(runData);
ClientResult evaluationRun = await evaluationClient.CreateEvaluationRunAsync(
    evaluationId: evaluationId,
    content: runContent);
Console.WriteLine($"Evaluation run created: {GetString(evaluationRun, "id")}");

參考資料: EvaluationClient 協定方法

下一步