從你的 azd 專案目錄中,確認代理程式已部署並可調用:
azd ai agent show
發送測試提示:
azd ai agent invoke "Write a haiku about deploying cloud applications."
你應該會在幾秒內看到回覆。
- 打開 Foundry 入口 ,進入你的專案。
- 選擇你的代理人,然後選擇 Playground 標籤。
- 傳送一個測試提示,例如
Write a haiku about deploying cloud applications.
你應該會在幾秒內看到回覆。
安裝 Foundry SDK:
pip install "azure-ai-projects>=2.0.0" azure-identity
設定兩個環境變數,然後建立專案客戶端。 將FOUNDRY_PROJECT_ENDPOINT設為您的專案端點,並將FOUNDRY_MODEL_NAME設為可作為評審模型使用的聊天完成部署。 以下範例假設你在此情境下執行:
import os
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
endpoint = os.environ["FOUNDRY_PROJECT_ENDPOINT"]
model_deployment = os.environ["FOUNDRY_MODEL_NAME"]
credential = DefaultAzureCredential()
project_client = AIProjectClient(endpoint=endpoint, credential=credential)
client = project_client.get_openai_client()
確認你部署的代理已註冊且可用。 請將 <your-agent-name> 替換為您託管的代理程式名稱:
agent = project_client.agents.get("<your-agent-name>")
print(f"Found agent: {agent.name}")
如果代理存在,呼叫會回傳;若名稱錯誤或代理未部署,則會產生錯誤。
安裝 Foundry SDK 與 OpenAI 評估客戶端:
dotnet add package Azure.AI.Projects --prerelease
dotnet add package OpenAI
dotnet add package Azure.Identity
設定兩個環境變數,然後建立客戶端。 將FOUNDRY_PROJECT_ENDPOINT設為您的專案端點,並將FOUNDRY_MODEL_NAME設為可作為評審模型使用的聊天完成部署。 以下範例假設你在此情境下執行:
using System.ClientModel;
using System.Text.Json;
using Azure.AI.Projects;
using Azure.AI.Projects.Agents;
using Azure.Core;
using Azure.Identity;
using OpenAI;
using OpenAI.Evals;
#pragma warning disable AAIP001, OPENAI001
var endpoint = Environment.GetEnvironmentVariable("FOUNDRY_PROJECT_ENDPOINT")!;
var modelDeployment = Environment.GetEnvironmentVariable("FOUNDRY_MODEL_NAME")!;
var credential = new DefaultAzureCredential();
AIProjectClient projectClient = new(new Uri(endpoint), credential);
// OpenAI-compatible evals client bound to the Foundry project endpoint.
// A Microsoft Entra token works as the credential because both use "Authorization: Bearer".
var token = credential.GetToken(new TokenRequestContext(["https://ai.azure.com/.default"])).Token;
EvaluationClient evalClient = new(
new ApiKeyCredential(token),
new OpenAIClientOptions { Endpoint = new Uri($"{endpoint}/openai/v1") });
確認你部署的代理已註冊且可用。 請將 <your-agent-name> 替換為您託管的代理程式名稱:
ProjectsAgentRecord agent = projectClient.AgentAdministrationClient.GetAgent("<your-agent-name>");
Console.WriteLine($"Found agent: {agent.Name}");
如果代理存在,呼叫會回傳;若名稱錯誤或代理未部署,則會產生錯誤。
安裝 Foundry SDK:
npm install @azure/ai-projects @azure/identity dotenv
設定兩個環境變數,然後建立專案客戶端。 將FOUNDRY_PROJECT_ENDPOINT設為您的專案端點,並將FOUNDRY_MODEL_NAME設為可作為評審模型使用的聊天完成部署。 以下範例假設你在此情境下執行:
import { DefaultAzureCredential } from "@azure/identity";
import { AIProjectClient } from "@azure/ai-projects";
import "dotenv/config";
const endpoint = process.env["FOUNDRY_PROJECT_ENDPOINT"] || "";
const modelDeployment = process.env["FOUNDRY_MODEL_NAME"] || "";
const projectClient = new AIProjectClient(
endpoint,
new DefaultAzureCredential(),
);
const client = projectClient.getOpenAIClient();
確認你部署的代理已註冊且可用。 請將 <your-agent-name> 替換為您託管的代理程式名稱:
const agent = await projectClient.agents.get("<your-agent-name>");
console.log(`Found agent: ${agent.name}`);
如果代理存在,呼叫會回傳;若名稱錯誤或代理未部署,則會產生錯誤。
參考資料: AIProjectClient 類別
首先,為你的代理建立一個測試查詢的 JSONL 檔案。 每一行都是帶有 query 欄位的 JSON 物件。 將它儲存在你代理人的來源資料夾中,格式為 src/<your-agent-name>/tests/queries.jsonl:
{"query": "Write a haiku about deploying cloud applications."}
接著在同一個代理程式原始碼資料夾中建立一個 eval.yaml 檔案,命名為 src/<your-agent-name>/eval.yaml。 它會指向你的資料集,並列出可套用的內建評估器。
dataset.local_uri路徑是相對於這個資料夾的。 請將<your-agent-name>替換成您託管之 Agent 的名稱,並將<your-chat-completion-deployment>替換成裁判模型部署:
name: agent-eval
agent:
name: <your-agent-name>
kind: hosted
dataset:
local_uri: tests/queries.jsonl
evaluators:
- builtin.intent_resolution
- builtin.task_adherence
options:
eval_model: <your-chat-completion-deployment>
max_samples: 15
eval_model值為用來評分回覆的裁判模型;您可以重新使用您的 Agent 已使用的部署。
- 在 Foundry 入口網站 中,開啟您的代理程式,然後選取 「評估」 索引標籤,再選取 「建立」。
- 選擇 評估目標時,請選擇 代理人。
- 在 選取評估範圍 中,請選取 個別輪次。
- 選擇 資料來源時,選擇 現有資料集 ,並從專案資料資產中選擇 CSV 或 JSONL 測試查詢檔。
- 如果出現「配置代理」步驟,請檢視代理並接受預設的使用者提示。
{{item.query}} 僅在您的 Agent 預期不同的輸入格式時,才進行調整。
- 選擇 測試標準時,選擇一個或多個代理人評估器,如 任務依從 性與 意圖解決。
保持精靈開啟。 你在下一步提交評估。
首先,為你的代理建立一個測試查詢的 JSONL 檔案。 每一行都是帶有 query 欄位的 JSON 物件。 儲存為 queries.jsonl:
{"query": "Write a haiku about deploying cloud applications."}
將檔案上傳為你的專案資料集:
dataset = project_client.datasets.upload_file(
name="agent-test-queries",
version="1",
file_path="./queries.jsonl",
)
接著,選擇內建的評估器並映射其輸入。 參數 data_mapping 告訴每個評估者查詢及代理回應的位置。 AI 輔助的評估者在initialization_parameters中需要一個評審模型;值必須是您專案中的聊天完成部署。
from azure.ai.projects.models import TestingCriterionAzureAIEvaluator
testing_criteria = [
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="Intent Resolution",
evaluator_name="builtin.intent_resolution",
initialization_parameters={"model": model_deployment},
data_mapping={
"query": "{{item.query}}",
"response": "{{sample.output_items}}",
},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="Task Adherence",
evaluator_name="builtin.task_adherence",
initialization_parameters={"model": model_deployment},
data_mapping={
"query": "{{item.query}}",
"response": "{{sample.output_items}}",
},
),
]
建立評估。 它定義了測試資料結構與測試標準,並作為一個或多個執行的容器:
from openai.types.eval_create_params import DataSourceConfigCustom
data_source_config = DataSourceConfigCustom(
type="custom",
item_schema={
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
include_sample_schema=True,
)
evaluation = client.evals.create(
name="Agent Quality Evaluation",
data_source_config=data_source_config,
testing_criteria=testing_criteria,
)
print(f"Evaluation created: {evaluation.id}")
首先,為你的代理建立一個測試查詢的 JSONL 檔案。 每一行都是帶有 query 欄位的 JSON 物件。 儲存為 queries.jsonl:
{"query": "Write a haiku about deploying cloud applications."}
將檔案上傳為你的專案資料集:
AIProjectDataset dataset = projectClient.Datasets.UploadFile(
name: "agent-test-queries",
version: "1",
filePath: "./queries.jsonl");
接著,選擇內建的評估器並映射其輸入。 參數 data_mapping 告訴每個評估者查詢及代理回應的位置。 AI 輔助的評估者在initialization_parameters中需要一個評審模型;值必須是您專案中的聊天完成部署。
var testingCriteria = new object[]
{
new
{
type = "azure_ai_evaluator",
name = "Intent Resolution",
evaluator_name = "builtin.intent_resolution",
initialization_parameters = new { model = modelDeployment },
data_mapping = new { query = "{{item.query}}", response = "{{sample.output_items}}" },
},
new
{
type = "azure_ai_evaluator",
name = "Task Adherence",
evaluator_name = "builtin.task_adherence",
initialization_parameters = new { model = modelDeployment },
data_mapping = new { query = "{{item.query}}", response = "{{sample.output_items}}" },
},
};
建立評估。 它定義了測試資料結構與測試標準,並作為一個或多個執行的容器:
var createEvaluation = new
{
name = "Agent Quality Evaluation",
data_source_config = new
{
type = "custom",
item_schema = new
{
type = "object",
properties = new { query = new { type = "string" } },
required = new[] { "query" },
},
include_sample_schema = true,
},
testing_criteria = testingCriteria,
};
ClientResult evaluationResult = evalClient.CreateEvaluation(
BinaryContent.Create(BinaryData.FromObjectAsJson(createEvaluation)));
string evaluationId = JsonDocument.Parse(evaluationResult.GetRawResponse().Content.ToString())
.RootElement.GetProperty("id").GetString()!;
Console.WriteLine($"Evaluation created: {evaluationId}");
首先,為你的代理建立一個測試查詢的 JSONL 檔案。 每一行都是帶有 query 欄位的 JSON 物件。 儲存為 queries.jsonl:
{"query": "Write a haiku about deploying cloud applications."}
將檔案上傳為你的專案資料集:
const dataset = await projectClient.datasets.uploadFile(
"agent-test-queries",
"1",
"./queries.jsonl",
);
接著,選擇內建的評估器並映射其輸入。 參數 data_mapping 告訴每個評估者查詢及代理回應的位置。 AI 輔助的評估者在initialization_parameters中需要一個評審模型;值必須是您專案中的聊天完成部署。
const testingCriteria = [
{
type: "azure_ai_evaluator",
name: "Intent Resolution",
evaluator_name: "builtin.intent_resolution",
initialization_parameters: { model: modelDeployment },
data_mapping: {
query: "{{item.query}}",
response: "{{sample.output_items}}",
},
},
{
type: "azure_ai_evaluator",
name: "Task Adherence",
evaluator_name: "builtin.task_adherence",
initialization_parameters: { model: modelDeployment },
data_mapping: {
query: "{{item.query}}",
response: "{{sample.output_items}}",
},
},
];
建立評估。 它定義了測試資料結構與測試標準,並作為一個或多個執行的容器:
const dataSourceConfig = {
type: "custom",
item_schema: {
type: "object",
properties: { query: { type: "string" } },
required: ["query"],
},
include_sample_schema: true,
};
const evaluation = await client.evals.create({
name: "Agent Quality Evaluation",
data_source_config: dataSourceConfig,
testing_criteria: testingCriteria,
});
console.log(`Evaluation created: ${evaluation.id}`);
從 azd 工作區根目錄執行評估:
azd ai agent eval run --config eval.yaml
注意
azd ai agent eval run 會將 --config 路徑解析為相對於 src/ 底下代理程式的來源資料夾(例如 src/<your-agent-name>/eval.yaml),而不是目前目錄。 請將 eval.yaml 及其 local_uri 所指向的資料集保留在該資料夾中。
指令讀 eval.yaml,將每個查詢傳送給你的客服人員,評分回應,完成後會列印摘要:
Eval run started
Eval: eval_b36748dede424e4ba3f8e6c99ca2cf27
Run: evalrun_5f72ef189ad24790a32128e6f230b131
(✓) Done Eval run
Results: 1 total, 1 passed, 0 failed, 0 errored
Per-criteria results:
intent_resolution: 1 passed, 0 failed, 0 errored
task_adherence: 1 passed, 0 failed, 0 errored
- 在 檢閱並提交 步驟中,輸入此評估的 名稱。
- 檢視目標、範圍、資料來源及所選評估者。
- 選擇 提交 開始跑步。
建立一個執行作業,將每個測試查詢傳送給你的代理程式,並套用評估器。 請將 <your-agent-name> 替換為您託管的代理程式名稱:
eval_run = client.evals.runs.create(
eval_id=evaluation.id,
name="Agent Evaluation Run",
data_source={
"type": "azure_ai_target_completions",
"source": {"type": "file_id", "id": dataset.id},
"input_messages": {
"type": "template",
"template": [
{
"type": "message",
"role": "user",
"content": {"type": "input_text", "text": "{{item.query}}"},
}
],
},
"target": {
"type": "azure_ai_agent",
"name": "<your-agent-name>",
# "version": "1", # Optional; omit to use the latest version
},
},
)
print(f"Evaluation run started: {eval_run.id}")
建立一個執行作業,將每個測試查詢傳送給你的代理程式,並套用評估器。 請將 <your-agent-name> 替換為您託管的代理程式名稱:
var createRun = new
{
name = "Agent Evaluation Run",
data_source = new
{
type = "azure_ai_target_completions",
source = new { type = "file_id", id = dataset.Id },
input_messages = new
{
type = "template",
template = new object[]
{
new { type = "message", role = "user", content = new { type = "input_text", text = "{{item.query}}" } },
},
},
// Add a "version" property to the target to pin a specific agent version; omit to use the latest.
target = new { type = "azure_ai_agent", name = "<your-agent-name>" },
},
};
ClientResult runResult = evalClient.CreateEvaluationRun(
evaluationId, BinaryContent.Create(BinaryData.FromObjectAsJson(createRun)));
string runId = JsonDocument.Parse(runResult.GetRawResponse().Content.ToString())
.RootElement.GetProperty("id").GetString()!;
Console.WriteLine($"Evaluation run started: {runId}");
建立一個執行作業,將每個測試查詢傳送給你的代理程式,並套用評估器。 請將 <your-agent-name> 替換為您託管的代理程式名稱:
const evalRun = await client.evals.runs.create(evaluation.id, {
name: "Agent Evaluation Run",
data_source: {
type: "azure_ai_target_completions",
source: { type: "file_id", id: dataset.id },
input_messages: {
type: "template",
template: [
{
type: "message",
role: "user",
content: { type: "input_text", text: "{{item.query}}" },
},
],
},
target: {
type: "azure_ai_agent",
name: "<your-agent-name>",
// version: "1", // Optional; omit to use the latest version
},
},
});
console.log(`Evaluation run started: ${evalRun.id}`);
列出近期評價:
azd ai agent eval list
Eval ID Name Status of last run Runs
------- ---- ------------------ ----
* eval_b36748dede424e4ba3f8e6c99ca2cf27 agent-eval Completed 1
* = active eval in current environment
顯示最新的評估及其執行結果:
azd ai agent eval show
Eval: eval_b36748dede424e4ba3f8e6c99ca2cf27
Name: agent-eval
Agent: <your-agent-name>
Runs: 1
Recent runs:
Run ID Status Passed Failed Created
------ ------ ------ ------ -------
evalrun_5f72ef189ad24790a32128e6f230b131 Completed 1/1 0 2026-06-17 14:52 UTC
使用這些結果來確認受評估的是哪個代理程式版本,以及產生了哪些評估器分數。 欲查看每位評估者詳細資料及 Foundry 入口網站報告連結,請執行 azd ai agent eval show <eval-id> --eval-run-id <run-id>。
- 詳情頁面顯示目標、資料集、狀態、代幣使用情況,以及每位評估者的總分。
- 選擇執行名稱以查看列級結果:每個查詢、代理回應、評估者分數及分數說明。
進行完成投票,然後列印狀態和報告網址,即可在 Foundry 入口開啟結果:
import time
while True:
run = client.evals.runs.retrieve(run_id=eval_run.id, eval_id=evaluation.id)
if run.status in ["completed", "failed"]:
break
time.sleep(5)
print(f"Status: {run.status}")
print(f"Report URL: {run.report_url}")
在執行層級,你可以看到每個評估器彙總的通過與失敗次數:
print(run.result_counts)
for criteria in run.per_testing_criteria_results:
print(criteria.testing_criteria, "passed:", criteria.passed, "failed:", criteria.failed)
ResultCounts(errored=0, failed=0, passed=1, total=1, skipped=0)
Intent Resolution passed: 1 failed: 0
Task Adherence passed: 1 failed: 0
對於列級詳細資料,請列出輸出項目。 每個結果都包含評審人員姓名、通過或不通過,以及一個分數:
for item in client.evals.runs.output_items.list(run_id=eval_run.id, eval_id=evaluation.id):
for result in item.results:
print(item.id, result.name, "passed:", result.passed, "score:", result.score)
進行完成投票,然後列印狀態和報告網址,這樣就能在 Foundry 入口網站開啟結果:
JsonElement run = default;
while (true)
{
ClientResult runStatus = evalClient.GetEvaluationRun(evaluationId, runId, options: null);
run = JsonDocument.Parse(runStatus.GetRawResponse().Content.ToString()).RootElement;
string status = run.GetProperty("status").GetString()!;
if (status is "completed" or "failed") break;
Thread.Sleep(TimeSpan.FromSeconds(5));
}
Console.WriteLine($"Status: {run.GetProperty("status").GetString()}");
Console.WriteLine($"Report URL: {run.GetProperty("report_url").GetString()}");
在執行層級,你可以看到每個評估器彙總的通過與失敗次數:
Console.WriteLine(run.GetProperty("result_counts").GetRawText());
foreach (JsonElement criteria in run.GetProperty("per_testing_criteria_results").EnumerateArray())
{
Console.WriteLine(
$"{criteria.GetProperty("testing_criteria").GetString()} " +
$"passed: {criteria.GetProperty("passed").GetInt32()} " +
$"failed: {criteria.GetProperty("failed").GetInt32()}");
}
對於列級詳細資料,請列出輸出項目。 每個結果都包含評審人員姓名、通過或不通過,以及一個分數:
ClientResult outputItems = evalClient.GetEvaluationRunOutputItems(
evaluationId, runId, limit: 100, order: null, after: null, outputItemStatus: null, options: null);
foreach (JsonElement item in JsonDocument.Parse(outputItems.GetRawResponse().Content.ToString())
.RootElement.GetProperty("data").EnumerateArray())
{
foreach (JsonElement result in item.GetProperty("results").EnumerateArray())
{
Console.WriteLine(
$"{item.GetProperty("id").GetString()} {result.GetProperty("name").GetString()} " +
$"passed: {result.GetProperty("passed")} score: {result.GetProperty("score")}");
}
}
進行完成投票,然後列印狀態和報告網址,這樣就能在 Foundry 入口網站開啟結果:
let run = evalRun;
while (!["completed", "failed"].includes(run.status)) {
run = await client.evals.runs.retrieve(run.id, {
eval_id: evaluation.id,
});
await new Promise((resolve) => setTimeout(resolve, 5000));
}
console.log(`Status: ${run.status}`);
console.log(`Report URL: ${run.report_url}`);
在執行層級,你可以看到每個評估器彙總的通過與失敗次數:
console.log(JSON.stringify(run.result_counts));
for (const criteria of run.per_testing_criteria_results) {
console.log(
criteria.testing_criteria,
"passed:",
criteria.passed,
"failed:",
criteria.failed,
);
}
{"errored":0,"failed":0,"passed":1,"total":1,"skipped":0}
Intent Resolution passed: 1 failed: 0
Task Adherence passed: 1 failed: 0
對於列級詳細資料,請列出輸出項目。 每個結果都包含評審人員姓名、通過或不通過,以及一個分數:
for await (const item of client.evals.runs.outputItems.list(run.id, {
eval_id: evaluation.id,
})) {
for (const result of item.results) {
console.log(item.id, result.name, "passed:", result.passed, "score:", result.score);
}
}