Important
本文中標示為預覽的項目目前仍在預覽中。 此預覽版未簽訂服務等級協議,Microsoft 不建議用於生產工作負載。 某些功能可能不被支援或功能受限。 欲了解更多資訊,請參閱Microsoft Azure預覽補充使用條款。
內建的評估器提供一種簡便的方式來監控應用程式世代的品質。 為了自訂你的評估,你可以建立自己的程式碼、提示或端點式評估器。
自訂評估器讓你能定義領域專屬的品質指標,超越 內建的評估器目錄。 當你需要衡量應用獨有的標準時,例如品牌語調、領域特定準確性或輸出格式合規性,請使用自訂評估器。
你可以建立三種類型的客製化評估器:
| 以程式碼為基礎 | 提示型 | 端點式 | |
|---|---|---|---|
| 運作原理 | Python grade() 函數以確定性邏輯評分每一項。 |
評審提示指示大型語言模型(LLM)為每項項目評分。 | 外部 HTTP 端點接收評估資料並回傳分數。 |
| 適用對象 | 規則式檢查、關鍵字匹配、格式驗證、長度限制。 | 主觀品質判斷、語意相似性、語氣分析。 | 自訂評分邏輯託管於您自己的基礎設施、專有模型,或需要網路存取的複雜管線。 |
| 計分方法 | 連續:浮動從 0.0 到 1.0(越高越好)。 | 序數、連續或二元。 你定義了序數分數和連續分數的最小/最大範圍。 分數越高越好。 | 由你的終端點定義。 回傳一個符合標準評估結果架構的 JSON 物件。 |
| 輸出合約 | 一個介於 0.0 到 1.0 之間的單一浮點數值。 | 一個 JSON 物件,且 resultreason。 的 result 類型取決於評分方法:序數用整數,連續用 float,二進位用布林。 |
一個包含 score、 reason、 status和 可選 properties的 JSON 物件。 請參見 端點回應架構。 |
建立自訂評估器後,你可以將其加入 Foundry 專案的評估器目錄,並用於 批次評估執行。
程式碼評估工具
基於程式碼的評估器是一個名為 grade 的Python函式,接收兩個字典參數(sample 和 item),並回傳浮點分數介於 0.0 到 1.0(越高越好)。 實務上,所有資料皆透過 item以下方式存取:
-
Dataset evaluation:輸入欄位如
response或ground_truth可在Python程式碼中取得,如item.get("response")或item.get("ground_truth")。 -
模型或代理人目標評估:要擷取產生的回應文字,請使用
item.get("sample", {}).get("output_text")。
Note
目前,模型或代理人目標產生的回應文字是透過 item.get("sample", {}).get("output_text")存取的。 此存取模式可能會在未來的 API 更新中改變。
以下範例根據回答長度評分,偏好 50 到 500 字元之間的回答:
def grade(sample: dict, item: dict) -> float:
"""Score based on response length (prefer 50-500 chars)."""
# For dataset evaluation, access fields directly from item:
response = item.get("response", "")
# For model/agent target evaluation, use item.get("sample") instead:
# response = item.get("sample", {}).get("output_text", "")
if not response:
return 0.0
length = len(response)
if length < 50:
return 0.2
elif length > 500:
return 0.5
return 1.0
Note
若函 grade() 式產生例外或逾時,服務會將該項目的結果記錄為 , 0.0 並在評估報告中標記為錯誤。 設計你的功能時採取防禦性——用於 try/except 風險操作並回傳備用分數,而非讓例外事件擴散。
支援的套件與限制
基於程式碼的評估器運行於沙盒化的 Python 環境中,具有以下限制條件:
- 程式碼大小必須小於 256 KB。
- 每次評分電話執行時間限制為2分鐘。
- 執行時無法使用網路存取。
- 記憶體限制為 2 GB,磁碟限制為 1 GB,CPU 限制為 2 核心。
以下第三方方案可供選擇:
| Package | 版本 |
|---|---|
numpy |
2.2.4 |
scipy |
1.15.2 |
pandas |
2.2.3 |
scikit-learn |
1.6.1 |
rapidfuzz |
3.10.1 |
sympy |
1.13.3 |
jsonschema |
4.23.0 |
pydantic |
2.10.6 |
deepdiff |
8.4.2 |
nltk |
3.9.1 |
rouge-score |
0.1.2 |
pyyaml |
6.0.2 |
NLTK 語料庫punkt、stopwords、wordnetomw-1.4names皆為預先載入。
執行時參數
pass_threshold
deployment_name和 是建立基於程式碼的評估器時的初始化參數。 即使基於程式碼的評估器不會呼叫 LLM,服務 API 架構仍要求 deployment_name 評估執行的編排。 你可以從你的專案中傳遞任何有效的模型部署名稱。
提示評估工具
基於提示的評估器使用法官提示範本,由大型語言模型(LLM)對每個項目進行評估。 模板變數使用雙大括號(例如 {{query}}),並映射到你的輸入資料欄位。
基於提示的評估器支援三種評分方法:
- 序數:整數分數,在你定義的離散量表上(例如1到5)。 越高越好。
- 連續:對你定義的範圍(例如 0.0–1.0)進行細緻測量的浮動分數。 越高越好。
- 二元 (真/假):基於閾值的檢查布林結果。
評估器必須回傳一個 JSON 物件,且 resultreason。 與你的計分方法相符的類型 result :序數用整數,連續用浮點數,二進位用布林值。
以下範例提示使用序數評分(1–5)來評估回應的友善度:
Friendliness assesses the warmth and approachability of the response.
Rate the friendliness of the response between one and five using the following scale:
1 - Unfriendly or hostile
2 - Mostly unfriendly
3 - Neutral
4 - Mostly friendly
5 - Very friendly
Assign a rating based on the tone and demeanor of the response.
Response:
{{response}}
Output Format (JSON):
{
"result": <integer from 1 to 5>,
"reason": "<brief explanation for the score>"
}
執行時參數
在建立基於提示的評估器時,這兩個初始deployment_namethreshold化參數都是必不可少的。
基於端點的評估器
端點式評估器會將評分委派給你擁有並操作的外部 HTTP 端點。 評估服務會呼叫你的端點,針對每個項目(或一組項目),並將映射的輸入資料以 JSON 有效載荷傳遞。 你的端點會用你選擇的任何邏輯處理資料,並回傳一個帶有分數的 JSON 回應。
當你需要時,可以使用端點式評估器:
- 評分時可網路存取外部服務或資料庫。
- 專有模型或機器學習管線託管在你自己的基礎設施上。
- 複雜的評分邏輯,超出沙盒式程式碼評估器的限制。
- 與現有評估服務或 API 整合。
運作原理
- 你部署一個 HTTP 端點,接受帶有評估資料的 POST 請求。
- 你在 Foundry 專案中建立一個連線,儲存端點 URL 和認證憑證。
- 你註冊一個端點式的評估器,該評估器會參考該連線。
- 當評估執行時,服務會解決連線,呼叫你的端點並輸入資料,並將回應記錄為評估結果。
端點請求結構
評估服務會向你的端點發送一個 POST 請求,並附帶包含評估元資料和映射輸入欄位的 JSON 正文。
下表描述你的端點會接收到的欄位:
| Field | 類型 | Description |
|---|---|---|
schema_version |
string |
請求架構的版本。 目前。"0.0.1" |
evaluator_name |
string |
被執行的評估員註冊名稱。 |
evaluator_version |
string |
評估者定義的版本。 |
evaluation_level |
string |
評估細緻度: "turn" 針對每項目或 "conversation" 完整對話。 |
data |
object |
包含評估輸入資料。 詳見 data.item 及 data.sample 以下。 |
data.item |
object |
評估資料集的輸入欄位,透過 data_mapping 配置映射。 |
data.sample |
object |
由模型或代理人目標產生的輸出。 只在評估目標時出現。 |
範例請求:
{
"schema_version": "0.0.1",
"evaluator_name": "my_endpoint_evaluator",
"evaluator_version": "1",
"evaluation_level": "turn",
"data": {
"item": {
"query": "What is the capital of France?"
},
"sample": {
"response": "Paris"
}
}
}
端點回應結構
下表描述了端點可回傳的欄位:
| Field | 類型 | Description |
|---|---|---|
score |
double 或 bool,可為零 |
評估分數。 類型依評估者而異。 跳過或錯誤時會顯示為 null。 |
reason |
string,可空 |
樂譜說明。 非LLM評估者則無效。 |
status |
string |
執行狀態: "completed"、、 "error"或 "skipped"。 |
properties |
object,可空 |
用於評估者專用資料的鍵值袋,標準欄位未捕捉。 |
threshold |
integer,可空 |
通過/不通過門檻。 不使用閾值的評估者則為空。 |
passed |
bool,可空 |
分數是否達到門檻。 當評估者出錯或被跳過時,則為 null。 |
schema_version |
string |
回應模式的版本。 請使用 "0.0.1"。 |
error |
object |
錯誤 status 細節 當 是 "error"。 包含 code 和 message。 未包含在成功回應中。 |
成功回應:
你的端點必須回傳一個符合標準評估結果結構的 JSON 物件:
{
"schema_version": "0.0.1",
"score": 0.95,
"reason": "The response accurately answers the question using the provided context.",
"status": "completed",
"properties": {
"confidence": 0.87,
"source_coverage": "full"
},
"threshold": 3,
"passed": true
}
失敗回應:
若發生錯誤,您的端點需回傳符合以下結構的 JSON 物件:
{
"schema_version": "0.0.1",
"status": "error",
"error": {
"code": "500",
"message": "Model inference failed"
}
}
Authentication
端點式評估器透過專案連線支援兩種認證方法:
| 方法 | 運作原理 | 最適合用於 |
|---|---|---|
| API 金鑰 | 服務在呼叫端點時會透過請求標頭傳遞金鑰。 | 簡單端點、Azure Functions 搭配函數層鍵、第三方 API。 |
| Microsoft Entra 身份識別 | Azure 函式會取得一個受管理身份憑證,並將其作為持有人憑證傳遞。 | Azure Functions with role-based access control, Azure Functions with Easy Auth. |
建立端點連線
連線會儲存端點網址和認證憑證。 使用 Azure Cognitive Services 管理客戶端建立連線:
API 金鑰連接
from azure.mgmt.cognitiveservices import CognitiveServicesManagementClient
from azure.mgmt.cognitiveservices.models import ConnectionPropertiesV2BasicResource
mgmt_client = CognitiveServicesManagementClient(
credential=credential,
subscription_id=subscription_id,
)
connection = ConnectionPropertiesV2BasicResource(
properties={
"category": "ApiKey",
"target": "https://your-endpoint.azurewebsites.net/api/evaluate",
"authType": "ApiKey",
"credentials": {
"key": "<your-api-key>",
},
},
)
mgmt_client.account_connections.create(
resource_group_name=resource_group,
account_name=account_name,
connection_name="my-endpoint-connection",
connection=connection,
)
Microsoft Entra ID connection
connection = ConnectionPropertiesV2BasicResource(
properties={
"category": "CustomKeys",
"target": "https://your-endpoint.azurewebsites.net/api/evaluate",
"authType": "AAD",
"credentials": {
"Audience": "api://<your-app-registration-client-id>",
},
},
)
mgmt_client.account_connections.create(
resource_group_name=resource_group,
account_name=account_name,
connection_name="my-endpoint-entra-connection",
connection=connection,
)
為了 Entra ID 認證,你的端點必須設定接受專案管理身份所發出的憑證。 這通常涉及:
- 在 Microsoft Entra ID 中註冊應用程式作為你的端點。
- 在您的端點啟用 Easy Auth(或等效的令牌驗證)。
- 將專案的管理身份賦予目標應用程式的應用程式角色指派。
註冊評估員
建立連線後,註冊一個基於端點的評估器以參考該連線:
endpoint_evaluator = project_client.beta.evaluators.create_version(
name="my-endpoint-evaluator",
evaluator_version={
"name": "my-endpoint-evaluator",
"categories": [EvaluatorCategory.QUALITY],
"display_name": "My Endpoint Evaluator",
"description": "Scores responses using a custom evaluation endpoint",
"definition": {
"type": "endpoint",
"connection_name": "my-endpoint-connection",
},
},
)
使用端點基礎評估器執行評估
使用欄位 data_mapping 指定哪些輸入資料欄位會傳送到你的端點:
testing_criteria = [
{
"type": "azure_ai_evaluator",
"name": "endpoint_eval",
"evaluator_name": "my-endpoint-evaluator",
"data_mapping": {
"query": "{{item.query}}",
"response": "{{item.response}}",
"context": "{{item.context}}",
},
},
]
這些 data_mapping 金鑰會成為你的端點收到的 JSON 欄位。 用 {{item.<field_name>}} 語法將它們映射到評估資料集中的欄位。
部署你的端點
你的評估端點可以是任何接受 POST 請求並回傳 JSON 的 HTTP 服務。 常見的主機選項包括:
- Azure Functions:輕量級、無伺服器主機,用於簡單的評分邏輯。
- Azure App 服務:用於複雜評估流程的完整網頁應用託管。
- Azure 容器應用程式: Container-based hosting for ML model inference.
端點必須在評估服務逾時(30 秒)內回應,並對每個請求回傳有效的 JSON 回應。
用 SDK 建立自訂評估器
前提與設置
安裝 SDK 並設定你的客戶端:
pip install "azure-ai-projects>=2.0.0"
import os
import time
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import EvaluatorCategory, EvaluatorDefinitionType
from openai.types.eval_create_params import DataSourceConfigCustom
from openai.types.evals.create_eval_jsonl_run_data_source_param import (
CreateEvalJSONLRunDataSourceParam,
SourceFileContent,
SourceFileContentContent,
)
# Azure AI Project endpoint
# Example: https://<account_name>.services.ai.azure.com/api/projects/<project_name>
endpoint = os.environ["AZURE_AI_PROJECT_ENDPOINT"]
# Model deployment name (required for prompt-based evaluators)
# Example: gpt-5-mini
model_deployment_name = os.environ.get("AZURE_AI_MODEL_DEPLOYMENT_NAME", "")
# Create the project client
project_client = AIProjectClient(
endpoint=endpoint,
credential=DefaultAzureCredential(),
)
# Get the OpenAI client for evaluation API
client = project_client.get_openai_client()
建立一個基於程式碼的評估器
將函式以欄位中的grade()字串傳遞code_text。 定義 宣 data_schema 告函式預期的輸入欄位,並 metrics 定義 來描述函式回傳的分數。 基於程式碼的評估器使用 continuous 範圍為 0.0 至 1.0 的度量類型。
首先,定義評估器版本架構:
code_evaluator = project_client.beta.evaluators.create_version(
name="response_length_scorer",
evaluator_version={
"name": "response_length_scorer",
"categories": [EvaluatorCategory.QUALITY],
"display_name": "Response Length Scorer",
"description": "Scores responses based on length, preferring 50-500 characters",
"definition": {
"type": EvaluatorDefinitionType.CODE,
"code_text": (
'def grade(sample: dict, item: dict) -> float:\n'
' """Score based on response length (prefer 50-500 chars)."""\n'
' response = item.get("response", "")\n'
' if not response:\n'
' return 0.0\n'
' length = len(response)\n'
' if length < 50:\n'
' return 0.2\n'
' elif length > 500:\n'
' return 0.5\n'
' return 1.0\n'
),
"init_parameters": {
"type": "object",
"properties": {
"deployment_name": {"type": "string"},
"pass_threshold": {"type": "number"},
},
"required": ["deployment_name", "pass_threshold"],
},
"metrics": {
"result": {
"type": "continuous",
"desirable_direction": "increase",
"min_value": 0.0,
"max_value": 1.0,
}
},
"data_schema": {
"type": "object",
"required": ["item"],
"properties": {
"item": {
"type": "object",
"properties": {
"response": {"type": "string"},
},
},
},
},
},
},
)
完整範例請參見基於 code 的評估器Python SDK 範例。
建立以提示為基礎的評估器
在現場傳遞裁判提示 prompt_text 。 定義 宣 data_schema 告提示所期望的輸入欄位,以及 描述 metrics 評分方法與範圍。 宣 init_parameters 告模型部署及評估器執行時所需的閾值。
prompt_evaluator = project_client.beta.evaluators.create_version(
name="friendliness_evaluator",
evaluator_version={
"name": "friendliness_evaluator",
"categories": [EvaluatorCategory.QUALITY],
"display_name": "Friendliness Evaluator",
"description": "Evaluates the warmth and approachability of a response",
"definition": {
"type": EvaluatorDefinitionType.PROMPT,
"prompt_text": (
"Friendliness assesses the warmth and approachability of the response.\n"
"Rate the friendliness of the response between one and five "
"using the following scale:\n\n"
"1 - Unfriendly or hostile\n"
"2 - Mostly unfriendly\n"
"3 - Neutral\n"
"4 - Mostly friendly\n"
"5 - Very friendly\n\n"
"Assign a rating based on the tone and demeanor of the response.\n\n"
"Response:\n{{response}}\n\n"
"Output Format (JSON):\n"
'{\n "result": <integer from 1 to 5>,\n'
' "reason": "<brief explanation for the score>"\n}\n'
),
"init_parameters": {
"type": "object",
"properties": {
"deployment_name": {"type": "string"},
"threshold": {"type": "number"},
},
"required": ["deployment_name", "threshold"],
},
"data_schema": {
"type": "object",
"properties": {
"response": {"type": "string"},
},
"required": ["response"],
},
"metrics": {
"custom_prompt": {
"type": "ordinal",
"desirable_direction": "increase",
"min_value": 1,
"max_value": 5,
}
},
},
},
)
完整範例請參考基於 prompt-based evaluator Python SDK 範例。
用自訂評估器執行評估
建立自訂評估器後,在評估執行中使用它們,就像使用內建評估器一樣。 你可以在一次運行中包含多位評估者。
以下範例同時執行基於 response_length_scorer 程式碼與提示 friendliness_evaluator 的程式碼。
定義並執行評估
# Define the data schema
data_source_config = DataSourceConfigCustom(
type="custom",
item_schema={
"type": "object",
"properties": {
"response": {"type": "string"},
},
"required": ["response"],
},
)
# Reference both custom evaluators in testing criteria
testing_criteria = [
{
"type": "azure_ai_evaluator",
"name": "response_length_scorer",
"evaluator_name": "response_length_scorer",
"initialization_parameters": {
"deployment_name": model_deployment_name,
"pass_threshold": 0.5,
},
},
{
"type": "azure_ai_evaluator",
"name": "friendliness_evaluator",
"evaluator_name": "friendliness_evaluator",
"data_mapping": {
"response": "{{item.response}}",
},
"initialization_parameters": {
"deployment_name": model_deployment_name,
"threshold": 3,
},
},
]
# Create the evaluation
eval_object = client.evals.create(
name="custom-eval-test",
data_source_config=data_source_config,
testing_criteria=testing_criteria,
)
# Run the evaluation with inline data
eval_run = client.evals.runs.create(
eval_id=eval_object.id,
name="custom-eval-run-01",
data_source=CreateEvalJSONLRunDataSourceParam(
type="jsonl",
source=SourceFileContent(
type="file_content",
content=[
SourceFileContentContent(
item={
"response": "I'm sorry this watch isn't working for you. I'd be happy to help you with a replacement!",
}
),
SourceFileContentContent(
item={
"response": "I will not apologize for my behavior!",
}
),
],
),
),
)
取得成果
輪詢評估執行直到結束,然後取得每個項目的結果和回報網址。
while True:
run = client.evals.runs.retrieve(run_id=eval_run.id, eval_id=eval_object.id)
if run.status in ("completed", "failed"):
break
time.sleep(5)
# Get per-item results
output_items = list(
client.evals.runs.output_items.list(run_id=run.id, eval_id=eval_object.id)
)
print(f"Status: {run.status}")
print(f"Report: {run.report_url}")
清理資源
刪除自訂評估器版本及不再需要時的評估:
# Delete the custom evaluator version
project_client.beta.evaluators.delete_version(
name="response_length_scorer",
version=code_evaluator.version,
)
# Delete the evaluation
client.evals.delete(eval_id=eval_object.id)
欲了解更多資料來源選項、評估器映射及進階情境,請參閱 從 SDK 執行評估。
如需更多範例,包括列出、更新及刪除評估器,請參閱 evaluator 目錄管理 Python SDK 範例。
在入口網站建立自訂評估器
你可以直接在 Azure AI Foundry 入口網站建立自訂評估器,無需撰寫 SDK 程式碼。
- 在你的 Foundry 專案中,請前往 Evaluation>Evaluator 目錄。
- 選擇 自訂評估器>建立。
- 請填寫以下欄位:
| Field | Description |
|---|---|
| Name | 評估器的唯一識別碼(例如, response_length_scorer)。 |
| 顯示名稱 | 評估器目錄中顯示的人類可讀名稱。 |
| Description | 評估者所衡量的簡要摘要。 |
| Type | 基於程式碼或提示為基礎。 決定你提供Python grade()函式還是法官提示。 |
| 計分方法 | 基於程式碼的評估器使用連續(0.0–1.0)。 基於提示的評估者可以使用序數、連續或二進位評分,並設定自訂範圍。 |
| 程式碼或提示 | 針對程式碼,請在程式碼編輯器中撰寫函 grade() 式。 針對提示型,請在提示編輯器中撰寫評審提示。 請參閱本文前述的程式碼與提示評估器章節,了解範例與需求。 |
在入口網站評估中使用自訂評估器
建立自訂評估器後,請在入口網站進行評估執行:
- 在你的 Foundry 專案中,進入 評估 並選擇 建立。
- 請依照評估建立向導操作。 在 「標準 」步驟中,選擇 「新增評估器」。
- 從評估者目錄中選擇你的客製化評估者。
- 提供所需的初始化參數。 對於基於提示的評估器,請提供 模型部署 與 閾值。 對於基於程式碼的評估器,請提供 通過門檻。
- 完成巫師並開始評估跑。
關於從入口網站執行評估的詳細步驟,請參見 「從入口網站執行評估」。
對話層級自訂評估器
自訂評估者可以評分整段對話,而非單回合。 為了實現對話層級的評估:
- 評估跑開始
evaluation_level="conversation" - 設計你的
grade()函式時,應該是對話item["messages"]陣列
在對話層級執行時, item dict 會接收完整的對話訊息陣列,而非單一查詢/回應對。 這讓你能建立自訂指標,評估整個使用者互動。
範例:會話層級合規檢查
以下範例檢查代理人在對話過程中是否透露了必要的免責聲明:
def grade(sample: dict, item: dict) -> float:
"""Check if agent disclosed required disclaimer during conversation."""
messages = item.get("messages", [])
for msg in messages:
if msg.get("role") == "assistant":
content = msg.get("content", "")
if isinstance(content, str) and "not financial advice" in content.lower():
return 1.0
return 0.0 # Disclaimer never provided
範例:對話長度評分器
此範例根據對話是否在目標回合數內解決來評分:
def grade(sample: dict, item: dict) -> float:
"""Score based on conversation length (prefer shorter resolutions)."""
messages = item.get("messages", [])
# Count user turns (excludes system messages)
user_turns = sum(1 for msg in messages if msg.get("role") == "user")
if user_turns <= 2:
return 1.0 # Resolved quickly
elif user_turns <= 4:
return 0.7 # Reasonable length
elif user_turns <= 6:
return 0.4 # Getting long
else:
return 0.2 # Too many turns