Custom evaluators (preview)

Important

本文中標示為預覽的項目目前仍在預覽中。 此預覽版未簽訂服務等級協議,Microsoft 不建議用於生產工作負載。 某些功能可能不被支援或功能受限。 欲了解更多資訊,請參閱Microsoft Azure預覽補充使用條款。

內建的評估器提供一種簡便的方式來監控應用程式世代的品質。 為了自訂你的評估,你可以建立自己的程式碼、提示或端點式評估器。

自訂評估器讓你能定義領域專屬的品質指標,超越 內建的評估器目錄。 當你需要衡量應用獨有的標準時,例如品牌語調、領域特定準確性或輸出格式合規性,請使用自訂評估器。

你可以建立三種類型的客製化評估器:

以程式碼為基礎 提示型 端點式
運作原理 Python grade() 函數以確定性邏輯評分每一項。 評審提示指示大型語言模型(LLM)為每項項目評分。 外部 HTTP 端點接收評估資料並回傳分數。
適用對象 規則式檢查、關鍵字匹配、格式驗證、長度限制。 主觀品質判斷、語意相似性、語氣分析。 自訂評分邏輯託管於您自己的基礎設施、專有模型,或需要網路存取的複雜管線。
計分方法 連續:浮動從 0.0 到 1.0(越高越好)。 序數、連續或二元。 你定義了序數分數和連續分數的最小/最大範圍。 分數越高越好。 由你的終端點定義。 回傳一個符合標準評估結果架構的 JSON 物件。
輸出合約 一個介於 0.0 到 1.0 之間的單一浮點數值。 一個 JSON 物件,且 resultreason。 的 result 類型取決於評分方法:序數用整數,連續用 float,二進位用布林。 一個包含 score、 reason、 status和 可選 properties的 JSON 物件。 請參見 端點回應架構。

建立自訂評估器後,你可以將其加入 Foundry 專案的評估器目錄,並用於 批次評估執行。

程式碼評估工具

基於程式碼的評估器是一個名為 grade 的Python函式,接收兩個字典參數(sample 和 item),並回傳浮點分數介於 0.0 到 1.0(越高越好)。 實務上,所有資料皆透過 item以下方式存取:

  • Dataset evaluation:輸入欄位如response或ground_truth可在Python程式碼中取得,如item.get("response")或item.get("ground_truth")。
  • 模型或代理人目標評估:要擷取產生的回應文字,請使用 item.get("sample", {}).get("output_text")。

Note

目前,模型或代理人目標產生的回應文字是透過 item.get("sample", {}).get("output_text")存取的。 此存取模式可能會在未來的 API 更新中改變。

以下範例根據回答長度評分,偏好 50 到 500 字元之間的回答:

def grade(sample: dict, item: dict) -> float:
    """Score based on response length (prefer 50-500 chars)."""
    # For dataset evaluation, access fields directly from item:
    response = item.get("response", "")

    # For model/agent target evaluation, use item.get("sample") instead:
    # response = item.get("sample", {}).get("output_text", "")

    if not response:
        return 0.0

    length = len(response)
    if length < 50:
        return 0.2
    elif length > 500:
        return 0.5
    return 1.0

Note

若函 grade() 式產生例外或逾時,服務會將該項目的結果記錄為 , 0.0 並在評估報告中標記為錯誤。 設計你的功能時採取防禦性——用於 try/except 風險操作並回傳備用分數,而非讓例外事件擴散。

支援的套件與限制

基於程式碼的評估器運行於沙盒化的 Python 環境中,具有以下限制條件:

  • 程式碼大小必須小於 256 KB。
  • 每次評分電話執行時間限制為2分鐘。
  • 執行時無法使用網路存取。
  • 記憶體限制為 2 GB,磁碟限制為 1 GB,CPU 限制為 2 核心。

以下第三方方案可供選擇:

Package 版本
numpy 2.2.4
scipy 1.15.2
pandas 2.2.3
scikit-learn 1.6.1
rapidfuzz 3.10.1
sympy 1.13.3
jsonschema 4.23.0
pydantic 2.10.6
deepdiff 8.4.2
nltk 3.9.1
rouge-score 0.1.2
pyyaml 6.0.2

NLTK 語料庫punkt、stopwords、wordnetomw-1.4names皆為預先載入。

執行時參數

pass_threshold deployment_name和 是建立基於程式碼的評估器時的初始化參數。 即使基於程式碼的評估器不會呼叫 LLM,服務 API 架構仍要求 deployment_name 評估執行的編排。 你可以從你的專案中傳遞任何有效的模型部署名稱。

提示評估工具

基於提示的評估器使用法官提示範本,由大型語言模型(LLM)對每個項目進行評估。 模板變數使用雙大括號(例如 {{query}}),並映射到你的輸入資料欄位。

基於提示的評估器支援三種評分方法:

  • 序數:整數分數,在你定義的離散量表上(例如1到5)。 越高越好。
  • 連續:對你定義的範圍(例如 0.0–1.0)進行細緻測量的浮動分數。 越高越好。
  • 二元 (真/假):基於閾值的檢查布林結果。

評估器必須回傳一個 JSON 物件,且 resultreason。 與你的計分方法相符的類型 result :序數用整數,連續用浮點數,二進位用布林值。

以下範例提示使用序數評分(1–5)來評估回應的友善度:

Friendliness assesses the warmth and approachability of the response.
Rate the friendliness of the response between one and five using the following scale:

1 - Unfriendly or hostile
2 - Mostly unfriendly
3 - Neutral
4 - Mostly friendly
5 - Very friendly

Assign a rating based on the tone and demeanor of the response.

Response:
{{response}}

Output Format (JSON):
{
  "result": <integer from 1 to 5>,
  "reason": "<brief explanation for the score>"
}

執行時參數

在建立基於提示的評估器時,這兩個初始deployment_namethreshold化參數都是必不可少的。

基於端點的評估器

端點式評估器會將評分委派給你擁有並操作的外部 HTTP 端點。 評估服務會呼叫你的端點,針對每個項目(或一組項目),並將映射的輸入資料以 JSON 有效載荷傳遞。 你的端點會用你選擇的任何邏輯處理資料,並回傳一個帶有分數的 JSON 回應。

當你需要時,可以使用端點式評估器:

  • 評分時可網路存取外部服務或資料庫。
  • 專有模型或機器學習管線託管在你自己的基礎設施上。
  • 複雜的評分邏輯,超出沙盒式程式碼評估器的限制。
  • 與現有評估服務或 API 整合。

運作原理

  1. 你部署一個 HTTP 端點,接受帶有評估資料的 POST 請求。
  2. 你在 Foundry 專案中建立一個連線,儲存端點 URL 和認證憑證。
  3. 你註冊一個端點式的評估器,該評估器會參考該連線。
  4. 當評估執行時,服務會解決連線,呼叫你的端點並輸入資料,並將回應記錄為評估結果。

端點請求結構

評估服務會向你的端點發送一個 POST 請求,並附帶包含評估元資料和映射輸入欄位的 JSON 正文。

下表描述你的端點會接收到的欄位:

Field 類型 Description
schema_version string 請求架構的版本。 目前。"0.0.1"
evaluator_name string 被執行的評估員註冊名稱。
evaluator_version string 評估者定義的版本。
evaluation_level string 評估細緻度: "turn" 針對每項目或 "conversation" 完整對話。
data object 包含評估輸入資料。 詳見 data.item 及 data.sample 以下。
data.item object 評估資料集的輸入欄位,透過 data_mapping 配置映射。
data.sample object 由模型或代理人目標產生的輸出。 只在評估目標時出現。

範例請求:

{
  "schema_version": "0.0.1",
  "evaluator_name": "my_endpoint_evaluator",
  "evaluator_version": "1",
  "evaluation_level": "turn",
  "data": {
    "item": {
      "query": "What is the capital of France?"
    },
    "sample": {
      "response": "Paris"
    }
  }
}

端點回應結構

下表描述了端點可回傳的欄位:

Field 類型 Description
score double 或 bool,可為零 評估分數。 類型依評估者而異。 跳過或錯誤時會顯示為 null。
reason string,可空 樂譜說明。 非LLM評估者則無效。
status string 執行狀態: "completed"、、 "error"或 "skipped"。
properties object,可空 用於評估者專用資料的鍵值袋,標準欄位未捕捉。
threshold integer,可空 通過/不通過門檻。 不使用閾值的評估者則為空。
passed bool,可空 分數是否達到門檻。 當評估者出錯或被跳過時,則為 null。
schema_version string 回應模式的版本。 請使用 "0.0.1"。
error object 錯誤 status 細節 當 是 "error"。 包含 code 和 message。 未包含在成功回應中。

成功回應:

你的端點必須回傳一個符合標準評估結果結構的 JSON 物件:

{
  "schema_version": "0.0.1",
  "score": 0.95,
  "reason": "The response accurately answers the question using the provided context.",
  "status": "completed",
  "properties": {
    "confidence": 0.87,
    "source_coverage": "full"
  },
  "threshold": 3,
  "passed": true
}

失敗回應:

若發生錯誤,您的端點需回傳符合以下結構的 JSON 物件:

 {
   "schema_version": "0.0.1",
   "status": "error",
   "error": {
     "code": "500",
     "message": "Model inference failed"
   }
 }

Authentication

端點式評估器透過專案連線支援兩種認證方法:

方法 運作原理 最適合用於
API 金鑰 服務在呼叫端點時會透過請求標頭傳遞金鑰。 簡單端點、Azure Functions 搭配函數層鍵、第三方 API。
Microsoft Entra 身份識別 Azure 函式會取得一個受管理身份憑證,並將其作為持有人憑證傳遞。 Azure Functions with role-based access control, Azure Functions with Easy Auth.

建立端點連線

連線會儲存端點網址和認證憑證。 使用 Azure Cognitive Services 管理客戶端建立連線:

API 金鑰連接

from azure.mgmt.cognitiveservices import CognitiveServicesManagementClient
from azure.mgmt.cognitiveservices.models import ConnectionPropertiesV2BasicResource

mgmt_client = CognitiveServicesManagementClient(
    credential=credential,
    subscription_id=subscription_id,
)

connection = ConnectionPropertiesV2BasicResource(
    properties={
        "category": "ApiKey",
        "target": "https://your-endpoint.azurewebsites.net/api/evaluate",
        "authType": "ApiKey",
        "credentials": {
            "key": "<your-api-key>",
        },
    },
)

mgmt_client.account_connections.create(
    resource_group_name=resource_group,
    account_name=account_name,
    connection_name="my-endpoint-connection",
    connection=connection,
)

Microsoft Entra ID connection

connection = ConnectionPropertiesV2BasicResource(
    properties={
        "category": "CustomKeys",
        "target": "https://your-endpoint.azurewebsites.net/api/evaluate",
        "authType": "AAD",
        "credentials": {
            "Audience": "api://<your-app-registration-client-id>",
        },
    },
)

mgmt_client.account_connections.create(
    resource_group_name=resource_group,
    account_name=account_name,
    connection_name="my-endpoint-entra-connection",
    connection=connection,
)

為了 Entra ID 認證,你的端點必須設定接受專案管理身份所發出的憑證。 這通常涉及:

  • 在 Microsoft Entra ID 中註冊應用程式作為你的端點。
  • 在您的端點啟用 Easy Auth(或等效的令牌驗證)。
  • 將專案的管理身份賦予目標應用程式的應用程式角色指派。

註冊評估員

建立連線後,註冊一個基於端點的評估器以參考該連線:

endpoint_evaluator = project_client.beta.evaluators.create_version(
    name="my-endpoint-evaluator",
    evaluator_version={
        "name": "my-endpoint-evaluator",
        "categories": [EvaluatorCategory.QUALITY],
        "display_name": "My Endpoint Evaluator",
        "description": "Scores responses using a custom evaluation endpoint",
        "definition": {
            "type": "endpoint",
            "connection_name": "my-endpoint-connection",
        },
    },
)

使用端點基礎評估器執行評估

使用欄位 data_mapping 指定哪些輸入資料欄位會傳送到你的端點:

testing_criteria = [
    {
        "type": "azure_ai_evaluator",
        "name": "endpoint_eval",
        "evaluator_name": "my-endpoint-evaluator",
        "data_mapping": {
            "query": "{{item.query}}",
            "response": "{{item.response}}",
            "context": "{{item.context}}",
        },
    },
]

這些 data_mapping 金鑰會成為你的端點收到的 JSON 欄位。 用 {{item.<field_name>}} 語法將它們映射到評估資料集中的欄位。

部署你的端點

你的評估端點可以是任何接受 POST 請求並回傳 JSON 的 HTTP 服務。 常見的主機選項包括:

  • Azure Functions:輕量級、無伺服器主機,用於簡單的評分邏輯。
  • Azure App 服務:用於複雜評估流程的完整網頁應用託管。
  • Azure 容器應用程式: Container-based hosting for ML model inference.

端點必須在評估服務逾時(30 秒)內回應,並對每個請求回傳有效的 JSON 回應。

用 SDK 建立自訂評估器

前提與設置

安裝 SDK 並設定你的客戶端:

pip install "azure-ai-projects>=2.0.0"
import os
import time
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import EvaluatorCategory, EvaluatorDefinitionType
from openai.types.eval_create_params import DataSourceConfigCustom
from openai.types.evals.create_eval_jsonl_run_data_source_param import (
    CreateEvalJSONLRunDataSourceParam,
    SourceFileContent,
    SourceFileContentContent,
)

# Azure AI Project endpoint
# Example: https://<account_name>.services.ai.azure.com/api/projects/<project_name>
endpoint = os.environ["AZURE_AI_PROJECT_ENDPOINT"]

# Model deployment name (required for prompt-based evaluators)
# Example: gpt-5-mini
model_deployment_name = os.environ.get("AZURE_AI_MODEL_DEPLOYMENT_NAME", "")

# Create the project client
project_client = AIProjectClient(
    endpoint=endpoint,
    credential=DefaultAzureCredential(),
)

# Get the OpenAI client for evaluation API
client = project_client.get_openai_client()

建立一個基於程式碼的評估器

將函式以欄位中的grade()字串傳遞code_text。 定義 宣 data_schema 告函式預期的輸入欄位,並 metrics 定義 來描述函式回傳的分數。 基於程式碼的評估器使用 continuous 範圍為 0.0 至 1.0 的度量類型。

首先,定義評估器版本架構:

code_evaluator = project_client.beta.evaluators.create_version(
    name="response_length_scorer",
    evaluator_version={
        "name": "response_length_scorer",
        "categories": [EvaluatorCategory.QUALITY],
        "display_name": "Response Length Scorer",
        "description": "Scores responses based on length, preferring 50-500 characters",
        "definition": {
            "type": EvaluatorDefinitionType.CODE,
            "code_text": (
                'def grade(sample: dict, item: dict) -> float:\n'
                '    """Score based on response length (prefer 50-500 chars)."""\n'
                '    response = item.get("response", "")\n'
                '    if not response:\n'
                '        return 0.0\n'
                '    length = len(response)\n'
                '    if length < 50:\n'
                '        return 0.2\n'
                '    elif length > 500:\n'
                '        return 0.5\n'
                '    return 1.0\n'
            ),
            "init_parameters": {
                "type": "object",
                "properties": {
                    "deployment_name": {"type": "string"},
                    "pass_threshold": {"type": "number"},
                },
                "required": ["deployment_name", "pass_threshold"],
            },
            "metrics": {
                "result": {
                    "type": "continuous",
                    "desirable_direction": "increase",
                    "min_value": 0.0,
                    "max_value": 1.0,
                }
            },
            "data_schema": {
                "type": "object",
                "required": ["item"],
                "properties": {
                    "item": {
                        "type": "object",
                        "properties": {
                            "response": {"type": "string"},
                        },
                    },
                },
            },
        },
    },
)

完整範例請參見基於 code 的評估器Python SDK 範例。

建立以提示為基礎的評估器

在現場傳遞裁判提示 prompt_text 。 定義 宣 data_schema 告提示所期望的輸入欄位,以及 描述 metrics 評分方法與範圍。 宣 init_parameters 告模型部署及評估器執行時所需的閾值。

prompt_evaluator = project_client.beta.evaluators.create_version(
    name="friendliness_evaluator",
    evaluator_version={
        "name": "friendliness_evaluator",
        "categories": [EvaluatorCategory.QUALITY],
        "display_name": "Friendliness Evaluator",
        "description": "Evaluates the warmth and approachability of a response",
        "definition": {
            "type": EvaluatorDefinitionType.PROMPT,
            "prompt_text": (
                "Friendliness assesses the warmth and approachability of the response.\n"
                "Rate the friendliness of the response between one and five "
                "using the following scale:\n\n"
                "1 - Unfriendly or hostile\n"
                "2 - Mostly unfriendly\n"
                "3 - Neutral\n"
                "4 - Mostly friendly\n"
                "5 - Very friendly\n\n"
                "Assign a rating based on the tone and demeanor of the response.\n\n"
                "Response:\n{{response}}\n\n"
                "Output Format (JSON):\n"
                '{\n  "result": <integer from 1 to 5>,\n'
                '  "reason": "<brief explanation for the score>"\n}\n'
            ),
            "init_parameters": {
                "type": "object",
                "properties": {
                    "deployment_name": {"type": "string"},
                    "threshold": {"type": "number"},
                },
                "required": ["deployment_name", "threshold"],
            },
            "data_schema": {
                "type": "object",
                "properties": {
                    "response": {"type": "string"},
                },
                "required": ["response"],
            },
            "metrics": {
                "custom_prompt": {
                    "type": "ordinal",
                    "desirable_direction": "increase",
                    "min_value": 1,
                    "max_value": 5,
                }
            },
        },
    },
)

完整範例請參考基於 prompt-based evaluator Python SDK 範例。

用自訂評估器執行評估

建立自訂評估器後,在評估執行中使用它們,就像使用內建評估器一樣。 你可以在一次運行中包含多位評估者。

以下範例同時執行基於 response_length_scorer 程式碼與提示 friendliness_evaluator 的程式碼。

定義並執行評估

# Define the data schema
data_source_config = DataSourceConfigCustom(
    type="custom",
    item_schema={
        "type": "object",
        "properties": {
            "response": {"type": "string"},
        },
        "required": ["response"],
    },
)

# Reference both custom evaluators in testing criteria
testing_criteria = [
    {
        "type": "azure_ai_evaluator",
        "name": "response_length_scorer",
        "evaluator_name": "response_length_scorer",
        "initialization_parameters": {
            "deployment_name": model_deployment_name,
            "pass_threshold": 0.5,
        },
    },
    {
        "type": "azure_ai_evaluator",
        "name": "friendliness_evaluator",
        "evaluator_name": "friendliness_evaluator",
        "data_mapping": {
            "response": "{{item.response}}",
        },
        "initialization_parameters": {
            "deployment_name": model_deployment_name,
            "threshold": 3,
        },
    },
]

# Create the evaluation
eval_object = client.evals.create(
    name="custom-eval-test",
    data_source_config=data_source_config,
    testing_criteria=testing_criteria,
)

# Run the evaluation with inline data
eval_run = client.evals.runs.create(
    eval_id=eval_object.id,
    name="custom-eval-run-01",
    data_source=CreateEvalJSONLRunDataSourceParam(
        type="jsonl",
        source=SourceFileContent(
            type="file_content",
            content=[
                SourceFileContentContent(
                    item={
                        "response": "I'm sorry this watch isn't working for you. I'd be happy to help you with a replacement!",
                    }
                ),
                SourceFileContentContent(
                    item={
                        "response": "I will not apologize for my behavior!",
                    }
                ),
            ],
        ),
    ),
)

取得成果

輪詢評估執行直到結束,然後取得每個項目的結果和回報網址。

while True:
    run = client.evals.runs.retrieve(run_id=eval_run.id, eval_id=eval_object.id)
    if run.status in ("completed", "failed"):
        break
    time.sleep(5)

# Get per-item results
output_items = list(
    client.evals.runs.output_items.list(run_id=run.id, eval_id=eval_object.id)
)

print(f"Status: {run.status}")
print(f"Report: {run.report_url}")

清理資源

刪除自訂評估器版本及不再需要時的評估:

# Delete the custom evaluator version
project_client.beta.evaluators.delete_version(
    name="response_length_scorer",
    version=code_evaluator.version,
)

# Delete the evaluation
client.evals.delete(eval_id=eval_object.id)

欲了解更多資料來源選項、評估器映射及進階情境,請參閱 從 SDK 執行評估。

如需更多範例,包括列出、更新及刪除評估器,請參閱 evaluator 目錄管理 Python SDK 範例。

在入口網站建立自訂評估器

你可以直接在 Azure AI Foundry 入口網站建立自訂評估器,無需撰寫 SDK 程式碼。

  1. 在你的 Foundry 專案中,請前往 Evaluation>Evaluator 目錄。
  2. 選擇 自訂評估器>建立。
  3. 請填寫以下欄位:
Field Description
Name 評估器的唯一識別碼(例如, response_length_scorer)。
顯示名稱 評估器目錄中顯示的人類可讀名稱。
Description 評估者所衡量的簡要摘要。
Type 基於程式碼或提示為基礎。 決定你提供Python grade()函式還是法官提示。
計分方法 基於程式碼的評估器使用連續(0.0–1.0)。 基於提示的評估者可以使用序數、連續或二進位評分,並設定自訂範圍。
程式碼或提示 針對程式碼,請在程式碼編輯器中撰寫函 grade() 式。 針對提示型,請在提示編輯器中撰寫評審提示。 請參閱本文前述的程式碼與提示評估器章節,了解範例與需求。

在入口網站評估中使用自訂評估器

建立自訂評估器後,請在入口網站進行評估執行:

  1. 在你的 Foundry 專案中,進入 評估 並選擇 建立。
  2. 請依照評估建立向導操作。 在 「標準 」步驟中,選擇 「新增評估器」。
  3. 從評估者目錄中選擇你的客製化評估者。
  4. 提供所需的初始化參數。 對於基於提示的評估器,請提供 模型部署 與 閾值。 對於基於程式碼的評估器,請提供 通過門檻。
  5. 完成巫師並開始評估跑。

關於從入口網站執行評估的詳細步驟,請參見 「從入口網站執行評估」。

對話層級自訂評估器

自訂評估者可以評分整段對話,而非單回合。 為了實現對話層級的評估:

  1. 評估跑開始evaluation_level="conversation"
  2. 設計你的 grade() 函式時,應該是對話 item["messages"] 陣列

在對話層級執行時, item dict 會接收完整的對話訊息陣列,而非單一查詢/回應對。 這讓你能建立自訂指標,評估整個使用者互動。

範例:會話層級合規檢查

以下範例檢查代理人在對話過程中是否透露了必要的免責聲明:

def grade(sample: dict, item: dict) -> float:
    """Check if agent disclosed required disclaimer during conversation."""
    messages = item.get("messages", [])
    
    for msg in messages:
        if msg.get("role") == "assistant":
            content = msg.get("content", "")
            if isinstance(content, str) and "not financial advice" in content.lower():
                return 1.0
    
    return 0.0  # Disclaimer never provided

範例:對話長度評分器

此範例根據對話是否在目標回合數內解決來評分:

def grade(sample: dict, item: dict) -> float:
    """Score based on conversation length (prefer shorter resolutions)."""
    messages = item.get("messages", [])
    
    # Count user turns (excludes system messages)
    user_turns = sum(1 for msg in messages if msg.get("role") == "user")
    
    if user_turns <= 2:
        return 1.0  # Resolved quickly
    elif user_turns <= 4:
        return 0.7  # Reasonable length
    elif user_turns <= 6:
        return 0.4  # Getting long
    else:
        return 0.2  # Too many turns