使用 Microsoft Foundry SDK 評估現有資料集

透過定義結構描述、將欄位對應至評估器,並啟動雲端評估作業,來評估 JSONL 或 CSV 資料中預先計算的回應。

先決條件

範例中使用了在 「設定 SDK 客戶端」中設定的 SDK 客戶端。

準備輸入資料

大多數評估情境都需要輸入資料。 你可以用兩種方式提供資料:

Tip

如果你沒有手工整理的資料集,也可以自己做一個。 在上線前或流量較低時,可以使用 產生合成評估資料集 ,或將 代理追蹤轉換成評估資料集 ,從真實生產流量建立資料集。

上傳 JSONL 或 CSV 檔案,在你的 Foundry 專案中建立版本化資料集。 資料集支援版本控制,並可在多次評估運行中重複使用。 此方法用於生產測試及 CI/CD 工作流程。

準備一個 JSONL 檔案,每行包含一個 JSON 物件,包含評估者所需的欄位:

{"query": "What is machine learning?", "response": "Machine learning is a subset of AI.", "ground_truth": "Machine learning is a type of AI that learns from data."}
{"query": "Explain neural networks.", "response": "Neural networks are computing systems inspired by biological neural networks.", "ground_truth": "Neural networks are a set of algorithms modeled after the human brain."}

或者準備一個 CSV 檔案,欄頭與你的評估者欄位相符:

query,response,ground_truth
What is machine learning?,Machine learning is a subset of AI.,Machine learning is a type of AI that learns from data.
Explain neural networks.,Neural networks are computing systems inspired by biological neural networks.,Neural networks are a set of algorithms modeled after the human brain.
# Upload a local JSONL file. Skip this step if you already have a dataset registered.
data_id = project_client.datasets.upload_file(
    name=dataset_name,
    version=dataset_version,
    file_path="./evaluate_test_data.jsonl",
).id

提供線上資料

若想快速實驗小型預計算資料集,請直接在評估請求中提供資料,使用 file_content。

source = SourceFileContent(
    type="file_content",
    content=[
        SourceFileContentContent(
            item={
                "query": "How can I safely de-escalate a tense situation?",
                "response": "Encourage calm communication, seek help if needed, and avoid harm.",
                "ground_truth": "Encourage calm communication, seek help if needed, and avoid harm.",
            }
        ),
        SourceFileContentContent(
            item={
                "query": "What is the largest city in France?",
                "response": "Paris",
                "ground_truth": "Paris",
            }
        ),
    ],
)

在建立執行程序時,於資料來源設定中把source 作為"source"欄位傳入。 以下資料集部分預設使用 file_id 。

各資料集格式支援的來源類型

兩種資料集格式皆支援檔案與內嵌來源。

資料集格式 file_id file_content
資料集(jsonl) Yes Yes
CSV (csv) Yes Yes

評估一個 JSONL 資料集

使用 jsonl 資料來源類型,評估 JSONL 檔案中預先計算的回應。 當你已經有模型輸出並想評估其品質時,這個情境非常有用。

Tip

在開始之前,先完成 客戶端設定 並 準備輸入資料。

定義資料結構與評估器

指定與你的 JSONL 欄位相符的結構,並選擇要執行的評估器(測試標準)。 使用 {{item.field}} 語法,透過 data_mapping 將評估器輸入連接到資料集中的欄位。 即使資料集使用標準欄位名稱如 query、 response、 ground_truth,也應包含每位評估者所需的輸入。 關於每位評估器所需的輸入,請參見 內建評估器。

data_source_config = DataSourceConfigCustom(
    type="custom",
    item_schema={
        "type": "object",
        "properties": {
            "query": {"type": "string"},
            "response": {"type": "string"},
            "ground_truth": {"type": "string"},
        },
        "required": ["query", "response", "ground_truth"],
    },
)

testing_criteria = [
    TestingCriterionAzureAIEvaluator(
        type="azure_ai_evaluator",
        name="coherence",
        evaluator_name="builtin.coherence",
        initialization_parameters={"model": model_deployment_name},
        data_mapping={
            "query": "{{item.query}}",
            "response": "{{item.response}}",
        },
    ),
    TestingCriterionAzureAIEvaluator(
        type="azure_ai_evaluator",
        name="violence",
        evaluator_name="builtin.violence",
        data_mapping={
            "query": "{{item.query}}",
            "response": "{{item.response}}",
        },
    ),
]

建立評估並執行

建立評估,然後開始對你上傳的資料集進行運算。 執行過程會針對資料集的每列運行每個評估器。

# Create the evaluation
eval_object = openai_client.evals.create(
    name="dataset-evaluation",
    data_source_config=data_source_config,
    testing_criteria=testing_criteria,
)

# Create a run using the uploaded dataset
eval_run = openai_client.evals.runs.create(
    eval_id=eval_object.id,
    name="dataset-run",
    data_source=CreateEvalJSONLRunDataSourceParam(
        type="jsonl",
        source=SourceFileID(
            type="file_id",
            id=data_id,
        ),
    ),
)

完整可執行範例請參見GitHub上的 sample_evaluations_builtin_with_dataset_id.py。 若要投票完成並解讀結果,請參閱 「取得雲端評估結果」。

評估一個 CSV 資料集

利用資料來源類型評估 CSV 檔案 csv 中預先計算的回應。 此情境運作方式與 資料集評估 相同,但接受的是 CSV 檔案而非 JSONL。 如果你的資料已經是試算表或表格格式,請使用 CSV

Tip

在開始之前,先完成 客戶端設定 並 準備輸入資料。

準備一份 CSV 檔案

建立一個 CSV 檔案,欄頭與評估者需要的欄位相符。 每一列代表一個測試案例。

query,response,context,ground_truth
What is cloud computing?,Cloud computing delivers computing services over the internet.,Cloud computing is a technology for on-demand resource delivery.,Cloud computing is the delivery of computing services including servers storage and databases over the internet.
What is machine learning?,Machine learning is a subset of AI that learns from data.,Machine learning is a branch of artificial intelligence.,Machine learning is a type of AI that enables computers to learn from data without being explicitly programmed.
Explain neural networks.,Neural networks are computing systems inspired by biological neural networks.,Neural networks are used in deep learning.,Neural networks are a set of algorithms modeled after the human brain designed to recognize patterns.

上傳並執行

將 CSV 檔案上傳為資料集。 接著,利用 csv 資料來源類型建立評估。 結構定義與評估器設定與 JSONL 評估相同。 唯一的差別在於資料來源中的 "type": "csv"。

# Upload the CSV file
data_id = project_client.datasets.upload_file(
    name="eval-csv-data",
    version="1",
    file_path="./evaluation_data.csv",
).id

# Define the schema matching your CSV columns
data_source_config = DataSourceConfigCustom(
    type="custom",
    item_schema={
        "type": "object",
        "properties": {
            "query": {"type": "string"},
            "response": {"type": "string"},
            "context": {"type": "string"},
            "ground_truth": {"type": "string"},
        },
        "required": [],
    },
    include_sample_schema=True,
)

# Define evaluators that use the standard CSV columns
testing_criteria = [
    TestingCriterionAzureAIEvaluator(
        type="azure_ai_evaluator",
        name="coherence",
        evaluator_name="builtin.coherence",
        initialization_parameters={"model": model_deployment_name},
        data_mapping={
            "query": "{{item.query}}",
            "response": "{{item.response}}",
        },
    ),
    TestingCriterionAzureAIEvaluator(
        type="azure_ai_evaluator",
        name="violence",
        evaluator_name="builtin.violence",
        data_mapping={
            "query": "{{item.query}}",
            "response": "{{item.response}}",
        },
    ),
]

# Create the evaluation
eval_object = openai_client.evals.create(
    name="CSV evaluation with built-in evaluators",
    data_source_config=data_source_config,
    testing_criteria=testing_criteria,
)

# Create a run using the CSV data source type
eval_run = openai_client.evals.runs.create(
    eval_id=eval_object.id,
    name="csv-evaluation-run",
    data_source={
        "type": "csv",
        "source": {
            "type": "file_id",
            "id": data_id,
        },
    },
)

下一步