透過定義結構描述、將欄位對應至評估器,並啟動雲端評估作業,來評估 JSONL 或 CSV 資料中預先計算的回應。
先決條件
範例中使用了在 「設定 SDK 客戶端」中設定的 SDK 客戶端。
大多數評估情境都需要輸入資料。 你可以用兩種方式提供資料:
上傳資料集(建議)
上傳 JSONL 或 CSV 檔案,在你的 Foundry 專案中建立版本化資料集。 資料集支援版本控制,並可在多次評估運行中重複使用。 此方法用於生產測試及 CI/CD 工作流程。
準備一個 JSONL 檔案,每行包含一個 JSON 物件,包含評估者所需的欄位:
{"query": "What is machine learning?", "response": "Machine learning is a subset of AI.", "ground_truth": "Machine learning is a type of AI that learns from data."}
{"query": "Explain neural networks.", "response": "Neural networks are computing systems inspired by biological neural networks.", "ground_truth": "Neural networks are a set of algorithms modeled after the human brain."}
或者準備一個 CSV 檔案,欄頭與你的評估者欄位相符:
query,response,ground_truth
What is machine learning?,Machine learning is a subset of AI.,Machine learning is a type of AI that learns from data.
Explain neural networks.,Neural networks are computing systems inspired by biological neural networks.,Neural networks are a set of algorithms modeled after the human brain.
# Upload a local JSONL file. Skip this step if you already have a dataset registered.
data_id = project_client.datasets.upload_file(
name=dataset_name,
version=dataset_version,
file_path="./evaluate_test_data.jsonl",
).id
// Upload a local JSONL file. Skip this step if you already have a dataset registered.
FileDataset dataset = await projectClient.Datasets.UploadFileAsync(
name: datasetName,
version: datasetVersion,
filePath: "./evaluate_test_data.jsonl");
string dataId = dataset.Id;
參考資料: AIProjectDatasetsOperations.UploadFileAsync
// Upload a local JSONL file. Skip this step if you already have a
// dataset registered.
const dataset = await projectClient.datasets.uploadFile(
datasetName,
datasetVersion,
"./evaluate_test_data.jsonl",
);
const dataId = dataset.id;
參考資料: datasets.uploadFile
cURL 範例使用現有的資料集 ID 或內嵌內容。 使用 Python 或 JavaScript/TypeScript 標籤來上傳本地檔案,然後在 cURL 請求中提供其資料集 ID。
提供線上資料
若想快速實驗小型預計算資料集,請直接在評估請求中提供資料,使用 file_content。
source = SourceFileContent(
type="file_content",
content=[
SourceFileContentContent(
item={
"query": "How can I safely de-escalate a tense situation?",
"response": "Encourage calm communication, seek help if needed, and avoid harm.",
"ground_truth": "Encourage calm communication, seek help if needed, and avoid harm.",
}
),
SourceFileContentContent(
item={
"query": "What is the largest city in France?",
"response": "Paris",
"ground_truth": "Paris",
}
),
],
)
object source = new
{
type = "file_content",
content = new[]
{
new
{
item = new
{
query = "How can I safely de-escalate a tense situation?",
ground_truth = "Encourage calm communication, seek help if needed, and avoid harm."
}
},
new
{
item = new
{
query = "What is the largest city in France?",
ground_truth = "Paris"
}
}
}
};
const source = {
type: "file_content",
content: [
{
item: {
query: "How can I safely de-escalate a tense situation?",
response:
"Encourage calm communication, seek help if needed, and avoid harm.",
ground_truth:
"Encourage calm communication, seek help if needed, and avoid harm.",
},
},
{
item: {
query: "What is the largest city in France?",
response: "Paris",
ground_truth: "Paris",
},
},
],
};
你不需要另外寄出申請。 將 file_content 物件直接包含在對應的 建立評估並執行 區段所示的 cURL 請求主體中。
在建立執行程序時,於資料來源設定中把source 作為"source"欄位傳入。 以下資料集部分預設使用 file_id 。
兩種資料集格式皆支援檔案與內嵌來源。
| 資料集格式 |
file_id |
file_content |
資料集(jsonl) |
Yes |
Yes |
CSV (csv) |
Yes |
Yes |
評估一個 JSONL 資料集
使用 jsonl 資料來源類型,評估 JSONL 檔案中預先計算的回應。 當你已經有模型輸出並想評估其品質時,這個情境非常有用。
定義資料結構與評估器
指定與你的 JSONL 欄位相符的結構,並選擇要執行的評估器(測試標準)。 使用 {{item.field}} 語法,透過 data_mapping 將評估器輸入連接到資料集中的欄位。 即使資料集使用標準欄位名稱如 query、 response、 ground_truth,也應包含每位評估者所需的輸入。 關於每位評估器所需的輸入,請參見 內建評估器。
data_source_config = DataSourceConfigCustom(
type="custom",
item_schema={
"type": "object",
"properties": {
"query": {"type": "string"},
"response": {"type": "string"},
"ground_truth": {"type": "string"},
},
"required": ["query", "response", "ground_truth"],
},
)
testing_criteria = [
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="coherence",
evaluator_name="builtin.coherence",
initialization_parameters={"model": model_deployment_name},
data_mapping={
"query": "{{item.query}}",
"response": "{{item.response}}",
},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="violence",
evaluator_name="builtin.violence",
data_mapping={
"query": "{{item.query}}",
"response": "{{item.response}}",
},
),
]
object dataSourceConfig = new
{
type = "custom",
item_schema = new
{
type = "object",
properties = new
{
query = new { type = "string" },
response = new { type = "string" },
ground_truth = new { type = "string" }
},
required = new[] { "query", "response", "ground_truth" }
}
};
object[] testingCriteria =
[
new
{
type = "azure_ai_evaluator",
name = "coherence",
evaluator_name = "builtin.coherence",
initialization_parameters = new { model = modelDeploymentName },
data_mapping = new
{
query = "{{item.query}}",
response = "{{item.response}}"
}
},
new
{
type = "azure_ai_evaluator",
name = "violence",
evaluator_name = "builtin.violence",
data_mapping = new
{
query = "{{item.query}}",
response = "{{item.response}}"
}
},
new
{
type = "azure_ai_evaluator",
name = "f1",
evaluator_name = "builtin.f1_score",
data_mapping = new
{
response = "{{item.response}}",
ground_truth = "{{item.ground_truth}}"
}
}
];
const dataSourceConfig = {
type: "custom",
item_schema: {
type: "object",
properties: {
query: { type: "string" },
response: { type: "string" },
ground_truth: { type: "string" },
},
required: ["query", "response", "ground_truth"],
},
};
const testingCriteria = [
{
type: "azure_ai_evaluator",
name: "coherence",
evaluator_name: "builtin.coherence",
initialization_parameters: { model: modelDeploymentName },
data_mapping: {
query: "{{item.query}}",
response: "{{item.response}}",
},
},
{
type: "azure_ai_evaluator",
name: "violence",
evaluator_name: "builtin.violence",
data_mapping: {
query: "{{item.query}}",
response: "{{item.response}}",
},
},
{
type: "azure_ai_evaluator",
name: "f1",
evaluator_name: "builtin.f1_score",
data_mapping: {
response: "{{item.response}}",
ground_truth: "{{item.ground_truth}}",
},
},
];
curl --request POST \
--url "https://${ACCOUNT}.services.ai.azure.com/api/projects/${PROJECT}/openai/v1/evals" \
--header "Authorization: Bearer ${TOKEN}" \
--header "Content-Type: application/json" \
--data '{
"name": "dataset-evaluation",
"data_source_config": {
"type": "custom",
"item_schema": {
"type": "object",
"properties": {
"query": { "type": "string" },
"response": { "type": "string" },
"ground_truth": { "type": "string" }
},
"required": ["query", "response", "ground_truth"]
}
},
"testing_criteria": [
{
"type": "azure_ai_evaluator",
"name": "coherence",
"evaluator_name": "builtin.coherence",
"initialization_parameters": {"model": "gpt-5-mini"},
"data_mapping": {
"query": "{{item.query}}",
"response": "{{item.response}}"
}
},
{
"type": "azure_ai_evaluator",
"name": "violence",
"evaluator_name": "builtin.violence",
"data_mapping": {
"query": "{{item.query}}",
"response": "{{item.response}}"
}
},
{
"type": "azure_ai_evaluator",
"name": "f1",
"evaluator_name": "builtin.f1_score",
"data_mapping": {
"response": "{{item.response}}",
"ground_truth": "{{item.ground_truth}}"
}
}
]
}'
建立評估並執行
建立評估,然後開始對你上傳的資料集進行運算。 執行過程會針對資料集的每列運行每個評估器。
# Create the evaluation
eval_object = openai_client.evals.create(
name="dataset-evaluation",
data_source_config=data_source_config,
testing_criteria=testing_criteria,
)
# Create a run using the uploaded dataset
eval_run = openai_client.evals.runs.create(
eval_id=eval_object.id,
name="dataset-run",
data_source=CreateEvalJSONLRunDataSourceParam(
type="jsonl",
source=SourceFileID(
type="file_id",
id=data_id,
),
),
)
BinaryData evaluationData = BinaryData.FromObjectAsJson(new
{
name = "dataset-evaluation",
data_source_config = dataSourceConfig,
testing_criteria = testingCriteria
});
using BinaryContent evaluationContent = BinaryContent.Create(evaluationData);
ClientResult evaluation = await evaluationClient.CreateEvaluationAsync(
evaluationContent);
string evaluationId = GetString(evaluation, "id");
object dataSource = new
{
type = "jsonl",
source = new { type = "file_id", id = dataId }
};
BinaryData runData = BinaryData.FromObjectAsJson(new
{
name = "dataset-run",
data_source = dataSource
});
using BinaryContent runContent = BinaryContent.Create(runData);
ClientResult evaluationRun = await evaluationClient.CreateEvaluationRunAsync(
evaluationId: evaluationId,
content: runContent);
Console.WriteLine($"Evaluation run created: {GetString(evaluationRun, "id")}");
參考資料: EvaluationClient 協定方法
// Create the evaluation
const evalObject = await openaiClient.evals.create({
name: "dataset-evaluation",
data_source_config: dataSourceConfig,
testing_criteria: testingCriteria,
});
// Create a run using the uploaded dataset
const evalRun = await openaiClient.evals.runs.create(evalObject.id, {
name: "dataset-run",
data_source: {
type: "jsonl",
source: {
type: "file_id",
id: dataId,
},
},
});
# Step 1: Create the evaluation
EVAL_ID=$(curl --silent --request POST \
--url "https://${ACCOUNT}.services.ai.azure.com/api/projects/${PROJECT}/openai/v1/evals" \
--header "Authorization: Bearer ${TOKEN}" \
--header "Content-Type: application/json" \
--data '{
"name": "dataset-evaluation",
"data_source_config": {
"type": "custom",
"item_schema": {
"type": "object",
"properties": {
"query": { "type": "string" },
"response": { "type": "string" },
"ground_truth": { "type": "string" }
},
"required": ["query", "response", "ground_truth"]
}
},
"testing_criteria": [
{
"type": "azure_ai_evaluator",
"name": "coherence",
"evaluator_name": "builtin.coherence",
"initialization_parameters": { "model": "gpt-5-mini" },
"data_mapping": {
"query": "{{item.query}}",
"response": "{{item.response}}"
}
},
{
"type": "azure_ai_evaluator",
"name": "violence",
"evaluator_name": "builtin.violence",
"data_mapping": {
"query": "{{item.query}}",
"response": "{{item.response}}"
}
},
{
"type": "azure_ai_evaluator",
"name": "f1",
"evaluator_name": "builtin.f1_score",
"data_mapping": {
"response": "{{item.response}}",
"ground_truth": "{{item.ground_truth}}"
}
}
]
}' | jq -r '.id')
# Step 2: Create a run against your dataset
curl --request POST \
--url "https://${ACCOUNT}.services.ai.azure.com/api/projects/${PROJECT}/openai/v1/evals/${EVAL_ID}/runs" \
--header "Authorization: Bearer ${TOKEN}" \
--header "Content-Type: application/json" \
--data '{
"name": "dataset-run",
"data_source": {
"type": "jsonl",
"source": {
"type": "file_id",
"id": "YOUR_DATASET_ID"
}
}
}'
完整可執行範例請參見GitHub上的 sample_evaluations_builtin_with_dataset_id.py。 若要投票完成並解讀結果,請參閱 「取得雲端評估結果」。
評估一個 CSV 資料集
利用資料來源類型評估 CSV 檔案 csv 中預先計算的回應。 此情境運作方式與 資料集評估 相同,但接受的是 CSV 檔案而非 JSONL。 如果你的資料已經是試算表或表格格式,請使用 CSV
準備一份 CSV 檔案
建立一個 CSV 檔案,欄頭與評估者需要的欄位相符。 每一列代表一個測試案例。
query,response,context,ground_truth
What is cloud computing?,Cloud computing delivers computing services over the internet.,Cloud computing is a technology for on-demand resource delivery.,Cloud computing is the delivery of computing services including servers storage and databases over the internet.
What is machine learning?,Machine learning is a subset of AI that learns from data.,Machine learning is a branch of artificial intelligence.,Machine learning is a type of AI that enables computers to learn from data without being explicitly programmed.
Explain neural networks.,Neural networks are computing systems inspired by biological neural networks.,Neural networks are used in deep learning.,Neural networks are a set of algorithms modeled after the human brain designed to recognize patterns.
上傳並執行
將 CSV 檔案上傳為資料集。 接著,利用 csv 資料來源類型建立評估。 結構定義與評估器設定與 JSONL 評估相同。 唯一的差別在於資料來源中的 "type": "csv"。
# Upload the CSV file
data_id = project_client.datasets.upload_file(
name="eval-csv-data",
version="1",
file_path="./evaluation_data.csv",
).id
# Define the schema matching your CSV columns
data_source_config = DataSourceConfigCustom(
type="custom",
item_schema={
"type": "object",
"properties": {
"query": {"type": "string"},
"response": {"type": "string"},
"context": {"type": "string"},
"ground_truth": {"type": "string"},
},
"required": [],
},
include_sample_schema=True,
)
# Define evaluators that use the standard CSV columns
testing_criteria = [
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="coherence",
evaluator_name="builtin.coherence",
initialization_parameters={"model": model_deployment_name},
data_mapping={
"query": "{{item.query}}",
"response": "{{item.response}}",
},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="violence",
evaluator_name="builtin.violence",
data_mapping={
"query": "{{item.query}}",
"response": "{{item.response}}",
},
),
]
# Create the evaluation
eval_object = openai_client.evals.create(
name="CSV evaluation with built-in evaluators",
data_source_config=data_source_config,
testing_criteria=testing_criteria,
)
# Create a run using the CSV data source type
eval_run = openai_client.evals.runs.create(
eval_id=eval_object.id,
name="csv-evaluation-run",
data_source={
"type": "csv",
"source": {
"type": "file_id",
"id": data_id,
},
},
)
object csvDataSourceConfig = new
{
type = "custom",
item_schema = new
{
type = "object",
properties = new
{
query = new { type = "string" },
response = new { type = "string" },
context = new { type = "string" },
ground_truth = new { type = "string" }
},
required = Array.Empty<string>()
},
include_sample_schema = true
};
object[] csvTestingCriteria =
[
new
{
type = "azure_ai_evaluator",
name = "coherence",
evaluator_name = "builtin.coherence",
initialization_parameters = new { model = modelDeploymentName },
data_mapping = new
{
query = "{{item.query}}",
response = "{{item.response}}"
}
},
new
{
type = "azure_ai_evaluator",
name = "violence",
evaluator_name = "builtin.violence",
data_mapping = new
{
query = "{{item.query}}",
response = "{{item.response}}"
}
},
new
{
type = "azure_ai_evaluator",
name = "f1",
evaluator_name = "builtin.f1_score",
data_mapping = new
{
response = "{{item.response}}",
ground_truth = "{{item.ground_truth}}"
}
}
];
FileDataset csvDataset = await projectClient.Datasets.UploadFileAsync(
name: "eval-csv-data",
version: "1",
filePath: "./evaluation_data.csv");
BinaryData evaluationData = BinaryData.FromObjectAsJson(new
{
name = "CSV evaluation with built-in evaluators",
data_source_config = csvDataSourceConfig,
testing_criteria = csvTestingCriteria
});
using BinaryContent evaluationContent = BinaryContent.Create(evaluationData);
ClientResult evaluation = await evaluationClient.CreateEvaluationAsync(
evaluationContent);
string evaluationId = GetString(evaluation, "id");
BinaryData runData = BinaryData.FromObjectAsJson(new
{
name = "csv-evaluation-run",
data_source = new
{
type = "csv",
source = new { type = "file_id", id = csvDataset.Id }
}
});
using BinaryContent runContent = BinaryContent.Create(runData);
ClientResult evaluationRun = await evaluationClient.CreateEvaluationRunAsync(
evaluationId: evaluationId,
content: runContent);
Console.WriteLine($"Evaluation run created: {GetString(evaluationRun, "id")}");
參考: AIProjectDatasetsOperations.UploadFileAsync 及 EvaluationClient 協議方法。
目前的 JavaScript/TypeScript SDK 範例並未展示 CSV 評估。 此流程請使用 Python 或 C# 索引標籤。
請使用 Python 或 C# 分頁上傳 CSV 檔案。 接著,你可以使用 Evals REST 端點,並搭配 csv 資料來源類型和已上傳的資料集 ID。
下一步