僅適用於:Foundry(經典)入口。 這篇文章無法在新的 Foundry 入口網站中提供。
了解更多關於新入口網站的資訊。
註
本文中的連結可能會開啟新版 Microsoft Foundry 文件的內容,而非您目前正在瀏覽的 Foundry(經典版)文件。
重要
本文中標示為預覽的項目目前仍在預覽中。 此預覽版未簽訂服務等級協議,Microsoft 不建議用於生產工作負載。 某些功能可能不被支援或功能受限。 欲了解更多資訊,請參閱Microsoft Azure預覽補充使用條款。
你可以透過將生成式 AI 應用應用到大量資料集中,徹底評估其效能。 在你的開發環境中,使用 Azure AI 評估 SDK 評估應用程式。
當您提供測試資料集或目標時,生成式 AI 應用的輸出會以數學指標及 AI 輔助的品質與安全評估器進行量化衡量。 內建或客製化的評估器能為您提供對應用程式功能與限制的全面洞察。
在本文中,你將學習如何在單一資料列上執行評估器,並在應用目標上執行較大的測試資料集。 你可以使用內建的評估器,這些評估器在本地使用 Azure AI 評估 SDK。 接著,你要學會追蹤 Foundry 專案中的結果和評估日誌。
開始
首先,安裝 Azure AI 評估 SDK 中的 evaluators 套件:
pip install azure-ai-evaluation
註
欲了解更多資訊,請參閱 Azure AI 評估客戶端函式庫Python。
內建評估器
內建品質與安全指標接受查詢與回應對,並提供針對特定評估者的額外資訊。
| 分類 | 評估員 |
|---|---|
| 通用用途 |
CoherenceEvaluator, FluencyEvaluator, QAEvaluator |
| 文本相似性 |
SimilarityEvaluator,F1ScoreEvaluator,BleuScoreEvaluator,GleuScoreEvaluator,RougeScoreEvaluator,MeteorScoreEvaluator |
| 檢索增強生成(RAG) |
RetrievalEvaluator,DocumentRetrievalEvaluator,GroundednessEvaluator,GroundednessProEvaluator,RelevanceEvaluator,ResponseCompletenessEvaluator |
| 風險與安全 |
ViolenceEvaluator, SexualEvaluator, SelfHarmEvaluator, HateUnfairnessEvaluatorIndirectAttackEvaluator, ProtectedMaterialEvaluatorUngroundedAttributesEvaluatorCodeVulnerabilityEvaluator,ContentSafetyEvaluator |
| 代理式 |
IntentResolutionEvaluator, ToolCallAccuracyEvaluator, TaskAdherenceEvaluator |
| Azure OpenAI |
AzureOpenAILabelGrader, AzureOpenAIStringCheckGrader, , AzureOpenAITextSimilarityGraderAzureOpenAIGrader |
內建評估器的資料需求
內建評估工具可接受查詢與回應配對、JSON Lines (JSONL) 格式的對話清單,或同時接受兩者。
| 評估員 | 對話與單回合文字支援 | 對話與單回合文字與圖片支援 | 單回合僅支援文字 | 需求 ground_truth |
支援 代理輸入 |
|---|---|---|---|---|---|
| 品質評估員 | |||||
IntentResolutionEvaluator |
✓ | ||||
ToolCallAccuracyEvaluator |
✓ | ||||
TaskAdherenceEvaluator |
✓ | ||||
GroundednessEvaluator |
✓ | ✓ | |||
GroundednessProEvaluator |
✓ | ||||
RetrievalEvaluator |
✓ | ||||
DocumentRetrievalEvaluator |
✓ | ✓ | |||
RelevanceEvaluator |
✓ | ✓ | |||
CoherenceEvaluator |
✓ | ||||
FluencyEvaluator |
✓ | ||||
ResponseCompletenessEvaluator |
✓ | ✓ | |||
QAEvaluator |
✓ | ✓ | |||
| 自然語言處理(NLP)評估員 | |||||
SimilarityEvaluator |
✓ | ✓ | |||
F1ScoreEvaluator |
✓ | ✓ | |||
RougeScoreEvaluator |
✓ | ✓ | |||
GleuScoreEvaluator |
✓ | ✓ | |||
BleuScoreEvaluator |
✓ | ✓ | |||
MeteorScoreEvaluator |
✓ | ✓ | |||
| 安全評估員 | |||||
ViolenceEvaluator |
✓ | ||||
SexualEvaluator |
✓ | ||||
SelfHarmEvaluator |
✓ | ||||
HateUnfairnessEvaluator |
✓ | ||||
ProtectedMaterialEvaluator |
✓ | ||||
ContentSafetyEvaluator |
✓ | ||||
UngroundedAttributesEvaluator |
✓ | ||||
CodeVulnerabilityEvaluator |
✓ | ||||
IndirectAttackEvaluator |
✓ | ||||
| Azure OpenAI 評分器 | |||||
AzureOpenAILabelGrader |
✓ | ||||
AzureOpenAIStringCheckGrader |
✓ | ||||
AzureOpenAITextSimilarityGrader |
✓ | ✓ | |||
AzureOpenAIGrader |
✓ |
註
除 SimilarityEvaluator外,AI 輔助品質評估器包含一個理由欄位。 他們會使用像是思緒鏈推理等技巧來產生分數的解釋。
由於評估品質提升,生成時會消耗更多代幣使用量。 具體來說,用於評估器生成的 max_token 在大多數 AI 輔助評估器中被設為 800。 為容納較長輸入,RetrievalEvaluator 的值為 1600,ToolCallAccuracyEvaluator 的值為 3000。
Azure OpenAI 評分器需要一個模板,描述他們的輸入欄位如何轉換成評分器使用的 real 輸入。 例如,如果你有兩個輸入,分別是 query 和 response,以及格式為 {{item.query}}的範本,則只使用查詢。 同樣地,您可以使用像 {{item.conversation}} 這樣的寫法來接收對話輸入,但系統是否能正確處理,取決於您如何設定評分器的其他部分,以讓它預期該輸入。
欲了解更多關於代理評估者資料需求的資訊,請參閱 「評估您的 AI 代理」。
單回合支援文字
所有內建的評估器都將單回合輸入視為字串中的查詢與回應對。 例如:
from azure.ai.evaluation import RelevanceEvaluator
query = "What is the capital of life?"
response = "Paris."
# Initialize an evaluator:
relevance_eval = RelevanceEvaluator(model_config)
relevance_eval(query=query, response=response)
若要使用 本地評估 來執行批次評估,或上傳 資料集以執行雲端評估,請以 JSONL 格式表示資料集。 先前的單回合資料為查詢與回應配對,等同於下列資料集中的一行;以下範例顯示三行資料:
{"query":"What is the capital/major city of France?","response":"Paris."}
{"query":"What atoms compose water?","response":"Hydrogen and oxygen."}
{"query":"What color is my shirt?","response":"Blue."}
評估測試資料集可依各內建評估器的需求包含以下元素:
- 查詢:發送給生成式 AI 應用程式的查詢。
- 回應:生成式 AI 應用程式對查詢的回應。
- 情境:所產生回應的來源。 也就是說,用於提供依據的文件。
- 真實標準:由使用者或人工提供的回應,作為正確答案。
想了解每位評估者的需求,請參閱 內建評估器。
文字對話支援
對於支持文本對話的評估工具,你可以提供 conversation 作為輸入。 此輸入包含一個Python字典,其中含有一個列表,該列表包括messages、content、role,以及可選的context。
請參考以下 Python 兩回合對話:
conversation = {
"messages": [
{
"content": "Which tent is the most waterproof?",
"role": "user"
},
{
"content": "The Alpine Explorer Tent is the most waterproof",
"role": "assistant",
"context": "From the our product list the alpine explorer tent is the most waterproof. The Adventure Dining Table has higher weight."
},
{
"content": "How much does it cost?",
"role": "user"
},
{
"content": "The Alpine Explorer Tent is $120.",
"role": "assistant",
"context": None
}
]
}
若要使用 本地評估 來執行批次評估,或 上傳資料集來執行雲端評估,你需要用 JSONL 格式來表示資料集。 前述對話相當於 JSONL 檔案中的一行資料集,如下範例:
{"conversation":
{
"messages": [
{
"content": "Which tent is the most waterproof?",
"role": "user"
},
{
"content": "The Alpine Explorer Tent is the most waterproof",
"role": "assistant",
"context": "From the our product list the alpine explorer tent is the most waterproof. The Adventure Dining Table has higher weight."
},
{
"content": "How much does it cost?",
"role": "user"
},
{
"content": "The Alpine Explorer Tent is $120.",
"role": "assistant",
"context": null
}
]
}
}
我們的評估工具了解對話的第一回合提供來自 query 來自 user、context 來自 assistant、response 來自 assistant,並以查詢-回應格式進行處理。 每回合評估對話,結果會匯總至所有回合以產生對話分數。
註
在第二回合中,即使 context 為 null 或缺少該鍵值,評估工具也會將該回合解讀為空字串,而不會因錯誤而失敗,這可能會導致評估結果產生誤導。
我們強烈建議您驗證評估資料,以符合資料要求。
關於對話模式,這裡有一個範例:GroundednessEvaluator
# Conversation mode:
import json
import os
from azure.ai.evaluation import GroundednessEvaluator, AzureOpenAIModelConfiguration
model_config = AzureOpenAIModelConfiguration(
azure_endpoint=os.environ.get("AZURE_ENDPOINT"),
api_key=os.environ.get("AZURE_API_KEY"),
azure_deployment=os.environ.get("AZURE_DEPLOYMENT_NAME"),
api_version=os.environ.get("AZURE_API_VERSION"),
)
# Initialize the Groundedness evaluator:
groundedness_eval = GroundednessEvaluator(model_config)
conversation = {
"messages": [
{ "content": "Which tent is the most waterproof?", "role": "user" },
{ "content": "The Alpine Explorer Tent is the most waterproof", "role": "assistant", "context": "From the our product list the alpine explorer tent is the most waterproof. The Adventure Dining Table has higher weight." },
{ "content": "How much does it cost?", "role": "user" },
{ "content": "$120.", "role": "assistant", "context": "The Alpine Explorer Tent is $120."}
]
}
# Alternatively, you can load the same content from a JSONL file.
groundedness_conv_score = groundedness_eval(conversation=conversation)
print(json.dumps(groundedness_conv_score, indent=4))
對於對話輸出,每一回合的結果會儲存在一個清單中,而整體對話分數 'groundedness': 4.0 則是對各回合結果取平均而得:
{
"groundedness": 5.0,
"gpt_groundedness": 5.0,
"groundedness_threshold": 3.0,
"evaluation_per_turn": {
"groundedness": [
5.0,
5.0
],
"gpt_groundedness": [
5.0,
5.0
],
"groundedness_reason": [
"The response accurately and completely answers the query by stating that the Alpine Explorer Tent is the most waterproof, which is directly supported by the context. There are no irrelevant details or incorrect information present.",
"The RESPONSE directly answers the QUERY with the exact information provided in the CONTEXT, making it fully correct and complete."
],
"groundedness_result": [
"pass",
"pass"
],
"groundedness_threshold": [
3,
3
]
}
}
註
若要支援更多評估器模型,請使用無前綴的鍵。 例如,使用 groundedness.groundedness。
支援映像及多模態文字與映像的對話功能
對於支援影像與多模態影像與文字對話的評估器,你可以將圖片網址或 Base64 編碼的圖片傳送入 conversation。
支援的情境包括:
- 可以輸入多張圖片和文字來進行影像或文字生成。
- 僅文字輸入的影像生成。
- 僅輸入圖片以生成文字。
from pathlib import Path
from azure.ai.evaluation import ContentSafetyEvaluator
import base64
# Create an instance of an evaluator with image and multi-modal support.
safety_evaluator = ContentSafetyEvaluator(credential=azure_cred, azure_ai_project=project_scope)
# Example of a conversation with an image URL:
conversation_image_url = {
"messages": [
{
"role": "system",
"content": [
{"type": "text", "text": "You are an AI assistant that understands images."}
],
},
{
"role": "user",
"content": [
{"type": "text", "text": "Can you describe this image?"},
{
"type": "image_url",
"image_url": {
"url": "https://cdn.britannica.com/68/178268-050-5B4E7FB6/Tom-Cruise-2013.jpg"
},
},
],
},
{
"role": "assistant",
"content": [
{
"type": "text",
"text": "The image shows a man with short brown hair smiling, wearing a dark-colored shirt.",
}
],
},
]
}
# Example of a conversation with base64 encoded images:
base64_image = ""
with Path.open("Image1.jpg", "rb") as image_file:
base64_image = base64.b64encode(image_file.read()).decode("utf-8")
conversation_base64 = {
"messages": [
{"content": "create an image of a branded apple", "role": "user"},
{
"content": [{"type": "image_url", "image_url": {"url": f"data:image/jpg;base64,{base64_image}"}}],
"role": "assistant",
},
]
}
# Run the evaluation on the conversation to output the result.
safety_score = safety_evaluator(conversation=conversation_image_url)
目前,影像與多模態評估器支援:
- 僅限單回合:對話中只能有一則用戶訊息和一則助理訊息。
- 僅包含一則系統訊息的對話。
- 對話資料負載小於 10 MB,包括圖像。
- 絕對網址和 Base64 編碼的圖片。
- 單一回合中包含多張影像。
- JPG/JPEG、PNG 和 GIF 檔案格式。
設定
對於 AI 輔助品質評估者,除GroundednessProEvaluator預覽外,你必須在你的gpt-35-turbo中指定一個 GPT 模型(gpt-4、gpt-4-turbo、gpt-4o、gpt-4o-mini、或model_config)。 GPT 模型作為評審,對評估資料進行評分。 我們支援 Azure OpenAI 或 OpenAI 模型配置架構。 為了獲得最佳效能與可解析的回應,我們建議使用非預覽版的 GPT 模型搭配評估工具。
註
將gpt-3.5-turbo替換為gpt-4o-mini以應用於評估模型。 根據 OpenAI 的說法,它 gpt-4o-mini 更便宜、更強大、速度也相當。
要用 API 金鑰進行推論呼叫,請確保你至少有 Azure OpenAI 資源的 Cognitive Services OpenAI User 角色。 欲了解更多權限資訊,請參閱 一個 Azure OpenAI 資源的權限。
對於所有風險與安全評估者以及(預覽),而非在中的 GPT 部署,你必須提供你的資訊。
對於所有風險與安全評估工具,以及 GroundednessProEvaluator (預覽版),您必須提供 model_config 資訊,而不是在 azure_ai_project 中提供 GPT 部署。 這會透過你的 Foundry 專案存取後端評估服務。
AI 輔助內建評估器的使用提示
為了透明起見,我們將品質評估者的提示開源於我們的評估者庫及 Azure AI Evaluation Python SDK 存儲庫,唯獨安全評估者與GroundednessProEvaluator,這些是由 Azure AI 內容安全 提供支持。 這些提示作為語言模型執行評估任務的指令,該任務需要對衡量指標及其相應評分標準有友善的定義。 我們強烈建議您根據情境細節客製化定義與評分標準。 欲了解更多資訊,請參閱 自訂評估器。
綜合評估者
複合評估器是內建的評估器,結合個別品質或安全指標。 它們現成可用,提供多種評估計量,適用於查詢與回應配對或聊天訊息。
| 綜合評估員 | 包含 | 描述 |
|---|---|---|
QAEvaluator |
GroundednessEvaluator,RelevanceEvaluator,CoherenceEvaluator,FluencyEvaluator,SimilarityEvaluator,F1ScoreEvaluator |
將所有品質評估器合併,輸出查詢與回應對的綜合指標 |
ContentSafetyEvaluator |
ViolenceEvaluator, SexualEvaluator, , SelfHarmEvaluatorHateUnfairnessEvaluator |
將所有安全評估器合併,輸出查詢與回應對的綜合指標 |
使用 evaluate() 在測試資料集上進行本地評估
在你對單一資料列進行內建或自訂評估器抽查後,你可以將多個評估器與 evaluate() API 合併到整個測試資料集。
Microsoft Foundry 專案的前置設置步驟
如果這次是你第一次執行評估並記錄到 Foundry 專案,可能需要執行以下設定步驟:
- 在資源層級建立並連結您的儲存帳號至 Foundry 專案。 這個 bicep 範本會使用金鑰驗證,佈建和連接儲存體帳戶到您的 Foundry 專案。
- 確保連接的儲存帳號能存取所有專案。
- 如果您將儲存帳戶與 Microsoft Entra ID 連結,請確保在 Microsoft Azure 入口網站中對您的帳戶和 Foundry 專案資源同時賦予Storage Blob Data Owner的 Microsoft 身份權限。
在資料集上進行評估,並將結果記錄到 Foundry。
為了確保 evaluate() API 能正確解析資料,你必須指定欄位映射,將欄位從資料集映射到評估者接受的關鍵字。 此範例指定了 、 queryresponse、 和 context的資料映射。
from azure.ai.evaluation import evaluate
result = evaluate(
data="data.jsonl", # Provide your data here:
evaluators={
"groundedness": groundedness_eval,
"answer_length": answer_length
},
# Column mapping:
evaluator_config={
"groundedness": {
"column_mapping": {
"query": "${data.queries}",
"context": "${data.context}",
"response": "${data.response}"
}
}
},
# Optionally, provide your Foundry project information to track your evaluation results in your project portal.
azure_ai_project = azure_ai_project,
# Optionally, provide an output path to dump a JSON file of metric summary, row-level data, and the metric and Foundry project URL.
output_path="./myevalresults.json"
)
提示
取得 result.studio_url 屬性的內容,以獲得查看您在 Foundry 專案中已記錄的評估結果的連結。
評估工具會以字典形式輸出結果,其中包含彙總的 metrics,以及資料列層級的資料與計量。 請參考以下範例輸出:
{'metrics': {'answer_length.value': 49.333333333333336,
'groundedness.gpt_groundeness': 5.0, 'groundedness.groundeness': 5.0},
'rows': [{'inputs.response': 'Paris is the capital/major city of France.',
'inputs.context': 'Paris has been the capital/major city of France since '
'the 10th century and is known for its '
'cultural and historical landmarks.',
'inputs.query': 'What is the capital/major city of France?',
'outputs.answer_length.value': 31,
'outputs.groundeness.groundeness': 5,
'outputs.groundeness.gpt_groundeness': 5,
'outputs.groundeness.groundeness_reason': 'The response to the query is supported by the context.'},
{'inputs.response': 'Albert Einstein developed the theory of '
'relativity.',
'inputs.context': 'Albert Einstein developed the theory of '
'relativity, with his special relativity '
'published in 1905 and general relativity in '
'1915.',
'inputs.query': 'Who developed the theory of relativity?',
'outputs.answer_length.value': 51,
'outputs.groundeness.groundeness': 5,
'outputs.groundeness.gpt_groundeness': 5,
'outputs.groundeness.groundeness_reason': 'The response to the query is supported by the context.'},
{'inputs.response': 'The speed of light is approximately 299,792,458 '
'meters per second.',
'inputs.context': 'The exact speed of light in a vacuum is '
'299,792,458 meters per second, a constant '
"used in physics to represent 'c'.",
'inputs.query': 'What is the speed of light?',
'outputs.answer_length.value': 66,
'outputs.groundeness.groundeness': 5,
'outputs.groundeness.gpt_groundeness': 5,
'outputs.groundeness.groundeness_reason': 'The response to the query is supported by the context.'}],
'traces': {}}
要求 evaluate()
API evaluate() 需要特定的資料格式和評估器參數的鍵名,才能在你的 Foundry 專案中正確顯示評估結果圖表。
資料格式
API evaluate() 僅接受 JSONL 格式的資料。 對於所有內建評估器, evaluate() 需要以下格式的資料及所需的輸入欄位。 請參見 前一節關於內建評估器所需資料輸入的部分。 以下程式碼片段是一行可能長什麼樣子的範例:
{
"query":"What is the capital/major city of France?",
"context":"France is in Europe",
"response":"Paris is the capital/major city of France.",
"ground_truth": "Paris"
}
評估器參數格式
當你輸入內建的評估器時,請在 evaluators 參數列表中指定正確的關鍵字映射。 下表顯示了您內建評估器的結果在登錄到您的 Foundry 專案時顯示在 UI 中所需的關鍵字映射。
| 評估員 | 關鍵字參數 |
|---|---|
GroundednessEvaluator |
"groundedness" |
GroundednessProEvaluator |
"groundedness_pro" |
RetrievalEvaluator |
"retrieval" |
RelevanceEvaluator |
"relevance" |
CoherenceEvaluator |
"coherence" |
FluencyEvaluator |
"fluency" |
SimilarityEvaluator |
"similarity" |
F1ScoreEvaluator |
"f1_score" |
RougeScoreEvaluator |
"rouge" |
GleuScoreEvaluator |
"gleu" |
BleuScoreEvaluator |
"bleu" |
MeteorScoreEvaluator |
"meteor" |
ViolenceEvaluator |
"violence" |
SexualEvaluator |
"sexual" |
SelfHarmEvaluator |
"self_harm" |
HateUnfairnessEvaluator |
"hate_unfairness" |
IndirectAttackEvaluator |
"indirect_attack" |
ProtectedMaterialEvaluator |
"protected_material" |
CodeVulnerabilityEvaluator |
"code_vulnerability" |
UngroundedAttributesEvaluator |
"ungrounded_attributes" |
QAEvaluator |
"qa" |
ContentSafetyEvaluator |
"content_safety" |
這裡有一個設定 evaluators 參數的範例:
result = evaluate(
data="data.jsonl",
evaluators={
"sexual":sexual_evaluator,
"self_harm":self_harm_evaluator,
"hate_unfairness":hate_unfairness_evaluator,
"violence":violence_evaluator
}
)
對目標進行本機評估
如果您有一組查詢要執行並進行評估,evaluate() API 也支援 target 參數。 這個參數會將查詢傳送給應用程式收集答案,然後用評估器執行該查詢與回應。
目標可以是目錄中任何可呼叫的類別。 在這個例子中,有一個Python腳本 askwiki.py,且可呼叫的類別 askwiki() 被設定為目標。 如果您有一組查詢資料集,可以送入簡單的 askwiki 應用程式,就可以評估輸出內容的依據性。 請確認您設定的資料在 "column_mapping" 中具有正確的欄位映射。 你可以用來 "default" 指定所有評估器的欄位映射。
以下是內容:"data.jsonl"
{"query":"When was United States found ?", "response":"1776"}
{"query":"What is the capital/major city of France?", "response":"Paris"}
{"query":"Who is the best tennis player of all time ?", "response":"Roger Federer"}
from askwiki import askwiki
result = evaluate(
data="data.jsonl",
target=askwiki,
evaluators={
"groundedness": groundedness_eval
},
evaluator_config={
"default": {
"column_mapping": {
"query": "${data.queries}",
"context": "${outputs.context}",
"response": "${outputs.response}"
}
}
}
)
相關內容
- Azure AI 評估用於 Python 的客戶端函式庫
- AI 評估 SDK 問題排除
- 生成式人工智慧中的可觀察性
使用 Microsoft Foundry SDK - 產生合成與模擬資料以供評估
- 請參閱 Foundry 入口網站的評估結果
- 開始使用 Foundry
- 開始使用評估樣本