ターンレベルまたは会話レベルでデータセットからの完全な会話を評価します。
前提条件
- クラウド評価の前提条件とクライアントのセットアップを完了します。
-
messages配列を使用した会話データ。 - 選択した評価レベルをサポートするエバリュエーター (一貫性や接地など)。
この例では、「SDK クライアントのセットアップ」で構成された SDK クライアントを使用します。
会話データセットを評価する
完全な会話を評価して、個々の応答だけでなく、ユーザーの操作全体にわたってエージェントの品質を評価します。 会話レベルの評価を用いて、タスクの不完全な解決、ユーザーのフラストレーション、ターンレベルの評価では見逃されるツール呼び出しの回帰といった品質上の問題を特定します。
たとえば、ユーザーが複数のターンで不満を感じるサポート エージェントを考えてみましょう。
ターン 1 — ユーザー: "パスワードをリセットする必要があります。"エージェント: "アカウントが見つかりました。 リセット リンクを送信します。
ターン 2 — ユーザー: "私は電子メールを受け取りませんでした。エージェント: "リンクを再送信しました。 スパムを確認してください。
ターン 3 - ユーザー: "まだ何もありません。 直接リセットできますか?エージェント: "別のリセット リンクを送信しました。"
ターンレベルのエバリュエーターは、最後の応答 (丁寧でアクションを実行する) のみをスコア付けするので、スコアが良くなります。 会話レベルのエバリュエーターは、エージェントが別の方法を試さずに同じ失敗したアクションを 3 回繰り返し、ユーザーの問題を未解決のままにした、会話フラグ全体で 顧客満足度 を評価します。
会話レベルの評価は、ターンレベルの評価といくつかの点で異なります。
| 特徴 | ターン レベル | 会話レベル |
|---|---|---|
| スコープ | 個々のクエリ応答ペア | 複数の交換で会話を完了する |
| Metrics | 応答ごとの品質と安全性 | 会話レベルの結果とユーザーの満足度 |
| データ形式 |
queryフィールドとresponse フィールドを含む JSONL |
完全な会話を含む messages 配列を持つ JSONL |
| 使用例 | 個々のモデル応答のテスト | エンド ツー エンドのエージェント エクスペリエンスのテスト |
データ ソースに一致する会話ワークフローを選択します。
| Workflow | いつ使用するか | データ ソースの種類 |
|---|---|---|
| データセットまたはインラインから | ローカルの会話トレースまたはテストデータがある場合 |
jsonl で file_id または file_content |
| デプロイされた会話 | Application Insights から特定の会話またはサンプリングされた運用トラフィックを評価する場合 |
azure_ai_trace_data_source_preview と trace_source |
| シミュレートされた会話 | 合成テスト用の会話を生成したい |
azure_ai_target_completions と conversation_gen_preview |
評価レベルを選択する
実行の evaluation_level パラメーターは、エバリュエーターが個々のターンをスコア付けするか、会話を完了するかを決定します。
| 価値 | Behavior |
|---|---|
"turn" |
エバリュエーターは、各ターンを個別にスコア付けします。 |
"conversation" |
エバリュエーターは、会話全体をスコア付けします。 |
| (省略) | 既定値は "turn" です。 |
Important
エバリュエーターの互換性: 各エバリュエーターは、特定の評価レベルをサポートします。
supported_evaluation_levels内のエバリュエーターのフィールドを確認します。
-
ターン専用エバリュエーター (
fluency、relevanceなど) は、evaluation_level="conversation"では使用できません。 - 現在、すべての会話レベルエバリュエーターは、
"turn"レベルと"conversation"レベルの両方をサポートしています。
一般的なエラー
| Error | 原因 | ソリューション |
|---|---|---|
| 互換性のない評価レベル | ターン専用エバリュエーターでの evaluation_level="conversation" の使用 |
ターン専用評価子を削除するか、evaluation_level="turn" に変更する |
会話データを準備する
各行の messages フィールドに完全な会話が含まれる JSONL ファイルを作成します。 各メッセージには、 role (ユーザー、アシスタント、またはシステム) と contentが含まれている必要があります。 完全な例については、SDK の conversation 評価サンプルを参照してください。
{"messages": [{"role": "user", "content": "What's my account balance?"}, {"role": "assistant", "content": "Your current balance is $1,234.56."}, {"role": "user", "content": "Thanks!"}, {"role": "assistant", "content": "You're welcome! Is there anything else?"}]}
エージェントがツールを使用している場合は、ツール定義とツール呼び出しを含めることもできます。
{"messages": [{"role": "user", "content": "What is the capital/major city of France?"}, {"role": "assistant", "content": "Paris"}]}
{"messages": [{"role": "user", "content": "How do I reverse a string in Python?"}, {"role": "assistant", "content": "You can reverse a string in Python by using slicing: string[::-1]"}]}
{"messages": [{"role": "user", "content": "What are the main causes of climate change?"}, {"role": "assistant", "content": "The main causes of climate change are the increase in greenhouse gases in the atmosphere, primarily due to human activities such as burning fossil fuels and deforestation."}]}
{"messages": [{"role": "user", "content": "What's my account balance?"}, {"role": "assistant", "content": null, "tool_calls": [{"id": "call_abc123", "type": "function", "function": {"name": "get_account_balance", "arguments": "{\"account_id\": \"ACCT-7890\"}"}}]}, {"role": "tool", "tool_call_id": "call_abc123", "content": "{ \"balance\": 1234.56, \"currency\": \"USD\" }"}, {"role": "assistant", "content": "Your current balance is 1,234.56."}, {"role": "user", "content": "Thanks!"}, {"role": "assistant", "content": "You're welcome! Is there anything else?"}], "tool_definitions": [{"name": "get_account_balance", "description": "Retrieves the current balance for a customer account", "parameters": {"type": "object", "properties": {"account_id": {"type": "string"}}, "required": ["account_id"]}}]}
{"messages": [{"role": "user", "content": "Explain the theory of relativity in simple terms."}, {"role": "assistant", "content": "Einstein's theory of relativity shows that space and time are interconnected and relative to the observer's frame of reference."}]}
{"messages": [{"role": "user", "content": "What's the weather in Seattle?"}, {"role": "assistant", "content": null, "tool_calls": [{"id": "call_002", "type": "function", "function": {"name": "get_weather", "arguments": "{\"location\": \"Seattle, WA\"}"}}]}, {"role": "tool", "tool_call_id": "call_002", "content": "{ \"temperature\": 55, \"condition\": \"Cloudy\" }"}, {"role": "assistant", "content": "It's currently 55F and cloudy in Seattle."}], "tool_definitions": [{"name": "get_weather", "description": "Get the current weather for a location", "parameters": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}}]}
{"messages": [{"role": "user", "content": "What is the tallest mountain in the world?"}, {"role": "assistant", "content": "Mount Everest is the tallest mountain in the world."}]}
{"messages": [{"role": "user", "content": "Is 4 x 2 = 16?"}, {"role": "assistant", "content": "No, 4 x 2 = 8."}]}
{"messages": [{"role": "user", "content": "What is the best Italian desert?"}, {"role": "assistant", "content": "Tiramisu is a popular Italian dessert."}]}
{"messages": [{"role": "user", "content": "What is the chemical formula for water?"}, {"role": "assistant", "content": "The chemical formula for water is H2O."}]}
データ スキーマとエバリュエーターを定義する
会話データのスキーマ "messages" を指定し、会話レベルの評価用に設計されたエバリュエーターを選択します。 会話レベルのエバリュエーターは、個々のターンではなく、相互作用全体を評価します。
pip install "azure-ai-projects>=2.2.0"
import os
from openai.types.eval_create_params import DataSourceConfigCustom
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import TestingCriterionAzureAIEvaluator
endpoint = os.environ["AZURE_AI_PROJECT_ENDPOINT"]
model_deployment_name = os.environ["AZURE_AI_MODEL_DEPLOYMENT_NAME"]
with (
DefaultAzureCredential() as credential,
AIProjectClient(endpoint=endpoint, credential=credential) as project_client,
project_client.get_openai_client() as openai_client,
):
data_source_config = DataSourceConfigCustom(
type="custom",
item_schema={
"type": "object",
"properties": {
"messages": {"type": "array"},
"tool_definitions": {"type": "array"},
},
"required": ["messages"],
},
include_sample_schema=False,
)
testing_criteria = [
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="conversation_coherence",
evaluator_name="builtin.coherence",
initialization_parameters={"model": model_deployment_name},
data_mapping={"messages": "{{item.messages}}"},
),
TestingCriterionAzureAIEvaluator(
type="azure_ai_evaluator",
name="groundedness",
evaluator_name="builtin.groundedness",
initialization_parameters={"model": model_deployment_name},
data_mapping={"messages": "{{item.messages}}"},
),
]
評価を作成して実行する
準備: sample_data_multiturn_conversations.jsonl をダウンロードします
from openai.types.evals.create_eval_jsonl_run_data_source_param import (
CreateEvalJSONLRunDataSourceParam,
SourceFileID,
)
# Upload conversation data
data_id = project_client.datasets.upload_file(
name="multiturn-conversation-data",
version="1",
file_path="./sample_data_multiturn_conversations.jsonl",
).id
# Create the evaluation
eval_object = openai_client.evals.create(
name="Multi-turn Conversation Evaluation",
data_source_config=data_source_config,
testing_criteria=testing_criteria,
)
# Create a run with evaluation_level set to "conversation"
eval_run = openai_client.evals.runs.create(
eval_id=eval_object.id,
name="multiturn-conversation-run",
data_source=CreateEvalJSONLRunDataSourceParam(
type="jsonl",
source=SourceFileID(
type="file_id",
id=data_id,
),
),
extra_body={"evaluation_level": "conversation"},
)
完了するまでポーリングして結果を解釈するには、「クラウド評価結果を取得する」を参照してください。
実行可能な完全な例については、GitHubのsample_multiturn_conversation_evaluation.pyを参照してください。
次のステップ
- 完了するまでポーリングして結果を解釈するには、「クラウド評価結果を取得する」を参照してください。
- 運用トレースを評価するには、「 デプロイされたモデルとエージェントの会話を評価する」を参照してください。
- 合成会話を生成するには、「 エージェントの会話をシミュレートする」を参照してください。