Microsoft Foundry SDK を使用して会話データセットを評価する (プレビュー)

ターンレベルまたは会話レベルでデータセットからの完全な会話を評価します。

前提条件

この例では、「SDK クライアントのセットアップ」で構成された SDK クライアントを使用します

会話データセットを評価する

完全な会話を評価して、個々の応答だけでなく、ユーザーの操作全体にわたってエージェントの品質を評価します。 会話レベルの評価を用いて、タスクの不完全な解決、ユーザーのフラストレーション、ターンレベルの評価では見逃されるツール呼び出しの回帰といった品質上の問題を特定します。

たとえば、ユーザーが複数のターンで不満を感じるサポート エージェントを考えてみましょう。

ターン 1 — ユーザー: "パスワードをリセットする必要があります。"エージェント: "アカウントが見つかりました。 リセット リンクを送信します。

ターン 2 — ユーザー: "私は電子メールを受け取りませんでした。エージェント: "リンクを再送信しました。 スパムを確認してください。

ターン 3 - ユーザー: "まだ何もありません。 直接リセットできますか?エージェント: "別のリセット リンクを送信しました。"

ターンレベルのエバリュエーターは、最後の応答 (丁寧でアクションを実行する) のみをスコア付けするので、スコアが良くなります。 会話レベルのエバリュエーターは、エージェントが別の方法を試さずに同じ失敗したアクションを 3 回繰り返し、ユーザーの問題を未解決のままにした、会話フラグ全体で 顧客満足度 を評価します。

会話レベルの評価は、ターンレベルの評価といくつかの点で異なります。

特徴 ターン レベル 会話レベル
スコープ 個々のクエリ応答ペア 複数の交換で会話を完了する
Metrics 応答ごとの品質と安全性 会話レベルの結果とユーザーの満足度
データ形式 queryフィールドとresponse フィールドを含む JSONL 完全な会話を含む messages 配列を持つ JSONL
使用例 個々のモデル応答のテスト エンド ツー エンドのエージェント エクスペリエンスのテスト

データ ソースに一致する会話ワークフローを選択します。

Workflow いつ使用するか データ ソースの種類
データセットまたはインラインから ローカルの会話トレースまたはテストデータがある場合 jsonlfile_id または file_content
デプロイされた会話 Application Insights から特定の会話またはサンプリングされた運用トラフィックを評価する場合 azure_ai_trace_data_source_previewtrace_source
シミュレートされた会話 合成テスト用の会話を生成したい azure_ai_target_completionsconversation_gen_preview

評価レベルを選択する

実行の evaluation_level パラメーターは、エバリュエーターが個々のターンをスコア付けするか、会話を完了するかを決定します。

価値 Behavior
"turn" エバリュエーターは、各ターンを個別にスコア付けします。
"conversation" エバリュエーターは、会話全体をスコア付けします。
(省略) 既定値は "turn" です。

Important

エバリュエーターの互換性: 各エバリュエーターは、特定の評価レベルをサポートします。 supported_evaluation_levels内のエバリュエーターのフィールドを確認します。

  • ターン専用エバリュエーター ( fluencyrelevanceなど) は、 evaluation_level="conversation"では使用できません。
  • 現在、すべての会話レベルエバリュエーターは、 "turn" レベルと "conversation" レベルの両方をサポートしています。

一般的なエラー

Error 原因 ソリューション
互換性のない評価レベル ターン専用エバリュエーターでの evaluation_level="conversation" の使用 ターン専用評価子を削除するか、evaluation_level="turn" に変更する

会話データを準備する

各行の messages フィールドに完全な会話が含まれる JSONL ファイルを作成します。 各メッセージには、 role (ユーザー、アシスタント、またはシステム) と contentが含まれている必要があります。 完全な例については、SDK の conversation 評価サンプルを参照してください。

 {"messages": [{"role": "user", "content": "What's my account balance?"}, {"role": "assistant", "content": "Your current balance is $1,234.56."}, {"role": "user", "content": "Thanks!"}, {"role": "assistant", "content": "You're welcome! Is there anything else?"}]}

エージェントがツールを使用している場合は、ツール定義とツール呼び出しを含めることもできます。

{"messages": [{"role": "user", "content": "What is the capital/major city of France?"}, {"role": "assistant", "content": "Paris"}]}
{"messages": [{"role": "user", "content": "How do I reverse a string in Python?"}, {"role": "assistant", "content": "You can reverse a string in Python by using slicing: string[::-1]"}]}
{"messages": [{"role": "user", "content": "What are the main causes of climate change?"}, {"role": "assistant", "content": "The main causes of climate change are the increase in greenhouse gases in the atmosphere, primarily due to human activities such as burning fossil fuels and deforestation."}]}
{"messages": [{"role": "user", "content": "What's my account balance?"}, {"role": "assistant", "content": null, "tool_calls": [{"id": "call_abc123", "type": "function", "function": {"name": "get_account_balance", "arguments": "{\"account_id\": \"ACCT-7890\"}"}}]}, {"role": "tool", "tool_call_id": "call_abc123", "content": "{ \"balance\": 1234.56, \"currency\": \"USD\" }"}, {"role": "assistant", "content": "Your current balance is 1,234.56."}, {"role": "user", "content": "Thanks!"}, {"role": "assistant", "content": "You're welcome! Is there anything else?"}], "tool_definitions": [{"name": "get_account_balance", "description": "Retrieves the current balance for a customer account", "parameters": {"type": "object", "properties": {"account_id": {"type": "string"}}, "required": ["account_id"]}}]}
{"messages": [{"role": "user", "content": "Explain the theory of relativity in simple terms."}, {"role": "assistant", "content": "Einstein's theory of relativity shows that space and time are interconnected and relative to the observer's frame of reference."}]}
{"messages": [{"role": "user", "content": "What's the weather in Seattle?"}, {"role": "assistant", "content": null, "tool_calls": [{"id": "call_002", "type": "function", "function": {"name": "get_weather", "arguments": "{\"location\": \"Seattle, WA\"}"}}]}, {"role": "tool", "tool_call_id": "call_002", "content": "{ \"temperature\": 55, \"condition\": \"Cloudy\" }"}, {"role": "assistant", "content": "It's currently 55F and cloudy in Seattle."}], "tool_definitions": [{"name": "get_weather", "description": "Get the current weather for a location", "parameters": {"type": "object", "properties": {"location": {"type": "string"}}, "required": ["location"]}}]}
{"messages": [{"role": "user", "content": "What is the tallest mountain in the world?"}, {"role": "assistant", "content": "Mount Everest is the tallest mountain in the world."}]}
{"messages": [{"role": "user", "content": "Is 4 x 2 = 16?"}, {"role": "assistant", "content": "No, 4 x 2 = 8."}]}
{"messages": [{"role": "user", "content": "What is the best Italian desert?"}, {"role": "assistant", "content": "Tiramisu is a popular Italian dessert."}]}
{"messages": [{"role": "user", "content": "What is the chemical formula for water?"}, {"role": "assistant", "content": "The chemical formula for water is H2O."}]}

データ スキーマとエバリュエーターを定義する

会話データのスキーマ "messages" を指定し、会話レベルの評価用に設計されたエバリュエーターを選択します。 会話レベルのエバリュエーターは、個々のターンではなく、相互作用全体を評価します。

pip install "azure-ai-projects>=2.2.0"
import os
from openai.types.eval_create_params import DataSourceConfigCustom
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import TestingCriterionAzureAIEvaluator

endpoint = os.environ["AZURE_AI_PROJECT_ENDPOINT"]
model_deployment_name = os.environ["AZURE_AI_MODEL_DEPLOYMENT_NAME"]

with (
    DefaultAzureCredential() as credential,
    AIProjectClient(endpoint=endpoint, credential=credential) as project_client,
    project_client.get_openai_client() as openai_client,
):
    data_source_config = DataSourceConfigCustom(
        type="custom",
        item_schema={
            "type": "object",
            "properties": {
                "messages": {"type": "array"},
                "tool_definitions": {"type": "array"},
            },
            "required": ["messages"],
        },
        include_sample_schema=False,
    )

    testing_criteria = [
        TestingCriterionAzureAIEvaluator(
            type="azure_ai_evaluator",
            name="conversation_coherence",
            evaluator_name="builtin.coherence",
            initialization_parameters={"model": model_deployment_name},
            data_mapping={"messages": "{{item.messages}}"},
        ),
        TestingCriterionAzureAIEvaluator(
            type="azure_ai_evaluator",
            name="groundedness",
            evaluator_name="builtin.groundedness",
            initialization_parameters={"model": model_deployment_name},
            data_mapping={"messages": "{{item.messages}}"},
        ),
    ]

評価を作成して実行する

準備: sample_data_multiturn_conversations.jsonl をダウンロードします

from openai.types.evals.create_eval_jsonl_run_data_source_param import (
    CreateEvalJSONLRunDataSourceParam,
    SourceFileID,
)

# Upload conversation data
data_id = project_client.datasets.upload_file(
    name="multiturn-conversation-data",
    version="1",
    file_path="./sample_data_multiturn_conversations.jsonl",
).id

# Create the evaluation
eval_object = openai_client.evals.create(
    name="Multi-turn Conversation Evaluation",
    data_source_config=data_source_config,
    testing_criteria=testing_criteria,
)

# Create a run with evaluation_level set to "conversation"
eval_run = openai_client.evals.runs.create(
    eval_id=eval_object.id,
    name="multiturn-conversation-run",
    data_source=CreateEvalJSONLRunDataSourceParam(
        type="jsonl",
        source=SourceFileID(
            type="file_id",
            id=data_id,
        ),
    ),
    extra_body={"evaluation_level": "conversation"},
)

完了するまでポーリングして結果を解釈するには、「クラウド評価結果を取得する」を参照してください。

実行可能な完全な例については、GitHubのsample_multiturn_conversation_evaluation.pyを参照してください。

次のステップ