用 Azure Machine Learning CLI、SDK 和 REST API 訓練模型

適用於:Azure CLI ml extension v2 (current)Python SDK azure-ai-ml v2 (current)

Azure Machine Learning 提供多種方式提交機器學習訓練工作。 在本文中,您將學習如何透過以下方法提交職缺:

  • Azure CLI 機器學習擴充功能:ml 擴充,亦稱為 CLI v2。
  • Python SDK v2 for Azure Machine Learning.
  • REST API:CLI 和 SDK 所建立的 API。

先決條件

若要使用 SDK:

複製範例庫

本文中的程式碼片段基於 Azure Machine Learning 範例 GitHub 儲存庫。 要將該儲存庫複製到你的開發環境,請使用以下指令:

git clone --depth 1 https://github.com/Azure/azureml-examples
cd azureml-examples

提示

使用 --depth 1 僅複製存放庫的最新認可,如此可縮短完成作業的時間。

本文剩下的指令假設你是從 azureml-examples 目錄執行。

範例工作

本文範例使用鳶尾花資料集來訓練 MLFlow 模型。

雲端訓練

當你在雲端訓練時,必須連接到你的 Azure Machine Learning 工作空間,並選擇一個運算資源來執行訓練工作。

連線到工作區

提示

請使用以下分頁選擇你想用來訓練模型的方法。 選擇分頁會自動將本文中所有分頁切換到同一個分頁。你隨時可以選擇其他分頁。

要連接工作區,你需要識別碼參數——訂閱、資源群組和工作區名稱。 使用 MLClient 命名空間中的 azure.ai.ml 中的這些詳細資訊,以獲取所需的 Azure Machine Learning 工作區的句柄。 要驗證,請使用 default Azure authentication。 欲了解更多如何設定憑證及連接工作區的資訊,請參閱此 example。

#import required libraries
from azure.ai.ml import MLClient
from azure.identity import DefaultAzureCredential

#Enter details of your Azure Machine Learning workspace
subscription_id = '<SUBSCRIPTION_ID>'
resource_group = '<RESOURCE_GROUP>'
workspace = '<AZUREML_WORKSPACE_NAME>'

#connect to the workspace
ml_client = MLClient(DefaultAzureCredential(), subscription_id, resource_group, workspace)

請列印工作區名稱以驗證連結:

print(ml_client.workspace_name)

建立一個訓練用的運算資源

註

若想嘗試 無伺服器運算,跳過此步驟,直接提交 訓練工作即可。

Azure Machine Learning 運算叢集是一個完全管理的運算資源,可以用來執行訓練工作。 在以下範例中,你會建立一個名為 cpu-cluster的運算叢集。

from azure.ai.ml.entities import AmlCompute

# specify aml compute name.
cpu_compute_target = "cpu-cluster"

try:
    ml_client.compute.get(cpu_compute_target)
except Exception:
    print("Creating a new cpu compute target...")
    compute = AmlCompute(
        name=cpu_compute_target, size="STANDARD_D2_V2", min_instances=0, max_instances=4
    )
    ml_client.compute.begin_create_or_update(compute).result()

確認計算叢集是否存在:

cpu_cluster = ml_client.compute.get("cpu-cluster")
print(f"Compute '{cpu_cluster.name}' provisioning state: {cpu_cluster.provisioning_state}")

提交訓練作業

執行此腳本時,使用一個 command,執行位於 ./sdk/python/jobs/single-step/lightgbm/iris/src/ 下的 main.py Python 腳本。 你把指令以 job 提交給 Azure Machine Learning。

註

若要使用 無伺服器運算,請在這段程式碼中刪除 compute="cpu-cluster" 。

from azure.ai.ml import command, Input

# define the command
command_job = command(
    code="./src",
    command="python main.py --iris-csv ${{inputs.iris_csv}} --learning-rate ${{inputs.learning_rate}} --boosting ${{inputs.boosting}}",
    environment="AzureML-lightgbm-3.2-ubuntu18.04-py37-cpu@latest",
    inputs={
        "iris_csv": Input(
            type="uri_file",
            path="https://azuremlexamples.blob.core.windows.net/datasets/iris.csv",
        ),
        "learning_rate": 0.9,
        "boosting": "gbdt",
    },
    compute="cpu-cluster",
)

在同一場 Python 會議中提交職缺:

# submit the command
returned_job = ml_client.jobs.create_or_update(command_job)
# get a URL for the status of the job
returned_job.studio_url

在前面的例子中,你設定了:

  • code - 執行指令程式碼所在的路徑。
  • command - 需要執行的指令。
  • environment - 執行訓練腳本所需的環境。 在此範例中,使用由Azure Machine Learning提供的精選或現成環境,稱為 AzureML-lightgbm-3.2-ubuntu18.04-py37-cpu@latest。 你也可以透過指定一個基礎 docker 映像檔,並在其上設定一個 conda yaml,來使用自訂環境。
  • inputs - 使用名稱值對作為指令的輸入字典。 鍵是工作上下文中輸入的名稱,而值則是輸入值。 使用command表達式參考${{inputs.<input_name>}}中的輸入。 若要使用檔案或資料夾作為輸入,請使用該 Input 類別。 欲了解更多資訊,請參閱 SDK 與 CLI v2 表達式。

欲了解更多資訊,請參閱 參考文件。

當你提交工作時,服務會回傳一個 Azure Machine Learning 工作室 中工作狀態的 URL。 使用工作室介面查看工作進度。 你也可以用來 returned_job.status 查詢該職缺的當前狀態。

print(f"Studio URL: {returned_job.studio_url}")

重要

Azure Machine Learning 的訓練與指令工作不支援使用自訂網域名稱標籤的 Azure 容器登錄檔(ACR)。 參考此類登錄檔的工作可能在啟動時因映像拉取或環境解析錯誤而失敗。 為了避免這個問題:

  • 針對您的 ACR,使用預設登入伺服器格式 (<registry-name>.azurecr.io)。
  • 建立登錄檔時,將 網域名稱標籤範圍 設為 不安全。

監控訓練作業

請等訓練工作完成後再註冊模型。 工作狀態會依 Starting → Preparing → Running → Completed轉換。

用 ml_client.jobs.stream() 來即時監控工作輸出:

ml_client.jobs.stream(returned_job.name)

或者,也可以用程式化方式查看職缺狀態:

returned_job = ml_client.jobs.get(returned_job.name)
print(f"Job status: {returned_job.status}")

註冊訓練好的模型

以下範例示範如何在您的 Azure Machine Learning 工作空間中註冊模型。

提示

訓練作業回傳一個 name 屬性。 將這個名稱作為通往模型路徑的一部分。

from azure.ai.ml.entities import Model
from azure.ai.ml.constants import AssetTypes

run_model = Model(
    path="azureml://jobs/{}/outputs/artifacts/paths/model/".format(returned_job.name),
    name="run-model-example",
    description="Model created from run.",
    type=AssetTypes.MLFLOW_MODEL
)

ml_client.models.create_or_update(run_model)

確認該型號有註冊:

registered_model = ml_client.models.get("run-model-example", version="1")
print(f"Model '{registered_model.name}' version {registered_model.version} registered successfully.")

清理資源

如果你不打算用運算叢集做更多訓練工作,刪除它以避免產生費用。 只要叢集存在,就會持續計費,即使沒有任何節點在執行。

ml_client.compute.begin_delete("cpu-cluster").wait()

解決常見錯誤

錯誤 原因 Resolution
ImportError: No module named 'azure.identity' 遺失包裹azure-identity pip install azure-identity執行
DefaultAzureCredential failed 未登入 Azure 先執行 az login ,或設定環境變數以進行 服務主體認證
ComputeNotFound 叢集名稱不符或叢集刪除 確認叢集名稱並檢查配置狀態
EnvironmentNotFound 已淘汰或無法使用的精選環境 列出可用 ml_client.environments.list() 環境,並使用目前版本
QuotaExceeded 虛擬機容量的 vCPU 配額不足 請求增加配額 或使用較小的虛擬機容量

關於環境特定問題,請參見 「排除環境映像建置問題」。

下一步

既然你已經有訓練好模型,請學習 如何用線上端點部署它。

更多範例請參見 Azure Machine Learning 範例 GitHub 資料庫。

欲了解更多關於本文中使用的 Azure CLI 指令、Python SDK 類別或 REST API 的資訊,請參閱以下參考文件: