整合 OneLake 與 Azure Databricks

本文說明如何從Azure Databricks存取OneLake 資料。 這兩種方式都使用服務主體認證及 OneLake ABFS 端點。 選擇與你 Databricks 計算類型相符的區塊:

  • 標準或工作叢集:使用 Spark ABFS 驅動程式搭配 OAuth 設定,直接透過 Spark DataFrames 讀寫資料。
  • 無伺服器運算:無伺服器執行時不允許你設定自訂的 Spark 設定屬性。 相反地,請使用 Microsoft 驗證資源庫(MSAL)和 Python deltalake 函式庫來驗證並讀取或寫入 Delta 資料表。

有關相關的 Databricks 整合情境,請參閱以下資源:

Scenario 文件
從 Unity 目錄查詢 OneLake 資料,但不複製 啟用 OneLake 目錄聯盟
從 Fabric 存取 Databricks Unity 目錄資料 同步 Azure Databricks Unity 目錄

必要條件

在你連線之前,請確保你具備:

  • Fabric 工作區和湖倉。
  • 進階版 Azure Databricks 工作區。
  • 一個至少具備Contributor角色指派的服務主體。
  • Databricks secrets 或 Azure Key Vault(AKV)來儲存和檢索機密資料。 本文中的範例使用 Databricks 的秘密。

使用標準叢集連接至 OneLake

請使用正確的 OneLake ABFS 路徑格式

請使用以下其中一種 URI 格式:

  • abfss://<workspace_id_or_name>@onelake.dfs.fabric.microsoft.com/<lakehouse_id_or_name>.lakehouse/Files/<path>
  • abfss://<workspace_id_or_name>@onelake.dfs.fabric.microsoft.com/<lakehouse_id_or_name>.lakehouse/Tables/<path>

你可以使用識別碼或名字。 如果使用名稱,請避免工作區和湖屋名稱中的特殊字元和空白。

使用服務主體驗證

將此選項用於自動化工作和集中式秘密輪替。

workspace_name = "<workspace_name>"
lakehouse_name = "<lakehouse_name>"
tenant_id = dbutils.secrets.get(scope="<scope-name>", key="<tenant-id-key>")
service_principal_id = dbutils.secrets.get(scope="<scope-name>", key="<client-id-key>")
service_principal_secret = dbutils.secrets.get(scope="<scope-name>", key="<client-secret-key>")

spark.conf.set("fs.azure.account.auth.type", "OAuth")
spark.conf.set(
   "fs.azure.account.oauth.provider.type",
   "org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider",
)
spark.conf.set("fs.azure.account.oauth2.client.id", service_principal_id)
spark.conf.set("fs.azure.account.oauth2.client.secret", service_principal_secret)
spark.conf.set(
   "fs.azure.account.oauth2.client.endpoint",
   f"https://login.microsoftonline.com/{tenant_id}/oauth2/token",
)

# Read
df = spark.read.format("parquet").load(
   f"abfss://{workspace_name}@onelake.dfs.fabric.microsoft.com/{lakehouse_name}.lakehouse/Files/data"
)
df.show(10)

# Write
df.write.format("delta").mode("overwrite").save(
   f"abfss://{workspace_name}@onelake.dfs.fabric.microsoft.com/{lakehouse_name}.lakehouse/Tables/dbx_delta_spn"
)

使用無伺服器運算連接 OneLake

Databricks 的無伺服器運算 讓你可以在不配置叢集的情況下執行工作負載,但它只允許支援的 Spark 屬性的子集。 你無法設定 fs.azure.* 標準叢集使用的 Spark 設定。

注意

此限制並非 Azure Databricks 獨有。 Amazon Web Services(AWS) 和 Google Cloud 上的 Databricks 無伺服器實作也有相同行為。

如果你嘗試在無伺服器筆記本中設定不支援的 Spark 設定,系統會回傳 CONFIG_NOT_AVAILABLE 錯誤。

螢幕擷取畫面,如果使用者嘗試在無伺服器計算中修改不支援的 Spark 組態,則顯示錯誤訊息。

相反地,使用 MSAL 取得 OAuth 令牌及 Python deltalake 函式庫,用該令牌讀寫 Delta 資料表。

設置一台無伺服器的筆記本

  1. 在你的 Databricks 工作空間建立一個筆記本,並連接到無伺服器運算。

    螢幕擷取畫面顯示如何將 Databricks 筆記本與無伺服器計算連線。

  2. 匯入 Python 模組。 在此範例中,請使用兩個模組:

    • msal 與Microsoft 身分識別平台進行驗證。
    • deltalake 讀取與寫入帶有 Python 的 Delta Lake 資料表。
    from msal import ConfidentialClientApplication
    from deltalake import DeltaTable, write_deltalake
    
  3. 宣告 Microsoft Entra 租戶的變數,包括應用程式識別碼。 使用已部署 Fabric 的該租用戶 ID。

    # Fetch from Databricks secrets.
    tenant_id = dbutils.secrets.get(scope="<replace-scope-name>",key="<replace value with key value for tenant_id>")
    client_id = dbutils.secrets.get(scope="<replace-scope-name>",key="<replace value with key value for client_id>")
    client_secret = dbutils.secrets.get(scope="<replace-scope-name>",key="<replace value with key value for secret>")
    
  4. 宣告 Fabric 工作區變數。

    workspace_id = "<replace with workspace name>"
    lakehouse_id = "<replace with lakehouse name>"
    table_to_read = "<name of lakehouse table to read>"
    onelake_uri = f"abfss://{workspace_id}@onelake.dfs.fabric.microsoft.com/{lakehouse_id}.lakehouse/Tables/{table_to_read}"
    
  5. 初始化客戶端以取得令牌。

    authority = f"https://login.microsoftonline.com/{tenant_id}"
    
    app = ConfidentialClientApplication(
        client_id,
        authority=authority,
        client_credential=client_secret
    )
    
    result = app.acquire_token_for_client(scopes=["https://onelake.fabric.microsoft.com/.default"])
    
    if "access_token" in result:
        print("Access token acquired.")
        token_val = result['access_token']
    else:
        raise Exception(f"Failed to acquire token: {result.get('error_description', result)}")
    
  6. 閱讀OneLake的三角洲表格。

    dt = DeltaTable(onelake_uri, storage_options={"bearer_token": f"{token_val}", "use_fabric_endpoint": "true"})
    df = dt.to_pandas()
    print(df.head())
    
  7. 寫入 Delta 表格到 OneLake。

    target_uri = f"abfss://{workspace_id}@onelake.dfs.fabric.microsoft.com/{lakehouse_id}.lakehouse/Tables/<target_table_name>"
    write_deltalake(
        target_uri,
        df,
        mode="overwrite",
        storage_options={"bearer_token": f"{token_val}", "use_fabric_endpoint": "true"}
    )
    

設計考量

  • 盡可能在每條資料表路徑中使用一個寫入模式。 從多個運算引擎或執行時版本寫入相同的儲存路徑可能會造成衝突。
  • 使用密匙管理來管理服務帳戶憑證。
  • 當你需要虛擬存取時,使用 OneLake 捷徑,替代實體寫入另一個湖倉位置。
  • 若要限制 OneLake 的公共網路存取,請將你的 Databricks 存取連接器的資源 ID 加入工作區的 資源實例規則 中,讓 OneLake 能在每次請求時驗證其管理身份。