本文說明如何從Azure Databricks存取
- 標準或工作叢集:使用 Spark ABFS 驅動程式搭配 OAuth 設定,直接透過 Spark DataFrames 讀寫資料。
-
無伺服器運算:無伺服器執行時不允許你設定自訂的 Spark 設定屬性。 相反地,請使用 Microsoft 驗證資源庫(MSAL)和 Python
deltalake函式庫來驗證並讀取或寫入 Delta 資料表。
有關相關的 Databricks 整合情境,請參閱以下資源:
| Scenario | 文件 |
|---|---|
| 從 Unity 目錄查詢 OneLake 資料,但不複製 | 啟用 OneLake 目錄聯盟 |
| 從 Fabric 存取 Databricks Unity 目錄資料 | 同步 Azure Databricks Unity 目錄 |
必要條件
在你連線之前,請確保你具備:
- Fabric 工作區和湖倉。
- 進階版 Azure Databricks 工作區。
- 一個至少具備Contributor角色指派的服務主體。
- Databricks secrets 或 Azure Key Vault(AKV)來儲存和檢索機密資料。 本文中的範例使用 Databricks 的秘密。
使用標準叢集連接至 OneLake
請使用正確的 OneLake ABFS 路徑格式
請使用以下其中一種 URI 格式:
abfss://<workspace_id_or_name>@onelake.dfs.fabric.microsoft.com/<lakehouse_id_or_name>.lakehouse/Files/<path>abfss://<workspace_id_or_name>@onelake.dfs.fabric.microsoft.com/<lakehouse_id_or_name>.lakehouse/Tables/<path>
你可以使用識別碼或名字。 如果使用名稱,請避免工作區和湖屋名稱中的特殊字元和空白。
使用服務主體驗證
將此選項用於自動化工作和集中式秘密輪替。
workspace_name = "<workspace_name>"
lakehouse_name = "<lakehouse_name>"
tenant_id = dbutils.secrets.get(scope="<scope-name>", key="<tenant-id-key>")
service_principal_id = dbutils.secrets.get(scope="<scope-name>", key="<client-id-key>")
service_principal_secret = dbutils.secrets.get(scope="<scope-name>", key="<client-secret-key>")
spark.conf.set("fs.azure.account.auth.type", "OAuth")
spark.conf.set(
"fs.azure.account.oauth.provider.type",
"org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider",
)
spark.conf.set("fs.azure.account.oauth2.client.id", service_principal_id)
spark.conf.set("fs.azure.account.oauth2.client.secret", service_principal_secret)
spark.conf.set(
"fs.azure.account.oauth2.client.endpoint",
f"https://login.microsoftonline.com/{tenant_id}/oauth2/token",
)
# Read
df = spark.read.format("parquet").load(
f"abfss://{workspace_name}@onelake.dfs.fabric.microsoft.com/{lakehouse_name}.lakehouse/Files/data"
)
df.show(10)
# Write
df.write.format("delta").mode("overwrite").save(
f"abfss://{workspace_name}@onelake.dfs.fabric.microsoft.com/{lakehouse_name}.lakehouse/Tables/dbx_delta_spn"
)
使用無伺服器運算連接 OneLake
Databricks 的無伺服器運算 讓你可以在不配置叢集的情況下執行工作負載,但它只允許支援的 Spark 屬性的子集。 你無法設定 fs.azure.* 標準叢集使用的 Spark 設定。
注意
此限制並非 Azure Databricks 獨有。 Amazon Web Services(AWS) 和 Google Cloud 上的 Databricks 無伺服器實作也有相同行為。
如果你嘗試在無伺服器筆記本中設定不支援的 Spark 設定,系統會回傳 CONFIG_NOT_AVAILABLE 錯誤。
相反地,使用 MSAL 取得 OAuth 令牌及 Python deltalake 函式庫,用該令牌讀寫 Delta 資料表。
設置一台無伺服器的筆記本
在你的 Databricks 工作空間建立一個筆記本,並連接到無伺服器運算。
匯入 Python 模組。 在此範例中,請使用兩個模組:
- msal 與Microsoft 身分識別平台進行驗證。
- deltalake 讀取與寫入帶有 Python 的 Delta Lake 資料表。
from msal import ConfidentialClientApplication from deltalake import DeltaTable, write_deltalake宣告 Microsoft Entra 租戶的變數,包括應用程式識別碼。 使用已部署 Fabric 的該租用戶 ID。
# Fetch from Databricks secrets. tenant_id = dbutils.secrets.get(scope="<replace-scope-name>",key="<replace value with key value for tenant_id>") client_id = dbutils.secrets.get(scope="<replace-scope-name>",key="<replace value with key value for client_id>") client_secret = dbutils.secrets.get(scope="<replace-scope-name>",key="<replace value with key value for secret>")宣告 Fabric 工作區變數。
workspace_id = "<replace with workspace name>" lakehouse_id = "<replace with lakehouse name>" table_to_read = "<name of lakehouse table to read>" onelake_uri = f"abfss://{workspace_id}@onelake.dfs.fabric.microsoft.com/{lakehouse_id}.lakehouse/Tables/{table_to_read}"初始化客戶端以取得令牌。
authority = f"https://login.microsoftonline.com/{tenant_id}" app = ConfidentialClientApplication( client_id, authority=authority, client_credential=client_secret ) result = app.acquire_token_for_client(scopes=["https://onelake.fabric.microsoft.com/.default"]) if "access_token" in result: print("Access token acquired.") token_val = result['access_token'] else: raise Exception(f"Failed to acquire token: {result.get('error_description', result)}")閱讀OneLake的三角洲表格。
dt = DeltaTable(onelake_uri, storage_options={"bearer_token": f"{token_val}", "use_fabric_endpoint": "true"}) df = dt.to_pandas() print(df.head())寫入 Delta 表格到 OneLake。
target_uri = f"abfss://{workspace_id}@onelake.dfs.fabric.microsoft.com/{lakehouse_id}.lakehouse/Tables/<target_table_name>" write_deltalake( target_uri, df, mode="overwrite", storage_options={"bearer_token": f"{token_val}", "use_fabric_endpoint": "true"} )
設計考量
- 盡可能在每條資料表路徑中使用一個寫入模式。 從多個運算引擎或執行時版本寫入相同的儲存路徑可能會造成衝突。
- 使用密匙管理來管理服務帳戶憑證。
- 當你需要虛擬存取時,使用 OneLake 捷徑,替代實體寫入另一個湖倉位置。
- 若要限制 OneLake 的公共網路存取,請將你的 Databricks 存取連接器的資源 ID 加入工作區的 資源實例規則 中,讓 OneLake 能在每次請求時驗證其管理身份。
相關內容
- 使用 OneLake 搭配 Azure Databricks
- 在 Azure Databricks 上掛載雲端物件儲存