Important
Lakebase 的宣告式自動化套件支援目前仍處於 測試階段。
本頁展示了完整的宣告式自動化套件,適用於生產準備的 Lakebase 專案,包含最常用的功能:
- 受保護的生產分支
- 具備可讀次級節點的高可用性(HA)讀寫端點
- 服務主體的內嵌工作區層級
CAN_MANAGE權限 - 來自 Unity Catalog 的連續已同步資料表串流
- Lakebase 資料庫的 Unity Catalog 繫結
- 連接 Lakebase 專案的 Databricks 應用程式
關於使用 Lakebase 的宣告式自動化套件的逐步介紹,請參見 「管理 Lakebase with Declarative Automation Bundles」。
先決條件
在開始之前,您需要:
- Databricks CLI v1.0.0 或更新版本。 若要檢查您的版本,請執行
databricks --version。 要安裝或升級,請參見 「安裝或更新 Databricks CLI」。 - 一個啟用 Lakebase 的 Azure Databricks 工作空間。
- 一個為 OAuth 機器對機器(M2M)認證設定的服務主體。 該套件賦予此主要工作區
CAN_MANAGE專案權限。 請參閱 使用 OAuth 授權 Service Principal 存取 Azure Databricks 及 管理專案權限。 - 啟用變更資料饋送(CDF)的 Unity Catalog Delta 資料表,可作為同步來源。 如果不需要資料同步,請移除
postgres_synced_tablesandpostgres_catalogs區塊。
完整套件組態
該套件使用所有工作區特定值的變數。 將它們設定在 .databricks/bundle/<target>/variables.json 檔案中,或在部署時透過 --var 傳入。
當你建立專案時,Azure Databricks 會自動建立一個production分支、一個primary讀寫端點、一個與你身份綁定的擁有者 Postgres 角色,以及databricks_postgres一個資料庫。 要配置這些隱含建立的資源,請將 宣告為 replace_existing: true。
bundle:
name: lakebase-typical-project
variables:
project_id:
description: 'Lakebase project ID (lowercase, hyphen-delimited)'
default: 'my-lakebase-project'
display_name:
description: 'Human-readable project name shown in the UI'
default: 'My Lakebase project'
pg_version:
description: 'Postgres major version'
default: 17
min_cu:
description: 'Minimum compute units on the default endpoint'
default: 0.5
max_cu:
description: 'Maximum compute units on the default endpoint'
default: 4.0
suspend_timeout:
description: 'Idle time before the default endpoint suspends. Ignored when no_suspension is true.'
default: '300s'
admin_sp_app_id:
description: 'Application ID of the service principal to grant CAN_MANAGE on the project'
default: '<your-sp-application-id>'
source_table:
description: 'Unity Catalog three-part name of the Delta table to sync (catalog.schema.table)'
default: '<catalog>.<schema>.<table>'
primary_key_column:
description: 'Primary key column of the source Delta table'
default: '<pk>'
storage_catalog:
description: 'Unity Catalog catalog where the sync pipeline stores its metadata'
default: '<catalog>'
storage_schema:
description: 'Unity Catalog schema where the sync pipeline stores its metadata'
default: '<schema>'
app_name:
description: 'Databricks App name (must be unique in the workspace)'
default: 'my-lakebase-app'
uc_catalog_id:
description: 'Name to register the Lakebase database in Unity Catalog'
default: 'my_lakebase_uc_catalog'
database_name:
description: 'Postgres-internal name for the app database'
default: 'app_database'
targets:
prod:
default: true
workspace:
host: https://<your-workspace>.cloud.databricks.com
resources:
# Project — top-level container for branches, endpoints, and databases.
# The permissions block grants workspace-level CAN_MANAGE to the service principal.
postgres_projects:
lakebase_project:
project_id: ${var.project_id}
# purge_on_delete: true # Uncomment to permanently delete on destroy (default: soft delete, 7-day retention).
pg_version: ${var.pg_version}
display_name: ${var.display_name}
default_endpoint_settings:
autoscaling_limit_min_cu: ${var.min_cu}
autoscaling_limit_max_cu: ${var.max_cu}
suspend_timeout_duration: ${var.suspend_timeout}
permissions:
- service_principal_name: ${var.admin_sp_app_id}
level: CAN_MANAGE
# Configure the implicitly created production branch as protected.
postgres_branches:
production:
branch_id: production
parent: ${resources.postgres_projects.lakebase_project.name}
no_expiry: true
is_protected: true
replace_existing: true
# Configure the implicitly created primary endpoint with HA.
# HA requires no_suspension: true. group.min: 2 adds a standby for automatic failover.
postgres_endpoints:
primary:
endpoint_id: primary
parent: ${resources.postgres_branches.production.name}
endpoint_type: ENDPOINT_TYPE_READ_WRITE
autoscaling_limit_min_cu: ${var.min_cu}
autoscaling_limit_max_cu: ${var.max_cu}
no_suspension: true
group:
min: 2
max: 2
enable_readable_secondaries: true
replace_existing: true
# Postgres role that owns the app database.
postgres_roles:
app_role:
role_id: app-role # Resource ID: lowercase letters, digits, and hyphens.
parent: ${resources.postgres_branches.production.name}
postgres_role: app_role # Postgres identifier: lowercase letters, digits, and underscores.
# Named Postgres database for the app.
postgres_databases:
app_db:
database_id: app-database
parent: ${resources.postgres_branches.production.name}
postgres_database: ${var.database_name}
role: ${resources.postgres_roles.app_role.id}
# Sync a Unity Catalog Delta table into the project continuously.
postgres_synced_tables:
orders_sync:
synced_table_id: '${var.storage_catalog}.${var.storage_schema}.orders_synced'
branch: ${resources.postgres_branches.production.name}
postgres_database: ${var.database_name}
source_table_full_name: ${var.source_table}
primary_key_columns:
- ${var.primary_key_column}
scheduling_policy: CONTINUOUS
create_database_objects_if_missing: true
new_pipeline_spec:
storage_catalog: ${var.storage_catalog}
storage_schema: ${var.storage_schema}
# Bind the Lakebase database into Unity Catalog so it is queryable as UC data.
postgres_catalogs:
lakebase_uc_catalog:
catalog_id: ${var.uc_catalog_id}
postgres_database: ${var.database_name}
branch: ${resources.postgres_branches.production.name}
create_database_if_missing: true
# Databricks App connected to the project.
# Update source_code_path to point to your app source directory.
apps:
lakebase_app:
name: ${var.app_name}
description: 'App backed by Lakebase autoscaling'
source_code_path: ./app_src
config:
command:
- flask
- run
- --host=0.0.0.0
- --port=8000
resources:
- name: lakebase-db
postgres:
branch: ${resources.postgres_branches.production.name}
database: ${resources.postgres_databases.app_db.name}
permission: CAN_CONNECT_AND_CREATE
注意事項
每個 Lakebase 專案都會自動建立一個 databricks_postgres 資料庫,其擁有者是與你的身分連結的 Postgres 角色。 此套件會建立一個獨立命名的資料庫(),${var.database_name}由專用應用程式角色擁有,用以隔離應用程式資料。 若要直接使用隱含資料庫與角色,移除 postgres_roles 和 資源區塊,直接設定 postgres_databases 和 postgres_database: databricks_postgrespostgres_synced_tables,並將應用程式資源更新為 postgres_catalogsdatabase: ${resources.postgres_branches.production.name}/databases/databricks-postgres 。
若要改為將隱含的擁有者角色和 databricks_postgres 資料庫納入 bundle 管理,請使用它們現有的 ID 搭配 replace_existing: true 進行宣告。 資料庫的 ID 總是 databricks-postgres。 角色 ID 是根據你的 Databricks 身份衍生,而非固定名稱,所以先查查看:
databricks postgres list-roles projects/<project-id>/branches/production
然後宣告這兩個資源,並使其與該角色上已設定的每個欄位一致。 省略 membership_roles 會在角色被採用時移除 DATABRICKS_SUPERUSER 成員資格,請明確聲明:
postgres_roles:
owner:
role_id: <role-id-from-list-roles>
parent: ${resources.postgres_branches.production.name}
postgres_role: user@databricks.com # Or the service principal application ID.
identity_type: USER # Or SERVICE_PRINCIPAL.
membership_roles:
- DATABRICKS_SUPERUSER
replace_existing: true
postgres_databases:
databricks_postgres:
database_id: databricks-postgres
parent: ${resources.postgres_branches.production.name}
postgres_database: databricks_postgres
role: ${resources.postgres_roles.owner.id}
replace_existing: true
注意事項
要拆除這個組合所產生的資源,執行 databricks bundle destroy -t prod。 預設情況下,專案會被軟刪除並保留7天,然後才永久刪除,因此你可以在保留期間恢復。 若要立即僅刪除該專案,請搭配 --purge 使用 Databricks CLI,或取消註解上方專案資源中的 purge_on_delete: true,以便在每次執行 destroy 時將其永久刪除:
databricks postgres delete-project projects/<project-id> --purge
套用此套件
驗證與部署:
databricks bundle validate -t prod
databricks bundle deploy -t prod
如果 databricks bundle deploy 第一次跑沒完成,就重跑一次。
部署內容
該組合會產生以下資源:
- 一個 Lakebase 專案,裡面有你指定的計算預設值。
- 受保護的
production分支。 - 一個主要的讀寫端點,具備 HA 和可讀的次要節點。
- 一個持續同步管線,將 Unity 目錄 Delta 資料表串流到專案資料庫。
- 以 Lakebase 資料庫為後端的 Unity Catalog 目錄,可作為 Unity Catalog 資料進行查詢。
- 一個連接到專案資料庫的 Databricks 應用程式。
- 您所指定之服務主體的工作區
CAN_MANAGE權限。
其他資源
- 高可用性 涵蓋了 HA 模式以及何時在生產環境中使用它們。
- 提供 Lakehouse 資料並同步資料表, 涵蓋排程選項及管線管理。
- 管理專案權限 涵蓋工作區層級及資料庫層級的存取控制。
- 「將自訂 Databricks 應用程式連接到 Lakebase 」展示了如何將 Databricks 應用程式與自動擴展專案連結。
- 宣告式自動化套件 資源提供完整的宣告式自動化套件資源參考。