Nota
L'accesso a questa pagina richiede l'autorizzazione. È possibile provare ad accedere o modificare le directory.
L'accesso a questa pagina richiede l'autorizzazione. È possibile provare a modificare le directory.
Importante
Il supporto di Terraform per Lakebase è in fase Beta.
Questa pagina mostra una configurazione Terraform completa per un progetto di scalabilità automatica di Lakebase pronto per la produzione con le funzionalità più usate:
- Ramo di produzione protetto
- Endpoint di lettura/scrittura a disponibilità elevata con repliche secondarie leggibili
- Entità servizio con
DATABRICKS_SUPERUSERprivilegi di database - Database Postgres di proprietà dell'app
- Database Postgres registrato in Unity Catalog per query di federazione Lakehouse da Databricks SQL e notebook
- Streaming continuo di tabelle sincronizzate da Unity Catalog
- App Databricks connessa al progetto Lakebase
Per un'introduzione dettagliata a Terraform con Lakebase, vedere Introduzione a Terraform per Lakebase.
Prerequisiti
Prima di iniziare, devi avere il seguente:
- Terraform installato (versione 1.0 o versione successiva). Vedere Installare Terraform.
- Entità servizio configurata per l'autenticazione da computer a computer OAuth (M2M) con
CAN_MANAGEautorizzazione per il progetto Lakebase. Consulta Autorizzare l'accesso dell'entità servizio ad Azure Databricks con OAuth e Gestisci le autorizzazioni del progetto. - Tabella Delta del catalogo Unity con CDF (Change Data Feed) abilitata per l'uso come origine di sincronizzazione.
Configurazione completa
Quando si crea un progetto, Azure Databricks crea automaticamente un production ramo, un primary endpoint di lettura/scrittura, un ruolo Postgres proprietario associato all'identità e un databricks_postgres database. Per configurare queste risorse create in modo implicito, dichiararle in Terraform con replace_existing = true. Per altri dettagli, vedere databricks_postgres_branch, databricks_postgres_endpoint, databricks_postgres_rolee databricks_postgres_database.
Avvertimento
Questa configurazione imposta is_protected = true sul ramo production e include una variabile unprotect_for_destroy integrata nella specifica del ramo. Terraform non può eliminare un progetto contenente rami protetti e il ramo non può essere eliminato direttamente perché il production relativo ciclo di vita è controllato dal progetto. Per eliminare le risorse in modo pulito, usare un'eliminazione definitiva in due passaggi:
# Step 1: unprotect the branch
terraform apply -var="unprotect_for_destroy=true"
# Step 2: destroy all resources
terraform destroy -var="unprotect_for_destroy=true"
Dopo aver eseguito terraform destroy, il progetto viene eliminato in modo non definitivo e conservato per 7 giorni prima dell'eliminazione permanente. Per eliminarla immediatamente in modo permanente, impostare purge_on_delete = true sulla risorsa databricks_postgres_project prima di eseguire destroy.
variable "admin_sp_app_id" {
description = "Application ID of the service principal to grant admin access"
type = string
}
variable "unprotect_for_destroy" {
description = "Set to true before destroy to unprotect the production branch"
type = bool
default = false
}
# Project — top-level container for branches, endpoints, databases, and roles.
resource "databricks_postgres_project" "this" {
project_id = "my-lakebase-project"
# purge_on_delete = true # Uncomment to permanently delete on destroy (default: soft delete, 7-day retention).
spec = {
pg_version = 17
display_name = "My Lakebase Project"
default_endpoint_settings = {
autoscaling_limit_min_cu = 0.5
autoscaling_limit_max_cu = 4.0
suspend_timeout_duration = "300s"
}
}
}
# Configure the implicitly created production branch as protected.
resource "databricks_postgres_branch" "production" {
branch_id = "production"
parent = databricks_postgres_project.this.name
spec = {
no_expiry = true
is_protected = var.unprotect_for_destroy ? false : true
}
replace_existing = true
}
# Configure the implicitly created primary endpoint with HA.
# HA requires no_suspension = true. group.min = 2 adds a standby for automatic failover.
resource "databricks_postgres_endpoint" "primary" {
endpoint_id = "primary"
parent = databricks_postgres_branch.production.name
spec = {
endpoint_type = "ENDPOINT_TYPE_READ_WRITE"
autoscaling_limit_min_cu = 0.5
autoscaling_limit_max_cu = 4.0
no_suspension = true
group = {
min = 2
max = 2
enable_readable_secondaries = true
}
}
replace_existing = true
}
# Grant workspace-level CAN_MANAGE on the project to the service principal.
# Use status.project_id (bare ID) not .name (full resource path) — the permissions
# API rejects the full path with a "resource type not found" error.
resource "databricks_permissions" "project" {
database_project_name = databricks_postgres_project.this.status.project_id
access_control {
service_principal_name = var.admin_sp_app_id
permission_level = "CAN_MANAGE"
}
}
# Create a Postgres role backed by the service principal with full database privileges.
# depends_on serializes creation — Lakebase processes one branch operation at a time.
resource "databricks_postgres_role" "admin_sp" {
role_id = "admin-sp"
parent = databricks_postgres_branch.production.name
spec = {
identity_type = "SERVICE_PRINCIPAL"
postgres_role = var.admin_sp_app_id
auth_method = "LAKEBASE_OAUTH_V1"
membership_roles = ["DATABRICKS_SUPERUSER"]
attributes = {
createdb = true
createrole = true
bypassrls = true
}
}
depends_on = [databricks_postgres_endpoint.primary]
}
# Create a Postgres database owned by the admin SP role.
resource "databricks_postgres_database" "app" {
database_id = "app"
parent = databricks_postgres_branch.production.name
spec = {
postgres_database = "app"
role = databricks_postgres_role.admin_sp.name
}
}
# Register the Postgres database in Unity Catalog. This makes the database queryable
# from Databricks SQL and notebooks through Lakehouse Federation, and serves as the
# parent namespace for synced tables that live inside the Lakebase Catalog.
# create_database_if_missing is set explicitly because the database is managed by
# the databricks_postgres_database resource above.
resource "databricks_postgres_catalog" "app_catalog" {
catalog_id = "app_catalog"
spec = {
postgres_database = databricks_postgres_database.app.status.postgres_database
branch = databricks_postgres_branch.production.name
create_database_if_missing = false
}
}
# Sync a Unity Catalog Delta table into the Lakebase database continuously.
# Prefixing synced_table_id with the Lakebase Catalog name places the synced table
# inside the catalog so it's discoverable alongside the rest of the catalog's contents.
# postgres_database references the catalog's status, which implicitly orders this
# resource after the catalog without an explicit depends_on.
resource "databricks_postgres_synced_table" "orders" {
synced_table_id = "app_catalog.default.orders_synced"
spec = {
branch = databricks_postgres_branch.production.name
postgres_database = databricks_postgres_catalog.app_catalog.status.postgres_database
source_table_full_name = "my_catalog.default.orders"
primary_key_columns = ["order_id"]
scheduling_policy = "CONTINUOUS"
create_database_objects_if_missing = true
new_pipeline_spec = {
storage_catalog = "my_catalog"
storage_schema = "default"
}
}
}
# Databricks App connected to the Lakebase project.
# database must be the full resource name (databricks_postgres_database.app.name),
# not the Postgres database name. permission must be "CAN_CONNECT_AND_CREATE".
resource "databricks_app" "this" {
name = "my-lakebase-app"
description = "App backed by Lakebase autoscaling project"
depends_on = [databricks_postgres_database.app]
resources = [{
name = "lakebase-db"
postgres = {
branch = databricks_postgres_branch.production.name
database = databricks_postgres_database.app.name
permission = "CAN_CONNECT_AND_CREATE"
}
}]
}
Annotazioni
Questa configurazione crea un ruolo separato () e un database (admin_spapp) anziché gestire il ruolo proprietario implicito e databricks_postgres il database. Per portare invece tali risorse implicite sotto la gestione di Terraform, dichiaratele con replace_existing = true utilizzando i rispettivi ID esistenti. L'ID del database è sempre databricks-postgres. L'ID del ruolo viene ricavato dall'identità che ha creato il ramo: la parte dell'indirizzo email prima di @ (in minuscolo, con i caratteri non alfanumerici sostituiti da trattini) per un utente, oppure sp-<application-id> per un'entità servizio. Se non sei sicuro del valore esatto, leggilo dall'app Lakebase o dall'API Postgres invece di derivarlo a mano.
spec.membership_roles sovrascrive le appartenenze del ruolo a ogni applicazione anziché unirle. Mantieni DATABRICKS_SUPERUSER nell'elenco; se lo si omette, vengono rimosse tutte le appartenenze associate al ruolo.
resource "databricks_postgres_role" "owner" {
role_id = "jane-doe" # normalized login of the creating identity
parent = databricks_postgres_branch.production.name
spec = {
postgres_role = "jane.doe@databricks.com" # the raw login
membership_roles = ["DATABRICKS_SUPERUSER"]
attributes = {
createdb = true
createrole = true
bypassrls = true
}
}
replace_existing = true
}
resource "databricks_postgres_database" "databricks_postgres" {
database_id = "databricks-postgres"
parent = databricks_postgres_branch.production.name
spec = {
postgres_database = "databricks_postgres"
# spec.role is omitted, so the database keeps its existing owner.
}
replace_existing = true
}
Risorse aggiuntive
- L'alta disponibilità tratta i pattern di alta disponibilità e quando usarli in produzione.
- Le tabelle di sincronizzazione illustrano le opzioni di pianificazione e la gestione delle pipeline.
- Gestire le autorizzazioni del progetto copre i controlli di accesso a livello di area di lavoro e a livello di database.
- Databricks Apps con Lakebase illustra come connettere le app ai progetti di scalabilità automatica.
- Il Registro Terraform fornisce il riferimento completo alle risorse.