Configuración típica del proyecto de Lakebase con Terraform

Importante

El soporte de Terraform para Lakebase está en Beta.

Esta página muestra una configuración completa de Terraform para un proyecto Lakebase listo para producción con las características más usadas:

  • Rama de producción protegida
  • Punto de conexión de lectura y escritura de alta disponibilidad (HA) con secundarias legibles
  • Entidad de servicio con privilegios de base de datos DATABRICKS_SUPERUSER
  • Base de datos Postgres de la aplicación
  • Base de datos Postgres registrada en Unity Catalog para consultas de Lakehouse Federation desde Databricks SQL y cuadernos
  • Transmisión en continuo de tabla sincronizada desde Unity Catalog
  • Aplicación de Databricks conectada al proyecto lakebase

Para obtener una introducción paso a paso a Terraform con Lakebase, consulte Introducción a Terraform for Lakebase.

Prerequisites

Antes de empezar, debe disponer de lo siguiente:

Configuración completa

Al crear un proyecto, Azure Databricks crea automáticamente una production rama, un primary punto de conexión de lectura y escritura, un rol de Postgres propietario vinculado a la identidad y una databricks_postgres base de datos. Para configurar estos recursos creados implícitamente, declarelos en Terraform con replace_existing = true. Para obtener más información, vea databricks_postgres_branch, databricks_postgres_endpoint, databricks_postgres_roley databricks_postgres_database.

Advertencia

Esta configuración establece is_protected = true en la production rama e incluye una unprotect_for_destroy variable conectada a la especificación de rama. Terraform no puede eliminar un proyecto que contenga ramas protegidas y la production rama no se puede eliminar directamente porque el proyecto controla su ciclo de vida. Para destruir los recursos de forma limpia, utilice un proceso de destrucción en dos pasos:

# Step 1: unprotect the branch
terraform apply -var="unprotect_for_destroy=true"

# Step 2: destroy all resources
terraform destroy -var="unprotect_for_destroy=true"

Después de ejecutar terraform destroy, el proyecto se elimina temporalmente y se conserva durante 7 días antes de la eliminación permanente. Para eliminarlo de inmediato y de forma permanente, establezca purge_on_delete = true en el recurso databricks_postgres_project antes de ejecutar destroy.

variable "admin_sp_app_id" {
  description = "Application ID of the service principal to grant admin access"
  type        = string
}

variable "unprotect_for_destroy" {
  description = "Set to true before destroy to unprotect the production branch"
  type        = bool
  default     = false
}

# Project — top-level container for branches, endpoints, databases, and roles.
resource "databricks_postgres_project" "this" {
  project_id = "my-lakebase-project"
  # purge_on_delete = true  # Uncomment to permanently delete on destroy (default: soft delete, 7-day retention).
  spec = {
    pg_version   = 17
    display_name = "My Lakebase Project"
    default_endpoint_settings = {
      autoscaling_limit_min_cu = 0.5
      autoscaling_limit_max_cu = 4.0
      suspend_timeout_duration = "300s"
    }
  }
}

# Configure the implicitly created production branch as protected.
resource "databricks_postgres_branch" "production" {
  branch_id = "production"
  parent    = databricks_postgres_project.this.name
  spec = {
    no_expiry    = true
    is_protected = var.unprotect_for_destroy ? false : true
  }
  replace_existing = true
}

# Configure the implicitly created primary endpoint with HA.
# HA requires no_suspension = true. group.min = 2 adds a standby for automatic failover.
resource "databricks_postgres_endpoint" "primary" {
  endpoint_id = "primary"
  parent      = databricks_postgres_branch.production.name
  spec = {
    endpoint_type            = "ENDPOINT_TYPE_READ_WRITE"
    autoscaling_limit_min_cu = 0.5
    autoscaling_limit_max_cu = 4.0
    no_suspension            = true
    group = {
      min                         = 2
      max                         = 2
      enable_readable_secondaries = true
    }
  }
  replace_existing = true
}

# Grant workspace-level CAN_MANAGE on the project to the service principal.
# Use status.project_id (bare ID) not .name (full resource path) — the permissions
# API rejects the full path with a "resource type not found" error.
resource "databricks_permissions" "project" {
  database_project_name = databricks_postgres_project.this.status.project_id
  access_control {
    service_principal_name = var.admin_sp_app_id
    permission_level       = "CAN_MANAGE"
  }
}

# Create a Postgres role backed by the service principal with full database privileges.
# depends_on serializes creation — Lakebase processes one branch operation at a time.
resource "databricks_postgres_role" "admin_sp" {
  role_id = "admin-sp"
  parent  = databricks_postgres_branch.production.name
  spec = {
    identity_type    = "SERVICE_PRINCIPAL"
    postgres_role    = var.admin_sp_app_id
    auth_method      = "LAKEBASE_OAUTH_V1"
    membership_roles = ["DATABRICKS_SUPERUSER"]
    attributes = {
      createdb   = true
      createrole = true
      bypassrls  = true
    }
  }
  depends_on = [databricks_postgres_endpoint.primary]
}

# Create a Postgres database owned by the admin SP role.
resource "databricks_postgres_database" "app" {
  database_id = "app"
  parent      = databricks_postgres_branch.production.name
  spec = {
    postgres_database = "app"
    role              = databricks_postgres_role.admin_sp.name
  }
}

# Register the Postgres database in Unity Catalog. This makes the database queryable
# from Databricks SQL and notebooks through Lakehouse Federation, and serves as the
# parent namespace for synced tables that live inside the Lakebase Catalog.
# create_database_if_missing is set explicitly because the database is managed by
# the databricks_postgres_database resource above.
resource "databricks_postgres_catalog" "app_catalog" {
  catalog_id = "app_catalog"
  spec = {
    postgres_database          = databricks_postgres_database.app.status.postgres_database
    branch                     = databricks_postgres_branch.production.name
    create_database_if_missing = false
  }
}

# Sync a Unity Catalog Delta table into the Lakebase database continuously.
# Prefixing synced_table_id with the Lakebase Catalog name places the synced table
# inside the catalog so it's discoverable alongside the rest of the catalog's contents.
# postgres_database references the catalog's status, which implicitly orders this
# resource after the catalog without an explicit depends_on.
resource "databricks_postgres_synced_table" "orders" {
  synced_table_id = "app_catalog.default.orders_synced"
  spec = {
    branch                             = databricks_postgres_branch.production.name
    postgres_database                  = databricks_postgres_catalog.app_catalog.status.postgres_database
    source_table_full_name             = "my_catalog.default.orders"
    primary_key_columns                = ["order_id"]
    scheduling_policy                  = "CONTINUOUS"
    create_database_objects_if_missing = true
    new_pipeline_spec = {
      storage_catalog = "my_catalog"
      storage_schema  = "default"
    }
  }
}

# Databricks App connected to the Lakebase project.
# database must be the full resource name (databricks_postgres_database.app.name),
# not the Postgres database name. permission must be "CAN_CONNECT_AND_CREATE".
resource "databricks_app" "this" {
  name        = "my-lakebase-app"
  description = "App backed by Lakebase autoscaling project"
  depends_on  = [databricks_postgres_database.app]
  resources = [{
    name = "lakebase-db"
    postgres = {
      branch     = databricks_postgres_branch.production.name
      database   = databricks_postgres_database.app.name
      permission = "CAN_CONNECT_AND_CREATE"
    }
  }]
}

Note

Esta configuración crea un rol independiente (admin_sp) y una base de datos (app) en lugar de gestionar el rol de propietario implícito y la base de datos databricks_postgres. Para poner esos recursos implícitos bajo la gestión de Terraform, declárelos mediante replace_existing = true usando sus identificadores existentes. El identificador de la base de datos es siempre databricks-postgres. El ID de rol se deriva de la identidad que creó la rama: la parte del correo electrónico anterior a @ (en minúsculas, con los caracteres no alfanuméricos sustituidos por guiones) para un usuario, o sp-<application-id> para una entidad de servicio. Si no tienes claro el valor exacto, léelo desde la app de Lakebase o la API de Postgres en lugar de derivarlo a mano.

spec.membership_roles sobrescribe las pertenencias del rol en cada aplicación en lugar de combinarse con ellas. Mantenga DATABRICKS_SUPERUSER en la lista; dejarlo fuera elimina todas las pertenencias del rol.

resource "databricks_postgres_role" "owner" {
  role_id = "jane-doe" # normalized login of the creating identity
  parent  = databricks_postgres_branch.production.name
  spec = {
    postgres_role    = "jane.doe@databricks.com" # the raw login
    membership_roles = ["DATABRICKS_SUPERUSER"]
    attributes = {
      createdb   = true
      createrole = true
      bypassrls  = true
    }
  }
  replace_existing = true
}

resource "databricks_postgres_database" "databricks_postgres" {
  database_id = "databricks-postgres"
  parent      = databricks_postgres_branch.production.name
  spec = {
    postgres_database = "databricks_postgres"
    # spec.role is omitted, so the database keeps its existing owner.
  }
  replace_existing = true
}

Recursos adicionales