Schedule GPU workloads and compose tasks

Important

This feature is in Public Preview. To use it, a workspace admin must enable the AI Runtime preview from the Previews page. See Manage Azure Databricks previews.

Use Declarative Automation Bundles to schedule an AI Runtime GPU workload and combine it with other work, such as upstream data preparation. Define the workload and its schedule in YAML, then deploy them as a job. For a training pipeline, run preprocessing on CPU compute and start GPU training after the data is ready.

This guide starts with a workload that prints "Hello world", then builds a scheduled preprocessing and training pipeline. If you already have an AI Runtime workload YAML, see Convert an AI Runtime workload to a bundle.

How it works

  • A bundle contains your code and a databricks.yml configuration.
  • A job groups tasks and defines when they run.
  • An ai_runtime_task runs your command on serverless GPU compute.

The depends_on field connects tasks into a directed acyclic graph (DAG), and the schedule field schedules your job to run on a cadence.

databricks bundle deploy uploads the code and creates or updates the job. databricks bundle run starts a job run immediately. On each run, AI Runtime provisions the GPU compute, runs your command, and records the run in the named MLflow experiment.

Requirements

Hello world example with ai_runtime_task

This example runs a shell command on one A10 GPU.

  1. Create a directory for the bundle. In that directory, create command.sh with the following contents:

    #!/usr/bin/env bash
    set -euo pipefail
    echo "Hello world"
    
  2. Create databricks.yml in the same directory:

    bundle:
      name: hello-ai-runtime
    
    resources:
      jobs:
        hello:
          name: hello-ai-runtime
          tasks:
            - task_key: hello
              environment_key: default
              ai_runtime_task:
                experiment: hello-ai-runtime
                deployments:
                  - command_path: ${workspace.file_path}/command.sh
                    compute:
                      accelerator_type: GPU_1xA10
                      accelerator_count: 1
          environments:
            - environment_key: default
              spec:
                environment_version: '6'
    
    targets:
      dev:
        mode: development
        default: true
    
  3. From the bundle directory, validate, deploy, and run the job:

    databricks bundle validate --target dev
    databricks bundle deploy --target dev
    databricks bundle run hello --target dev
    
  4. Open the run URL printed by the CLI and view the hello task's output. It contains Hello world. The run also appears in the hello-ai-runtime MLflow experiment.

More complex example: Schedule a preprocessing and training pipeline

This example prepares a small dataset on CPU compute, then trains a linear model on one A10 GPU. Both tasks use the same volume file to pass data between them. The job is scheduled for 09:00 UTC daily, with the schedule paused until you test it.

Create the project

  1. Create a separate directory with the following layout:

    scheduled-training/
    ├── databricks.yml
    ├── prep.py
    ├── command.sh
    └── src/
        └── train.py
    
  2. Create prep.py with the following contents. Replace <catalog>, <schema>, and <volume> with your volume's names. This source notebook normalizes the input values and writes the prepared data to the volume:

    # Databricks notebook source
    import json
    from pathlib import Path
    
    data_path = Path("/Volumes/<catalog>/<schema>/<volume>/scheduled-training/data.json")
    inputs = [value / 10 for value in range(10)]
    data = {"x": inputs, "y": [2 * value + 1 for value in inputs]}
    data_path.parent.mkdir(parents=True, exist_ok=True)
    data_path.write_text(json.dumps(data))
    print(f"Prepared {len(inputs)} rows at {data_path}")
    
  3. Create src/train.py with the following contents. Use the same volume path as in prep.py:

    import json
    from pathlib import Path
    
    import torch
    
    data_path = Path("/Volumes/<catalog>/<schema>/<volume>/scheduled-training/data.json")
    data = json.loads(data_path.read_text())
    x = torch.tensor(data["x"], device="cuda").reshape(-1, 1)
    y = torch.tensor(data["y"], device="cuda").reshape(-1, 1)
    model = torch.nn.Linear(1, 1).to("cuda")
    optimizer = torch.optim.SGD(model.parameters(), lr=0.1)
    
    for _ in range(200):
        optimizer.zero_grad()
        loss = torch.nn.functional.mse_loss(model(x), y)
        loss.backward()
        optimizer.step()
    
    print(f"Trained on {x.device}: loss={loss.item():.4f}")
    
  4. Create command.sh to run the training code. The bundle below packages src/, so CODE_SOURCE_PATH points to the extracted src directory:

    #!/usr/bin/env bash
    set -euo pipefail
    cd "$CODE_SOURCE_PATH"
    python train.py
    
  5. Create databricks.yml with the following contents. The bundle packages src/ for the GPU task. The prep notebook runs on serverless CPU compute, and depends_on starts train only after prep succeeds. max_concurrent_runs: 1 prevents runs of this job from overwriting each other's input file:

    bundle:
      name: scheduled-training
    
    artifacts:
      code:
        type: tgz
        path: .
        include: [src]
        files:
          - source: ./dist/code.tgz
    
    resources:
      jobs:
        train_pipeline:
          name: scheduled-training
          max_concurrent_runs: 1
          schedule:
            quartz_cron_expression: '0 0 9 * * ?'
            timezone_id: UTC
            pause_status: PAUSED
          tasks:
            - task_key: prep
              notebook_task:
                notebook_path: ./prep.py
            - task_key: train
              depends_on:
                - task_key: prep
              environment_key: training
              ai_runtime_task:
                experiment: scheduled-training
                code_source_path: ./dist/code.tgz
                deployments:
                  - command_path: ${workspace.file_path}/command.sh
                    compute:
                      accelerator_type: GPU_1xA10
                      accelerator_count: 1
          environments:
            - environment_key: training
              spec:
                environment_version: '6'
                dependencies:
                  - torch
    
    targets:
      dev:
        mode: development
        default: true
    

Test and enable the schedule

  1. From scheduled-training/, deploy the bundle and start a manual run:

    databricks bundle validate --target dev
    databricks bundle deploy --target dev
    databricks bundle run train_pipeline --target dev
    
  2. Open the run URL printed by the CLI. Confirm that prep succeeds before train starts. The prep output reports 10 prepared rows, and the train output reports Trained on cuda:0 and the loss.

  3. In databricks.yml, change the schedule's pause_status from PAUSED to UNPAUSED. The schedule block is now:

    schedule:
      quartz_cron_expression: '0 0 9 * * ?'
      timezone_id: UTC
      pause_status: UNPAUSED
    
  4. Deploy the change to activate the daily schedule:

    databricks bundle deploy --target dev
    

    Setting pause_status: UNPAUSED explicitly enables this schedule even for a development target. In Jobs & Pipelines, open the deployed job and confirm that its schedule is active for 09:00 UTC. See Run jobs on a schedule.

To pause scheduled runs, set pause_status: PAUSED and deploy again. To remove the example job and the files uploaded by the bundle, run the following from its bundle directory:

databricks bundle destroy --target dev

The prepared data in your volume remains. Delete the scheduled-training data directory when you no longer need it.

Additional resources