Workload YAML reference

Important

This feature is in Public Preview.

This page describes the workload YAML fields accepted by databricks air run. For command syntax and flags, see databricks air run. To submit a workload, see Databricks CLI quickstart for AI Runtime.

Every workload requires experiment_name, compute, and command.

Example

The following example shows a sample workload configuration:

experiment_name: my-training
compute:
  accelerator_type: GPU_1xA10
  num_accelerators: 1
environment:
  version: 6
  dependencies:
    - torch
code_source:
  type: snapshot
  snapshot:
    root_path: .
command: python $CODE_SOURCE_PATH/train.py

experiment_name

Type: String

Required. The experiment_name key is the MLflow experiment name.

compute

Type: Map

Required. The compute object defines the accelerator configuration.

Key Type Description
accelerator_type string Required. An accelerator type accepted by the CLI and available in the workspace. Matching is case-sensitive.
num_accelerators integer Required. A positive accelerator count supported for the selected accelerator type.
pool_id string The ID for provisioned GPU capacity.
priority_class string The scheduling priority for workloads using provisioned GPU capacity. Accepted values are BEST_EFFORT, NORMAL, and CRITICAL. Requires pool_id.

Accelerator types and counts depend on the workspace, cloud, and region. For hardware selection and availability, see Hardware options.

command

Type: String

Required. The command key contains a nonempty shell command or script. Paths in the command refer to remote file paths in the workload container. When code_source is set, $CODE_SOURCE_PATH contains the absolute path to the extracted code.

environment

Type: Map

The environment object selects a managed environment with optional dependencies or a custom Unity Catalog image.

Key Type Description
version string or integer A numeric environment version, such as 6, or a Databricks AI environment identifier, such as databricks_ai_v6. A databricks_ai_v identifier requires version 5 or above. If omitted, the CLI uses the current AI Runtime default.
dependencies list of strings Package specifications passed to the serverless environment. Defaults to no additional dependencies.
unity_catalog_image string A custom image in the format <catalog>.<schema>.<image>:<tag>. Mutually exclusive with version and dependencies. For setup instructions, see Use custom Docker images with AI Runtime.

For guidance about selecting and extending an environment, see Set up your environment.

Managed environment example
environment:
  version: 6
  dependencies:
    - torch

To use a custom image instead of a managed environment, specify only unity_catalog_image:

Custom image example
environment:
  unity_catalog_image: main.ml.training:v1

code_source

Type: Map

The code_source object packages local code for the workload.

Key Type Description
type string Required when code_source is set. Must be snapshot.
snapshot object Required when code_source is set. Defines the local code to package.

snapshot

Type: Map

The snapshot object supports the following keys:

Key Type Description
root_path string Required. The local directory to package. The directory must exist when the CLI stages the run.
remote_volume string A Unity Catalog volume location for the uploaded archive. Must start with /Volumes/. Defaults to a workspace location.
include_paths list of strings Paths to include relative to root_path. Entries must be nonempty and relative, and cannot contain ... Omit the key to include all selected files. An empty list is not valid.
git object Pins the snapshot to a local Git branch or commit.

If you omit git, the CLI packages the working tree, including uncommitted and untracked files that are not excluded by Git ignore rules. A relative root_path resolves from the directory that contains the workload YAML file.

git

Type: Map

The git object supports the following keys:

Key Type Description
branch string Required when commit is omitted. Packages the local HEAD of the branch. Selected files cannot contain uncommitted changes.
commit string Required when branch is omitted. Packages a commit from the local Git object store. Mutually exclusive with branch.
Example

The following example packages selected paths from the local main branch:

code_source:
  type: snapshot
  snapshot:
    root_path: ./src
    include_paths:
      - training
      - configs
    git:
      branch: main

env_variables

Type: Map of strings

The env_variables key contains plain environment variable names and string values. A name cannot also appear in secrets.

Example
env_variables:
  BATCH_SIZE: '32'
  LEARNING_RATE: '0.001'

idempotency_token

Type: String

The idempotency_token key identifies repeated submissions. A repeated submission with the same token returns the existing run. The CLI generates a UUID when this key is omitted. The --idempotency-key command flag takes precedence.

max_retries

Type: Integer

The max_retries key specifies the number of retries after the initial attempt. The value must be nonnegative. Default: 3.

mlflow_artifact_location

Type: String

The mlflow_artifact_location key contains a dbfs:/ URI or /Volumes/ path.

mlflow_experiment_directory

Type: String

The mlflow_experiment_directory key contains the workspace directory for the experiment. The path must start with /Workspace.

mlflow_run_name

Type: String

The mlflow_run_name key contains the run name.

parameters

Type: Map

The parameters key contains free-form nested values. The CLI serializes the values to hyperparameters.yaml with the other launch artifacts.

Example
parameters:
  learning_rate: 0.001
  model:
    hidden_size: 1024

permissions

Type: List of maps

The permissions key contains grants applied after submission. If a grant fails, the run remains created and the CLI prints a warning.

Each grant supports the following keys:

Key Type Description
level string Required. A permission level, such as CAN_VIEW or CAN_MANAGE. The workspace validates the value.
user_name string The user email. Specify exactly one of user_name, group_name, or service_principal_name for each grant.
group_name string The group name. Specify exactly one of user_name, group_name, or service_principal_name for each grant.
service_principal_name string The service principal name. Specify exactly one of user_name, group_name, or service_principal_name for each grant.
Example
permissions:
  - user_name: trainer@example.com
    level: CAN_MANAGE
  - group_name: training-team
    level: CAN_VIEW

secrets

Type: Map of strings

The secrets key contains environment variable names mapped to secret references in the exact scope/key format. The scope and key must be nonempty. A name cannot also appear in env_variables.

Example
secrets:
  HF_TOKEN: my-scope/hf-token

For information about creating and managing secrets, see Secret management.

timeout_minutes

Type: Integer

The timeout_minutes key sets the wall-clock limit for the complete run, including retries. The value must be positive. If omitted, the workload uses the AI Runtime default.

usage_policy_id

Type: String

The usage_policy_id key contains the canonical UUID of an existing serverless usage policy. Mutually exclusive with usage_policy_name.

usage_policy_name

Type: String

The usage_policy_name key contains the name of an existing serverless usage policy. Mutually exclusive with usage_policy_id.

For information about creating usage policies, see Attribute usage with serverless usage policies.

Additional resources