Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Important
This feature is in Beta. To use it, a workspace admin must enable the AI Runtime Beta Features and Databricks Artifact Registry previews from the workspace Previews page.
AI Runtime can run a custom Docker container image stored in Artifact Registry. Use a custom image when you need:
- System libraries or complex dependencies that you cannot install with
environment.dependencies. - A reproducible environment across development, research, and production.
- Private, organization-approved images built by your platform or security team.
Prerequisites
- Install or update to Databricks CLI version 1.19.0 or above. See Install or update the Databricks CLI.
- Ask a workspace admin to enable the AI Runtime Beta Features preview. For instructions, see Manage workspace-level previews.
- Set up Artifact Registry in the same workspace and obtain permission to push and read images. See Get started with Artifact Registry.
- Install and start Docker on your local machine.
Push an image to Artifact Registry
Before using a custom image with AI Runtime, store it in Artifact Registry in the same workspace that you use to submit the workload.
Create or select a Unity Catalog catalog and schema for the image.
Follow Get started with Artifact Registry to configure Docker authentication, grant the required privileges, and push the image.
Note the image's Unity Catalog name in the following format:
<catalog>.<schema>.<image>:<tag>For example,
main.ml.training:v1. Do not include the registry hostname in the workload configuration.
Tip
Alternatively, use the databricks air images push helper command in Databricks CLI.
Use a Docker image in a workload
Specify the image's Unity Catalog name in your workload YAML under environment.unity_catalog_image:
experiment_name: my-dcs-training
environment:
unity_catalog_image: main.ml.training:v1
compute:
num_accelerators: 1
accelerator_type: GPU_1xA10
command: python /app/train.py
This example runs train.py from /app in the image. To upload application code separately without rebuilding the image, see Run uploaded application code.
When bringing your own Docker image, environment.dependencies and environment.version are not supported. Specifying environment.unity_catalog_image with either field triggers an error. If you have additional dependencies, install the packages in the Dockerfile instead.
Submit the workload:
databricks air run -f workload.yaml -p my-databricks-profile
The profile must authenticate to the same workspace where the image is stored.
Environment variables injected into your container
AI Runtime injects the following environment variables into every container at runtime:
CODE_SOURCE_PATH: path to uploaded application code, whencode_sourceis configured.NUM_NODES: total number of nodes.LOCAL_WORLD_SIZE: GPUs per node.WORLD_SIZE: total number of processes.POD_RANK: current node rank (0-indexed). Also injected asNODE_RANK.LOCAL_ADDR: local node IP (multi-node only).MASTER_ADDR: rank-0 coordination address (multi-node only).MASTER_PORT: rank-0 coordination port (multi-node only).
Examples
The following examples show how to run uploaded application code and distributed training with a custom image.
Run uploaded application code
Use code_source to upload application code separately from the image. You can edit and resubmit your code without rebuilding the image. Install the code's Python and system dependencies in the image.
Place train.py in a local src directory next to workload.yaml. The following configuration uploads src and runs its train.py inside the custom image:
experiment_name: my-dcs-uploaded-code
environment:
unity_catalog_image: main.ml.training:v1
compute:
num_accelerators: 1
accelerator_type: GPU_1xA10
code_source:
type: snapshot
snapshot:
root_path: ./src
command: |-
cd "$CODE_SOURCE_PATH"
python3 train.py
root_path resolves relative to workload.yaml. AI Runtime sets $CODE_SOURCE_PATH to the uploaded directory's path in the container. See code_source for snapshot options.
Multi-node H100 with RDMA
For multi-node H100 jobs that need full network bandwidth on AWS p5 instances, base your image on one of the Databricks base images with NCCL and EFA preconfigured:
experiment_name: my-dcs-distributed
environment:
unity_catalog_image: main.ml.training:v1
compute:
num_accelerators: 16 # 2 nodes × 8 H100
accelerator_type: GPU_8xH100
command: |-
torchrun \
--nnodes="${NUM_NODES}" \
--nproc_per_node="${LOCAL_WORLD_SIZE}" \
--node_rank="${POD_RANK}" \
--rdzv_endpoint="${MASTER_ADDR}:${MASTER_PORT}" \
/app/train.py
Build your own image
When building your own image, Databricks recommends using databricks-ai-runtime skill with a coding agent or starting from a Databricks base image.
Use a coding agent
Install the databricks-ai-runtime Claude Code skill for step-by-step Dockerfile guidance, including building from scratch, CUDA/NCCL/EFA compatibility, common problems, and a pre-build checklist. This skill requires Databricks CLI version 1.0.0 or newer.
databricks aitools install --skills databricks-ai-runtime --experimental
Databricks base images
Databricks publishes base images on Docker Hub at databricksruntime/air with CUDA, NCCL, and cloud-specific networking (AWS EFA or Azure InfiniBand) preconfigured.
| Tag | Variant | CUDA | Use when |
|---|---|---|---|
dcs-base-azure-runtime |
Runtime | 12 | Installing pre-built wheels only |
dcs-base-azure-devel |
Devel | 12 | Compiling CUDA extensions (requires nvcc) |
The following Dockerfile adds PyTorch to a Databricks base image. The base images provide Python at /opt/venv, managed by uv. uv pip install targets that environment by default. To use a different environment, create and activate a venv before running uv pip install.
To include your training script in the image, place train.py next to the Dockerfile. The Dockerfile copies it to /app/train.py. If you upload application code with code_source, omit the COPY instruction and keep train.py in your snapshot directory instead.
FROM databricksruntime/air:dcs-base-azure-runtime
RUN uv pip install --no-cache \
torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0
RUN uv pip install --no-cache \
transformers==4.45.0 \
accelerate==0.34.0 \
'mlflow>=3.6'
COPY ./train.py /app/train.py
Build the image locally:
docker build --platform linux/amd64 -t my-training-image:v1 .
Then follow Get started with Artifact Registry to tag and push the image to Artifact Registry. Use the resulting <catalog>.<schema>.<image>:<tag> name as environment.unity_catalog_image in the workload YAML.
Tip
Alternatively, use the databricks air images push helper command in Databricks CLI and follow the interactive prompts.
Limitations
- Images must be stored in Artifact Registry in the workspace where you submit the workload.
- Image size must be under 20 GB.
WORKDIRis not honored at runtime. Use absolute paths for files baked into the image. For example, usepython /app/train.py, notpython train.py.- You cannot use
environment.dependenciesorenvironment.versionwithenvironment.unity_catalog_image. If you need extra packages beyond what is in the image, you must add them to the Dockerfile.
Troubleshooting
For registry-related authentication, permission, image push, or image discovery errors, see Troubleshoot Artifact Registry.
ssl.SSLError when loading dependencies
A custom image can fail at runtime with an OpenSSL error when a library tries to create an SSL context, for example:
ssl.SSLError: [CRYPTO] unknown error (_ssl.c:3076)
The error appears while importing libraries that open network connections, such as huggingface_hub, and prevents them from loading.
This occurs because AI Runtime workloads run on hosts with Federal Information Processing Standards (FIPS) enabled. When the image's cryptographic libraries are not FIPS-compliant, OpenSSL fails to initialize in FIPS mode, so creating an SSL context fails.
Recommended solution:
Enterprise, government, healthcare, and finance workloads often depend on FIPS 140-2 or 140-3 compliance for FedRAMP, CMMC, or HIPAA audits. If your workload must remain FIPS-compliant, build your image with FIPS-compliant cryptographic libraries.
If your workload does not require FIPS compliance, you can disable FIPS mode by setting the OPENSSL_FORCE_FIPS_MODE environment variable to 0. Doing so can silently break compliance requirements.
To disable FIPS mode, set it in your workload YAML under env_variables:
env_variables:
OPENSSL_FORCE_FIPS_MODE: '0'
Alternatively, set the variable in your Dockerfile so it applies to every workload that uses the image:
ENV OPENSSL_FORCE_FIPS_MODE=0
Resubmit the workload and confirm the SSL error no longer appears when dependencies load.