Serve custom LLMs with Custom Model Serving

Important

This feature is in Public Preview. Workspace admins can control access to this feature from the Previews page. See Manage Azure Databricks previews.

This page shows how to serve your own LLM on a Model Serving GPU endpoint. Any model that vLLM runs becomes a production endpoint with an OpenAI-compatible API. This is how it works:

  • You run the inference server, such as vLLM, and choose its version and settings.
  • You test the server in a serverless GPU notebook. A wrong flag or an out-of-memory error shows up in seconds instead of after a deployment.
  • You deploy the same command and environment that you built and know works to a production serving endpoint.

Quickstart

Import the following notebook into your workspace and click Run all on AI Runtime with an A10 GPU. In about 15 minutes, you have a Qwen3.5-4B endpoint that answers OpenAI-compatible chat requests.

Serve Qwen3.5-4B with vLLM

Get notebook

When to use custom LLM serving

Use custom LLM serving in the following cases:

  • You fine-tuned an LLM on AI Runtime and want to serve it. See the section below for an example.
  • You want to serve an open model that Foundation Model APIs (FMAPI) doesn't support, like a new LLM, a speech-to-text model, or an embedding model.
  • You want to change the runtime or architecture of an LLM and need control of the runtime and environment.

Don't use custom LLM serving in the following cases:

  • FMAPI serves the model you need, unchanged. FMAPI is simpler to use and highly optimized.
  • The model doesn't fit on a single GPU with about 80 GB of memory. For example, Qwen3.8-27B and gpt-oss-120b fit, but Kimi K3 and GLM 5.3 don't.
  • You need frontier-level throughput or price per token without tuning. Performance matches open source vLLM on the same GPU.

Requirements

Custom LLM serving has the following requirements:

  • Your workspace must have serverless GPU compute.
  • You must log the model from a serverless GPU notebook. A model logged from a CPU environment packages CPU dependencies, and the GPU endpoint fails to start.
  • You must have permission to create models in a Unity Catalog schema.
  • You must register the model with env_pack="databricks_model_serving". Custom LLM serving is built on express deployments, which package the notebook's environment with the model.
  • You must use MLflow 3.12 or above and databricks-sdk 0.102.0 or above. The AI v6 environment includes both.

Starter notebooks

Each notebook takes one model from Hugging Face to a queried endpoint, like the quickstart. Run one as is, or start from one to serve your own model.

Model Task Environment Notebook GPU Endpoint GPU
Qwen3.5-4B Chat AI v6 or Standard v6 A10 A10G (AWS), A100 (Azure)
Qwen3.8-27B Chat with tool calls AI v6 H100 H100 (AWS), A100 (Azure)
Gemma 4 26B-A4B Chat with tool calls AI v6 H100 H100 (AWS), A100 (Azure)
Muse Glimmer 30B Chat Standard v6 H100 H100 (AWS), A100 (Azure)
Whisper large-v3-turbo Speech to text AI v6 A10 A10G (AWS), A100 (Azure)
Qwen3-Embedding-0.6B Embeddings AI v6 A10 A10G (AWS), A100 (Azure)

AI v6 and Standard v6 are AI Runtime environments. AI v6 comes with vLLM, PyTorch, Transformers, and other common machine learning packages pre-installed and ready for use. See the AI v6 package list. A Standard v6 notebook installs vLLM itself, so you choose its version. Use a Standard v6 notebook when your model needs a newer vLLM or transformers than AI v6 includes, as with Muse Glimmer.

Serve your own model

To serve another model, pick the starter notebook with the same task and the GPU your model needs. Change MODEL_REPO_ID and the flags in vllm_command, then run the notebook. Every notebook runs the following steps:

  1. Download the model weights from Hugging Face, or get them from a training checkpoint.
  2. Start vLLM in the notebook and query it.
  3. Log the model with the vLLM command as its entrypoint, and register it to Unity Catalog as an express deployment.
  4. Create a serving endpoint, which starts the same command.
  5. Query the endpoint with the OpenAI client, the Databricks SDK, or SQL ai_query.

In the following example, the model's metadata holds the task and the entrypoint:

import mlflow
from mlflow.pyfunc.model import ChatCompletionResponse, ChatModel

# The endpoint runs the entrypoint and never calls predict, but MLflow needs a model class to log.
class Placeholder(ChatModel):
    def predict(self, context, messages, params):
        return ChatCompletionResponse.from_dict({"choices": []})

model_info = mlflow.pyfunc.log_model(
    name="my-model",
    python_model=Placeholder(),
    artifacts={"model_dir": "my-model"},  # the weights folder, which --model names
    metadata={
        "task": "llm/v1/chat",
        "entrypoint": (
            "python -u -m vllm.entrypoints.openai.api_server "
            "--model my-model --served-model-name my-model "
            "--host 0.0.0.0 --port 8080 --max-model-len 16384"
        ),
    },
)
mlflow.register_model(model_info.model_uri, "<catalog>.<schema>.my_model", env_pack="databricks_model_serving")

Serve a fine-tuned model

Fine-tune a model on AI Runtime and serve it from the same notebook. For an example, see Supervised fine-tuning (Full) and serving of Qwen3.5-0.8B. The tutorial fine-tunes Qwen3.5-0.8B on a single H100, compares answers before and after training, and serves the fine-tuned model.

Supported tasks

The task in the model's metadata sets the API the endpoint serves. Your server must expose the OpenAI-compatible API for that task. Other tasks, such as llm/v1/completions, aren't supported. The following table lists the supported tasks.

task Model type Query with
llm/v1/chat Chat models, including vision-language models chat.completions
llm/v1/embeddings Embedding models embeddings
llm/v1/audio/transcriptions Speech-to-text models audio.transcriptions
llm/v1/audio/translations Speech-to-English translation models audio.translations

Choose a GPU

Azure Databricks recommends developing on the same GPU that you serve on, so the settings you test are the settings you deploy. The following table lists the GPUs for custom LLM serving.

workload_type GPU Notes
GPU_SMALL 1x T4 (16 GB) Models up to about 5B parameters.
GPU_LARGE 1x A100 (80 GB) Large models. Generally available.
GPU_LARGE_RTX 1x RTX PRO 6000 (96 GB) More memory and speed than A100. Available in southeastasia and westus2.

Common issues

The following issues are the most common when you change a notebook:

  • The download fails in /Workspace, which doesn't accept multi-GB files. Download the weights to local disk, as the starter notebooks do.
  • The local server doesn't start. Serverless GPU notebooks allow only ports 3000 to 3999, so test on one of those. Only the entrypoint uses port 8080, and it must otherwise match the command you tested.
  • The endpoint can't find the weights. The entrypoint runs in the model's artifacts folder, so --model must name the weights folder you logged.
  • An embedding model doesn't serve embeddings. Start vLLM with --runner pooling.
  • Registration fails with TimeoutError('Timed out after 0:05:00') while uploading model_version.tar or model_environment.tar. Upgrade to databricks-sdk 0.102.0 or above and register the model again.

Create an endpoint

Create the endpoint from the Serving UI or with the Databricks SDK, as the starter notebooks do. workload_type picks the GPU, and workload_size (Small, Medium, or Large) sets the number of replicas.

ServedEntityInput(
    entity_name="<catalog>.<schema>.<model>",
    entity_version="<version>",
    workload_type=ServingModelWorkloadType.GPU_MEDIUM,
    workload_size="Small",
    scale_to_zero_enabled=False,
)

Scale-to-zero and capacity

Currently custom llm endpoints don't autoscale dynamically, so size workload_size for your peak traffic.

With scale-to-zero, an idle endpoint stops all replicas. The next request waits one to several minutes while vLLM loads the model again and all replicas start up. Azure Databricks recommends turning off scale-to-zero for production traffic.

Warning

Scale-up capacity is not guaranteed. Whenever Azure Databricks needs to acquire a new GPU for your endpoint, such as on creation, on a workload_size increase, or when an endpoint wakes up from zero, the request can stop responding if the cloud provider has no GPU capacity in your region. This applies to all GPU types. Databricks mitigates this with warm pools and prereservation, which keep GPU capacity available and ready.

Query your endpoint

A ready chat endpoint appears in the AI Playground. The following examples query an endpoint with the Databricks SDK, the OpenAI client, and REST.

Databricks SDK

from databricks.sdk import WorkspaceClient
from databricks.sdk.service.serving import ChatMessage, ChatMessageRole

w = WorkspaceClient()
w.serving_endpoints.query(
    name="<endpoint-name>",
    messages=[ChatMessage(role=ChatMessageRole.USER, content="Hello")],
)

OpenAI client

from openai import OpenAI

client = OpenAI(api_key=DATABRICKS_TOKEN, base_url=f"{DATABRICKS_HOST}/serving-endpoints")

client.chat.completions.create(model="<endpoint-name>", messages=[{"role": "user", "content": "Hello"}])
client.embeddings.create(model="<endpoint-name>", input=["The quick brown fox jumps over the lazy dog."])
with open("speech.flac", "rb") as audio:
    client.audio.transcriptions.create(model="<endpoint-name>", file=audio)

REST: chat

curl -X POST \
  -u "token:$DATABRICKS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Hello"}]}' \
  https://<workspace-url>/serving-endpoints/<endpoint-name>/invocations

REST: embeddings

curl -X POST \
  -u "token:$DATABRICKS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"input":["The quick brown fox jumps over the lazy dog."]}' \
  https://<workspace-url>/serving-endpoints/<endpoint-name>/invocations

Some embedding models expect a prefix on each input, such as search_query: or search_document:. Check the model card.

Monitor your endpoint

The endpoint's Logs tab shows your server's stdout and stderr live. The logs API returns the same output.

Azure Databricks forwards your vLLM server's metrics and charts them on the endpoint's Metrics tab:

  • Latency: time to first token, time per output token, request latency, and queue time.
  • Load: requests running and waiting.
  • KV cache: usage and hit rate.
  • Throughput: prompt and generation tokens per second.

The export metrics API returns the same metrics in Prometheus format, so you can scrape them into Prometheus or Datadog.

With telemetry turned on, Azure Databricks also saves the server logs and the vLLM Prometheus metrics to Unity Catalog tables, so you can query them over longer periods. See Persist custom model serving data to Unity Catalog.

Pricing

You pay per GPU instance hour, the same as for other GPU custom model serving. See Model Serving pricing.

Limitations and region availability

You log a custom LLM from AI Runtime, so custom LLM serving is available in the same regions as serverless GPU compute. Those are US regions on AWS and Azure. GCP isn't supported.

The following features are coming soon:

  • LoRA adapters.
  • Cross-region serving.
  • Logging models outside AI Runtime, for example from Databricks Runtime or CPU serverless.

The following features aren't supported:

  • KV-cache-aware routing.
  • Route optimization.