Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Important
This feature is in Public Preview. Workspace admins can control access to this feature from the Previews page. See Manage Azure Databricks previews.
This page shows how to serve your own LLM on a Model Serving GPU endpoint. Any model that vLLM runs becomes a production endpoint with an OpenAI-compatible API. This is how it works:
- You run the inference server, such as vLLM, and choose its version and settings.
- You test the server in a serverless GPU notebook. A wrong flag or an out-of-memory error shows up in seconds instead of after a deployment.
- You deploy the same command and environment that you built and know works to a production serving endpoint.
Quickstart
Import the following notebook into your workspace and click Run all on AI Runtime with an A10 GPU. In about 15 minutes, you have a Qwen3.5-4B endpoint that answers OpenAI-compatible chat requests.
Serve Qwen3.5-4B with vLLM
When to use custom LLM serving
Use custom LLM serving in the following cases:
- You fine-tuned an LLM on AI Runtime and want to serve it. See the section below for an example.
- You want to serve an open model that Foundation Model APIs (FMAPI) doesn't support, like a new LLM, a speech-to-text model, or an embedding model.
- You want to change the runtime or architecture of an LLM and need control of the runtime and environment.
Don't use custom LLM serving in the following cases:
- FMAPI serves the model you need, unchanged. FMAPI is simpler to use and highly optimized.
- The model doesn't fit on a single GPU with about 80 GB of memory. For example, Qwen3.8-27B and gpt-oss-120b fit, but Kimi K3 and GLM 5.3 don't.
- You need frontier-level throughput or price per token without tuning. Performance matches open source vLLM on the same GPU.
Requirements
Custom LLM serving has the following requirements:
- Your workspace must have serverless GPU compute.
- You must log the model from a serverless GPU notebook. A model logged from a CPU environment packages CPU dependencies, and the GPU endpoint fails to start.
- You must have permission to create models in a Unity Catalog schema.
- You must register the model with
env_pack="databricks_model_serving". Custom LLM serving is built on express deployments, which package the notebook's environment with the model. - You must use MLflow 3.12 or above and
databricks-sdk0.102.0 or above. The AI v6 environment includes both.
Starter notebooks
Each notebook takes one model from Hugging Face to a queried endpoint, like the quickstart. Run one as is, or start from one to serve your own model.
| Model | Task | Environment | Notebook GPU | Endpoint GPU |
|---|---|---|---|---|
| Qwen3.5-4B | Chat | AI v6 or Standard v6 | A10 | A10G (AWS), A100 (Azure) |
| Qwen3.8-27B | Chat with tool calls | AI v6 | H100 | H100 (AWS), A100 (Azure) |
| Gemma 4 26B-A4B | Chat with tool calls | AI v6 | H100 | H100 (AWS), A100 (Azure) |
| Muse Glimmer 30B | Chat | Standard v6 | H100 | H100 (AWS), A100 (Azure) |
| Whisper large-v3-turbo | Speech to text | AI v6 | A10 | A10G (AWS), A100 (Azure) |
| Qwen3-Embedding-0.6B | Embeddings | AI v6 | A10 | A10G (AWS), A100 (Azure) |
AI v6 and Standard v6 are AI Runtime environments. AI v6 comes with vLLM, PyTorch, Transformers, and other common machine learning packages pre-installed and ready for use. See the AI v6 package list. A Standard v6 notebook installs vLLM itself, so you choose its version. Use a Standard v6 notebook when your model needs a newer vLLM or transformers than AI v6 includes, as with Muse Glimmer.
Serve your own model
To serve another model, pick the starter notebook with the same task and the GPU your model needs. Change MODEL_REPO_ID and the flags in vllm_command, then run the notebook. Every notebook runs the following steps:
- Download the model weights from Hugging Face, or get them from a training checkpoint.
- Start vLLM in the notebook and query it.
- Log the model with the vLLM command as its entrypoint, and register it to Unity Catalog as an express deployment.
- Create a serving endpoint, which starts the same command.
- Query the endpoint with the OpenAI client, the Databricks SDK, or SQL
ai_query.
In the following example, the model's metadata holds the task and the entrypoint:
import mlflow
from mlflow.pyfunc.model import ChatCompletionResponse, ChatModel
# The endpoint runs the entrypoint and never calls predict, but MLflow needs a model class to log.
class Placeholder(ChatModel):
def predict(self, context, messages, params):
return ChatCompletionResponse.from_dict({"choices": []})
model_info = mlflow.pyfunc.log_model(
name="my-model",
python_model=Placeholder(),
artifacts={"model_dir": "my-model"}, # the weights folder, which --model names
metadata={
"task": "llm/v1/chat",
"entrypoint": (
"python -u -m vllm.entrypoints.openai.api_server "
"--model my-model --served-model-name my-model "
"--host 0.0.0.0 --port 8080 --max-model-len 16384"
),
},
)
mlflow.register_model(model_info.model_uri, "<catalog>.<schema>.my_model", env_pack="databricks_model_serving")
Serve a fine-tuned model
Fine-tune a model on AI Runtime and serve it from the same notebook. For an example, see Supervised fine-tuning (Full) and serving of Qwen3.5-0.8B. The tutorial fine-tunes Qwen3.5-0.8B on a single H100, compares answers before and after training, and serves the fine-tuned model.
Supported tasks
The task in the model's metadata sets the API the endpoint serves. Your server must expose the OpenAI-compatible API for that task. Other tasks, such as llm/v1/completions, aren't supported. The following table lists the supported tasks.
task |
Model type | Query with |
|---|---|---|
llm/v1/chat |
Chat models, including vision-language models | chat.completions |
llm/v1/embeddings |
Embedding models | embeddings |
llm/v1/audio/transcriptions |
Speech-to-text models | audio.transcriptions |
llm/v1/audio/translations |
Speech-to-English translation models | audio.translations |
Choose a GPU
Azure Databricks recommends developing on the same GPU that you serve on, so the settings you test are the settings you deploy. The following table lists the GPUs for custom LLM serving.
workload_type |
GPU | Notes |
|---|---|---|
GPU_SMALL |
1x T4 (16 GB) | Models up to about 5B parameters. |
GPU_LARGE |
1x A100 (80 GB) | Large models. Generally available. |
GPU_LARGE_RTX |
1x RTX PRO 6000 (96 GB) | More memory and speed than A100. Available in southeastasia and westus2. |
Common issues
The following issues are the most common when you change a notebook:
- The download fails in
/Workspace, which doesn't accept multi-GB files. Download the weights to local disk, as the starter notebooks do. - The local server doesn't start. Serverless GPU notebooks allow only ports 3000 to 3999, so test on one of those. Only the entrypoint uses port 8080, and it must otherwise match the command you tested.
- The endpoint can't find the weights. The entrypoint runs in the model's artifacts folder, so
--modelmust name the weights folder you logged. - An embedding model doesn't serve embeddings. Start vLLM with
--runner pooling. - Registration fails with
TimeoutError('Timed out after 0:05:00')while uploadingmodel_version.tarormodel_environment.tar. Upgrade todatabricks-sdk0.102.0 or above and register the model again.
Create an endpoint
Create the endpoint from the Serving UI or with the Databricks SDK, as the starter notebooks do. workload_type picks the GPU, and workload_size (Small, Medium, or Large) sets the number of replicas.
ServedEntityInput(
entity_name="<catalog>.<schema>.<model>",
entity_version="<version>",
workload_type=ServingModelWorkloadType.GPU_MEDIUM,
workload_size="Small",
scale_to_zero_enabled=False,
)
Scale-to-zero and capacity
Currently custom llm endpoints don't autoscale dynamically, so size workload_size for your peak traffic.
With scale-to-zero, an idle endpoint stops all replicas. The next request waits one to several minutes while vLLM loads the model again and all replicas start up. Azure Databricks recommends turning off scale-to-zero for production traffic.
Warning
Scale-up capacity is not guaranteed. Whenever Azure Databricks needs to acquire a new GPU for your endpoint, such as on creation, on a workload_size increase, or when an endpoint wakes up from zero, the request can stop responding if the cloud provider has no GPU capacity in your region. This applies to all GPU types. Databricks mitigates this with warm pools and prereservation, which keep GPU capacity available and ready.
Query your endpoint
A ready chat endpoint appears in the AI Playground. The following examples query an endpoint with the Databricks SDK, the OpenAI client, and REST.
Databricks SDK
from databricks.sdk import WorkspaceClient
from databricks.sdk.service.serving import ChatMessage, ChatMessageRole
w = WorkspaceClient()
w.serving_endpoints.query(
name="<endpoint-name>",
messages=[ChatMessage(role=ChatMessageRole.USER, content="Hello")],
)
OpenAI client
from openai import OpenAI
client = OpenAI(api_key=DATABRICKS_TOKEN, base_url=f"{DATABRICKS_HOST}/serving-endpoints")
client.chat.completions.create(model="<endpoint-name>", messages=[{"role": "user", "content": "Hello"}])
client.embeddings.create(model="<endpoint-name>", input=["The quick brown fox jumps over the lazy dog."])
with open("speech.flac", "rb") as audio:
client.audio.transcriptions.create(model="<endpoint-name>", file=audio)
REST: chat
curl -X POST \
-u "token:$DATABRICKS_TOKEN" \
-H "Content-Type: application/json" \
-d '{"messages":[{"role":"user","content":"Hello"}]}' \
https://<workspace-url>/serving-endpoints/<endpoint-name>/invocations
REST: embeddings
curl -X POST \
-u "token:$DATABRICKS_TOKEN" \
-H "Content-Type: application/json" \
-d '{"input":["The quick brown fox jumps over the lazy dog."]}' \
https://<workspace-url>/serving-endpoints/<endpoint-name>/invocations
Some embedding models expect a prefix on each input, such as search_query: or search_document:. Check the model card.
Monitor your endpoint
The endpoint's Logs tab shows your server's stdout and stderr live. The logs API returns the same output.
Azure Databricks forwards your vLLM server's metrics and charts them on the endpoint's Metrics tab:
- Latency: time to first token, time per output token, request latency, and queue time.
- Load: requests running and waiting.
- KV cache: usage and hit rate.
- Throughput: prompt and generation tokens per second.
The export metrics API returns the same metrics in Prometheus format, so you can scrape them into Prometheus or Datadog.
With telemetry turned on, Azure Databricks also saves the server logs and the vLLM Prometheus metrics to Unity Catalog tables, so you can query them over longer periods. See Persist custom model serving data to Unity Catalog.
Pricing
You pay per GPU instance hour, the same as for other GPU custom model serving. See Model Serving pricing.
Limitations and region availability
You log a custom LLM from AI Runtime, so custom LLM serving is available in the same regions as serverless GPU compute. Those are US regions on AWS and Azure. GCP isn't supported.
The following features are coming soon:
- LoRA adapters.
- Cross-region serving.
- Logging models outside AI Runtime, for example from Databricks Runtime or CPU serverless.
The following features aren't supported:
- KV-cache-aware routing.
- Route optimization.