Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
How you query an agent depends on the agent server that serves it. Find your agent in the following table, then follow the matching section. To learn about agent servers, see Agent Server.
| Your agent | Hosted on | How to query |
|---|---|---|
Uses DurableAgentServer |
Agent Runtime on Databricks Apps | Invocation API at /api/invocations |
Uses the MLflow AgentServer or LongRunningAgentServer (legacy) |
Databricks Apps | Databricks OpenAI client or the OpenAI Responses API at /responses |
| Deployed to Model Serving (legacy) | Model Serving endpoint | Databricks OpenAI client, REST API, ai_query, or AI Playground |
Agents hosted on Databricks Apps require a Databricks OAuth token. Personal access tokens don't work for Databricks Apps. To generate OAuth tokens from a script, a service principal, another app, or a notebook, see Connect to an API Databricks app using token authentication.
Query an agent that uses DurableAgentServer
Agents that use DurableAgentServer serve the invocation API. Agents that you create with the Agent Bricks CLI use DurableAgentServer, and agentbricks deploy deploys them to Agent Runtime as an app named agent-bricks-<name>. Each request to the API starts one run of your agent, called an invocation.
| Endpoint | Description |
|---|---|
POST /api/invocations |
Starts an invocation. By default, the request waits and returns the result. Set stream to receive events as they happen, or background to return immediately. |
GET /api/invocations/<id> |
Returns the status of an invocation and, after it completes, its output. |
GET /api/invocations/<id>/events?after=<event-id> |
Streams the stored events that come after <event-id>. Use this endpoint to reconnect to a stream. |
Request body
The request body for POST /api/invocations accepts the following fields. The server rejects requests that contain other fields.
| Field | Description |
|---|---|
id |
Required. A UUID that you generate for each invocation. The server treats the ID as an idempotency key: resending the same request with the same ID returns the existing invocation instead of running the agent again. Reusing an ID for a different request returns a 409 error. |
session_id |
The conversation that the invocation belongs to. Invocations that share a session ID run one at a time, in order. Agents generated from the CLI templates require this field. |
input |
The input for your agent. Agents generated from the CLI templates accept a list of messages, or an object with a messages list. |
stream |
Set to true to receive events as Server-Sent Events (SSE). |
background |
Set to true to return a 202 response immediately with a status URL, and then poll for the result. |
Your agent's handler defines the structure of input. Agents generated from the CLI templates read the following fields when input is an object:
| Field | Description |
|---|---|
messages |
The conversation turns to send to the agent. |
actor |
The identity whose long-term memory the agent reads and writes. If you don't pass an actor, the agent uses the session ID, so memories don't carry over to a new session. Set actor from your application's signed-in user, not from text that the user types. |
model |
The model to use for this invocation, instead of the model set in the agent code. |
resume |
The response to an agent that paused for human input, such as approval of a tool call. When an agent pauses, the invocation's status is interrupted. Send resume in a new invocation with the same session_id to continue. |
To let the agent recall what it learned about a user across sessions, pass the user's ID as actor:
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"session_id": "support-case-123",
"input": {
"messages": [{ "role": "user", "content": "What does Databricks do?" }],
"actor": "user-42"
}
}
Agent Bricks
To test a deployed agent from your terminal, use agentbricks endpoint invoke. The command looks up the app and authenticates with your CLI profile.
agentbricks --profile <profile> endpoint invoke agent-bricks-<name> \
--path /api/invocations \
--json "{\"id\":\"$(uuidgen)\",\"session_id\":\"$(uuidgen)\",\"input\":[{\"role\":\"user\",\"content\":\"Hello\"}]}"
To stream the response, add "stream":true to the JSON body and pass --sse. To test the agent while it runs locally with agentbricks dev, replace the app name with --url http://localhost:8000.
REST API
Get the app URL. The URL field in the output is the base URL for the invocations API.
agentbricks --profile <profile> deployments get agent-bricks-<name>Get an OAuth token for your profile. The output contains the token in the
access_tokenfield.databricks auth token --profile <profile>Send a request:
curl --request POST \ --url <app-url>/api/invocations \ --header 'Authorization: Bearer <OAuth token>' \ --header 'content-type: application/json' \ --data '{ "id": "550e8400-e29b-41d4-a716-446655440000", "session_id": "support-case-123", "input": [{ "role": "user", "content": "What does Databricks do?" }] }'
The response contains the invocation id, its status, and an output field with the value that your agent returns. The server doesn't require an output schema: your handler can return any JSON-serializable value. Agents generated from the CLI templates return an object with the following fields:
output: the messages that the agent produced in this invocation.status:completed, orinterruptedif the agent paused for human input.
Python
The following example uses the Databricks SDK to look up the app URL and generate an OAuth token, and then calls the invocations API. The WorkspaceClient must use OAuth authentication.
import uuid
import requests
from databricks.sdk import WorkspaceClient
w = WorkspaceClient()
app_url = w.apps.get("agent-bricks-<name>").url
session_id = str(uuid.uuid4())
response = requests.post(
f"{app_url}/api/invocations",
headers=w.config.authenticate(),
json={
"id": str(uuid.uuid4()),
"session_id": session_id,
"input": [{"role": "user", "content": "What does Databricks do?"}],
},
)
response.raise_for_status()
print(response.json()["output"])
To continue the conversation, send the next message with the same session_id and a new id.
Stream, run in the background, and reconnect
- Stream: Set
"stream": true. The response is an SSE stream that includesrun.startedandrun.completed(orrun.failed) events, plus the events your agent emits, such asdeltaevents with streamed text. Each event has an ID. - Run in the background: Set
"background": true. The server returns a202response with astatus_url. PollGET /api/invocations/<id>until the status iscompleted. If you also set"stream": true, the response includes anevents_urlthat you can read events from. - Reconnect: If a stream disconnects, call
GET /api/invocations/<id>/events?after=<event-id>with the ID of the last event that you received.
The following Python example streams a response:
with requests.post(
f"{app_url}/api/invocations",
headers=w.config.authenticate(),
json={
"id": str(uuid.uuid4()),
"session_id": session_id,
"input": [{"role": "user", "content": "Summarize our last conversation."}],
"stream": True,
},
stream=True,
) as response:
response.raise_for_status()
for line in response.iter_lines(decode_unicode=True):
if line.startswith("data: "):
print(line[len("data: "):])
If you deploy the agent with more than one instance, send the session ID in an X-Routing-Key header to route all requests in a session to the same instance.
Query an agent that uses the legacy MLflow AgentServer
Use this section for agents that you deploy on Databricks Apps with the legacy agent server: the MLflow AgentServer or LongRunningAgentServer, with the ResponsesAgent interface. These agents serve the OpenAI Responses API at /responses.
LongRunningAgentServer serves the same API, so the following examples also apply to it. It also supports background runs: set background to true in the request, and then retrieve the response with GET /responses/<response-id>?stream=true&starting_after=<sequence-number>, which streams the events after that sequence number.
Databricks OpenAI client
Databricks recommends the Databricks OpenAI client for these agents. Include the apps/ prefix in the model name.
from databricks.sdk import WorkspaceClient
from databricks_openai import DatabricksOpenAI
input_msgs = [{"role": "user", "content": "What does Databricks do?"}]
app_name = "<agent-app-name>"
# The WorkspaceClient must use OAuth authentication.
w = WorkspaceClient()
client = DatabricksOpenAI(workspace_client=w)
# Non-streaming request
response = client.responses.create(model=f"apps/{app_name}", input=input_msgs)
print(response)
# Streaming request
streaming_response = client.responses.create(
model=f"apps/{app_name}", input=input_msgs, stream=True
)
for chunk in streaming_response:
print(chunk)
To pass custom_inputs, use the extra_body parameter:
response = client.responses.create(
model=f"apps/{app_name}",
input=input_msgs,
extra_body={"custom_inputs": {"id": 5}},
)
To get the trace ID for a request, include the x-mlflow-return-trace-id header. Then use MLflow get_trace to retrieve the full trace.
response = client.responses.create(
model=f"apps/{app_name}",
input=input_msgs,
extra_headers={"x-mlflow-return-trace-id": "true"},
)
trace_id = response.metadata["trace_id"]
trace = client.get_trace(trace_id)
REST API
Send requests to the /responses path of the app URL. The request body follows the OpenAI Responses API, so you can use any HTTP client or tool that supports it.
curl --request POST \
--url <app-url>/responses \
--header 'Authorization: Bearer <OAuth token>' \
--header 'content-type: application/json' \
--data '{
"input": [{ "role": "user", "content": "hi" }],
"stream": true
}'
To pass custom_inputs, add them to the request body:
curl --request POST \
--url <app-url>/responses \
--header 'Authorization: Bearer <OAuth token>' \
--header 'content-type: application/json' \
--data '{
"input": [{ "role": "user", "content": "hi" }],
"custom_inputs": { "id": 5 }
}'
To get the trace ID, include the x-mlflow-return-trace-id: true header. The response body includes the trace ID in a metadata.trace_id field. For streaming requests, the trace ID arrives as a separate SSE event (data: {"trace_id": "tr-..."}) near the end of the stream.
Query a legacy agent on Model Serving
Use this section for legacy agents deployed to Model Serving endpoints. You can authenticate with a Databricks OAuth token or a personal access token. To move these agents to Databricks Apps, see Migrate an agent from Model Serving to Databricks Apps.
Databricks OpenAI client
For agents that use the ResponsesAgent interface, call responses.create with the endpoint name as the model:
from databricks_openai import DatabricksOpenAI
input_msgs = [{"role": "user", "content": "What does Databricks do?"}]
endpoint = "<agent-endpoint-name>"
client = DatabricksOpenAI()
# Non-streaming request. Calls predict.
response = client.responses.create(model=endpoint, input=input_msgs)
print(response)
# Streaming request. Calls predict_stream.
streaming_response = client.responses.create(model=endpoint, input=input_msgs, stream=True)
for chunk in streaming_response:
print(chunk)
For agents that use the legacy ChatAgent or ChatModel interfaces, use the chat completions client:
from databricks.sdk import WorkspaceClient
messages = [{"role": "user", "content": "What does Databricks do?"}]
endpoint = "<agent-endpoint-name>"
client = WorkspaceClient().serving_endpoints.get_open_ai_client()
response = client.chat.completions.create(model=endpoint, messages=messages)
print(response)
With either client, pass custom_inputs or databricks_options through the extra_body parameter. For example, extra_body={"databricks_options": {"return_trace": True}} returns the trace with the response.
REST API
For agents that use the ResponsesAgent interface, send a request to /serving-endpoints/responses with the endpoint name as the model:
curl --request POST \
--url https://<workspace-url>/serving-endpoints/responses \
--header 'Authorization: Bearer <token>' \
--header 'content-type: application/json' \
--data '{
"model": "<agent-endpoint-name>",
"input": [{ "role": "user", "content": "hi" }],
"stream": true
}'
For agents that use the ChatAgent or ChatModel interfaces, send a request to /serving-endpoints/chat/completions with a messages list instead of input. To pass custom_inputs or databricks_options, add them to the request body. You can also send requests to the endpoint's /serving-endpoints/<agent-endpoint-name>/invocations URL. See Query individual models behind an endpoint.
AI Playground
To chat with an agent on Model Serving without writing code, open AI Playground and select the agent's serving endpoint. To pass custom_inputs to the agent from AI Playground, see Provide custom_inputs in the AI Playground and review app.
SQL with
Use ai_query to query an agent on Model Serving from SQL. See ai_query function for syntax and parameters.
SELECT ai_query(
"<agent-endpoint-name>", question
) FROM (VALUES ('what is MLflow?'), ('how does MLflow work?')) AS t(question);
Additional resources
- Agent Server
- Agent Runtime
- Set up production monitoring
- Query foundation and embedding models: Query foundation models and other models directly, instead of an agent.