使用 Unity 目錄函式建立代理工具

使用 Unity Catalog 函式來建立執行自訂邏輯並執行特定任務的代理工具,這些任務將大型語言模型的能力擴展到語言生成之外。

何時使用 Unity 目錄函式與 MCP 伺服器

Databricks 建議在查詢事先已知且代理提供參數時,使用 Unity Catalog 函式作為代理工具,專門用於結構化資料檢索工具。 請參見 「連接代理與結構化資料」。

在大多數其他使用情境中,Databricks 建議使用 MCP 伺服器或直接在代理程式中定義邏輯,以加快執行速度、支援每用戶驗證及增加彈性。

Requirements

要建立並使用 Unity 目錄函式作為代理工具,你需要以下條件:

  • Databricks Runtime:使用 Databricks Runtime 15.0 和更新版本
  • Python 版本:安裝 Python 3.10 或更新版本

若要執行 Unity 目錄函式:

  • 在你的工作空間中必須啟用無伺服器運算,才能在生產環境中執行 Unity Catalog 函式作為代理工具。 請參閱 無伺服器運算需求。
    • Python 函式的本機模式執行不需要無伺服器泛型計算才能執行,不過本機模式只適用於開發和測試用途。

若要建立 Unity 目錄函式:

  • 您的工作區中必須啟用無伺服器泛型計算,才能使用 Databricks Workspace Client 或 SQL 主體語句來建立函式。
    • Python 函式可以在沒有伺服器運算的情況下建立。

建立 Unity 目錄函式工具

以下步驟說明如何建立並測試 Unity 目錄函式。 在 Databricks 筆記本中執行下列程式代碼。

Tip

告訴 精靈代碼 (代理模式)幫你做這件事:

Create a Unity Catalog Python function that an AI agent can use as a tool. It should take two floating point numbers and return their sum, with type hints and a Google-style docstring. Register it using the Databricks Function Client, then test calling it.

安裝依賴項

安裝 Unity Catalog AI 套件,並包含 [databricks] 額外功能。

# Install Unity Catalog AI integration packages with the Databricks extra
%pip install unitycatalog-ai[databricks]

dbutils.library.restartPython()

初始化 Databricks 函式用戶端

初始化 Databricks 函式用戶端,這是在 Databricks 中建立、管理及執行 Unity 目錄函式的特殊介面。

from unitycatalog.ai.core.databricks import DatabricksFunctionClient

client = DatabricksFunctionClient()

定義工具的邏輯

Unity Catalog 工具其實本質上只是 Unity Catalog 使用者定義函數 (UDF)。 當你定義 Unity 目錄工具時,你是在 Unity 目錄中註冊一個函式。 欲了解更多關於 Unity Catalog UDFs,請參閱 Unity Catalog 中的 SQL 與 Python 使用者定義函式(UDF)。

警告

在代理程式工具中執行任意程式碼可能會暴露代理程式有權存取的敏感或私人資訊。 客戶只需執行受信任的程式碼,並設定防護欄及適當權限,以防止資料被誤導存取。

您可以使用兩個 API 的其中一個來建立 Unity 目錄函式:

  • create_python_function 接受 Python 可調用對象。
  • create_function 接受 SQL 主體 create 函式語句。 請參閱 建立 Python 函式。

使用 create_python_function API 來建立函式。

若要讓 Python 函式能被 Unity Catalog 函數數據模型辨識,您的函式必須符合下列條件:

  • 類型提示:函式簽章必須定義有效的 Python 類型提示。 具名自變數和傳回值都必須定義其類型。
  • 請勿使用變數自變數:不支援 *args 和 **kwargs 等變數自變數。 所有參數都必須明確定義。
  • 類型相容性:SQL 中不支援所有 Python 類型。 請參閱 Spark 支援的數據類型。
  • 描述性檔字串:Unity 目錄函式工具組會從您的 docstring 讀取、剖析及擷取重要資訊。
    • Docstring 必須根據 Google docstring語法來格式化。
    • 撰寫函式及其自變數的清楚描述,以協助 LLM 瞭解如何使用函式的方式和時機。
  • 相依性匯入:連結庫必須在函式主體內匯入。 執行工具時,將不會解析函式外部的匯入。

使用以下程式碼片段中的create_python_function來註冊 Python 的可調用對象add_numbers:


CATALOG = "my_catalog"
SCHEMA = "my_schema"

def add_numbers(number_1: float, number_2: float) -> float:
  """
  A function that accepts two floating point numbers adds them,
  and returns the resulting sum as a float.

  Args:
    number_1 (float): The first of the two numbers to add.
    number_2 (float): The second of the two numbers to add.

  Returns:
    float: The sum of the two input numbers.
  """
  return number_1 + number_2

function_info = client.create_python_function(
  func=add_numbers,
  catalog=CATALOG,
  schema=SCHEMA,
  replace=True
)

測試函式

測試您的函式,以檢查它是否如預期般運作。 在 API execute_function 中指定完整的函式名稱以運行函式:

result = client.execute_function(
  function_name=f"{CATALOG}.{SCHEMA}.add_numbers",
  parameters={"number_1": 36939.0, "number_2": 8922.4}
)

result.value # OUTPUT: '45861.4'

在你的代理程式中新增 Unity 目錄函式

當你建立並測試完 Unity 目錄函式後,請選擇以下方法之一將其加入代理程式。

MCP 圖示。 使用 MCP(建議)

Databricks 建議使用 MCP 伺服器,將 Unity 目錄功能加入代理程式。 MCP 方法提供更簡單的整合,結合自動工具發現及內建認證支援。

Unity 目錄函式的受管理 MCP URL 為: https://<workspace-hostname>/api/2.0/mcp/functions/{catalog}/{schema}。 你可以選擇性地透過附加 /{function_name}來指定特定函式。

以下範例說明如何透過 MCP 將您的代理程式與 Unity 目錄函式連結。 將 <catalog> 和 <schema> 替換為您的函數位置。

OpenAI Agents SDK (Apps)

from agents import Agent, Runner
from databricks.sdk import WorkspaceClient
from databricks_openai.agents import McpServer

workspace_client = WorkspaceClient()

async with McpServer.from_uc_function(
    catalog="<catalog>",
    schema="<schema>",
    workspace_client=workspace_client,
    name="uc-functions",
) as uc_server:
    agent = Agent(
        name="Tool-using agent",
        instructions="You are a helpful assistant. Use the available tools to answer questions.",
        model="databricks-claude-sonnet-4-5",
        mcp_servers=[uc_server],
    )
    result = await Runner.run(agent, "Look up customer info for Acme Corp")
    print(result.final_output)

授權應用程式存取 Unity 目錄功能:databricks.yml

resources:
  apps:
    my_agent_app:
      resources:
        - name: 'my_uc_function'
          uc_securable:
            securable_full_name: '<catalog>.<schema>.<function-name>'
            securable_type: 'FUNCTION'
            permission: 'EXECUTE'

LangGraph(應用程式)

from databricks.sdk import WorkspaceClient
from databricks_langchain import ChatDatabricks, DatabricksMCPServer, DatabricksMultiServerMCPClient
from langgraph.prebuilt import create_react_agent

workspace_client = WorkspaceClient()
host = workspace_client.config.host

mcp_client = DatabricksMultiServerMCPClient([
    DatabricksMCPServer(
        name="uc-functions",
        url=f"{host}/api/2.0/mcp/functions/<catalog>/<schema>",
        workspace_client=workspace_client,
    ),
])

async with mcp_client:
    tools = await mcp_client.get_tools()
    agent = create_react_agent(
        ChatDatabricks(endpoint="databricks-claude-sonnet-4-5"),
        tools=tools,
    )
    result = await agent.ainvoke(
        {"messages": [{"role": "user", "content": "Look up customer info for Acme Corp"}]}
    )
    print(result["messages"][-1].content)

授權應用程式存取 Unity 目錄功能:databricks.yml

resources:
  apps:
    my_agent_app:
      resources:
        - name: 'my_uc_function'
          uc_securable:
            securable_full_name: '<catalog>.<schema>.<function-name>'
            securable_type: 'FUNCTION'
            permission: 'EXECUTE'

模型服務

from databricks.sdk import WorkspaceClient
from databricks_mcp import DatabricksMCPClient
import mlflow

workspace_client = WorkspaceClient()
host = workspace_client.config.host

# Connect to the UC functions MCP server
mcp_client = DatabricksMCPClient(
    server_url=f"{host}/api/2.0/mcp/functions/<catalog>/<schema>",
    workspace_client=workspace_client,
)

# List available tools
tools = mcp_client.list_tools()

# Log the agent with the required resources for deployment
mlflow.pyfunc.log_model(
    "agent",
    python_model=my_agent,
    resources=mcp_client.get_databricks_resources(),
)

要部署代理,請參見 「為 AI 應用部署代理(模型服務)」。 關於使用 MCP 資源的日誌代理的詳細資訊,請參見 Azure Databricks 管理的 MCP 伺服器。

功能圖示。 使用 UCFunctionToolkit

使用 UCFunctionToolkit

此範例使用 LangChain,但類似的方法也可以套用至其他函式庫。 請參考「 Use Unity Catalog 工具搭配其他 AI 框架」。

安裝額外的相依套件

安裝 UCFunctionToolkit 的 LangChain 整合套件。

%pip install unitycatalog-langchain[databricks]==0.2.0

# Install the Databricks LangChain integration package
%pip install databricks-langchain==0.5.0

dbutils.library.restartPython()

使用 UCFunctionToolKit 來包裝函式

使用 UCFunctionToolkit 封裝函式,以便代理程式開發庫存取。 該工具包確保不同 AI 函式庫間的一致性,並新增了如自動追蹤檢索器等實用功能。

from databricks_langchain import UCFunctionToolkit

# Create a toolkit with the Unity Catalog function
func_name = f"{CATALOG}.{SCHEMA}.add_numbers"
toolkit = UCFunctionToolkit(function_names=[func_name])

tools = toolkit.tools

在代理程式中使用工具

使用 tools 中的 UCFunctionToolkit 屬性將工具新增至 LangChain 代理程式。

Note

此範例使用 LangChain。 不過,您可以將 Unity 目錄工具與其他架構整合,例如 LlamaIndex、OpenAI、Anthropic 等。 請參考「 Use Unity Catalog 工具搭配其他 AI 框架」。

本範例使用 LangChain AgentExecutor API 撰寫簡單的代理程式,以求簡單。 對於生產工作負載,請使用「Author a agent」中提到的代理創作工作流程 ,並部署到 Databricks 應用程式中。

from langchain.agents import AgentExecutor, create_tool_calling_agent
from langchain.prompts import ChatPromptTemplate
from databricks_langchain import (
  ChatDatabricks,
  UCFunctionToolkit,
)
import mlflow

# Initialize the LLM (optional: replace with your LLM of choice)
LLM_ENDPOINT_NAME = "databricks-meta-llama-3-3-70b-instruct"
llm = ChatDatabricks(endpoint=LLM_ENDPOINT_NAME, temperature=0.1)

# Define the prompt
prompt = ChatPromptTemplate.from_messages(
  [
    (
      "system",
      "You are a helpful assistant. Make sure to use tools for additional functionality.",
    ),
    ("placeholder", "{chat_history}"),
    ("human", "{input}"),
    ("placeholder", "{agent_scratchpad}"),
  ]
)

# Enable automatic tracing
mlflow.langchain.autolog()

# Define the agent, specifying the tools from the toolkit above
agent = create_tool_calling_agent(llm, tools, prompt)

# Create the agent executor
agent_executor = AgentExecutor(agent=agent, tools=tools, verbose=True)
agent_executor.invoke({"input": "What is 36939.0 + 8922.4?"})

將 Unity 目錄工具與其他 AI 框架搭配使用

除了 LangChain(如上圖所示)之外,Unity Catalog 代理工具還能與其他熱門的 AI 函式庫如 LlamaIndex、OpenAI 和 Anthropic 協同運作。 這些整合結合了 Unity Catalog 工具治理與第三方代理創作框架的功能。 例如,在 OpenAI 或 Anthropic 整合中,這些函式會在執行時由 AI 模型直接呼叫。

在以下分頁選擇你的框架,建立 Unity 目錄工具並搭配該框架使用。 在 Azure Databricks 筆記本或 Python 腳本中執行程式碼。

LlamaIndex

使用 Azure Databricks Unity Catalog 將 SQL 與 Python 函式整合為 LlamaIndex 工作流程的工具。 此整合將 Unity Catalog 的數據管理功能與 LlamaIndex 的能力結合,目的是為大型語言模型(LLM)編製索引並查詢大型數據集。

  1. 安裝適用於 LlamaIndex 的 Databricks Unity 目錄整合套件。

    %pip install unitycatalog-llamaindex[databricks]
    dbutils.library.restartPython()
    
  2. 建立 Unity 目錄函式用戶端的實例。

    from unitycatalog.ai.core.base import get_uc_function_client
    
    client = get_uc_function_client()
    
  3. 建立以 Python 撰寫的 Unity 目錄函式。

    CATALOG = "your_catalog"
    SCHEMA = "your_schema"
    
    func_name = f"{CATALOG}.{SCHEMA}.code_function"
    
    def code_function(code: str) -> str:
      """
      Runs Python code.
    
      Args:
        code (str): The Python code to run.
      Returns:
        str: The result of running the Python code.
      """
      import sys
      from io import StringIO
      stdout = StringIO()
      sys.stdout = stdout
      exec(code)
      return stdout.getvalue()
    
    client.create_python_function(
      func=code_function,
      catalog=CATALOG,
      schema=SCHEMA,
      replace=True
    )
    
  4. 建立一個 Unity 目錄函式作為工具包實例,並執行以驗證工具是否正常運作。

    from unitycatalog.ai.llama_index.toolkit import UCFunctionToolkit
    import mlflow
    
    # Enable traces
    mlflow.llama_index.autolog()
    
    # Create a UCFunctionToolkit that includes the UC function
    toolkit = UCFunctionToolkit(function_names=[func_name])
    
    # Fetch the tools stored in the toolkit
    tools = toolkit.tools
    python_exec_tool = tools[0]
    
    # Run the tool directly
    result = python_exec_tool.call(code="print(1 + 1)")
    print(result)  # Outputs: {"format": "SCALAR", "value": "2\n"}
    
  5. 藉由將 Unity 目錄函式定義為 LlamaIndex 工具集合的一部分,以使用 LlamaIndex ReActAgent 中的工具。 然後呼叫 LlamaIndex 工具集合,確認代理程式的行為正確。

    from llama_index.llms.openai import OpenAI
    from llama_index.core.agent import ReActAgent
    
    llm = OpenAI()
    
    agent = ReActAgent.from_tools(tools, llm=llm, verbose=True)
    
    agent.chat("Please run the following python code: `print(1 + 1)`")
    

OpenAI

使用 Azure Databricks Unity Catalog,將 SQL 和 Python 函式整合為 OpenAI 工作流程的工具。 此整合結合了 Unity Catalog 與 OpenAI 的治理,打造強大的 AI 應用程式。

  1. 安裝適用於 OpenAI 的 Databricks Unity 目錄整合套件。

    %pip install unitycatalog-openai[databricks]
    %pip install mlflow -U
    dbutils.library.restartPython()
    
  2. 建立 Unity 目錄函式用戶端的實例。

    from unitycatalog.ai.core.base import get_uc_function_client
    
    client = get_uc_function_client()
    
  3. 建立以 Python 撰寫的 Unity 目錄函式。

    CATALOG = "your_catalog"
    SCHEMA = "your_schema"
    
    func_name = f"{CATALOG}.{SCHEMA}.code_function"
    
    def code_function(code: str) -> str:
      """
      Runs Python code.
    
      Args:
        code (str): The python code to run.
      Returns:
        str: The result of running the Python code.
      """
      import sys
      from io import StringIO
      stdout = StringIO()
      sys.stdout = stdout
      exec(code)
      return stdout.getvalue()
    
    client.create_python_function(
      func=code_function,
      catalog=CATALOG,
      schema=SCHEMA,
      replace=True
    )
    
  4. 建立 Unity Catalog 函式的實例作為工具組,並藉由執行函式來確認工具運作正常。

    from unitycatalog.ai.openai.toolkit import UCFunctionToolkit
    import mlflow
    
    # Enable tracing
    mlflow.openai.autolog()
    
    # Create a UCFunctionToolkit that includes the UC function
    toolkit = UCFunctionToolkit(function_names=[func_name])
    
    # Fetch the tools stored in the toolkit
    tools = toolkit.tools
    client.execute_function = tools[0]
    
  5. 將要求連同工具一起提交至 OpenAI 模型。

    import openai
    
    messages = [
      {
        "role": "system",
        "content": "You are a helpful customer support assistant. Use the supplied tools to assist the user.",
      },
      {"role": "user", "content": "What is the result of 2**10?"},
    ]
    response = openai.chat.completions.create(
      model="gpt-4o-mini",
      messages=messages,
      tools=tools,
    )
    # check the model response
    print(response)
    
  6. 在 OpenAI 傳回回應之後,叫用 Unity Catalog 函式,以產生回應並傳回給 OpenAI。

    import json
    
    # OpenAI sends only a single request per tool call
    tool_call = response.choices[0].message.tool_calls[0]
    # Extract arguments that the Unity Catalog function needs to run
    arguments = json.loads(tool_call.function.arguments)
    
    # Run the function based on the arguments
    result = client.execute_function(func_name, arguments)
    print(result.value)
    
  7. 傳回答案之後,您可以建構後續對OpenAI呼叫的響應承載。

    # Create a message containing the result of the function call
    function_call_result_message = {
      "role": "tool",
      "content": json.dumps({"content": result.value}),
      "tool_call_id": tool_call.id,
    }
    assistant_message = response.choices[0].message.to_dict()
    completion_payload = {
      "model": "gpt-4o-mini",
      "messages": [*messages, assistant_message, function_call_result_message],
    }
    
    # Generate final response
    openai.chat.completions.create(
      model=completion_payload["model"], messages=completion_payload["messages"]
    )
    

Utilities

為了簡化製作工具回應的程式, ucai-openai 套件具有公用程式 generate_tool_call_messages,可轉換 OpenAI ChatCompletion 回應訊息,以便用於產生回應。

from unitycatalog.ai.openai.utils import generate_tool_call_messages

messages = generate_tool_call_messages(response=response, client=client)
print(messages)

Note

如果回應包含多個選擇專案,您可以在呼叫 generate_tool_call_messages 時傳遞choice_index自變數,以選擇要使用哪一個選擇專案。 目前不支援處理多項選項。

Anthropic

使用 Azure Databricks Unity Catalog 來整合 SQL 和 Python 函式,作為 Anthropic SDK LLM 呼叫的工具。 此整合結合了 Unity Catalog 的治理與 Anthropic 模型,打造強大的 AI 應用程式。

Note

Anthropic 整合功能需使用 Databricks Runtime 15.0 或以上版本。

  1. 安裝適用於 Anthropic 的 Databricks Unity Catalog 集成包。

    %pip install unitycatalog-anthropic[databricks]
    dbutils.library.restartPython()
    
  2. 建立 Unity 目錄函式用戶端的實例。

    from unitycatalog.ai.core.base import get_uc_function_client
    
    client = get_uc_function_client()
    
  3. 建立以 Python 撰寫的 Unity 目錄函式。

    CATALOG = "your_catalog"
    SCHEMA = "your_schema"
    
    func_name = f"{CATALOG}.{SCHEMA}.weather_function"
    
    def weather_function(location: str) -> str:
      """
      Fetches the current weather from a given location in degrees Celsius.
    
      Args:
        location (str): The location to fetch the current weather from.
      Returns:
        str: The current temperature for the location provided in Celsius.
      """
      return f"The current temperature for {location} is 24.5 celsius"
    
    client.create_python_function(
      func=weather_function,
      catalog=CATALOG,
      schema=SCHEMA,
      replace=True
    )
    
  4. 將 Unity Catalog 函數的實體建立為工具包。

    from unitycatalog.ai.anthropic.toolkit import UCFunctionToolkit
    
    # Create an instance of the toolkit
    toolkit = UCFunctionToolkit(function_names=[func_name], client=client)
    
  5. 在 Anthropic 中使用工具呼叫。

    import anthropic
    
    # Initialize the Anthropic client with your API key
    anthropic_client = anthropic.Anthropic(api_key="YOUR_ANTHROPIC_API_KEY")
    
    # User's question
    question = [{"role": "user", "content": "What's the weather in New York City?"}]
    
    # Make the initial call to Anthropic
    response = anthropic_client.messages.create(
      model="claude-3-5-sonnet-20240620",  # Specify the model
      max_tokens=1024,  # Use 'max_tokens' instead of 'max_tokens_to_sample'
      tools=toolkit.tools,
      messages=question  # Provide the conversation history
    )
    
    # Print the response content
    print(response)
    
  6. 建立工具回應。 如果需要調用工具,則來自 Claude 模型的回應包含一個工具請求元數據塊。

    from unitycatalog.ai.anthropic.utils import generate_tool_call_messages
    
    # Call the UC function and construct the required formatted response
    tool_messages = generate_tool_call_messages(
      response=response,
      client=client,
      conversation_history=question
    )
    
    # Continue the conversation with Anthropic
    tool_response = anthropic_client.messages.create(
      model="claude-3-5-sonnet-20240620",
      max_tokens=1024,
      tools=toolkit.tools,
      messages=tool_messages,
    )
    
    print(tool_response)
    

該 unitycatalog.ai-anthropic 包包括一個消息處理程序實用程式,用於簡化對 Unity Catalog 函數的調用的解析和處理。 該實用程式執行以下作:

  1. 檢測工具調用要求。
  2. 從查詢中擷取工具呼叫資訊。
  3. 執行對 Unity Catalog 函數的調用。
  4. 分析來自 Unity Catalog 函數的回應。
  5. 製作下一個消息格式以繼續與 Claude 的對話。

Note

必須在 API 的conversation_history參數中提供generate_tool_call_messages整個對話歷史記錄。 Claude 模型需要初始化對話(原始使用者輸入問題)以及所有後續 LLM 生成的回應和多輪次工具調用結果。

用清晰的文件改善工具呼叫

良好的文件幫助你的員工瞭解何時及如何使用每個工具。 請遵循下列最佳做法來記錄您的工具:

  • 針對 Unity 目錄函式,請使用 COMMENT 子句來描述工具功能和參數。
  • 清楚定義預期的輸入和輸出。
  • 撰寫有意義的描述,讓代理程式和人類更容易使用工具。

範例:有效的工具文件

下列範例顯示用於查詢結構化表格的工具的清晰COMMENT字串。

CREATE OR REPLACE FUNCTION main.default.lookup_customer_info(
  customer_name STRING COMMENT 'Name of the customer whose info to look up.'
)
RETURNS STRING
COMMENT 'Returns metadata about a specific customer including their email and ID.'
RETURN SELECT CONCAT(
    'Customer ID: ', customer_id, ', ',
    'Customer Email: ', customer_email
  )
  FROM main.default.customer_data
  WHERE customer_name = customer_name
  LIMIT 1;

範例:無效的工具文件

下列範例缺少重要的詳細數據,讓代理程序難以有效地使用此工具:

CREATE OR REPLACE FUNCTION main.default.lookup_customer_info(
  customer_name STRING COMMENT 'Name of the customer.'
)
RETURNS STRING
COMMENT 'Returns info about a customer.'
RETURN SELECT CONCAT(
    'Customer ID: ', customer_id, ', ',
    'Customer Email: ', customer_email
  )
  FROM main.default.customer_data
  WHERE customer_name = customer_name
  LIMIT 1;

使用無伺服器或本地模式執行函式

當 AI 服務判斷需要工具呼叫時,整合套件(UCFunctionToolkit 實例)會執行該 DatabricksFunctionClient.execute_function API。

呼叫 execute_function 可以在兩種執行模式中執行函式:無伺服器或本機。 此模式會決定執行函式的資源。

生產環境的無伺服器模式

無伺服器模式是執行 Unity Catalog 功能作為代理工具時,生產環境的預設且推薦的選項。 此模式使用無伺服器的通用計算平台(Spark Connect 無伺服器模式)來遠端執行函式,而 Lakeguard 確保代理程式的安全,避免本地執行任意程式碼的風險。

Note

Unity Catalog 作為代理工具執行的功能需要無伺服器通用運算(Spark Connect serverless),而非無伺服器 SQL 倉庫。 嘗試在沒有無伺服器泛型計算的情況下執行工具會產生類似的錯誤 PERMISSION_DENIED: Cannot access Spark Connect。

# Defaults to serverless if `execution_mode` is not specified
client = DatabricksFunctionClient(execution_mode="serverless")

當您的代理程式要求 在無伺服器 模式中執行工具時,會發生下列情況:

  1. DatabricksFunctionClient 將要求傳送至 Unity Catalog,以擷取函式定義,如果尚未在本機快取該定義。
  2. DatabricksFunctionClient 會擷取函式定義,並驗證參數名稱和類型。
  3. 會將 DatabricksFunctionClient 執行以 UDF 的形式提交至無伺服器泛型計算。

用於開發的本機模式

本地模式會在本地子程序執行 Python 函式,而非向無伺服器的通用運算請求。 這可讓您藉由提供本地堆棧追蹤,更有效率地對工具調用進行疑難解答。 其設計目的是開發及偵錯 Python Unity 目錄函式。

當您的代理程式要求以 本機 模式執行工具時,會 DatabricksFunctionClient 執行下列動作:

  1. 將請求傳送至 Unity Catalog,以擷取函式定義,若此定義尚未在本機快取。
  2. 擷取 Python 可呼叫定義,將可呼叫的定義快取到本地,並驗證參數名稱與類型。
  3. 使用受限制的子進程中的指定參數叫用可呼叫者,並具有逾時保護。
# Defaults to serverless if `execution_mode` is not specified
client = DatabricksFunctionClient(execution_mode="local")

"local" 模式下執行提供以下功能:

  • CPU 時間限制: 限制可呼叫執行的總 CPU 運行時間,以防止過多的計算負載。

    CPU 時間限制是以實際的CPU使用量為基礎,而不是時鐘時間。 由於系統排程和並行程序,CPU 時間可能會超過現實情況中的時鐘時間。

  • 記憶體限制: 限制配置給進程的虛擬記憶體。

  • 逾時保護: 強制執行執行函式的總時鐘逾時。

使用環境變數自定義這些限制(進一步閱讀)。

本機模式限制

  • Python 僅支援函式:本地模式下不支援基於 SQL 的函式。
  • 不受信任的程式代碼的安全性考慮:雖然本機模式會在進程隔離的子進程中執行函式,但在執行 AI 系統所產生的任意程式碼時,可能會有安全性風險。 這主要是當函式執行未經審查的動態生成 Python 程式碼時,才會特別令人擔憂。
  • 連結庫版本差異:連結庫版本在無伺服器和本機執行環境之間可能會有所不同,這可能會導致不同的函式行為。

環境變數

使用下列環境變數設定 函式在 中 DatabricksFunctionClient 執行的方式:

環境變數 預設值 Description
EXECUTOR_MAX_CPU_TIME_LIMIT 10 秒 允許的 CPU 執行時間上限(僅限本機模式)。
EXECUTOR_MAX_MEMORY_LIMIT 100 MB 進程允許的虛擬記憶體配置上限(僅限本機模式)。
EXECUTOR_TIMEOUT 20 秒 總壁時計時間上限(僅限本地模式)。
UCAI_DATABRICKS_SESSION_RETRY_MAX_ATTEMPTS 5 在令牌到期時,重試重新整理會話用戶端的嘗試次數上限。
UCAI_DATABRICKS_SERVERLESS_EXECUTION_RESULT_ROW_LIMIT 100 使用無伺服器計算和 databricks-connect執行函式時要傳回的數據列數目上限。

使用 http_request 呼叫外部 API(舊版)

你可以建立一個 Unity Catalog 函式,繞行 http_request() 呼叫基於 SQL 的工具定義中的外部服務。 此方法仍被支援,但不再建議用於新整合。 如需逐步解說(包括 SQL 範例和連線類型限制),請參閱 將 http_request() 包裝在 Unity Catalog 函式中。

筆記本範例

以下筆記本示範如何建立使用 Unity 目錄函式連接外部服務的代理工具。

Slack 傳訊代理程式工具

拿筆記本

Microsoft 圖形 API 代理工具

拿筆記本

Azure AI 搜索代理工具

拿筆記本

下一步