適用於:
Databricks SQL
Databricks Runtime
Important
這項功能位於 測試版 (Beta) 中。 工作區管理員可以從 「預覽 」頁面控制對此功能的存取。 請參閱 管理 Azure Databricks 預覽。
該 ai_search() 函式從一個或多個知識來源檢索資訊。 知識來源可以包括 AI 搜尋索引 或啟用 Lakebase 搜尋的同步表格。
給定一個自然語言查詢及為呼叫配置的知識來源,該函式會產生優化的搜尋查詢,檢索並去重,根據相關性重新排序,並回傳最相關的文件。 預設情況下,它也會在檢索的文件上綜合出一個紮實的自然語言答案。
用於 ai_search 大規模豐富操作資料與相關脈絡、建構批次檢索增強生成(RAG)流程,或將檢索作為複合 AI 系統的工具暴露——全部由單一 SQL 函式呼叫完成。
數據安全性
您的文件數據會在 Databricks 安全性周邊內處理。 Databricks 不儲存傳遞給 AI 函式呼叫的參數,但會保留元資料執行細節,例如 Databricks 執行時版本。
要求
- Databricks 執行環境 18.2 或以上。
- 一個或多個可搜尋的知識來源: AI 搜尋索引,或啟用 Lakebase 搜尋時同步的表格。
- 要使用 Lakebase 同步表格,請先準備搜尋。 參見 湖基配置。
- 若使用無伺服器運算,無伺服器環境版本必須設定為 3 或以上,這樣才能啟用像
VARIANT。 - 此功能
ai_search可透過 Databricks 筆記本、SQL 編輯器、Databricks 工作流程、工作或 Lakeflow 上的 Spark 宣告式管線使用。
Syntax
ai_search(query, knowledge_sources [, instructions] [, options])
引數
-
querySTRING:或VARIANT表達式。 自然語言搜尋查詢。VARIANT輸入,例如另一個 AI 函式的輸出,內部會序列化成 JSON 字串。 -
knowledge_sources: 一個VARIANT包含 JSON 知識來源配置陣列的 ORSTRING表達式,供搜尋。 請參見 知識來源配置。 你可以指定最多 10 個知識來源。 -
instructions:一個最多可選STRING的4,000字元表達式。 以自然語言指引查詢產生、元資料篩選生成及重新排序。 例如,'Prefer official documentation over internal articles when both cover the same topic.' -
options:一個可選MAP<STRING, STRING>的 。 支援的按鍵:-
'version':使用的功能版本。 -
'generate_answer':'true'(預設值)或'false'。 當'true'時,函數會從檢索的文件中綜合一個有基礎的自然語言答案,並在現場回傳。answer設定為'false'僅歸還文件。 -
'generate_citations':'true'或'false'(預設)。 當'true'時,函數會回傳引用,顯示模型用來支援其生成答案所擷取的區塊。 此選項需要'generate_answer'為'true'。
-
知識來源配置
這個 knowledge_sources 參數是一個 JSON 陣列。 每個元素都是一個 {type, config} 信封。 該 type 欄位識別如何 ai_search 連接來源,且與面向客戶的資源名稱是分開的。 設定 type 為 AI vector_search 搜尋索引,或 lakebase_search 啟用 Lakebase 搜尋的同步表格。 欄位 config 包含來源特定的配置。
| Key | Required | Description |
|---|---|---|
type |
Yes | 知識來源類型。 其中一個 vector_search (AI 搜尋索引)或 lakebase_search (啟用 Lakebase 搜尋的 Lakebase 同步表格)。 |
config |
Yes | 一個包含特定來源配置的物件。 請參閱 AI 搜尋索引配置 的 vector_search 和 的 Lakebase 配置。lakebase_search |
單一通話可合併 vector_search 來源, lakebase_search 最多可達 10 個來源限制。
AI 搜尋索引配置
對於設定 type 為 vector_search的 config AI 搜尋索引,接受以下鍵數:
| Key | Required | Description |
|---|---|---|
index_name |
Yes | Unity catalog.schema.my_index目錄中 AI 搜尋索引的三層名稱,例如 。 |
text_col |
Yes | 索引中包含文件文本的欄位則以 回傳為 page_content。 |
doc_uri_col |
Yes | 索引中包含文件 URI 的欄位被回傳為 doc_uri。 |
filter_columns |
No | 一個逗號分隔的字串或 JSON 欄位陣列,可用於元資料過濾。 若省略,則該清單由索引結構衍生,排除保留、文字及文件 URI 欄位。 |
以下範例將一個 AI 搜尋索引配置為知識來源:
[
{
"type": "vector_search",
"config": {
"index_name": "prod_catalog.docs.support_articles",
"text_col": "article_body",
"doc_uri_col": "article_url",
"filter_columns": "product,language"
}
}
]
若要從原始文件建立 AI 搜尋索引,請使用 ai_parse_document 並在 ai_prep_search Delta 表格中建立可搜尋的區塊。 接著從該表格 建立 AI 搜尋索引 。 索引上線後,使用其三層級名稱為 index_name。
湖基配置
在你使用Lakebase同步表格作為知識來源之前,請先準備好搜尋:
- 啟用 Lakebase Search 並安裝擴充功能。 請依照 Lakebase Search 在專案設定中啟用 Lakebase Search,然後連接到專案的 Postgres 資料庫,並
lakebase_vector透過執行CREATE EXTENSION IF NOT EXISTS lakebase_vector CASCADE;安裝擴充功能。 它提供近似的最近鄰向量搜尋。ai_search啟用 Lakebase Search 需要測試版存取權,且不可逆。 - 把你的 Unity Catalog 資料表同步到專案裡。 使用 同步表,將嵌入欄位映射到欄位
vector(n),並在上面建立lakebase_ann索引。 尺寸n必須與你的嵌入模型輸出相符。 - 離線產生嵌入,並用一個服務端點的模型。 將嵌入欄位填入嵌入模型——例如,透過查詢 的
ai_query嵌入模型——然後將該端點傳遞為embedding_model(見下文)。ai_search在搜尋時將每個查詢嵌入其中,因此查詢與儲存的文件嵌入共用一個向量空間。
對於設 type 為 lakebase_search的 config Lakebase 知識來源,接受以下鍵數:
| Key | Required | Description |
|---|---|---|
index_name |
Yes | Unity Catalog 中 Lakebase 同步表 的三層名稱,該表包含嵌入的區塊,例如 catalog.schema.my_synced_table。 |
text_col |
Yes | 表格中包含文件文本的欄位回傳為 page_content。 |
doc_uri_col |
Yes | 表格中包含文件 URI 的欄位回傳為 doc_uri。 |
embedding_column |
Yes | 表格中的嵌入欄位用於搜尋。 |
embedding_model |
Yes | 服務端點的模型過去用於離線嵌入 embedding_column 。
ai_search 在搜尋時會用相同的端點嵌入每個查詢,因此必須與產生儲存嵌入的模型相符。 請參考下方說明。 |
lakebase_project_id |
Yes | 例如 my-lakebase-project,承載同步資料表的 Lakebase 資料庫專案(實例)名稱。 這是專案名稱,不是 UUID。 |
filter_columns |
No | 一個逗號分隔的字串或 JSON 欄位陣列,可用於元資料過濾。 若省略,則由表格結構衍生,排除保留欄位、文字、文件 URI 及嵌入欄位。 設為空陣列([]),即使資料表有其他欄位,也可停用元資料過濾。 |
Important
embedding_model 必須是你用來產生離線嵌入 embedding_column 的同一個服務端點模型。
ai_search 在查詢時會用這個端點嵌入每個查詢;如果它與產生儲存向量的模型不同,查詢與文件嵌入會佔據不同的向量空間,檢索品質會大幅下降。
以下範例將 Lakebase 表格配置為知識來源。 替換 embedding_model 為你用來離線嵌入資料表的端點:
[
{
"type": "lakebase_search",
"config": {
"index_name": "prod_catalog.docs.support_articles_synced",
"text_col": "article_body",
"doc_uri_col": "article_url",
"embedding_column": "article_embedding",
"embedding_model": "databricks-gte-large-en",
"lakebase_project_id": "my-lakebase-project"
}
}
]
Returns
VARIANT A,其結構如下:
{
"document": [
{
"page_content": STRING, // Text content of the retrieved chunk
"doc_uri": STRING, // URI of the source document
"metadata": MAP // Additional metadata from the index
}
],
"answer": STRING, // Grounded answer synthesized from the retrieved
// documents, or null
"citations": [ // Present only when generate_citations is true
{
"document_index": INT // Position in the document array above, starting at 0
}
]
}
| Field | 類型 | Description |
|---|---|---|
document |
ARRAY |
依相關性排序的檢索文件陣列。 |
document[].page_content |
STRING |
取出的區塊文字內容。 |
document[].doc_uri |
STRING |
原始文件的 URI。 |
document[].metadata |
MAP |
索引中的額外元資料。 |
answer |
STRING |
這是從所取得文件中綜合而成的貼近自然語言的答案。
null 當答案產生被關閉或未取得任何文件時。 |
citations |
ARRAY |
只有在 generate_citations 為 true時,才會顯示 。 每個項目根據其在陣列中document以 0 為基礎的位置document_index指向回傳區塊,因此你可以看到模型用哪些區塊來支持其答案。 空陣列表示引用產生已執行但未產生引用。 當引用產生無法或未被調用query時,包含 SQL NULL時,值為 null 。 這些引用是模型中盡力挑選的支持片段,並非任何特定主張的證明。 |
當 generate_citations 是 true時,會一起讀 answer 和 citations 欄位。 每種數值組合的意義都不同:
- 非空
answer且陣列為citations空,表示已產生答案,但未被選取回傳區塊。 -
answer空citations值陣列表示引用產生已執行,但因未產生有根據的答案而產生引用。 - 空
citations值表示引用產生無法或未被調用,包括querySQL 時NULL。
Examples
基本搜尋
以下範例搜尋一個 AI 搜尋索引,並回傳排名文件及有根據的答案:
SELECT ai_search(
'How do I configure auto-scaling for my SQL warehouse?',
PARSE_JSON('[{
"type": "vector_search",
"config": {
"index_name": "prod_catalog.docs.support_articles",
"text_col": "article_body",
"doc_uri_col": "article_url",
"filter_columns": "product,language"
}
}]')
) AS result;
回傳產生答案的支援區塊
以下範例回傳模型所選擇支持答案的檢索區塊子集:
SELECT ai_search(
'How do I configure auto-scaling for my SQL warehouse?',
PARSE_JSON('[{
"type": "vector_search",
"config": {
"index_name": "prod_catalog.docs.support_articles",
"text_col": "article_body",
"doc_uri_col": "article_url"
}
}]'),
options => map('generate_citations', 'true')
) AS result;
陣 document 列仍包含完整的回傳區塊集合。 在每個引用中使用 , document_index 從該陣列中挑選支持的。 這些指標僅在單一回應中有效。 它們可以在通話間、重新排序變更或索引更新時改變。
多來源搜尋及說明
以下範例搜尋兩個 AI 搜尋索引,並用 instructions 來引導查詢產生與重新排序:
SELECT ai_search(
'What are the networking requirements for serverless SQL warehouses?',
PARSE_JSON('[
{
"type": "vector_search",
"config": {"index_name": "prod_catalog.docs.public_docs", "text_col": "content", "doc_uri_col": "doc_url"}
},
{
"type": "vector_search",
"config": {"index_name": "prod_catalog.docs.internal_kb", "text_col": "body", "doc_uri_col": "source_uri"}
}
]'),
'Focus on firewall rules and VPC/VNet configuration. Prefer official documentation over internal articles when both cover the same topic.'
) AS result;
搜尋湖底表格
以下範例搜尋 Lakebase 同步表格,並回傳排名文件及有根據的答案:
SELECT ai_search(
'How do I configure auto-scaling for my SQL warehouse?',
PARSE_JSON('[{
"type": "lakebase_search",
"config": {
"index_name": "prod_catalog.docs.support_articles_synced",
"text_col": "article_body",
"doc_uri_col": "article_url",
"embedding_column": "article_embedding",
"embedding_model": "databricks-gte-large-en",
"lakebase_project_id": "my-lakebase-project"
}
}]')
) AS result;
結合 AI 搜尋索引與 Lakebase 表格
以下範例在同一通話中搜尋 AI 搜尋索引與 Lakebase 資料表。 兩個來源的結果會被去重、合併重新排序,並以一個排名清單的形式回傳:
SELECT ai_search(
'What are the networking requirements for serverless SQL warehouses?',
PARSE_JSON('[
{
"type": "vector_search",
"config": {"index_name": "prod_catalog.docs.public_docs", "text_col": "content", "doc_uri_col": "doc_url"}
},
{
"type": "lakebase_search",
"config": {
"index_name": "prod_catalog.docs.internal_kb_synced",
"text_col": "body",
"doc_uri_col": "source_uri",
"embedding_column": "body_embedding",
"embedding_model": "databricks-gte-large-en",
"lakebase_project_id": "my-lakebase-project"
}
}
]')
) AS result;
豐富表格,提供檢索與生成答案
以下範例將以相關文件及建議解決方案豐富每張支援工單。 由於答案生成預設為開啟狀態,建議解析度可直接在 answer 現場取得,無需額外生成步驟。
SELECT
ticket_id,
customer_description,
ai_search(
customer_description,
PARSE_JSON('[{
"type": "vector_search",
"config": {
"index_name": "support.docs.product_documentation",
"text_col": "content",
"doc_uri_col": "doc_url"
}
}]'),
'Find product documentation, known issues, and troubleshooting guides relevant to this support ticket.'
):answer::STRING AS suggested_resolution
FROM support.tickets.open_tickets;
若要控制輸出格式或使用特定模型,請將擷取的文件設'generate_answer''false'為並串接至 。ai_query
Limitations
-
ai_search支援 AI 搜尋索引("type": "vector_search")與 Lakebase 同步表格("type": "lakebase_search")作為知識來源。 - 你可以在每次通話中指定最多 10 個知識來源,且可任意組合支援的類型。
-
instructions論證限於4,000字元。