快速入門:向量搜尋

註

Azure AI 搜尋服務 可透過 Azure 入口網站、REST API 及 Azure SDK 取得。 它同時也是 Foundry IQ 的基礎,這是一個管理式知識層,能將企業內容轉化為可重複使用、權限感知的知識庫,供 Microsoft Foundry 入口網站中的代理使用。

註

在向量搜尋工作流程中嵌入模型時,Azure AI 搜尋服務 支援 text-embedding-ada-002、 text-embedding-3-small,以及text-embedding-3-large來自 Azure OpenAI。 關於使用 AI 強化的聊天完成模型工作流程,請參閱 模型退休排程 以檢查 GPT-4 家族模型的棄用狀態。

在此快速入門中,您將使用 Azure AI 搜尋服務 .NET 用戶端函式庫 來建立、載入並查詢 向量索引。 .NET 用戶端函式庫提供對 REST API 的抽象化,用於索引操作。

在 Azure AI 搜尋服務 中,向量索引有一個定義向量與非向量場的索引結構、用於建立嵌入空間的演算法向量搜尋設定,以及查詢時評估的向量場定義設定。 索引 - 建立或更新 (REST API)建立向量索引。

提示

先決條件

  • 一個有有效訂閱的 Azure 帳號。 免費註冊帳號。

  • 一個Azure AI 搜尋服務服務。 你可以使用免費方案來完成大部分快速入門,但我們建議對於較大的資料檔案,使用基本版或更高等級。

  • .NET 8或更晚。

  • git 來複製樣本庫。

  • Azure CLI 用於與 Microsoft Entra ID 的無金鑰認證。

設定存取權限

在開始之前,請確認你有權限存取 Azure AI 搜尋服務 的內容和操作。 此快速入門使用 Microsoft Entra ID 進行驗證,並以角色為基礎的存取控制來授權。 您必須是 擁有者 或 使用者存取管理員 才能指派角色。 如果角色不可行,改用 金鑰驗證 。

要設定建議的以角色為基礎的存取權限:

  1. 為您的搜尋服務啟用基於角色的存取權限。

  2. 請將以下角色指派 到你的使用者帳號。

    • 搜尋服務貢獻者

    • 搜尋索引資料貢獻者

    • 搜尋索引資料閱讀器

取得端點

每個Azure AI 搜尋服務服務都有一個 endpoint,這是一個唯一用來識別並提供服務網路存取的網址。 在後面的章節中,你指定這個端點以程式化方式連接到你的搜尋服務。

要取得端點:

  1. 在Azure 入口網站中,進入您的搜尋服務。

  2. 從左側窗格選擇 「概覽」。

  3. 請記下該端點,應看起來像 https://my-service.search.windows.net。

設定環境

  1. 用 Git 複製樣本庫。

    git clone https://github.com/Azure-Samples/azure-search-dotnet-samples
    
  2. 進入快速啟動資料夾。

    cd azure-search-dotnet-samples/quickstart-vector-search
    
  3. 在 VectorSearchCreatePopulateIndex/appsettings.json 中,將 Endpoint 的佔位值替換成你在 取得端點獲得的 URL。

  4. 重複前一步,應用於VectorSearchExamples/appsettings.json。

  5. 若要使用 Microsoft Entra ID 進行無鑰匙認證,請登入您的 Azure 帳號。 如果你有多個訂閱,請選擇包含你 Azure AI 搜尋服務 服務的訂閱。

    az login
    

執行程式碼

  1. 執行第一個專案來建立並填充索引。

    cd VectorSearchCreatePopulateIndex
    dotnet run
    
  2. 在 VectorSearchExamples/Program.cs中,取消註解你想執行的查詢方法。

  3. 執行第二個專案,針對索引執行這些查詢。

    cd ..\VectorSearchExamples
    dotnet run
    

產出

第一個專案的產出包括索引建立的確認及文件上傳成功。

Creating or updating index 'hotels-vector-quickstart'...
Index 'hotels-vector-quickstart' updated.

Key: 1, Succeeded: True
Key: 2, Succeeded: True
Key: 3, Succeeded: True
Key: 4, Succeeded: True
Key: 48, Succeeded: True
Key: 49, Succeeded: True
Key: 13, Succeeded: True

第二個專案的輸出顯示每個啟用查詢方法的搜尋結果。 以下範例展示了單向量搜尋結果。

Single Vector Search Results:
Score: 0.6605852, HotelId: 48, HotelName: Nordick's Valley Motel
Score: 0.6333684, HotelId: 13, HotelName: Luxury Lion Resort
Score: 0.605672, HotelId: 4, HotelName: Sublime Palace Hotel
Score: 0.6026341, HotelId: 49, HotelName: Swirling Currents Hotel
Score: 0.57902366, HotelId: 2, HotelName: Old Century Hotel

了解程式碼

註

本節的程式碼片段可能已經過修改以提升可讀性。 完整工作範例請參考原始碼。

現在你已經執行過程式碼,讓我們來拆解幾個關鍵步驟:

  1. 建立向量索引
  2. 將文件上傳至索引
  3. 查詢索引

建立向量索引

在你將內容加入 Azure AI 搜尋服務 之前,必須建立索引來定義內容的儲存與結構。

索引架構是圍繞飯店內容組織的。 樣本資料包含虛構旅館的向量與非向量描述。 以下程式碼 建立 VectorSearchCreatePopulateIndex/Program.cs 索引結構,包括向量場 DescriptionVector。

static async Task CreateSearchIndex(string indexName, SearchIndexClient indexClient)
{
    var addressField = new ComplexField("Address");
    addressField.Fields.Add(new SearchableField("StreetAddress") { AnalyzerName = LexicalAnalyzerName.EnMicrosoft });
    addressField.Fields.Add(new SearchableField("City") { AnalyzerName = LexicalAnalyzerName.EnMicrosoft, IsFacetable = true, IsFilterable = true });
    addressField.Fields.Add(new SearchableField("StateProvince") { AnalyzerName = LexicalAnalyzerName.EnMicrosoft, IsFacetable = true, IsFilterable = true });
    addressField.Fields.Add(new SearchableField("PostalCode") { AnalyzerName = LexicalAnalyzerName.EnMicrosoft, IsFacetable = true, IsFilterable = true });
    addressField.Fields.Add(new SearchableField("Country") { AnalyzerName = LexicalAnalyzerName.EnMicrosoft, IsFacetable = true, IsFilterable = true });

    var allFields = new List<SearchField>()
    {
        new SimpleField("HotelId", SearchFieldDataType.String) { IsKey = true, IsFacetable = true, IsFilterable = true },
        new SearchableField("HotelName") { AnalyzerName = LexicalAnalyzerName.EnMicrosoft },
        new SearchableField("Description") { AnalyzerName = LexicalAnalyzerName.EnMicrosoft },
        new VectorSearchField("DescriptionVector", 1536, "my-vector-profile"),
        new SearchableField("Category") { AnalyzerName = LexicalAnalyzerName.EnMicrosoft, IsFacetable = true, IsFilterable = true },
        new SearchableField("Tags", collection: true) { AnalyzerName = LexicalAnalyzerName.EnMicrosoft, IsFacetable = true, IsFilterable = true },
        new SimpleField("ParkingIncluded", SearchFieldDataType.Boolean) { IsFacetable = true, IsFilterable = true },
        new SimpleField("LastRenovationDate", SearchFieldDataType.DateTimeOffset) { IsSortable = true },
        new SimpleField("Rating", SearchFieldDataType.Double) { IsFacetable = true, IsFilterable = true, IsSortable = true },
        addressField,
        new SimpleField("Location", SearchFieldDataType.GeographyPoint) { IsFilterable = true, IsSortable = true },
    };

    // Create the suggester configuration
    var suggester = new SearchSuggester("sg", new[] { "Address/City", "Address/Country" });

    // Create the semantic search
    var semanticSearch = new SemanticSearch()
    {
        Configurations =
        {
            new SemanticConfiguration(
                name: "semantic-config",
                prioritizedFields: new SemanticPrioritizedFields
                {
                    TitleField = new SemanticField("HotelName"),
                    KeywordsFields = { new SemanticField("Category") },
                    ContentFields = { new SemanticField("Description") }
                })
        }
    };

    // Add vector search configuration
    var vectorSearch = new VectorSearch();
    vectorSearch.Algorithms.Add(new HnswAlgorithmConfiguration(name: "my-hnsw-vector-config-1"));
    vectorSearch.Profiles.Add(new VectorSearchProfile(name: "my-vector-profile", algorithmConfigurationName: "my-hnsw-vector-config-1"));

    var definition = new SearchIndex(indexName)
    {
        Fields = allFields,
        Suggesters = { suggester },
        VectorSearch = vectorSearch,
        SemanticSearch = semanticSearch
    };

    // Create or update the index
    Console.WriteLine($"Creating or updating index '{indexName}'...");
    var result = await indexClient.CreateOrUpdateIndexAsync(definition);
    Console.WriteLine($"Index '{result.Value.Name}' updated.");
    Console.WriteLine();
}

重點摘要:

  • 你透過建立欄位清單來定義索引。

  • 此索引支援多種搜尋功能:

  • 第二個 VectorSearchField 參數指定 vectorSearchDimensions,必須與你的嵌入模型輸出大小相符。 這個快速入門使用 1,536 個維度,以符合 text-embedding-ada-002 模型。

  • 此 VectorSearch 配置定義了近似最近鄰(ANN)演算法。 支援的演算法包括 Hierarchical Navigable Small World (HNSW) 與詳盡 K-Nearest Neighbor (KNN)。 欲了解更多資訊,請參閱 向量搜尋中的相關性。

將文件上傳至索引

新建立的索引是空的。 要填充索引並使其可搜尋,您必須上傳符合索引結構的 JSON 文件。

在 Azure AI 搜尋服務 中,文件既是索引的輸入,也是查詢的輸出。 為簡化起見,此快速入門指南提供已預先計算好向量的飯店文件範例。 在生產環境中,內容常從連接的資料來源擷取,並透過 索引器轉換成 JSON。

以下程式碼會將文件 HotelData.json 上傳至您的搜尋服務。

static async Task UploadDocs(SearchClient searchClient)
{
    var jsonPath = Path.Combine(Directory.GetCurrentDirectory(), "HotelData.json");

    // Read and parse hotel data
    var json = await File.ReadAllTextAsync(jsonPath);
    List<Hotel> hotels = new List<Hotel>();
    try
    {
        using var doc = JsonDocument.Parse(json);
        if (doc.RootElement.ValueKind != JsonValueKind.Array)
        {
            Console.WriteLine("HotelData.json root is not a JSON array.");
        }
        // Deserialize all hotel objects
        hotels = doc.RootElement.EnumerateArray()
            .Select(e => JsonSerializer.Deserialize<Hotel>(e.GetRawText()))
            .Where(h => h != null)
            .ToList();
    }
    catch (JsonException ex)
    {
        Console.WriteLine($"Failed to parse HotelData.json: {ex.Message}");
    }

    try
    {
        // Upload hotel documents to Azure Search
        var result = await searchClient.UploadDocumentsAsync(hotels);
        foreach (var r in result.Value.Results)
        {
            Console.WriteLine($"Key: {r.Key}, Succeeded: {r.Succeeded}");
        }
    }
    catch (Exception ex)
    {
        Console.WriteLine("Failed to upload documents: " + ex);
    }
}

你的程式碼會透過 SearchClient 與 Azure AI 搜尋服務 服務中託管的特定搜尋索引互動,這是 Azure.Search.Documents 套件提供的主要物件。 SearchClient 提供索引操作的存取功能,例如:

  • 資料引入:UploadDocuments(),MergeDocuments(),DeleteDocuments()

  • 搜尋操作: Search(), Autocomplete(), Suggest()

查詢索引

這些 VectorSearchExamples 查詢展示了不同的搜尋模式。 範例向量查詢基於兩個字串:

  • 全文搜尋字串: "historic hotel walk to restaurants and shopping"

  • 向量查詢字串: "quintessential lodging near running trails, eateries, retail" (向量化成數學表示)

向量查詢字串在語意上與全文搜尋字串相似,但它包含索引中不存在的詞彙。 僅用關鍵字進行搜尋的向量查詢字串結果為零。 然而,向量搜尋是根據意義而非精確關鍵字尋找相關匹配。

以下範例從基本的向量查詢開始,逐步加入篩選器、關鍵字搜尋及語意重排序。

這個 SearchSingleVector 方法示範了一個基本情境,你想找到與向量查詢字串非常吻合的文件描述。 VectorizedQuery 配置向量搜尋:

  • KNearestNeighborsCount 限制根據向量相似度回傳的結果數量。
  • Fields 指定要搜尋的向量場。
public static async Task SearchSingleVector(SearchClient searchClient, ReadOnlyMemory<float> precalculatedVector)
{
    SearchResults<Hotel> response = await searchClient.SearchAsync<Hotel>(
        new SearchOptions
        {
            VectorSearch = new()
            {
                Queries = { new VectorizedQuery(precalculatedVector) { KNearestNeighborsCount = 5, Fields = { "DescriptionVector" } } }
            },
            Select = { "HotelId", "HotelName", "Description", "Category", "Tags" },
        });

    Console.WriteLine($"Single Vector Search Results:");
    await foreach (SearchResult<Hotel> result in response.GetResultsAsync())
    {
        Hotel doc = result.Document;
        Console.WriteLine($"Score: {result.Score}, HotelId: {doc.HotelId}, HotelName: {doc.HotelName}");
    }
    Console.WriteLine();
}

帶有濾波器的單向量搜尋

在Azure AI 搜尋服務中,filters 適用於索引中的非向量場。 這個 SearchSingleVectorWithFilter 方法會在 Tags 場地上篩選出不提供免費 Wi-Fi 的飯店。

public static async Task SearchSingleVectorWithFilter(SearchClient searchClient, ReadOnlyMemory<float> precalculatedVector)
{
    SearchResults<Hotel> responseWithFilter = await searchClient.SearchAsync<Hotel>(
        new SearchOptions
        {
            VectorSearch = new()
            {
                Queries = { new VectorizedQuery(precalculatedVector) { KNearestNeighborsCount = 5, Fields = { "DescriptionVector" } } }
            },
            Filter = "Tags/any(tag: tag eq 'free wifi')",
            Select = { "HotelId", "HotelName", "Description", "Category", "Tags" }
        });

    Console.WriteLine($"Single Vector Search With Filter Results:");
    await foreach (SearchResult<Hotel> result in responseWithFilter.GetResultsAsync())
    {
        Hotel doc = result.Document;
        Console.WriteLine($"Score: {result.Score}, HotelId: {doc.HotelId}, HotelName: {doc.HotelName}, Tags: {string.Join(String.Empty, doc.Tags)}");
    }
    Console.WriteLine();
}

使用地理濾波器的單向量搜尋

你可以指定地理 空間過濾器 ,限制結果只在特定地理區域內。 此 SingleSearchWithGeoFilter 方法指定一個地理點(華盛頓特區,使用經度與緯度座標),並返回300公里範圍內的飯店。 預設情況下,篩選器會在向量搜尋後執行。

public static async Task SingleSearchWithGeoFilter(SearchClient searchClient, ReadOnlyMemory<float> precalculatedVector)
{
    SearchResults<Hotel> responseWithGeoFilter = await searchClient.SearchAsync<Hotel>(
        new SearchOptions
        {
            VectorSearch = new()
            {
                Queries = { new VectorizedQuery(precalculatedVector) { KNearestNeighborsCount = 5, Fields = { "DescriptionVector" } } }
            },
            Filter = "geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300",
            Select = { "HotelId", "HotelName", "Description", "Address", "Category", "Tags" },
            Facets = { "Address/StateProvince" },
        });

    Console.WriteLine($"Vector query with a geo filter:");
    await foreach (SearchResult<Hotel> result in responseWithGeoFilter.GetResultsAsync())
    {
        Hotel doc = result.Document;
        Console.WriteLine($"HotelId: {doc.HotelId}");
        Console.WriteLine($"HotelName: {doc.HotelName}");
        Console.WriteLine($"Score: {result.Score}");
        Console.WriteLine($"City/State: {doc.Address.City}/{doc.Address.StateProvince}");
        Console.WriteLine($"Description: {doc.Description}");
        Console.WriteLine();
    }
    Console.WriteLine();
}

混合式搜尋 將全文與向量查詢合併於單一請求中。 此 SearchHybridVectorAndText 方法同時執行兩種查詢類型,然後使用互惠排名融合(Reciprocal Rank Fusion,RRF)將結果合併成統一的排名。 RRF 利用每個結果集結果排名的反向來產生合併排名。 請注意,混合搜尋分數普遍低於單一查詢分數。

public static async Task<SearchResults<Hotel>> SearchHybridVectorAndText(SearchClient searchClient, ReadOnlyMemory<float> precalculatedVector)
{
    SearchResults<Hotel> responseWithFilter = await searchClient.SearchAsync<Hotel>(
        "historic hotel walk to restaurants and shopping",
        new SearchOptions
        {
            VectorSearch = new()
            {
                Queries = { new VectorizedQuery(precalculatedVector) { KNearestNeighborsCount = 5, Fields = { "DescriptionVector" } } }
            },
            Select = { "HotelId", "HotelName", "Description", "Category", "Tags" },
            Size = 5,
        });

    Console.WriteLine($"Hybrid search results:");
    await foreach (SearchResult<Hotel> result in responseWithFilter.GetResultsAsync())
    {
        Hotel doc = result.Document;
        Console.WriteLine($"Score: {result.Score}");
        Console.WriteLine($"HotelId: {doc.HotelId}");
        Console.WriteLine($"HotelName: {doc.HotelName}");
        Console.WriteLine($"Description: {doc.Description}");
        Console.WriteLine($"Category: {doc.Category}");
        Console.WriteLine($"Tags: {string.Join(String.Empty, doc.Tags)}");
        Console.WriteLine();
    }
    Console.WriteLine();
    return responseWithFilter;
}

此 SearchHybridVectorAndSemantic 方法展示了語 意排序,即根據語言理解重新排序結果。

public static async Task SearchHybridVectorAndSemantic(SearchClient searchClient, ReadOnlyMemory<float> precalculatedVector)
{
    SearchResults<Hotel> responseWithFilter = await searchClient.SearchAsync<Hotel>(
        "historic hotel walk to restaurants and shopping",
        new SearchOptions
        {
            IncludeTotalCount = true,
            VectorSearch = new()
            {
                Queries = { new VectorizedQuery(precalculatedVector) { KNearestNeighborsCount = 5, Fields = { "DescriptionVector" } } }
            },
            Select = { "HotelId", "HotelName", "Description", "Category", "Tags" },
            SemanticSearch = new SemanticSearchOptions
            {
                SemanticConfigurationName = "semantic-config"
            },
            QueryType = SearchQueryType.Semantic,
            Size = 5
        });

    Console.WriteLine($"Hybrid search results:");
    await foreach (SearchResult<Hotel> result in responseWithFilter.GetResultsAsync())
    {
        Hotel doc = result.Document;
        Console.WriteLine($"Score: {result.Score}");
        Console.WriteLine($"HotelId: {doc.HotelId}");
        Console.WriteLine($"HotelName: {doc.HotelName}");
        Console.WriteLine($"Description: {doc.Description}");
        Console.WriteLine($"Category: {doc.Category}");
        Console.WriteLine();
    }
    Console.WriteLine();
}

將這些結果與前一個查詢的混合搜尋結果做比較。 若無語意重排序,Sublime Palace Hotel 會排名第一,因為互惠排名融合(Reciprocal Rank Fusion,RRF)將文字分數與向量分數合併,產生合併結果。 經過語意調整後,Swirling Currents Hotel 躍升至榜首。

語意排序器利用機器理解模型評估每個結果與查詢意圖的匹配程度。 Swirling Currents Hotel 的描述提到"walking access to shopping, dining, entertainment and the city center",這與搜尋關鍵字"walk to restaurants and shopping"高度吻合。 這種對附近餐飲與購物的語意比對,讓它的排名高於 Sublime Palace Hotel,因為後者在描述中未強調可步行的便利設施。

重點摘要:

  • 在混合式搜尋中,你可以將向量搜尋與關鍵字上的全文搜尋整合起來。 篩選器和語意排序僅適用於文本內容,不適用於向量。

  • 實際結果則包含更多細節,包括語意說明和重點。 此快速入門功能會修改結果以提升可讀性。 要取得完整的回應結構,請使用 REST 執行請求。

清理資源

當您在自己的訂用帳戶中工作時,建議您在完成專案後移除不再需要的資源。 若您讓資源繼續執行,則可能會產生費用。

在Azure入口網站中,從左側窗格選擇 所有資源或 資源群組以尋找並管理資源。 你可以單獨刪除資源,或是一次性刪除資源群組,移除所有資源。

在這個快速入門中,你會使用 Java 的 Azure AI 搜尋服務 用戶端函式庫來建立、載入並查詢 向量索引。 Java 用戶端函式庫在索引操作上提供 REST API 的抽象化。

在 Azure AI 搜尋服務 中,向量索引有一個定義向量與非向量場的索引結構、用於建立嵌入空間的演算法向量搜尋設定,以及查詢時評估的向量場定義設定。 索引 - 建立或更新 (REST API)建立向量索引。

提示

先決條件

設定存取權限

在開始之前,請確認你有權限存取 Azure AI 搜尋服務 的內容和操作。 此快速入門使用 Microsoft Entra ID 進行驗證,並以角色為基礎的存取控制來授權。 您必須是 擁有者 或 使用者存取管理員 才能指派角色。 如果角色不可行,改用 金鑰驗證 。

要設定建議的以角色為基礎的存取權限:

  1. 為您的搜尋服務啟用基於角色的存取權限。

  2. 請將以下角色指派 到你的使用者帳號。

    • 搜尋服務貢獻者

    • 搜尋索引資料貢獻者

    • 搜尋索引資料閱讀器

取得端點

每個Azure AI 搜尋服務服務都有一個 endpoint,這是一個唯一用來識別並提供服務網路存取的網址。 在後面的章節中,你指定這個端點以程式化方式連接到你的搜尋服務。

要取得端點:

  1. 在Azure 入口網站中,進入您的搜尋服務。

  2. 從左側窗格選擇 「概覽」。

  3. 請記下該端點,應看起來像 https://my-service.search.windows.net。

設定環境

  1. 用 Git 複製樣本庫。

    git clone https://github.com/Azure-Samples/azure-search-java-samples
    
  2. 進入快速啟動資料夾。

    cd azure-search-java-samples/quickstart-vector-search
    
  3. 在 src/main/resources/application.properties 中,將 azure.search.endpoint 的佔位值替換成你在 取得端點獲得的 URL。

  4. 安裝相依性。

    mvn clean dependency:copy-dependencies
    

    建置完成後,你應該會在專案目錄看到一個 target/dependency 資料夾。

  5. 若要使用 Microsoft Entra ID 進行無鑰匙認證,請登入您的 Azure 帳號。 如果你有多個訂閱,請選擇包含你 Azure AI 搜尋服務 服務的訂閱。

    az login
    

執行程式碼

  1. 建立向量索引。

    mvn compile exec:java "-Dexec.mainClass=com.example.search.CreateIndex"
    
  2. 載入包含預先計算嵌入的文件。

    mvn compile exec:java "-Dexec.mainClass=com.example.search.UploadDocuments"
    
  3. 執行向量搜尋查詢。

    mvn compile exec:java "-Dexec.mainClass=com.example.search.SearchSingle"
    
  4. (可選)執行額外的查詢變化。

    mvn compile exec:java "-Dexec.mainClass=com.example.search.SearchSingleWithFilter"
    mvn compile exec:java "-Dexec.mainClass=com.example.search.SearchSingleWithFilterGeo"
    mvn compile exec:java "-Dexec.mainClass=com.example.search.SearchHybrid"
    mvn compile exec:java "-Dexec.mainClass=com.example.search.SearchSemanticHybrid"
    

產出

輸出 CreateIndex.java 顯示索引名稱和確認。

Using Azure Search endpoint: https://<search-service-name>.search.windows.net
Using Azure Search index: hotels-vector-quickstart
Creating index...
hotels-vector-quickstart created

輸出 UploadDocuments.java 顯示每個索引文件的成功狀態。

Uploading documents...
Key: 1, Succeeded: true, ErrorMessage: none
Key: 2, Succeeded: true, ErrorMessage: none
Key: 3, Succeeded: true, ErrorMessage: none
Key: 4, Succeeded: true, ErrorMessage: none
Key: 48, Succeeded: true, ErrorMessage: none
Key: 49, Succeeded: true, ErrorMessage: none
Key: 13, Succeeded: true, ErrorMessage: none
Waiting for indexing... Current count: 0
All documents indexed successfully.

輸出 SearchSingle.java 結果顯示依相似度分數排名的向量搜尋結果。

Single Vector search found 5
- HotelId: 48, HotelName: Nordick's Valley Motel, Tags: ["continental breakfast","air conditioning","free wifi"], Score 0.6605852
- HotelId: 13, HotelName: Luxury Lion Resort, Tags: ["bar","concierge","restaurant"], Score 0.6333684
- HotelId: 4, HotelName: Sublime Palace Hotel, Tags: ["concierge","view","air conditioning"], Score 0.605672
- HotelId: 49, HotelName: Swirling Currents Hotel, Tags: ["air conditioning","laundry service","24-hour front desk service"], Score 0.6026341
- HotelId: 2, HotelName: Old Century Hotel, Tags: ["pool","free wifi","air conditioning","concierge"], Score 0.57902366

了解程式碼

註

本節的程式碼片段可能已經過修改以提升可讀性。 完整工作範例請參考原始碼。

現在你已經執行過程式碼,讓我們來拆解幾個關鍵步驟:

  1. 建立向量索引
  2. 將文件上傳至索引
  3. 查詢索引

建立向量索引

在你將內容加入 Azure AI 搜尋服務 之前,必須建立索引來定義內容的儲存與結構。

索引架構是圍繞飯店內容組織的。 樣本資料包含虛構旅館的向量與非向量描述。 以下程式碼 建立 CreateIndex.java 索引結構,包括向量場 DescriptionVector。

// Define fields
List<SearchField> fields = Arrays.asList(
    new SearchField("HotelId", SearchFieldDataType.STRING)
        .setKey(true)
        .setFilterable(true),
    new SearchField("HotelName", SearchFieldDataType.STRING)
        .setSortable(true)
        .setSearchable(true),
    new SearchField("Description", SearchFieldDataType.STRING)
        .setSearchable(true),
    new SearchField("DescriptionVector",
        SearchFieldDataType.collection(SearchFieldDataType.SINGLE))
        .setSearchable(true)
        .setVectorSearchDimensions(1536)
        .setVectorSearchProfileName("my-vector-profile"),
    new SearchField("Category", SearchFieldDataType.STRING)
        .setSortable(true)
        .setFilterable(true)
        .setFacetable(true)
        .setSearchable(true),
    new SearchField("Tags", SearchFieldDataType.collection(
        SearchFieldDataType.STRING))
        .setSearchable(true)
        .setFilterable(true)
        .setFacetable(true),
    // Additional fields: ParkingIncluded, LastRenovationDate, Rating, Address, Location
);

var searchIndex = new SearchIndex(indexName, fields);

// Define vector search configuration
var hnswParams = new HnswParameters()
    .setM(16)
    .setEfConstruction(200)
    .setEfSearch(128);
var hnsw = new HnswAlgorithmConfiguration("hnsw-vector-config");
hnsw.setParameters(hnswParams);

var vectorProfile = new VectorSearchProfile(
    "my-vector-profile",
    "hnsw-vector-config");
var vectorSearch = new VectorSearch()
    .setAlgorithms(Arrays.asList(hnsw))
    .setProfiles(Arrays.asList(vectorProfile));
searchIndex.setVectorSearch(vectorSearch);

// Define semantic configuration
var prioritizedFields = new SemanticPrioritizedFields()
    .setTitleField(new SemanticField("HotelName"))
    .setContentFields(Arrays.asList(new SemanticField("Description")))
    .setKeywordsFields(Arrays.asList(new SemanticField("Category")));
var semanticConfig = new SemanticConfiguration(
    "semantic-config",
    prioritizedFields);
var semanticSearch = new SemanticSearch()
    .setConfigurations(Arrays.asList(semanticConfig));
searchIndex.setSemanticSearch(semanticSearch);

// Define suggesters
var suggester = new SearchSuggester("sg", Arrays.asList("HotelName"));
searchIndex.setSuggesters(Arrays.asList(suggester));

// Create the search index
SearchIndex result = searchIndexClient.createOrUpdateIndex(searchIndex);

重點摘要:

  • 你透過建立欄位清單來定義索引。

  • 此索引支援多種搜尋功能:

  • 這個 setVectorSearchDimensions() 值必須與你的嵌入模型的輸出大小相符。 這個快速入門使用 1,536 個維度,以符合 text-embedding-ada-002 模型。

  • 此 VectorSearch 配置定義了近似最近鄰(ANN)演算法。 支援的演算法包括 Hierarchical Navigable Small World (HNSW) 與詳盡 K-Nearest Neighbor (KNN)。 欲了解更多資訊,請參閱 向量搜尋中的相關性。

將文件上傳至索引

新建立的索引是空的。 要填充索引並使其可搜尋,您必須上傳符合索引結構的 JSON 文件。

在 Azure AI 搜尋服務 中,文件既是索引的輸入,也是查詢的輸出。 為簡化起見,此快速入門指南提供已預先計算好向量的飯店文件範例。 在生產環境中,內容常從連接的資料來源擷取,並透過 索引器轉換成 JSON。

以下程式碼 UploadDocuments.java 將文件上傳至您的搜尋服務。

// Documents contain hotel data with 1536-dimension vectors for DescriptionVector
static final List<Map<String, Object>> DOCUMENTS = Arrays.asList(
    new HashMap<>() {{
        put("@search.action", "mergeOrUpload");
        put("HotelId", "1");
        put("HotelName", "Stay-Kay City Hotel");
        put("Description", "This classic hotel is fully-refurbished...");
        put("DescriptionVector", Arrays.asList(/* 1536 float values */));
        put("Category", "Boutique");
        put("Tags", Arrays.asList("view", "air conditioning", "concierge"));
        // Additional fields...
    }}
    // Additional hotel documents
);

// Upload documents to the index
IndexDocumentsResult result = searchClient.uploadDocuments(DOCUMENTS);
for (IndexingResult r : result.getResults()) {
    System.out.println("Key: %s, Succeeded: %s".formatted(r.getKey(), r.isSucceeded()));
}

你的程式碼會透過 SearchClient 與 Azure AI 搜尋服務 服務中託管的特定搜尋索引互動,這是 azure-search-documents 套件提供的主要物件。 SearchClient 會提供對作業的存取權,例如:

  • 資料引入:uploadDocuments,mergeDocuments,deleteDocuments

  • 搜尋操作: search, autocomplete, suggest

查詢索引

搜尋檔案中的查詢顯示出不同的搜尋模式。 範例向量查詢基於兩個字串:

  • 全文搜尋字串: "historic hotel walk to restaurants and shopping"

  • 向量查詢字串: "quintessential lodging near running trails, eateries, retail" (向量化成數學表示)

向量查詢字串在語意上與全文搜尋字串相似,但它包含索引中不存在的詞彙。 僅用關鍵字進行搜尋的向量查詢字串結果為零。 然而,向量搜尋是根據意義而非精確關鍵字尋找相關匹配。

以下範例從基本的向量查詢開始,逐步加入篩選器、關鍵字搜尋及語意重排序。

SearchSingle.java 示範一個基本情境,你想找到與向量查詢字串非常吻合的文件描述。 VectorizedQuery 配置向量搜尋:

  • setKNearestNeighborsCount() 限制根據向量相似度回傳的結果數量。
  • setFields() 指定要搜尋的向量場。
var vectorQuery = new VectorizedQuery(QueryVector.getVectorList())
    .setKNearestNeighborsCount(5)
    .setFields("DescriptionVector")
    .setExhaustive(true);

var vectorSearchOptions = new VectorSearchOptions()
    .setQueries(vectorQuery)
    .setFilterMode(VectorFilterMode.POST_FILTER);

var searchOptions = new SearchOptions()
    .setTop(7)
    .setIncludeTotalCount(true)
    .setSelect("HotelId", "HotelName", "Description", "Category", "Tags")
    .setVectorSearchOptions(vectorSearchOptions);

var results = searchClient.search("*", searchOptions, Context.NONE);

for (SearchResult result : results) {
    SearchDocument document = result.getDocument(SearchDocument.class);
    System.out.println("HotelId: %s, HotelName: %s, Score: %s".formatted(
        document.get("HotelId"), document.get("HotelName"), result.getScore()));
}

帶有濾波器的單向量搜尋

在Azure AI 搜尋服務中,filters 適用於索引中的非向量場。 SearchSingleWithFilter.java 在 Tags 欄位上設置過濾器,以過濾掉不提供免費 Wi-Fi 的飯店。

var vectorQuery = new VectorizedQuery(QueryVector.getVectorList())
    .setKNearestNeighborsCount(5)
    .setFields("DescriptionVector")
    .setExhaustive(true);

var vectorSearchOptions = new VectorSearchOptions()
    .setQueries(vectorQuery)
    .setFilterMode(VectorFilterMode.POST_FILTER);

// Add filter for "free wifi" tag
var searchOptions = new SearchOptions()
    .setTop(7)
    .setIncludeTotalCount(true)
    .setSelect("HotelId", "HotelName", "Description", "Category", "Tags")
    .setFilter("Tags/any(tag: tag eq 'free wifi')")
    .setVectorSearchOptions(vectorSearchOptions);

var results = searchClient.search("*", searchOptions, Context.NONE);

使用地理濾波器的單向量搜尋

你可以指定地理 空間過濾器 ,限制結果只在特定地理區域內。 SearchSingleWithGeoFilter.java 指定地理點(華盛頓特區,使用經度與緯度座標),並返回300公里範圍內的飯店。 setFilterMode方法在VectorSearchOptions被調用,以決定濾波器的運行時間。 在這種情況下,POST_FILTER 在向量搜索後運行過濾器。

var searchOptions = new SearchOptions()
    .setTop(5)
    .setIncludeTotalCount(true)
    .setSelect("HotelId", "HotelName", "Category", "Description",
               "Address/City", "Address/StateProvince")
    .setFacets("Address/StateProvince")
    .setFilter("geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300")
    .setVectorSearchOptions(vectorSearchOptions);

混合式搜尋 將全文與向量查詢合併於單一請求中。 SearchHybrid.java 同時執行兩種查詢類型,然後使用互惠排名融合(Reciprocal Rank Fusion,RRF)將結果合併成統一的排名。 RRF 利用每個結果集結果排名的反向來產生合併排名。 請注意,混合搜尋分數普遍低於單一查詢分數。

var vectorQuery = new VectorizedQuery(QueryVector.getVectorList())
    .setKNearestNeighborsCount(5)
    .setFields("DescriptionVector")
    .setExhaustive(true);

var vectorSearchOptions = new VectorSearchOptions()
    .setQueries(vectorQuery)
    .setFilterMode(VectorFilterMode.POST_FILTER);

var searchOptions = new SearchOptions()
    .setTop(5)
    .setIncludeTotalCount(true)
    .setSelect("HotelId", "HotelName", "Description", "Category", "Tags")
    .setVectorSearchOptions(vectorSearchOptions);

// Pass both text query and vector search options
var results = searchClient.search(
    "historic hotel walk to restaurants and shopping",
    searchOptions, Context.NONE);

SearchSemanticHybrid.java 示範語義排序,根據語言理解重新排序結果。

var vectorQuery = new VectorizedQuery(QueryVector.getVectorList())
    .setKNearestNeighborsCount(5)
    .setFields("DescriptionVector")
    .setExhaustive(true);

var vectorSearchOptions = new VectorSearchOptions()
    .setQueries(vectorQuery)
    .setFilterMode(VectorFilterMode.POST_FILTER);

SemanticSearchOptions semanticSearchOptions = new SemanticSearchOptions()
    .setSemanticConfigurationName("semantic-config");

var searchOptions = new SearchOptions()
    .setTop(5)
    .setIncludeTotalCount(true)
    .setSelect("HotelId", "HotelName", "Category", "Description")
    .setQueryType(QueryType.SEMANTIC)
    .setSemanticSearchOptions(semanticSearchOptions)
    .setVectorSearchOptions(vectorSearchOptions);

var results = searchClient.search(
    "historic hotel walk to restaurants and shopping",
    searchOptions, Context.NONE);

將這些結果與前一個查詢的混合搜尋結果做比較。 若無語意重排序,Sublime Palace Hotel 會排名第一,因為互惠排名融合(Reciprocal Rank Fusion,RRF)將文字分數與向量分數合併,產生合併結果。 經過語意調整後,Swirling Currents Hotel 躍升至榜首。

語意排序器利用機器理解模型評估每個結果與查詢意圖的匹配程度。 Swirling Currents Hotel 的描述提到"walking access to shopping, dining, entertainment and the city center",這與搜尋關鍵字"walk to restaurants and shopping"高度吻合。 這種對附近餐飲與購物的語意比對,讓它的排名高於 Sublime Palace Hotel,因為後者在描述中未強調可步行的便利設施。

重點摘要:

  • 在混合式搜尋中,你可以將向量搜尋與關鍵字上的全文搜尋整合起來。 篩選器和語意排序僅適用於文本內容,不適用於向量。

  • 實際結果則包含更多細節,包括語意說明和重點。 此快速入門功能會修改結果以提升可讀性。 要取得完整的回應結構,請使用 REST 執行請求。

清理資源

當您在自己的訂用帳戶中工作時,建議您在完成專案後移除不再需要的資源。 若您讓資源繼續執行,則可能會產生費用。

在Azure入口網站中,從左側窗格選擇 所有資源或 資源群組以尋找並管理資源。 你可以單獨刪除資源,或是一次性刪除資源群組,移除所有資源。

否則,請執行以下指令刪除你在此快速啟動中建立的索引。

mvn compile exec:java "-Dexec.mainClass=com.example.search.DeleteIndex"

在這個快速入門中,你會使用 Azure AI 搜尋服務 用戶端函式庫>建立、載入並查詢 向量索引。 JavaScript 用戶端函式庫提供對 REST API 的抽象化,用於索引操作。

在 Azure AI 搜尋服務 中,向量索引有一個定義向量與非向量場的索引結構、用於建立嵌入空間的演算法向量搜尋設定,以及查詢時評估的向量場定義設定。 索引 - 建立或更新 (REST API)建立向量索引。

提示

先決條件

設定存取權限

在開始之前,請確認你有權限存取 Azure AI 搜尋服務 的內容和操作。 此快速入門使用 Microsoft Entra ID 進行驗證,並以角色為基礎的存取控制來授權。 您必須是 擁有者 或 使用者存取管理員 才能指派角色。 如果角色不可行,改用 金鑰驗證 。

要設定建議的以角色為基礎的存取權限:

  1. 為您的搜尋服務啟用基於角色的存取權限。

  2. 請將以下角色指派 到你的使用者帳號。

    • 搜尋服務貢獻者

    • 搜尋索引資料貢獻者

    • 搜尋索引資料閱讀器

取得端點

每個Azure AI 搜尋服務服務都有一個 endpoint,這是一個唯一用來識別並提供服務網路存取的網址。 在後面的章節中,你指定這個端點以程式化方式連接到你的搜尋服務。

要取得端點:

  1. 在Azure 入口網站中,進入您的搜尋服務。

  2. 從左側窗格選擇 「概覽」。

  3. 請記下該端點,應看起來像 https://my-service.search.windows.net。

設定環境

  1. 用 Git 複製樣本庫。

    git clone https://github.com/Azure-Samples/azure-search-javascript-samples
    
  2. 進入快速啟動資料夾。

    cd azure-search-javascript-samples/quickstart-vector-js
    
  3. 在 sample.env 中,將 AZURE_SEARCH_ENDPOINT 的佔位值替換成你在 取得端點獲得的 URL。

  4. 重新命名 sample.env 為 .env。

    mv sample.env .env
    
  5. 安裝相依性。

    npm install
    

    安裝完成後,你應該會在專案目錄看到一個 node_modules 資料夾。

  6. 若要使用 Microsoft Entra ID 進行無鑰匙認證,請登入您的 Azure 帳號。 如果你有多個訂閱,請選擇包含你 Azure AI 搜尋服務 服務的訂閱。

    az login
    

執行程式碼

  1. 建立向量索引。

    node -r dotenv/config src/createIndex.js
    
  2. 載入包含預先計算嵌入的文件。

    node -r dotenv/config src/uploadDocuments.js
    
  3. 執行向量搜尋查詢。

    node -r dotenv/config src/searchSingle.js
    
  4. (可選)執行額外的查詢變化。

    node -r dotenv/config src/searchSingleWithFilter.js
    node -r dotenv/config src/searchSingleWithFilterGeo.js
    node -r dotenv/config src/searchHybrid.js
    node -r dotenv/config src/searchSemanticHybrid.js
    

產出

輸出 createIndex.js 顯示索引名稱和確認。

Using Azure Search endpoint: https://<search-service-name>.search.windows.net
Using Azure Search index: hotels-vector-quickstart
Creating index...
hotels-vector-quickstart created

輸出 uploadDocuments.js 顯示每個索引文件的成功狀態。

Uploading documents...
Key: 1, Succeeded: true, ErrorMessage: none
Key: 2, Succeeded: true, ErrorMessage: none
Key: 3, Succeeded: true, ErrorMessage: none
Key: 4, Succeeded: true, ErrorMessage: none
Key: 48, Succeeded: true, ErrorMessage: none
Key: 49, Succeeded: true, ErrorMessage: none
Key: 13, Succeeded: true, ErrorMessage: none
Waiting for indexing... Current count: 0
All documents indexed successfully.

輸出 searchSingle.js 結果顯示依相似度分數排名的向量搜尋結果。

Single Vector search found 5
- HotelId: 48, HotelName: Nordick's Valley Motel, Tags: ["continental breakfast","air conditioning","free wifi"], Score 0.6605852
- HotelId: 13, HotelName: Luxury Lion Resort, Tags: ["bar","concierge","restaurant"], Score 0.6333684
- HotelId: 4, HotelName: Sublime Palace Hotel, Tags: ["concierge","view","air conditioning"], Score 0.605672
- HotelId: 49, HotelName: Swirling Currents Hotel, Tags: ["air conditioning","laundry service","24-hour front desk service"], Score 0.6026341
- HotelId: 2, HotelName: Old Century Hotel, Tags: ["pool","free wifi","air conditioning","concierge"], Score 0.57902366

了解程式碼

註

本節的程式碼片段可能已經過修改以提升可讀性。 完整工作範例請參考原始碼。

現在你已經執行過程式碼,讓我們來拆解幾個關鍵步驟:

  1. 建立向量索引
  2. 將文件上傳至索引
  3. 查詢索引

建立向量索引

在你將內容加入 Azure AI 搜尋服務 之前,必須建立索引來定義內容的儲存與結構。

索引架構是圍繞飯店內容組織的。 樣本資料包含虛構旅館的向量與非向量描述。 以下程式碼 建立 createIndex.js 索引結構,包括向量場 DescriptionVector。

const searchFields = [
    { name: "HotelId", type: "Edm.String", key: true, sortable: true, filterable: true, facetable: true },
    { name: "HotelName", type: "Edm.String", searchable: true, filterable: true },
    { name: "Description", type: "Edm.String", searchable: true },
    {
        name: "DescriptionVector",
        type: "Collection(Edm.Single)",
        searchable: true,
        vectorSearchDimensions: 1536,
        vectorSearchProfileName: "vector-profile"
    },
    { name: "Category", type: "Edm.String", filterable: true, facetable: true },
    { name: "Tags", type: "Collection(Edm.String)", filterable: true },
    // Additional fields: ParkingIncluded, LastRenovationDate, Rating, Address, Location
];

const vectorSearch = {
    profiles: [
        {
            name: "vector-profile",
            algorithmConfigurationName: "vector-search-algorithm"
        }
    ],
    algorithms: [
        {
            name: "vector-search-algorithm",
            kind: "hnsw",
            parameters: { m: 4, efConstruction: 400, efSearch: 1000, metric: "cosine" }
        }
    ]
};

const semanticSearch = {
    configurations: [
        {
            name: "semantic-config",
            prioritizedFields: {
                contentFields: [{ name: "Description" }],
                keywordsFields: [{ name: "Category" }],
                titleField: { name: "HotelName" }
            }
        }
    ]
};

const searchIndex = {
    name: indexName,
    fields: searchFields,
    vectorSearch: vectorSearch,
    semanticSearch: semanticSearch,
    suggesters: [{ name: "sg", searchMode: "analyzingInfixMatching", sourceFields: ["HotelName"] }]
};

const result = await indexClient.createOrUpdateIndex(searchIndex);

重點摘要:

  • 你透過建立欄位清單來定義索引。

  • 此索引支援多種搜尋功能:

  • 屬性 vectorSearchDimensions 必須與你嵌入模型的輸出大小相符。 這個快速入門使用 1,536 個維度,以符合 text-embedding-ada-002 模型。

  • 此 vectorSearch 配置定義了近似最近鄰(ANN)演算法。 支援的演算法包括 Hierarchical Navigable Small World (HNSW) 與詳盡 K-Nearest Neighbor (KNN)。 欲了解更多資訊,請參閱 向量搜尋中的相關性。

將文件上傳至索引

新建立的索引是空的。 要填充索引並使其可搜尋,您必須上傳符合索引結構的 JSON 文件。

在 Azure AI 搜尋服務 中,文件既是索引的輸入,也是查詢的輸出。 為簡化起見,此快速入門指南提供已預先計算好向量的飯店文件範例。 在生產環境中,內容常從連接的資料來源擷取,並透過 索引器轉換成 JSON。

以下程式碼 uploadDocuments.js 將文件上傳至您的搜尋服務。

const DOCUMENTS = [
    // Array of hotel documents with embedded 1536-dimension vectors
    // Each document contains: HotelId, HotelName, Description, DescriptionVector,
    // Category, Tags, ParkingIncluded, LastRenovationDate, Rating, Address, Location
];

const searchClient = new SearchClient(searchEndpoint, indexName, credential);

const result = await searchClient.uploadDocuments(DOCUMENTS);
for (const r of result.results) {
    console.log(`Key: ${r.key}, Succeeded: ${r.succeeded}`);
}

你的程式碼會透過 SearchClient 與 Azure AI 搜尋服務 服務中託管的特定搜尋索引互動,這是 @azure/search-documents 套件提供的主要物件。 SearchClient 提供索引操作的存取功能,例如:

  • 資料引入:uploadDocuments,mergeDocuments,deleteDocuments

  • 搜尋操作: search, autocomplete, suggest

查詢索引

搜尋檔案中的查詢顯示出不同的搜尋模式。 範例向量查詢基於兩個字串:

  • 全文搜尋字串: "historic hotel walk to restaurants and shopping"

  • 向量查詢字串: "quintessential lodging near running trails, eateries, retail" (向量化成數學表示)

向量查詢字串在語意上與全文搜尋字串相似,但它包含索引中不存在的詞彙。 僅用關鍵字進行搜尋的向量查詢字串結果為零。 然而,向量搜尋是根據意義而非精確關鍵字尋找相關匹配。

以下範例從基本的向量查詢開始,逐步加入篩選器、關鍵字搜尋及語意重排序。

searchSingle.js 示範一個基本情境,你想找到與向量查詢字串非常吻合的文件描述。 物件 vectorQuery 配置向量搜尋:

  • kNearestNeighborsCount 限制根據向量相似度回傳的結果數量。
  • fields 指定要搜尋的向量場。
const vectorQuery = {
    vector: vector,
    kNearestNeighborsCount: 5,
    fields: ["DescriptionVector"],
    kind: "vector",
    exhaustive: true
};

const searchOptions = {
    top: 7,
    select: ["HotelId", "HotelName", "Description", "Category", "Tags"],
    includeTotalCount: true,
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

const results = await searchClient.search("*", searchOptions);

for await (const result of results.results) {
    const doc = result.document;
    console.log(`HotelId: ${doc.HotelId}, HotelName: ${doc.HotelName}, Score: ${result.score}`);
}

帶有濾波器的單向量搜尋

在Azure AI 搜尋服務中,filters 適用於索引中的非向量場。 searchSingleWithFilter.js 在 Tags 欄位上設置過濾器,以過濾掉不提供免費 Wi-Fi 的飯店。

const searchOptions = {
    top: 7,
    select: ["HotelId", "HotelName", "Description", "Category", "Tags"],
    includeTotalCount: true,
    filter: "Tags/any(tag: tag eq 'free wifi')", // Adding filter for "free wifi" tag
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

const results = await searchClient.search("*", searchOptions);

使用地理濾波器的單向量搜尋

你可以指定地理 空間過濾器 ,限制結果只在特定地理區域內。 searchSingleWithGeoFilter.js 指定地理點(華盛頓特區,使用經度與緯度座標),並返回300公里範圍內的飯店。 filterMode 屬性決定過濾器執行的時間。 在這種情況下,postFilter 在向量搜索後運行過濾器。

const searchOptions = {
    top: 5,
    includeTotalCount: true,
    select: ["HotelId", "HotelName", "Category", "Description", "Address/City", "Address/StateProvince"],
    facets: ["Address/StateProvince"],
    filter: "geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300",
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

混合式搜尋 將全文與向量查詢合併於單一請求中。 searchHybrid.js 同時執行兩種查詢類型,然後使用互惠排名融合(Reciprocal Rank Fusion,RRF)將結果合併成統一的排名。 RRF 利用每個結果集結果排名的反向來產生合併排名。 請注意,混合搜尋分數普遍低於單一查詢分數。

const vectorQuery = {
    vector: vector,
    kNearestNeighborsCount: 5,
    fields: ["DescriptionVector"],
    kind: "vector",
    exhaustive: true
};

const searchOptions = {
    top: 5,
    includeTotalCount: true,
    select: ["HotelId", "HotelName", "Description", "Category", "Tags"],
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

// Use search text for keyword search (hybrid search = vector + keyword)
const searchText = "historic hotel walk to restaurants and shopping";
const results = await searchClient.search(searchText, searchOptions);

searchSemanticHybrid.js 示範語義排序,根據語言理解重新排序結果。

const searchOptions = {
    top: 5,
    includeTotalCount: true,
    select: ["HotelId", "HotelName", "Category", "Description"],
    queryType: "semantic",
    semanticSearchOptions: {
        configurationName: "semantic-config"
    },
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

const searchText = "historic hotel walk to restaurants and shopping";
const results = await searchClient.search(searchText, searchOptions);

for await (const result of results.results) {
    console.log(`Score: ${result.score}, Re-ranker Score: ${result.rerankerScore}`);
}

將這些結果與前一個查詢的混合搜尋結果做比較。 若無語意重排序,Sublime Palace Hotel 會排名第一,因為互惠排名融合(Reciprocal Rank Fusion,RRF)將文字分數與向量分數合併,產生合併結果。 經過語意調整後,Swirling Currents Hotel 躍升至榜首。

語意排序器利用機器理解模型評估每個結果與查詢意圖的匹配程度。 Swirling Currents Hotel 的描述提到"walking access to shopping, dining, entertainment and the city center",這與搜尋關鍵字"walk to restaurants and shopping"高度吻合。 這種對附近餐飲與購物的語意比對,讓它的排名高於 Sublime Palace Hotel,因為後者在描述中未強調可步行的便利設施。

重點摘要:

  • 在混合式搜尋中,你可以將向量搜尋與關鍵字上的全文搜尋整合起來。 篩選器和語意排序僅適用於文本內容,不適用於向量。

  • 實際結果則包含更多細節,包括語意說明和重點。 此快速入門功能會修改結果以提升可讀性。 要取得完整的回應結構,請使用 REST 執行請求。

清理資源

當您在自己的訂用帳戶中工作時,建議您在完成專案後移除不再需要的資源。 若您讓資源繼續執行,則可能會產生費用。

在Azure入口網站中,從左側窗格選擇 所有資源或 資源群組以尋找並管理資源。 你可以單獨刪除資源,或是一次性刪除資源群組,移除所有資源。

否則,請執行以下指令刪除你在此快速啟動中建立的索引。

node -r dotenv/config src/deleteIndex.js

在這個快速入門中,你會使用 Python 的 Azure AI 搜尋服務 用戶端函式庫來建立、載入並查詢 向量索引。 Python 用戶端函式庫在索引操作上提供 REST API 的抽象化。

在 Azure AI 搜尋服務 中,向量索引有一個定義向量與非向量場的索引結構、用於建立嵌入空間的演算法向量搜尋設定,以及查詢時評估的向量場定義設定。 索引 - 建立或更新 (REST API)建立向量索引。

提示

先決條件

  • 一個有有效訂閱的 Azure 帳號。 免費註冊帳號。

  • 一個Azure AI 搜尋服務服務。 你可以使用免費方案來完成大部分快速入門,但我們建議對於較大的資料檔案,使用基本版或更高等級。

  • Python 3.8或更高。

  • Visual Studio Code,並搭配Python及Jupyter擴充功能。

  • git 來複製樣本庫。

  • Azure CLI 用於與 Microsoft Entra ID 的無金鑰認證。

設定存取權限

在開始之前,請確認你有權限存取 Azure AI 搜尋服務 的內容和操作。 此快速入門使用 Microsoft Entra ID 進行驗證,並以角色為基礎的存取控制來授權。 您必須是 擁有者 或 使用者存取管理員 才能指派角色。 如果角色不可行,改用 金鑰驗證 。

要設定建議的以角色為基礎的存取權限:

  1. 為您的搜尋服務啟用基於角色的存取權限。

  2. 請將以下角色指派 到你的使用者帳號。

    • 搜尋服務貢獻者

    • 搜尋索引資料貢獻者

    • 搜尋索引資料閱讀器

取得端點

每個Azure AI 搜尋服務服務都有一個 endpoint,這是一個唯一用來識別並提供服務網路存取的網址。 在後面的章節中,你指定這個端點以程式化方式連接到你的搜尋服務。

要取得端點:

  1. 在Azure 入口網站中,進入您的搜尋服務。

  2. 從左側窗格選擇 「概覽」。

  3. 請記下該端點,應看起來像 https://my-service.search.windows.net。

設定環境

  1. 用 Git 複製樣本庫。

    git clone https://github.com/Azure-Samples/azure-search-python-samples
    
  2. 進入快速啟動資料夾,在 Visual Studio Code 中開啟。

    cd azure-search-python-samples/Quickstart-Vector-Search
    code .
    
  3. 在 sample.env 中,將 AZURE_SEARCH_ENDPOINT 的佔位值替換成你在 取得端點獲得的 URL。

  4. 重新命名 sample.env 為 .env。

    mv sample.env .env
    
  5. 打開 vector-search-quickstart.ipynb。

  6. 按 Ctrl+Shift+P,選擇 筆記本:選擇筆記本核心,並依照指示建立虛擬環境。 選擇 requirements.txt 來指定相依性。

    完成後,你應該會在專案目錄中看到一個 .venv 資料夾。

  7. 若要使用 Microsoft Entra ID 進行無鑰匙認證,請登入您的 Azure 帳號。 如果你有多個訂閱,請選擇包含你 Azure AI 搜尋服務 服務的訂閱。

    az login
    

執行程式碼

  1. 執行 Install packages and set variables 儲存單元來安裝所需的套件並載入環境變數。

  2. 依序執行剩餘儲存格以建立向量索引、上傳文件,並執行不同類型的向量查詢。

產出

每個程式碼區塊會將其輸出顯示在筆記本上。 以下範例顯示了由 Single vector search 所產生的輸出,其依據相似度分數排列出向量搜尋結果。

Total results: 7
- HotelId: 48, HotelName: Nordick's Valley Motel, Category: Boutique
- HotelId: 13, HotelName: Luxury Lion Resort, Category: Luxury
- HotelId: 4, HotelName: Sublime Palace Hotel, Category: Boutique
- HotelId: 49, HotelName: Swirling Currents Hotel, Category: Suite
- HotelId: 2, HotelName: Old Century Hotel, Category: Boutique

了解程式碼

註

本節的程式碼片段可能已經過修改以提升可讀性。 完整工作範例請參考原始碼。

現在你已經執行過程式碼,讓我們來拆解幾個關鍵步驟:

  1. 建立向量索引
  2. 將文件上傳至索引
  3. 查詢索引

建立向量索引

在你將內容加入 Azure AI 搜尋服務 之前,必須建立索引來定義內容的儲存與結構。

索引架構是圍繞飯店內容組織的。 樣本資料包含虛構旅館的向量與非向量描述。 Create an index筆記本中的儲存格會建立索引結構,包括向量字段DescriptionVector。

fields = [
    SimpleField(name="HotelId", type=SearchFieldDataType.String, key=True, filterable=True),
    SearchableField(name="HotelName", type=SearchFieldDataType.String, sortable=True),
    SearchableField(name="Description", type=SearchFieldDataType.String),
    SearchField(
        name="DescriptionVector",
        type=SearchFieldDataType.Collection(SearchFieldDataType.Single),
        searchable=True,
        vector_search_dimensions=1536,
        vector_search_profile_name="my-vector-profile"
    ),
    SearchableField(name="Category", type=SearchFieldDataType.String, sortable=True, filterable=True, facetable=True),
    SearchField(name="Tags", type=SearchFieldDataType.Collection(SearchFieldDataType.String), searchable=True, filterable=True, facetable=True),
    # Additional fields omitted for brevity
]

重點摘要:

  • 你透過建立欄位清單來定義索引。 每個欄位皆透過一個輔助方法建立,該方法定義欄位類型及其設定。

  • 此索引支援多種搜尋功能:

  • 屬性 vector_search_dimensions 必須與你嵌入模型的輸出大小相符。 這個快速入門使用 1,536 個維度,以符合 text-embedding-ada-002 模型。

  • 此 VectorSearch 配置定義了近似最近鄰(ANN)演算法。 支援的演算法包括 Hierarchical Navigable Small World (HNSW) 與詳盡 K-Nearest Neighbor (KNN)。 欲了解更多資訊,請參閱 向量搜尋中的相關性。

將文件上傳至索引

新建立的索引是空的。 要填充索引並使其可搜尋,您必須上傳符合索引結構的 JSON 文件。

在 Azure AI 搜尋服務 中,文件既是索引的輸入,也是查詢的輸出。 為簡化起見,此快速入門指南提供已預先計算好向量的飯店文件範例。 在生產環境中,內容常從連接的資料來源擷取,並透過 索引器轉換成 JSON。

Create documents payload 和 Upload the documents 這兩個 cell 將文件載入索引中。

documents = [
    # List of hotel documents with embedded 1536-dimension vectors
    # Each document contains: HotelId, HotelName, Description, DescriptionVector,
    # Category, Tags, ParkingIncluded, LastRenovationDate, Rating, Address, Location
]

search_client = SearchClient(
    endpoint=search_endpoint,
    index_name=index_name,
    credential=credential
)

result = search_client.upload_documents(documents=documents)
for r in result:
    print(f"Key: {r.key}, Succeeded: {r.succeeded}, ErrorMessage: {r.error_message}")

你的程式碼會透過 SearchClient 與 Azure AI 搜尋服務 服務中託管的特定搜尋索引互動,這是 azure-search-documents 套件提供的主要物件。 SearchClient 提供索引操作的存取功能,例如:

  • 資料引入:upload_documents(),merge_documents(),delete_documents()

  • 搜尋操作: search(), autocomplete(), suggest()

查詢索引

筆記本中的查詢顯示出不同的搜尋模式。 範例向量查詢基於兩個字串:

  • 全文搜尋字串: "historic hotel walk to restaurants and shopping"

  • 向量查詢字串: "quintessential lodging near running trails, eateries, retail" (向量化成數學表示)

向量查詢字串在語意上與全文搜尋字串相似,但它包含索引中不存在的詞彙。 僅用關鍵字進行搜尋的向量查詢字串結果為零。 然而,向量搜尋是根據意義而非精確關鍵字尋找相關匹配。

以下範例從基本的向量查詢開始,逐步加入篩選器、關鍵字搜尋及語意重排序。

這個 Single vector search 儲存格展示了一個基本情境,你想找到與向量查詢字串非常吻合的文件描述。 VectorizedQuery 配置向量搜尋:

  • k_nearest_neighbors 限制根據向量相似度回傳的結果數量。
  • fields 指定要搜尋的向量場。
vector_query = VectorizedQuery(
    vector=vector,
    k_nearest_neighbors=5,
    fields="DescriptionVector",
    kind="vector",
    exhaustive=True
)

results = search_client.search(
    vector_queries=[vector_query],
    select=["HotelId", "HotelName", "Description", "Category", "Tags"],
    top=5,
    include_total_count=True
)

帶有濾波器的單向量搜尋

在Azure AI 搜尋服務中,filters 適用於索引中的非向量場。 Single vector search with filter 儲存格會篩選 Tags 欄位,以排除未提供免費 Wi-Fi 的所有旅館。

# vector_query omitted for brevity

results = search_client.search(
    vector_queries=[vector_query],
    filter="Tags/any(tag: tag eq 'free wifi')",
    select=["HotelId", "HotelName", "Description", "Category", "Tags"],
    top=7,
    include_total_count=True
)

使用地理濾波器的單向量搜尋

你可以指定地理 空間過濾器 ,限制結果只在特定地理區域內。 該 Single vector search with geo filter 儲存格指定地理點(華盛頓特區,使用經度與緯度座標),並返回300公里範圍內的飯店。 參數 vector_filter_mode 決定過濾器運行的時間。 在這種情況下,postFilter 在向量搜索後運行過濾器。

# vector_query omitted for brevity

results = search_client.search(
    include_total_count=True,
    top=5,
    select=[
        "HotelId", "HotelName", "Category", "Description", "Address/City", "Address/StateProvince"
    ],
    facets=["Address/StateProvince"],
    filter="geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300",
    vector_filter_mode="postFilter",
    vector_queries=[vector_query]
)

混合式搜尋 將全文與向量查詢合併於單一請求中。 該 Hybrid search 單元同時執行兩種查詢類型,然後使用互惠排名融合(Reciprocal Rank Fusion,RRF)將結果合併成統一的排名。 RRF 利用每個結果集結果排名的反向來產生合併排名。 請注意,混合搜尋分數普遍低於單一查詢分數。

# vector_query omitted for brevity

results = search_client.search(
    search_text="historic hotel walk to restaurants and shopping",
    vector_queries=[vector_query],
    select=["HotelId", "HotelName", "Description", "Category", "Tags"],
    top=5,
    include_total_count=True
)

該 Semantic hybrid search 單元展示了 語意排序,根據語言理解重新排序結果。

# vector_query omitted for brevity

results = search_client.search(
    search_text="historic hotel walk to restaurants and shopping",
    vector_queries=[vector_query],
    select=["HotelId", "HotelName", "Category", "Description"],
    query_type="semantic",
    semantic_configuration_name="my-semantic-config",
    top=5,
    include_total_count=True
)

將這些結果與前一個查詢的混合搜尋結果做比較。 若無語意重排序,Sublime Palace Hotel 會排名第一,因為互惠排名融合(Reciprocal Rank Fusion,RRF)將文字分數與向量分數合併,產生合併結果。 經過語意調整後,Swirling Currents Hotel 躍升至榜首。

語意排序器利用機器理解模型評估每個結果與查詢意圖的匹配程度。 Swirling Currents Hotel 的描述提到"walking access to shopping, dining, entertainment and the city center",這與搜尋關鍵字"walk to restaurants and shopping"高度吻合。 這種對附近餐飲與購物的語意比對,讓它的排名高於 Sublime Palace Hotel,因為後者在描述中未強調可步行的便利設施。

重點摘要:

  • 在混合式搜尋中,你可以將向量搜尋與關鍵字上的全文搜尋整合起來。 篩選器和語意排序僅適用於文本內容,不適用於向量。

  • 實際結果則包含更多細節,包括語意說明和重點。 此快速入門功能會修改結果以提升可讀性。 要取得完整的回應結構,請使用 REST 執行請求。

清理資源

當您在自己的訂用帳戶中工作時,建議您在完成專案後移除不再需要的資源。 若您讓資源繼續執行,則可能會產生費用。

在Azure入口網站中,從左側窗格選擇 所有資源或 資源群組以尋找並管理資源。 你可以單獨刪除資源,或是一次性刪除資源群組,移除所有資源。

否則,你可以執行 Clean up 程式碼儲存格刪除你在這個快速入門中建立的索引。

在這個快速入門中,你會使用 JavaScript 的 Azure AI 搜尋服務 用戶端函式庫(與 TypeScript 相容)來建立、載入並查詢 向量索引。 JavaScript 用戶端函式庫提供對 REST API 的抽象化,用於索引操作。

在 Azure AI 搜尋服務 中,向量索引有一個定義向量與非向量場的索引結構、用於建立嵌入空間的演算法向量搜尋設定,以及查詢時評估的向量場定義設定。 索引 - 建立或更新 (REST API)建立向量索引。

提示

先決條件

  • 一個有有效訂閱的 Azure 帳號。 免費註冊帳號。

  • 一個Azure AI 搜尋服務服務。 你可以使用免費方案來完成大部分快速入門,但我們建議對於較大的資料檔案,使用基本版或更高等級。

  • Node.js 20 LTS 或更晚的版本來執行編譯後的程式碼。

  • TypeScript 將 TypeScript 編譯成 JavaScript。

  • git 來複製樣本庫。

  • Azure CLI 用於與 Microsoft Entra ID 的無金鑰認證。

設定存取權限

在開始之前,請確認你有權限存取 Azure AI 搜尋服務 的內容和操作。 此快速入門使用 Microsoft Entra ID 進行驗證,並以角色為基礎的存取控制來授權。 您必須是 擁有者 或 使用者存取管理員 才能指派角色。 如果角色不可行,改用 金鑰驗證 。

要設定建議的以角色為基礎的存取權限:

  1. 為您的搜尋服務啟用基於角色的存取權限。

  2. 請將以下角色指派 到你的使用者帳號。

    • 搜尋服務貢獻者

    • 搜尋索引資料貢獻者

    • 搜尋索引資料閱讀器

取得端點

每個Azure AI 搜尋服務服務都有一個 endpoint,這是一個唯一用來識別並提供服務網路存取的網址。 在後面的章節中,你指定這個端點以程式化方式連接到你的搜尋服務。

要取得端點:

  1. 在Azure 入口網站中,進入您的搜尋服務。

  2. 從左側窗格選擇 「概覽」。

  3. 請記下該端點,應看起來像 https://my-service.search.windows.net。

設定環境

  1. 用 Git 複製樣本庫。

    git clone https://github.com/Azure-Samples/azure-search-javascript-samples
    
  2. 進入快速啟動資料夾。

    cd azure-search-javascript-samples/quickstart-vector-ts
    
  3. 在 sample.env 中,將 AZURE_SEARCH_ENDPOINT 的佔位值替換成你在 取得端點獲得的 URL。

  4. 重新命名 sample.env 為 .env。

    mv sample.env .env
    
  5. 安裝相依性。

    npm install
    

    安裝完成後,你應該會在專案目錄看到一個 node_modules 資料夾。

  6. 將 TypeScript 檔案編譯成 JavaScript。

    npm run build
    
  7. 若要使用 Microsoft Entra ID 進行無鑰匙認證,請登入您的 Azure 帳號。 如果你有多個訂閱,請選擇包含你 Azure AI 搜尋服務 服務的訂閱。

    az login
    

執行程式碼

  1. 建立向量索引。

    node -r dotenv/config dist/createIndex.js
    
  2. 載入包含預先計算嵌入的文件。

    node -r dotenv/config dist/uploadDocuments.js
    
  3. 執行向量搜尋查詢。

    node -r dotenv/config dist/searchSingle.js
    
  4. (可選)執行額外的查詢變化。

    node -r dotenv/config dist/searchSingleWithFilter.js
    node -r dotenv/config dist/searchSingleWithFilterGeo.js
    node -r dotenv/config dist/searchHybrid.js
    node -r dotenv/config dist/searchSemanticHybrid.js
    

    註

    這些指令會從 .js 資料夾中執行已編譯的 dist 檔案。 TypeScript 程式碼必須先轉譯成 JavaScript,Node.js 才能執行,這也是你之前執行 npm run build的原因。

產出

輸出 createIndex.ts 顯示索引名稱和確認。

Using Azure Search endpoint: https://<search-service-name>.search.windows.net
Using index name: hotels-vector-quickstart
Creating index...
hotels-vector-quickstart created

輸出 uploadDocuments.ts 顯示每個索引文件的成功狀態。

Uploading documents...
Key: 1, Succeeded: true, ErrorMessage: none
Key: 2, Succeeded: true, ErrorMessage: none
Key: 3, Succeeded: true, ErrorMessage: none
Key: 4, Succeeded: true, ErrorMessage: none
Key: 48, Succeeded: true, ErrorMessage: none
Key: 49, Succeeded: true, ErrorMessage: none
Key: 13, Succeeded: true, ErrorMessage: none
All documents indexed successfully.

輸出 searchSingle.ts 結果顯示依相似度分數排名的向量搜尋結果。

Single Vector search found 5
- HotelId: 48, HotelName: Nordick's Valley Motel, Tags: ["continental breakfast","air conditioning","free wifi"], Score 0.6605852
- HotelId: 13, HotelName: Luxury Lion Resort, Tags: ["bar","concierge","restaurant"], Score 0.6333684
- HotelId: 4, HotelName: Sublime Palace Hotel, Tags: ["concierge","view","air conditioning"], Score 0.605672
- HotelId: 49, HotelName: Swirling Currents Hotel, Tags: ["air conditioning","laundry service","24-hour front desk service"], Score 0.6026341
- HotelId: 2, HotelName: Old Century Hotel, Tags: ["pool","free wifi","air conditioning","concierge"], Score 0.57902366

了解程式碼

註

本節的程式碼片段可能已經過修改以提升可讀性。 完整工作範例請參考原始碼。

現在你已經執行過程式碼,讓我們來拆解幾個關鍵步驟:

  1. 建立向量索引
  2. 將文件上傳至索引
  3. 查詢索引

建立向量索引

在你將內容加入 Azure AI 搜尋服務 之前,必須建立索引來定義內容的儲存與結構。

索引架構是圍繞飯店內容組織的。 樣本資料包含虛構旅館的向量與非向量描述。 以下程式碼 建立 createIndex.ts 索引結構,包括向量場 DescriptionVector。

const searchFields: SearchField[] = [
    { name: "HotelId", type: "Edm.String", key: true, sortable: true, filterable: true, facetable: true },
    { name: "HotelName", type: "Edm.String", searchable: true, filterable: true },
    { name: "Description", type: "Edm.String", searchable: true },
    {
        name: "DescriptionVector",
        type: "Collection(Edm.Single)",
        searchable: true,
        vectorSearchDimensions: 1536,
        vectorSearchProfileName: "vector-profile"
    },
    { name: "Category", type: "Edm.String", filterable: true, facetable: true },
    { name: "Tags", type: "Collection(Edm.String)", filterable: true },
    // Additional fields: ParkingIncluded, LastRenovationDate, Rating, Address, Location
];

const vectorSearch: VectorSearch = {
    profiles: [
        {
            name: "vector-profile",
            algorithmConfigurationName: "vector-search-algorithm"
        }
    ],
    algorithms: [
        {
            name: "vector-search-algorithm",
            kind: "hnsw",
            parameters: { m: 4, efConstruction: 400, efSearch: 1000, metric: "cosine" }
        }
    ]
};

const semanticSearch: SemanticSearch = {
    configurations: [
        {
            name: "semantic-config",
            prioritizedFields: {
                contentFields: [{ name: "Description" }],
                keywordsFields: [{ name: "Category" }],
                titleField: { name: "HotelName" }
            }
        }
    ]
};

const searchIndex: SearchIndex = {
    name: indexName,
    fields: searchFields,
    vectorSearch: vectorSearch,
    semanticSearch: semanticSearch,
    suggesters: [{ name: "sg", searchMode: "analyzingInfixMatching", sourceFields: ["HotelName"] }]
};

const result = await indexClient.createOrUpdateIndex(searchIndex);

重點摘要:

  • 你透過建立欄位清單來定義索引。

  • 此索引支援多種搜尋功能:

  • 屬性 vectorSearchDimensions 必須與你嵌入模型的輸出大小相符。 這個快速入門使用 1,536 個維度,以符合 text-embedding-ada-002 模型。

  • 此 vectorSearch 配置定義了近似最近鄰(ANN)演算法。 支援的演算法包括 Hierarchical Navigable Small World (HNSW) 與詳盡 K-Nearest Neighbor (KNN)。 欲了解更多資訊,請參閱 向量搜尋中的相關性。

將文件上傳至索引

新建立的索引是空的。 要填充索引並使其可搜尋,您必須上傳符合索引結構的 JSON 文件。

在 Azure AI 搜尋服務 中,文件既是索引的輸入,也是查詢的輸出。 為簡化起見,此快速入門指南提供已預先計算好向量的飯店文件範例。 在生產環境中,內容常從連接的資料來源擷取,並透過 索引器轉換成 JSON。

以下程式碼 uploadDocuments.ts 將文件上傳至您的搜尋服務。

const DOCUMENTS = [
    // Array of hotel documents with embedded 1536-dimension vectors
    // Each document contains: HotelId, HotelName, Description, DescriptionVector,
    // Category, Tags, ParkingIncluded, LastRenovationDate, Rating, Address, Location
];

const searchClient = new SearchClient(searchEndpoint, indexName, credential);

const result = await searchClient.uploadDocuments(DOCUMENTS);
for (const r of result.results) {
    console.log(`Key: ${r.key}, Succeeded: ${r.succeeded}`);
}

你的程式碼會透過 SearchClient 與 Azure AI 搜尋服務 服務中託管的特定搜尋索引互動,這是 @azure/search-documents 套件提供的主要物件。 SearchClient 提供索引操作的存取功能,例如:

  • 資料引入:uploadDocuments,mergeDocuments,deleteDocuments

  • 搜尋操作: search, autocomplete, suggest

查詢索引

搜尋檔案中的查詢顯示出不同的搜尋模式。 範例向量查詢基於兩個字串:

  • 全文搜尋字串: "historic hotel walk to restaurants and shopping"

  • 向量查詢字串: "quintessential lodging near running trails, eateries, retail" (向量化成數學表示)

向量查詢字串在語意上與全文搜尋字串相似,但它包含索引中不存在的詞彙。 僅用關鍵字進行搜尋的向量查詢字串結果為零。 然而,向量搜尋是根據意義而非精確關鍵字尋找相關匹配。

以下範例從基本的向量查詢開始,逐步加入篩選器、關鍵字搜尋及語意重排序。

searchSingle.ts 示範一個基本情境,你想找到與向量查詢字串非常吻合的文件描述。 物件 VectorQuery 配置向量搜尋:

  • kNearestNeighborsCount 限制根據向量相似度回傳的結果數量。
  • fields 指定要搜尋的向量場。
const vectorQuery: VectorQuery<HotelDocument> = {
    vector: vector,
    kNearestNeighborsCount: 5,
    fields: ["DescriptionVector"],
    kind: "vector",
    exhaustive: true
};

const searchOptions: SearchOptions<HotelDocument> = {
    top: 7,
    select: ["HotelId", "HotelName", "Description", "Category", "Tags"] as const,
    includeTotalCount: true,
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

const results: SearchDocumentsResult<HotelDocument> = await searchClient.search("*", searchOptions);

for await (const result of results.results) {
    const doc = result.document;
    console.log(`HotelId: ${doc.HotelId}, HotelName: ${doc.HotelName}, Score: ${result.score}`);
}

帶有濾波器的單向量搜尋

在Azure AI 搜尋服務中,filters 適用於索引中的非向量場。 searchSingleWithFilter.ts 在 Tags 欄位上設置過濾器,以過濾掉不提供免費 Wi-Fi 的飯店。

const searchOptions: SearchOptions<HotelDocument> = {
    top: 7,
    select: ["HotelId", "HotelName", "Description", "Category", "Tags"] as const,
    includeTotalCount: true,
    filter: "Tags/any(tag: tag eq 'free wifi')", // Adding filter for "free wifi" tag
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

const results: SearchDocumentsResult<HotelDocument> = await searchClient.search("*", searchOptions);

使用地理濾波器的單向量搜尋

你可以指定地理 空間過濾器 ,限制結果只在特定地理區域內。 searchSingleWithGeoFilter.ts 指定地理點(華盛頓特區,使用經度與緯度座標),並返回300公里範圍內的飯店。 filterMode 屬性決定過濾器執行的時間。 在這種情況下,postFilter 在向量搜索後運行過濾器。

const searchOptions: SearchOptions<HotelDocument> = {
    top: 5,
    includeTotalCount: true,
    select: ["HotelId", "HotelName", "Category", "Description", "Address/City", "Address/StateProvince"] as const,
    facets: ["Address/StateProvince"],
    filter: "geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300",
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

混合式搜尋 將全文與向量查詢合併於單一請求中。 searchHybrid.ts 同時執行兩種查詢類型,然後使用互惠排名融合(Reciprocal Rank Fusion,RRF)將結果合併成統一的排名。 RRF 利用每個結果集結果排名的反向來產生合併排名。 請注意,混合搜尋分數普遍低於單一查詢分數。

const vectorQuery: VectorQuery<HotelDocument> = {
    vector: vector,
    kNearestNeighborsCount: 5,
    fields: ["DescriptionVector"],
    kind: "vector",
    exhaustive: true
};

const searchOptions: SearchOptions<HotelDocument> = {
    top: 5,
    includeTotalCount: true,
    select: ["HotelId", "HotelName", "Description", "Category", "Tags"] as const,
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

// Use search text for keyword search (hybrid search = vector + keyword)
const searchText = "historic hotel walk to restaurants and shopping";
const results: SearchDocumentsResult<HotelDocument> = await searchClient.search(searchText, searchOptions);

searchSemanticHybrid.ts 示範語義排序,根據語言理解重新排序結果。

const searchOptions: SearchOptions<HotelDocument> = {
    top: 5,
    includeTotalCount: true,
    select: ["HotelId", "HotelName", "Category", "Description"] as const,
    queryType: "semantic" as const,
    semanticSearchOptions: {
        configurationName: "semantic-config"
    },
    vectorSearchOptions: {
        queries: [vectorQuery],
        filterMode: "postFilter"
    }
};

const searchText = "historic hotel walk to restaurants and shopping";
const results: SearchDocumentsResult<HotelDocument> = await searchClient.search(searchText, searchOptions);

for await (const result of results.results) {
    console.log(`Score: ${result.score}, Re-ranker Score: ${result.rerankerScore}`);
}

將這些結果與前一個查詢的混合搜尋結果做比較。 若無語意重排序,Sublime Palace Hotel 會排名第一,因為互惠排名融合(Reciprocal Rank Fusion,RRF)將文字分數與向量分數合併,產生合併結果。 經過語意調整後,Swirling Currents Hotel 躍升至榜首。

語意排序器利用機器理解模型評估每個結果與查詢意圖的匹配程度。 Swirling Currents Hotel 的描述提到"walking access to shopping, dining, entertainment and the city center",這與搜尋關鍵字"walk to restaurants and shopping"高度吻合。 這種對附近餐飲與購物的語意比對,讓它的排名高於 Sublime Palace Hotel,因為後者在描述中未強調可步行的便利設施。

重點摘要:

  • 在混合式搜尋中,你可以將向量搜尋與關鍵字上的全文搜尋整合起來。 篩選器和語意排序僅適用於文本內容,不適用於向量。

  • 實際結果則包含更多細節,包括語意說明和重點。 此快速入門功能會修改結果以提升可讀性。 要取得完整的回應結構,請使用 REST 執行請求。

清理資源

當您在自己的訂用帳戶中工作時,建議您在完成專案後移除不再需要的資源。 若您讓資源繼續執行,則可能會產生費用。

在Azure入口網站中,從左側窗格選擇 所有資源或 資源群組以尋找並管理資源。 你可以單獨刪除資源,或是一次性刪除資源群組,移除所有資源。

否則,請執行以下指令刪除你在此快速啟動中建立的索引。

npm run build && node -r dotenv/config dist/deleteIndex.js

在這個快速入門中,你會使用 Azure AI 搜尋服務 REST API來建立、載入並查詢一個 向量索引。

在 Azure AI 搜尋服務 中,向量索引有一個定義向量與非向量場的索引結構、用於建立嵌入空間的演算法向量搜尋設定,以及查詢時評估的向量場定義設定。 索引 - 建立或更新 (REST API)建立向量索引。

提示

先決條件

設定存取權限

在開始之前,請確認你有權限存取 Azure AI 搜尋服務 的內容和操作。 此快速入門使用 Microsoft Entra ID 進行驗證,並以角色為基礎的存取控制來授權。 您必須是 擁有者 或 使用者存取管理員 才能指派角色。 如果角色不可行,改用 金鑰驗證 。

要設定建議的以角色為基礎的存取權限:

  1. 為您的搜尋服務啟用基於角色的存取權限。

  2. 請將以下角色指派 到你的使用者帳號。

    • 搜尋服務貢獻者

    • 搜尋索引資料貢獻者

    • 搜尋索引資料閱讀器

取得端點

每個Azure AI 搜尋服務服務都有一個 endpoint,這是一個唯一用來識別並提供服務網路存取的網址。 在後面的章節中,你指定這個端點以程式化方式連接到你的搜尋服務。

要取得端點:

  1. 在Azure 入口網站中,進入您的搜尋服務。

  2. 從左側窗格選擇 「概覽」。

  3. 請記下該端點,應看起來像 https://my-service.search.windows.net。

設定環境

  1. 用 Git 複製樣本庫。

    git clone https://github.com/Azure-Samples/azure-search-rest-samples
    
  2. 進入快速啟動資料夾,在 Visual Studio Code 中開啟。

    cd azure-search-rest-samples/Quickstart-vectors
    code .
    
  3. 在 az-search-quickstart-vectors.rest 中,將 @baseUrl 的佔位值替換成你在 取得端點獲得的 URL。

  4. 若要使用 Microsoft Entra ID 進行無鑰匙認證,請登入您的 Azure 帳號。 如果你有多個訂閱,請選擇包含你 Azure AI 搜尋服務 服務的訂閱。

    az login
    
  5. 使用 Microsoft Entra ID 進行無鑰匙認證時,請產生存取權杖。

    az account get-access-token --scope https://search.azure.com/.default --query accessToken -o tsv
    
  6. 將 的 @token 佔位值替換為前一步的標記。

執行程式碼

  1. 在 ### List existing indexes by name下選「 發送請求 」以驗證連線。

    回應應該會出現在相鄰的窗格。 如果您有現有索引,系統會列出這些索引。 否則,清單就是空白。 如果 HTTP 程式碼是 200 OK,你就可以繼續了。

  2. 依序發送剩餘請求,建立向量索引、上傳文件,並執行不同類型的向量查詢。

產出

每個查詢請求都會回傳 JSON 結果。 以下範例為請求的輸出 ### Run a single vector query ,顯示依相似度分數排名的向量搜尋結果。

{
  "@odata.count": 5,
  "value": [
    {
      "@search.score": 0.6605852,
      "HotelId": "48",
      "HotelName": "Nordick's Valley Motel",
      "Description": "Only 90 miles (about 2 hours) from the nation's capital and nearby most everything the historic valley has to offer. Hiking? Wine Tasting? Exploring the caverns? It's all nearby and we have specially priced packages to help make our B&B your home base for fun while visiting the valley.",
      "Category": "Boutique",
      "Tags": [
        "continental breakfast",
        "air conditioning",
        "free wifi"
      ]
    },
    {
      "@search.score": 0.6333684,
      "HotelId": "13",
      "HotelName": "Luxury Lion Resort",
      "Description": "Unmatched Luxury. Visit our downtown hotel to indulge in luxury accommodations. Moments from the stadium and transportation hubs, we feature the best in convenience and comfort.",
      "Category": "Luxury",
      "Tags": [
        "bar",
        "concierge",
        "restaurant"
      ]
    },
    {
      "@search.score": 0.605672,
      "HotelId": "4",
      "HotelName": "Sublime Palace Hotel",
      "Description": "Sublime Palace Hotel is located in the heart of the historic center of Sublime in an extremely vibrant and lively area within short walking distance to the sites and landmarks of the city and is surrounded by the extraordinary beauty of churches, buildings, shops and monuments. Sublime Cliff is part of a lovingly restored 19th century resort, updated for every modern convenience.",
      "Category": "Boutique",
      "Tags": [
        "concierge",
        "view",
        "air conditioning"
      ]
    },
    {
      "@search.score": 0.6026341,
      "HotelId": "49",
      "HotelName": "Swirling Currents Hotel",
      "Description": "Spacious rooms, glamorous suites and residences, rooftop pool, walking access to shopping, dining, entertainment and the city center. Each room comes equipped with a microwave, a coffee maker and a minifridge. In-room entertainment includes complimentary W-Fi and flat-screen TVs.",
      "Category": "Suite",
      "Tags": [
        "air conditioning",
        "laundry service",
        "24-hour front desk service"
      ]
    },
    {
      "@search.score": 0.57902366,
      "HotelId": "2",
      "HotelName": "Old Century Hotel",
      "Description": "The hotel is situated in a nineteenth century plaza, which has been expanded and renovated to the highest architectural standards to create a modern, functional and first-class hotel in which art and unique historical elements coexist with the most modern comforts. The hotel also regularly hosts events like wine tastings, beer dinners, and live music.",
      "Category": "Boutique",
      "Tags": [
        "pool",
        "free wifi",
        "air conditioning",
        "concierge"
      ]
    }
  ]
}

了解程式碼

註

本節的程式碼片段可能已經過修改以提升可讀性。 完整工作範例請參考原始碼。

現在你已經執行過程式碼,讓我們來拆解幾個關鍵步驟:

  1. 建立向量索引
  2. 將文件上傳至索引
  3. 查詢索引

建立向量索引

在你將內容加入 Azure AI 搜尋服務 之前,必須建立索引來定義內容的儲存與結構。 此快速入門將呼叫 Indexes - Create (REST API),以在您的搜尋服務中建立一個名為hotels-vector-quickstart的向量索引及其實體資料結構。

索引架構是圍繞飯店內容組織的。 樣本資料包含虛構旅館的向量與非向量描述。 以下摘錄展示了請求的 ### Create a new index 密鑰結構。

PUT {{baseUrl}}/indexes/hotels-vector-quickstart?api-version={{api-version}}  HTTP/1.1
Content-Type: application/json
Authorization: Bearer {{token}}

{
    "name": "hotels-vector-quickstart",
    "fields": [
        { "name": "HotelId", "type": "Edm.String", "key": true, "filterable": true },
        { "name": "HotelName", "type": "Edm.String", "searchable": true },
        { "name": "Description", "type": "Edm.String", "searchable": true },
        {
            "name": "DescriptionVector",
            "type": "Collection(Edm.Single)",
            "searchable": true,
            "dimensions": 1536,
            "vectorSearchProfile": "my-vector-profile"
        },
        { "name": "Category", "type": "Edm.String", "filterable": true, "facetable": true },
        { "name": "Tags", "type": "Collection(Edm.String)", "filterable": true, "facetable": true }
        // Additional fields omitted for brevity
    ],
    "vectorSearch": {
        "algorithms": [
            { "name": "hnsw-vector-config", "kind": "hnsw" }
        ],
        "profiles": [
            { "name": "my-vector-profile", "algorithm": "hnsw-vector-config" }
        ]
    },
    "semantic": {
        "configurations": [
            {
                "name": "semantic-config",
                "prioritizedFields": {
                    "titleField": { "fieldName": "HotelName" },
                    "prioritizedContentFields": [{ "fieldName": "Description" }]
                }
            }
        ]
    }
}

重點摘要:

  • 此索引支援多種搜尋功能:

  • 屬性 dimensions 必須與你嵌入模型的輸出大小相符。 這個快速入門使用 1,536 個維度,以符合 text-embedding-ada-002 模型。

  • 本 vectorSearch 節定義了近似最近鄰(ANN)演算法。 支援的演算法包括 Hierarchical Navigable Small World (HNSW) 與詳盡 K-Nearest Neighbor (KNN)。 欲了解更多資訊,請參閱 向量搜尋中的相關性。

將文件上傳至索引

新建立的索引是空的。 要填充索引並使其可搜尋,您必須上傳符合索引結構的 JSON 文件。

在 Azure AI 搜尋服務 中,文件既是索引的輸入,也是查詢的輸出。 為了簡化起見,這個快速入門指南提供飯店範例文件,格式為內嵌 JSON。 然而,在生產環境中,內容常從連接的資料來源擷取,並利用 索引器轉換成 JSON。

此快速入門程式呼叫 Documents - Index(REST API) 來將範例飯店文件加入索引中。 以下摘錄展示了該請求的 ### Upload 7 documents 結構。

POST {{baseUrl}}/indexes/hotels-vector-quickstart/docs/index?api-version={{api-version}}  HTTP/1.1
Content-Type: application/json
Authorization: Bearer {{token}}

{
    "value": [
        {
            "@search.action": "mergeOrUpload",
            "HotelId": "1",
            "HotelName": "Stay-Kay City Hotel",
            "Description": "This classic hotel is ideally located on the main commercial artery of the city...",
            "DescriptionVector": [-0.0347, 0.0289, ... ],  // 1536 floats
            "Category": "Boutique",
            "Tags": ["view", "air conditioning", "concierge"],
            "ParkingIncluded": false,
            "Rating": 3.60,
            "Address": { "City": "New York", "StateProvince": "NY" },
            "Location": { "type": "Point", "coordinates": [-73.975403, 40.760586] }
        }
        // Additional documents omitted for brevity
    ]
}

重點摘要:

  • 陣列中的 value 每份文件代表一家飯店,並包含與索引結構相符的欄位。 參數 @search.action 指定每份文件要執行的操作。 這個快速入門會使用 mergeOrUpload,如果文件不存在則會新增,如果已存在則會更新。

  • 有效載荷中的文件由索引結構中定義的欄位組成。

查詢索引

範例檔案中的查詢展示了不同的搜尋模式。 範例向量查詢基於兩個字串:

  • 全文搜尋字串: "historic hotel walk to restaurants and shopping"

  • 向量查詢字串: "quintessential lodging near running trails, eateries, retail" (向量化成數學表示)

向量查詢字串在語意上與全文搜尋字串相似,但它包含索引中不存在的詞彙。 僅用關鍵字進行搜尋的向量查詢字串結果為零。 然而,向量搜尋是根據意義而非精確關鍵字尋找相關匹配。

以下範例從基本的向量查詢開始,逐步加入篩選器、關鍵字搜尋及語意重排序。

這個 ### Run a single vector query 請求展示了一個基本情境,你希望找到與向量查詢字串非常吻合的文件描述。 陣列 vectorQueries 配置向量搜索。

  • k 限制根據向量相似度回傳的結果數量。
  • fields 指定要搜尋的向量場。
POST {{baseUrl}}/indexes/hotels-vector-quickstart/docs/search?api-version={{api-version}}  HTTP/1.1
Content-Type: application/json
Authorization: Bearer {{token}}

{
    "count": true,
    "select": "HotelId, HotelName, Description, Category, Tags",
    "vectorQueries": [
        {
            "vector": [ ... ],  // 1536-dimensional vector of "quintessential lodging near running trails, eateries, retail"
            "k": 5,
            "fields": "DescriptionVector",
            "kind": "vector",
            "exhaustive": true
        }
    ]
}

帶有濾波器的單向量搜尋

在Azure AI 搜尋服務中,filters 適用於索引中的非向量場。 ### Run a vector query with a filter請求會在Tags欄位中篩選出不提供免費 Wi-Fi 的飯店。

POST {{baseUrl}}/indexes/hotels-vector-quickstart/docs/search?api-version={{api-version}}  HTTP/1.1
Content-Type: application/json
Authorization: Bearer {{token}}

{
    "count": true,
    "select": "HotelId, HotelName, Description, Category, Tags",
    "filter": "Tags/any(tag: tag eq 'free wifi')",
    "vectorFilterMode": "postFilter",
    "vectorQueries": [
        {
            "vector": [ ... ],  // 1536-dimensional vector
            "k": 7,
            "fields": "DescriptionVector",
            "kind": "vector",
            "exhaustive": true
        }
    ]
}

使用地理濾波器的單向量搜尋

你可以指定地理 空間過濾器 ,限制結果只在特定地理區域內。 ### Run a vector query with a geo filter 要求會指定一個地理點 (華盛頓特區,使用經度與緯度座標),並傳回 300 公里內的飯店。 參數 vectorFilterMode 決定過濾器運行的時間。 在這種情況下,postFilter 在向量搜索後運行過濾器。

POST {{baseUrl}}/indexes/hotels-vector-quickstart/docs/search?api-version={{api-version}}  HTTP/1.1
Content-Type: application/json
Authorization: Bearer {{token}}

{
    "count": true,
    "select": "HotelId, HotelName, Address/City, Address/StateProvince, Description",
    "filter": "geo.distance(Location, geography'POINT(-77.03241 38.90166)') le 300",
    "vectorFilterMode": "postFilter",
    "top": 5,
    "facets": [ "Address/StateProvince"],
    "vectorQueries": [
        {
            "vector": [ ... ],  // 1536-dimensional vector
            "k": 5,
            "fields": "DescriptionVector",
            "kind": "vector",
            "exhaustive": true
        }
    ]
}

混合式搜尋 將全文與向量查詢合併於單一請求中。 請求 ### Run a hybrid query 會同時執行兩種查詢類型,然後使用互惠排名融合(Reciprocal Rank Fusion,RRF)將結果合併成統一的排名。 RRF 利用每個結果集結果排名的反向來產生合併排名。 請注意,混合搜尋分數普遍低於單一查詢分數。

POST {{baseUrl}}/indexes/hotels-vector-quickstart/docs/search?api-version={{api-version}}  HTTP/1.1
Content-Type: application/json
Authorization: Bearer {{token}}

{
    "count": true,
    "search": "historic hotel walk to restaurants and shopping",
    "select": "HotelId, HotelName, Category, Tags, Description",
    "top": 5,
    "vectorQueries": [
        {
            "vector": [ ... ],  // 1536-dimensional vector
            "k": 5,
            "fields": "DescriptionVector",
            "kind": "vector",
            "exhaustive": true
        }
    ]
}

此 ### Run a hybrid query with semantic reranking 請求展示了語 意排名,即根據語言理解對結果進行重新排序。

POST {{baseUrl}}/indexes/hotels-vector-quickstart/docs/search?api-version={{api-version}}  HTTP/1.1
Content-Type: application/json
Authorization: Bearer {{token}}

{
    "count": true,
    "search": "historic hotel walk to restaurants and shopping",
    "select": "HotelId, HotelName, Category, Description",
    "queryType": "semantic",
    "semanticConfiguration": "semantic-config",
    "top": 5,
    "vectorQueries": [
        {
            "vector": [ ... ],  // 1536-dimensional vector
            "k": 7,
            "fields": "DescriptionVector",
            "kind": "vector",
            "exhaustive": true
        }
    ]
}

將這些結果與前一個查詢的混合搜尋結果做比較。 若無語意重排序,Sublime Palace Hotel 會排名第一,因為互惠排名融合(Reciprocal Rank Fusion,RRF)將文字分數與向量分數合併,產生合併結果。 經過語意調整後,Swirling Currents Hotel 躍升至榜首。

語意排序器利用機器理解模型評估每個結果與查詢意圖的匹配程度。 Swirling Currents Hotel 的描述提到"walking access to shopping, dining, entertainment and the city center",這與搜尋關鍵字"walk to restaurants and shopping"高度吻合。 這種對附近餐飲與購物的語意比對,讓它的排名高於 Sublime Palace Hotel,因為後者在描述中未強調可步行的便利設施。

重點摘要:

  • 在混合式搜尋中,你可以將向量搜尋與關鍵字上的全文搜尋整合起來。 篩選器和語意排序僅適用於文本內容,不適用於向量。

清理資源

當您在自己的訂用帳戶中工作時,建議您在完成專案後移除不再需要的資源。 若您讓資源繼續執行,則可能會產生費用。

在Azure入口網站中,從左側窗格選擇 所有資源或 資源群組以尋找並管理資源。 你可以單獨刪除資源,或是一次性刪除資源群組,移除所有資源。

否則,您可以傳送 ### Delete an index 要求,刪除您在本快速入門中建立的索引。