搭配使用 Foundry 記憶體和 LangChain 及 LangGraph

使用 langchain-azure-ai 和 Foundry 記憶體為你的應用程式增加長期記憶。 在本文中,你會建立一個以記憶體為基礎的鏈條,儲存使用者偏好,在新會話中調用它們,並執行直接記憶體查詢。

此模式適用於 LangChain 與 LangGraph 應用。 核心概念是將短期聊天記錄保存在執行時,並使用 Foundry 記憶體作為使用者層級上下文的長期儲存。

Foundry Memory 專注於長期記憶。 將短期的回合逐回合狀態保留在 LangChain 或 LangGraph 執行時狀態中。

先決條件

  • 一個 Azure 訂閱。 免費創建一個。
  • Foundry 專案。
  • 一個已部署的 Microsoft Foundry 記憶體擷取聊天模型。
    • 這個教學使用了「gpt-4.1」。
  • 部署的聊天模型與用於記憶體庫的嵌入模型。
    • 這個教學使用 text-embedding-3-large。
  • Python 3.10 或更新版本。
  • Azure CLI登入(az login),這樣 DefaultAzureCredential 就能用角色 Azure AI Developer 進行認證。

配置你的環境

安裝本教學所需的套件。 用 langchain-azure-ai 於 LangChain 和 LangGraph 的整合,azure-ai-projects 於記憶體儲存管理,以及 azure-identity 於認證。

pip install -U "langchain-azure-ai" azure-ai-projects azure-identity

設定我們在這教學中使用的環境變數:

export AZURE_AI_PROJECT_ENDPOINT="https://<resource>.services.ai.azure.com/api/projects/<project>"

了解記憶體模型

Foundry Memory 儲存並檢索兩種長期記憶類型:

  • 使用者資料記憶:穩定的使用者事實與偏好,例如偏好名稱或飲食限制。
  • 聊天摘要記憶:對先前討論主題的精簡摘要。

記憶體利用「範圍」的概念來分割資訊,使其能夠一致地儲存與檢索。 範疇就像是用來組織資訊的識別碼或金鑰。

  • 你可以用 使用者 ID 作為長期記憶的穩定身份。 同一個使用者的通話時間保持不變。
  • 你可以用 會話 ID 作為短期對話的身份。 根據每次聊天會話進行調整。
  • 你可以將 資源 ID 作為跨多位使用者的長期記憶穩定識別碼。

這種分離讓你的應用程式能記住不同會話的使用者偏好,而不會混淆無關的對話。

建立記憶體儲存

在開始之前,你需要建立一個記憶體儲存區。 此操作請使用 Microsoft Foundry 專案 SDK azure-ai-projects。

import os

from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import (
	MemoryStoreDefaultDefinition,
	MemoryStoreDefaultOptions,
)
from azure.core.exceptions import ResourceNotFoundError
from azure.identity import DefaultAzureCredential

endpoint = os.environ["AZURE_AI_PROJECT_ENDPOINT"]
credential = DefaultAzureCredential()
client = AIProjectClient(endpoint=endpoint, credential=credential)

store_name = "lc-integration-test-store"
try:
    store = client.beta.memory_stores.get(store_name)
    print(f"✓ Memory store '{store_name}' already exists")
except ResourceNotFoundError:
    print(f"Creating memory store '{store_name}'...")
    definition = MemoryStoreDefaultDefinition(
        chat_model="gpt-4.1", 						# Change for your LLM model
        embedding_model="text-embedding-3-large",	# Change for your emebddings model
        options=MemoryStoreDefaultOptions(
            user_profile_enabled=True,
            chat_summary_enabled=True,
        ),
    )
    store = client.beta.memory_stores.create(
        name=store_name,
        description="Long-term memory store",
        definition=definition,
    )
    print(f"✓ Memory store '{store_name}' created successfully")
✓ Memory store 'lc-integration-test-store' created successfully

這段程式碼的功能是: 連線到你的 Foundry 專案,取得或建立記憶體儲存庫,並啟用使用者檔案功能以及聊天摘要擷取功能。

在 LangGraph 與 LangChain 中使用記憶體

Foundry Memory 透過引入兩個物件,整合 LangGraph 與 LangChain:

  • 這個類別 langchain_azure_ai.chat_message_history.AzureAIMemoryChatMessageHistory 會建立一個以記憶體為基礎的聊天記錄。
  • 這個類別 langchain_azure_ai.retrievers.AzureAIMemoryRetriever 允許從聊天訊息歷史中檢索記錄。

一般來說,你可以用以下實用的檢索策略來搭配它們:

  • 在對話初期擷取使用者資料記憶,以個人化回應。
  • 根據目前的聊天回合檢索聊天摘要記憶,以恢復相關的先前上下文。

範例:新增一個會話感知的記憶體層

在這個例子中,我們用 LangChain 建構一個單一可執行單元,擷取相關的長期記憶體,將其注入到模型的提示中,並結合短期聊天歷史與長期記憶來執行該模型。

讓我們來看看如何實作:

建立聊天訊息紀錄

這個例子使用穩定 user_id 檔作為記憶體範圍。 使用session_id來維護每個工作階段的交談上下文。

from langchain_azure_ai.chat_history import AzureAIMemoryChatMessageHistory
from langchain_azure_ai.retrievers import AzureAIMemoryRetriever
from langchain_core.chat_history import InMemoryChatMessageHistory

_session_histories: dict[tuple[str, str], AzureAIMemoryChatMessageHistory] = {}

def get_session_history(user_id: str, session_id: str) -> AzureAIMemoryChatMessageHistory:
    """Get or create a session history for a user and session.
    
    Args:
        user_id: Stable user identifier (used as scope in Foundry Memory)
        session_id: Ephemeral session identifier
        
    Returns:
        AzureAIMemoryChatMessageHistory instance
    """
    cache_key = (user_id, session_id)
    if cache_key not in _session_histories:
        _session_histories[cache_key] = AzureAIMemoryChatMessageHistory(
            project_endpoint=endpoint,
            credential=credential,
            store_name=store_name,
            scope=user_id,
            base_history=InMemoryChatMessageHistory(),
            update_delay=0,  # TEST MODE: process updates immediately (default ~300s)
        )
    return _session_histories[cache_key]


def get_foundry_retriever(user_id: str, session_id: str) -> AzureAIMemoryRetriever:
    """Get a retriever tied to the cached session history.
    
    This preserves incremental search state across turns.
    
    Args:
        user_id: Stable user identifier
        session_id: Ephemeral session identifier
        
    Returns:
        AzureAIMemoryRetriever instance
    """
    return get_session_history(user_id, session_id).get_retriever(k=5)

此程式碼片段的作用如下:為每個 ((user_id, session_id)) 配對建立以記憶體為基礎的歷程記錄與擷取器,並加以快取,使擷取狀態可在相同工作階段的多個回合之間持續保留。 在這個操作指南中,讓 update_delay=0 的記憶體更新即時可見。 在實際執行環境中,除非您特別需要即時擷取,否則請使用預設延遲。 session_histories 用來避免不斷重建物件。

用記憶體檢所來組合可執行的

讓我們建立一個可跑的程式來實作迴圈:

from typing import Any
import os

from langchain.chat_models import init_chat_model
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
from langchain_core.runnables import ConfigurableFieldSpec, RunnablePassthrough
from langchain_core.runnables.history import RunnableWithMessageHistory

llm = init_chat_model("azure_ai:gpt-4.1", temperature=0.7)

prompt = ChatPromptTemplate.from_messages(
    [
        ("system", "You are helpful and concise. Use prior memories when relevant."),
        MessagesPlaceholder("history"),
        ("system", "Memories:\n{memories}"),
        ("human", "{question}"),
    ]
)


def chain_for_session(user_id: str, session_id: str) -> RunnableWithMessageHistory:
    """Create a chain with message history for a specific user and session.
    
    Args:
        user_id: Stable user identifier
        session_id: Ephemeral session identifier
        
    Returns:
        Runnable chain with message history
    """
    retriever = get_foundry_retriever(user_id, session_id)

    def format_memories(x: dict[str, Any]) -> str:
        """Retrieve and format memories as text."""
        docs = retriever.invoke(x["question"])
        return (
            "\n".join([doc.page_content for doc in docs])
            if docs
            else "No relevant memories found."
        )

    # Use RunnablePassthrough.assign to add memories to the input dict
    # RunnableWithMessageHistory will inject history automatically
    chain = RunnablePassthrough.assign(memories=format_memories) | prompt | llm

    chain_with_history = RunnableWithMessageHistory(
        chain,
        get_session_history=get_session_history,
        input_messages_key="question",
        history_messages_key="history",
        history_factory_config=[
            ConfigurableFieldSpec(
                id="user_id",
                annotation=str,
                name="User ID",
                description="Unique identifier for the user.",
                default="",
                is_shared=True,
            ),
            ConfigurableFieldSpec(
                id="session_id",
                annotation=str,
                name="Session ID",
                description="Unique identifier for the session.",
                default="",
                is_shared=True,
            ),
        ],
    )
    return chain_with_history

這段程式碼的功能: 建立一個可執行程序,將擷取的記憶注入至提示詞,然後封裝 RunnableWithMessageHistory,使聊天歷史和長期記憶可以協同運作。

此模式可讓您的提示詞維持可預測性:每一回合都會在 Memories 區段中明確納入擷取到的記憶內容。

執行實際的跨工作階段情境

此情境展現了長期記憶的全部價值:

  1. 在 A 會話中,使用者會分享偏好。
  2. 在 B 階段,應用程式會自動回想這些偏好設定。
import time

user_id = "user_001"
session_id = "session_2026_02_10_001"
chain = chain_for_session(user_id, session_id)

# 4) Session A: seed preferences (long-term memory extraction happens async)
print(
	"\n=== Turn 1 (Session A): Introduce a preference "
	"(will be extracted into long-term memory) ==="
)
r1 = chain.invoke(
	{"question": "Hi! Call me JT. I prefer dark roast coffee and budget trips."},
	config={"configurable": {"user_id": user_id, "session_id": session_id}},
)
print("ASSISTANT:", r1.content)

print("\n=== Turn 2 (Session A): Add another preference ===")
r2 = chain.invoke(
	{
		"question": "Also, I usually drink green tea in the afternoon "
		"and I like staying in hostels."
	},
	config={"configurable": {"user_id": user_id, "session_id": session_id}},
)
print("ASSISTANT:", r2.content)

# Because we set update_delay=0, extraction should happen immediately for demo.
# If you use the default delay, you may need to wait before querying from new session.
time.sleep(60)

# 5) Cross-session test: same user_id, new session_id
session_id_b = "session_2026_02_10_002"
chain_b = chain_for_session(user_id, session_id_b)

print("\n=== Turn 3 (Session B): New session should recall coffee preference ===")
r4 = chain_b.invoke(
	{"question": "Remind me of my coffee preference and travel style."},
	config={"configurable": {"user_id": user_id, "session_id": session_id_b}},
)
print("ASSISTANT:", r4.content)

print("\n=== Turn 4 (Session B): Retrieve another preference ===")
r5 = chain_b.invoke(
	{
		"question": "What do I usually drink in the afternoon, "
		"and where do I like to stay?"
	},
	config={"configurable": {"user_id": user_id, "session_id": session_id_b}},
)
print("ASSISTANT:", r5.content)
=== Turn 1 (Session A) ===
ASSISTANT: Nice to meet you, JT. I noted that you prefer dark roast coffee and budget trips.

=== Turn 2 (Session A) ===
ASSISTANT: Got it. I also noted that you usually drink green tea in the afternoon and prefer hostels.

=== Turn 3 (Session B) ===
ASSISTANT: Your coffee preference is dark roast, and your travel style is budget trips.

=== Turn 4 (Session B) ===
ASSISTANT: You usually drink green tea in the afternoon, and you like staying in hostels.

這段內容 :在 A 會話中建立使用者偏好,為同一使用者啟動 B 會話,並顯示應用程式能在不同會話間回憶先前的偏好設定。

範例:直接查詢記憶體以處理非聊天用途

當您需要在交談管線以外直接讀取記憶(例如用於個人化中介層或設定檔檢查工具)時,可以使用臨時擷取工具。

adhoc = AzureAIMemoryRetriever(
	project_endpoint=endpoint,
	credential=credential,
	store_name=store_name,
	scope=user_id,
	k=5,
)
print("\n=== Turn 5 (Ad-hoc): Direct retriever query without session history ===")
adhoc_docs = adhoc.invoke("What are my drinking preferences?")
for i, doc in enumerate(adhoc_docs, start=1):
	print(f"MEMORY {i}:", doc.page_content)
MEMORY 1: Prefers dark roast coffee.
MEMORY 2: Prefers budget trips.
MEMORY 3: Usually drinks green tea in the afternoon.
MEMORY 4: Likes staying in hostels.

這段程式碼片段執行當前範圍的直接記憶體搜尋。 所有記憶都會被檢索(以 封頂 k),但依相關性排序。

當你需要直接記憶體讀取功能時,例如個人資料卡、個人化中介軟體或工作流程路由,請使用此模式。

範例:在圖表中使用記憶體

LangGraph 採用相同的概念模式:

  • 保持 user_id 穩定以維持長期記憶。
  • 使用 thread_id (或等效)用於短期討論串上下文。
  • 在呼叫模型節點前先取得記憶體。

如果您已有 StateGraph,請在模型節點中插入擷取作業,並將記憶體文字附加至模型輸入。 另一個典型策略是使用預模型鉤子。

from langgraph.graph import MessagesState


def call_model_with_foundry_memory(state: MessagesState, config: dict):
	user_id = config["configurable"]["user_id"]
	session_id = config["configurable"]["thread_id"]
	query = state["messages"][-1].content

	retriever = get_foundry_retriever(user_id, session_id)
	docs = retriever.invoke(query)
	memory_text = "\n".join(d.page_content for d in docs) if docs else ""

	response = llm.invoke(
		[
			{"role": "system", "content": "Use prior memories when relevant."},
			{"role": "system", "content": f"Memories:\n{memory_text}"},
			*state["messages"],
		]
	)
	return {"messages": [response]}

這段程式碼片段的作用:展示一個 LangGraph 節點模式,用於在目前回合中檢索 Foundry 記憶,並將其注入模型輸入中。

關於更廣泛的 LangGraph 記憶體概念,請參見:

了解預覽限制與操作指引

在進入生產環境前,請驗證以下限制條件:

  • 記憶還在預覽階段,行為可能會改變。
  • 記憶體需要相容的聊天與嵌入技術部署。
  • 依照每個存放區及範圍套用配額,包括搜尋與更新要求率。

同時也要規劃防禦性控制項,以應對記憶污染或提示注入攻擊嘗試。 在不受信任的輸入影響儲存記憶體之前,先驗證它們。

清理資源

執行樣本後,刪除作用域以避免測試資料洩漏到未來的執行中。

result = client.memory_stores.delete_scope(name=store_name, scope=user_id)
print(
	f"Deleted {getattr(result, 'deleted_count', 'all')} memories "
	f"for scope '{user_id}'."
)
Deleted 4 memories for scope 'user_001'.