備註
SharePoint 擷取支援在已啟用「配置強化安全性與合規」設定的工作空間中使用。
你可以將結構化、半結構化和非結構化檔案從 Microsoft SharePoint 匯入 Delta 表格。 SQL 與 Spark API 支援使用批次與串流 API(包括 Auto Loader、spark.read 及 COPY INTO),在 Unity Catalog 治理下進行 SharePoint 檔案的增量擷取。
選擇如何從 SharePoint 匯入
Lakeflow Connect 提供兩種從 SharePoint 匯入的方式。 兩者都在 SharePoint 中存取資料,但管理層級有所不同。
| Option | Description |
|---|---|
| 受管理的 SharePoint 連接器 | 一個全託管的企業應用連接器,能將資料匯入 Delta 表格,並使其與來源保持同步。 |
| SQL 與 Spark API | 透過 SQL、PySpark 或 Lakeflow 管線,運用批次和串流 API(例如 read_files、spark.read、COPY INTO 以及 Auto Loader)建構自訂資料擷取管線。 提供在攝取過程中執行複雜轉換的彈性,同時讓你對管線管理與維護有更大責任。 |
小提示
Databricks 建議大多數使用情境使用受管理型 SharePoint 連接器。 它提供完全管理的體驗,包含自動增量擷取、廣泛格式支援,以及靈活的檔案與資料夾選擇。
主要功能
SQL 與 Spark API 提供:
- 結構化、半結構化與非結構化檔案的匯入
- 細緻化資料擷取:擷取特定網站、子網站、文件庫、資料夾或單一檔案
- 在 Databricks Runtime 19 及更新版本中擷取 SharePoint 網站頁面(
.aspx)。 請參見 Ingest SharePoint 網站頁面。 - 批次與串流擷取,使用
spark.read、 自動載入器,以及COPY INTO - 結構化與半結構化格式(如 CSV 與 Excel)的自動結構推論與演進
- 透過 Unity 目錄連線的安全憑證儲存
- 使用
pathGlobFilter的模式匹配進行檔案選擇
需求
要從 SharePoint 匯入檔案,您必須具備以下條件:
- 啟用了 Unity Catalog 的工作區。
-
CREATE CONNECTION權限以建立 SharePoint 連線,或根據你的 叢集存取模式 給予使用現有連線的適當權限:- 專用存取模式:
MANAGE CONNECTION。 - 標準存取模式:
USE CONNECTION。
- 專用存取模式:
- 使用 Databricks Runtime 17.3 LTS 或以上版本的運算。
- OAuth 認證設定為
Sites.Read.AllORSites.Selected權限範圍。
建立連結
建立一個 Unity 目錄連線來儲存你的 SharePoint 憑證。 無論你使用 SQL 和 Spark API,還是管理式的 SharePoint 連接器,連線設定流程都是一樣的。
完整的連線設定說明,包括 OAuth 認證選項,請參見 SharePoint擷取設定概述。
從 SharePoint 讀取檔案
要讀取檔案,請用 databricks.connection 選項傳遞你建立的連線,並設定指向你想存取的SharePoint資源的 URL。 你提供的網址決定了資料擷取的範圍。
以下路徑類型在 Databricks Runtime 17.3 LTS 及以上版本中支援:
| 路徑類型 | Description |
|---|---|
| Site | 從網址列複製網站網址。https://mytenant.sharepoint.com/sites/test-site |
| 子網站 | 從網址列複製子站網址。https://mytenant.sharepoint.com/sites/test-site/test-subsite |
| 文件庫 | 從 網站內容 開啟圖書館,並從網址列複製網址。https://mytenant.sharepoint.com/sites/test-site/Shared%20Documentshttps://mytenant.sharepoint.com/sites/test-site/custom-drive |
| Folder | 從 網站內容 打開資料夾,並從網址列複製網址。 或者,在SharePoint開啟資料夾的Details面板,點擊Path旁的複製圖示。https://mytenant.sharepoint.com/sites/test-site/Shared%20Documents/Forms/AllItems.aspx?id=%2Fsites...https://mytenant.sharepoint.com/sites/test-site/custom-drive/test-folder |
| 檔案 | 選擇檔案,點選溢出選單(...),然後選擇 預覽。 從網址列複製 URL。 或者,在SharePoint開啟檔案的Details面板,點選Path旁的複製圖示。https://mytenant.sharepoint.com/sites/test-site/Shared%20Documents/Forms/AllItems.aspx?viewid=1a2b3c...https://mytenant.sharepoint.com/sites/test-site/custom-drive/test-folder/test.csv |
Databricks Runtime 18 LTS 及以上版本新增了以下路徑類型的支援:
| 路徑類型 | Description |
|---|---|
| 嵌套子網站 | 從網址列複製子站網址。https://mytenant.sharepoint.com/sites/test-site/subsite/nested-subsite/nested-nested-subsite |
| 分享連結 | 選擇檔案或資料夾,點擊溢出選單(...),然後選擇 複製連結。 Databricks 建議將分享連結設定為永不過期。https://mytenant.sharepoint.com/:i:/s/test-site/1A2B3C4D5E6F7G8H9I |
| Microsoft 365 網頁版(前稱 Office) | 在 Microsoft 365 網頁版中開啟該檔案,並從網址列複製網址。https://mytenant.sharepoint.com/:x:/r/sites/test-site/_layouts/15/Doc.aspx?sourcedoc=%1A2B... |
Databricks Runtime 19 及以上版本新增對以下路徑類型的支援:
| 路徑類型 | Description |
|---|---|
| 租戶 | 從地址列複製租戶根網址。 若要執行整個租戶範圍的擷取,您必須使用 OAuth 機器到機器(M2M)連線。 Microsoft Graph /sites/delta 端點需要應用程式令牌。https://mytenant.sharepoint.com |
| 網站頁面 | 擷取網站的 Site Pages(.aspx)文件庫,該文件庫儲存在其文件庫之外。 請參見 Ingest SharePoint 網站頁面。https://mytenant.sharepoint.com/sites/test-site/SitePages |
範例
有幾種方法可以透過 SQL 和 Spark API 從 SharePoint 讀取檔案。
使用自動載入器串流SharePoint檔案
Auto Loader 提供最有效率的方式,讓 SharePoint 以增量方式擷取結構化檔案。 它會自動偵測新檔案,並在新檔案到達時處理。 它也能自動擷取結構化與半結構化檔案,如 CSV 和 JSON,並進行結構化推論與演化。 如需自動載入器使用的詳細資訊,請參閱 常見的資料載入模式。
# Incrementally ingest new PDF files
df = (spark.readStream.format("cloudFiles")
.option("cloudFiles.format", "binaryFile")
.option("databricks.connection", "my_sharepoint_conn")
.option("cloudFiles.schemaLocation", "<path-to-schema-location>")
.option("pathGlobFilter", "*.pdf")
.load("https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents")
)
# Incrementally ingest CSV files with automatic schema inference and evolution
df = (spark.readStream.format("cloudFiles")
.option("cloudFiles.format", "csv")
.option("databricks.connection", "my_sharepoint_conn")
.option("cloudFiles.schemaLocation", "<path-to-schema-location>")
.option("pathGlobFilter", "*.csv")
.option("cloudFiles.inferColumnTypes", True)
.option("header", True)
.load("https://mytenant.sharepoint.com/sites/Engineering/Data/IoT_Logs")
)
使用 Spark 批次讀取 SharePoint 檔案
以下範例說明如何使用 spark.read 函式,在 Python 中擷取 SharePoint 檔案。 將 recursiveFileLookup 選項設為 true,即可從巢狀結構讀取檔案,例如網站中的文件庫、文件庫中的資料夾,或網站中的子網站。
# Read unstructured data as binary files
df = (spark.read
.format("binaryFile")
.option("databricks.connection", "my_sharepoint_conn")
.option("recursiveFileLookup", True)
.option("pathGlobFilter", "*.pdf") # optional. Example: only ingest PDFs
.load("https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents"))
# Read a batch of CSV files, infer the schema, and load the data into a DataFrame
df = (spark.read
.format("csv")
.option("databricks.connection", "my_sharepoint_conn")
.option("pathGlobFilter", "*.csv")
.option("recursiveFileLookup", True)
.option("inferSchema", True)
.option("header", True)
.load("https://mytenant.sharepoint.com/sites/Engineering/Data/IoT_Logs"))
# Read a specific Excel file from SharePoint, infer the schema, and load the data into a DataFrame
df = (spark.read
.format("excel")
.option("databricks.connection", "my_sharepoint_conn")
.option("headerRows", 1) # optional
.option("dataAddress", "Sheet1!A1:M20") # optional
.load("https://mytenant.sharepoint.com/sites/Finance/Shared%20Documents/Monthly/Report-Oct.xlsx"))
使用 Spark SQL 讀取SharePoint檔案
以下範例展示了如何使用 read_files 表值函式在 SQL 中擷取SharePoint檔案。 關於使用細節 read_files ,請參見 read_files 表值函數。
-- Read pdf files
CREATE TABLE my_table AS
SELECT * FROM read_files(
"https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents",
`databricks.connection` => "my_sharepoint_conn",
format => "binaryFile",
pathGlobFilter => "*.pdf", -- optional. Example: only ingest PDFs
schemaEvolutionMode => "none"
);
-- Read a specific Excel sheet and range
CREATE TABLE my_sheet_table AS
SELECT * FROM read_files(
"https://mytenant.sharepoint.com/sites/Finance/Shared%20Documents/Monthly/Report-Oct.xlsx",
`databricks.connection` => "my_sharepoint_conn",
format => "excel",
headerRows => 1, -- optional
dataAddress => "Sheet1!A2:D10", -- optional
schemaEvolutionMode => "none"
);
增量資料導入COPY INTO
COPY INTO 提供將檔案冪等且增量地載入到 Delta 表中。 關於使用情況的詳細資訊COPY INTO,請參見使用COPY INTO常見的資料載入模式。
CREATE TABLE IF NOT EXISTS sharepoint_pdf_table;
CREATE TABLE IF NOT EXISTS sharepoint_csv_table;
CREATE TABLE IF NOT EXISTS sharepoint_excel_table;
# Incrementally ingest new PDF files
COPY INTO sharepoint_pdf_table
FROM "https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents"
FILEFORMAT = BINARYFILE
PATTERN = '*.pdf'
FORMAT_OPTIONS ('databricks.connection' = 'my_sharepoint_conn')
COPY_OPTIONS ('mergeSchema' = 'true');
# Incrementally ingest CSV files with automatic schema inference and evolution
COPY INTO sharepoint_csv_table
FROM "https://mytenant.sharepoint.com/sites/Engineering/Data/IoT_Logs"
FILEFORMAT = CSV
PATTERN = '*.csv'
FORMAT_OPTIONS ('databricks.connection' = 'my_sharepoint_conn', 'header' = 'true', 'inferSchema' = 'true')
COPY_OPTIONS ('mergeSchema' = 'true');
# Ingest a single Excel file
COPY INTO sharepoint_excel_table
FROM "https://mytenant.sharepoint.com/sites/Finance/Shared%20Documents/Monthly/Report-Oct.xlsx"
FILEFORMAT = EXCEL
FORMAT_OPTIONS ('databricks.connection' = 'my_sharepoint_conn', 'headerRows' = '1')
COPY_OPTIONS ('mergeSchema' = 'true');
在 Lakeflow 管線中匯入 SharePoint 檔案
備註
要用 SQL 和 Spark API 匯入 SharePoint 檔案,需要 Databricks Runtime 17.3 或以上版本。
以下範例展示了如何在 Lakeflow pipelines 中使用 Auto Loader 讀取 SharePoint 檔案。
Python
from pyspark import pipelines as dp
# Incrementally ingest new PDF files
@dp.table
def sharepoint_pdf_table():
return (spark.readStream.format("cloudFiles")
.option("cloudFiles.format", "binaryFile")
.option("databricks.connection", "my_sharepoint_conn")
.option("pathGlobFilter", "*.pdf")
.load("https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents")
)
# Incrementally ingest CSV files with automatic schema inference and evolution
@dp.table
def sharepoint_csv_table():
return (spark.readStream.format("cloudFiles")
.option("cloudFiles.format", "csv")
.option("databricks.connection", "my_sharepoint_conn")
.option("pathGlobFilter", "*.csv")
.option("cloudFiles.inferColumnTypes", True)
.option("header", True)
.load("https://mytenant.sharepoint.com/sites/Engineering/Data/IoT_Logs")
)
# Read a specific Excel file from SharePoint in a materialized view
@dp.table
def sharepoint_excel_table():
return (spark.read.format("excel")
.option("databricks.connection", "my_sharepoint_conn")
.option("headerRows", 1) # optional
.option("inferColumnTypes", True) # optional
.option("dataAddress", "Sheet1!A1:M20") # optional
.load("https://mytenant.sharepoint.com/sites/Finance/Shared%20Documents/Monthly/Report-Oct.xlsx")
)
SQL
-- Incrementally ingest new PDF files
CREATE OR REFRESH STREAMING TABLE sharepoint_pdf_table
AS SELECT * FROM STREAM read_files(
"https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents",
format => "binaryFile",
`databricks.connection` => "my_sharepoint_conn",
pathGlobFilter => "*.pdf");
-- Incrementally ingest CSV files with automatic schema inference and evolution
CREATE OR REFRESH STREAMING TABLE sharepoint_csv_table
AS SELECT * FROM STREAM read_files(
"https://mytenant.sharepoint.com/sites/Engineering/Data/IoT_Logs",
format => "csv",
`databricks.connection` => "my_sharepoint_conn",
pathGlobFilter => "*.csv",
header => "true");
-- Read a specific Excel file from SharePoint in a materialized view
CREATE OR REFRESH MATERIALIZED VIEW sharepoint_excel_table
AS SELECT * FROM read_files(
"https://mytenant.sharepoint.com/sites/Finance/Shared%20Documents/Monthly/Report-Oct.xlsx",
`databricks.connection` => "my_sharepoint_conn",
format => "excel",
headerRows => 1, -- optional
dataAddress => "Sheet1!A2:D10", -- optional
`cloudFiles.schemaEvolutionMode` => "none"
);
解析非結構化檔案
當使用採用 binaryFile 格式的 SQL 與 Spark API 從 SharePoint 擷取非結構化檔案(如 PDF、Word 文件或 PowerPoint 檔案)時,檔案內容會以原始二進位資料儲存。 為了準備這些檔案以供 AI 工作負載使用,如檢索增強生成(RAG)、搜尋、分類或文件理解,你可以使用 ai_parse_document 將二進位內容解析為結構化且可查詢的輸出。
以下範例展示了如何解析儲存在名為 documents的青銅 Delta 表中的非結構化文件,並新增一個帶有解析內容的欄位:
CREATE TABLE documents AS
SELECT * FROM read_files(
"https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents",
`databricks.connection` => "my_sharepoint_conn",
format => "binaryFile",
pathGlobFilter => "*.{pdf,docx}",
schemaEvolutionMode => "none"
);
SELECT *, ai_parse_document(content, map('version', '2.0')) AS parsed_content
FROM documents;
該 parsed_content 欄位包含可直接用於下游 AI 流程的擷取文字、表格、版面資訊及元資料。
使用 Lakeflow 管線進行增量剖析
你也可以在 Lakeflow pipelines 內使用 ai_parse_document ,來啟用增量式解析。 隨著新檔案從 SharePoint 流入,隨著你的管線更新,它們會被自動解析。
例如,你可以定義一個 bronze 串流資料表來擷取原始二進位檔案,接著再定義第二個串流資料表,將解析後的內容寫入新欄位:
CREATE OR REFRESH STREAMING TABLE sharepoint_documents_table
AS SELECT * FROM STREAM read_files(
"https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents",
format => "binaryFile",
`databricks.connection` => "my_sharepoint_conn",
pathGlobFilter => "*.{pdf,docx}");
CREATE OR REFRESH STREAMING TABLE documents_parsed
AS SELECT
*,
ai_parse_document(content, map('version', '2.0')) AS parsed_content
FROM STREAM sharepoint_documents_table;
此方法可確保:
- 新匯入的 SharePoint 檔案會在管線更新時自動解析。
- 解析後的輸出會與輸入資料保持同步。
- 下游 AI 流程始終基於最新的文件表示形式運作。
請參閱 ai_parse_document 功能 以了解支援格式及進階選項,包括 version 將輸出結構釘選的選項。
SharePoint 元資料欄位
_sharepoint_metadata 欄位是一個隱藏的元資料欄位,提供存取SharePoint特定屬性的擷取檔案,來源為 Microsoft Graph driveItem 資源。 需要 Databricks Runtime 18 LTS 或更新版本,且從 SharePoint 讀取時支援所有檔案格式。 要將該 _sharepoint_metadata 欄位包含在回傳的資料框架中,您必須在讀取查詢中明確選取該欄位。
若資料來源包含名為 _sharepoint_metadata 的欄位,SharePoint 的元資料欄位會重新命名為 __sharepoint_metadata(加上一個前導底線)以避免重複。 會不斷添加額外底線,這個過程會持續到名稱唯一為止。
常見的檔案元資料如檔案路徑或大小,可以透過欄位 _metadata 查詢。 請參閱 檔案資料資料列。
架構
欄位 _sharepoint_metadata 為包含以下欄位的 a STRUCT 。 所有欄位皆可為空。
| 名稱 | 類型 | Description | Example |
|---|---|---|---|
| item_id | STRING |
該物品的 driveItem ID。 | 01OMQ3MNLH42C5J675CBEI5CRK7SPKQUTZ |
| site_id | STRING |
包含該項目的 SharePoint 網站的 ID。 | mytenant.sharepoint.com,69dc7b12-f92c-498d-9514-596b793a1f77,c6c1db8d-2b8d-48a1-a549-394b63d74725 |
| drive_id | STRING |
包含該物品的硬碟 ID。 | b!EnvcaSz5jUmVFFlreTofd43bwcaNK6FIpUk5S2PXRyWTvQraaWQkSpwQEgThHDS- |
| 磁碟機類型 | STRING |
硬碟類型,例如SharePoint函式庫用documentLibrary,商務用 OneDrive用business。 |
documentLibrary |
| parent_id | STRING |
父資料夾的 driveItem ID。 | 01OMQ3MNN6Y2GOVW7725BZO354PWSELRRZ |
| parent_name | STRING |
父資料夾的名稱。 | Shared Documents |
| parent_path | STRING |
父資料夾的磁碟相對路徑。 | /drives/b!EnvcaSz5.../root: |
| web_url | STRING |
SharePoint 上該項目的瀏覽器網址。 | https://mytenant.sharepoint.com/sites/TestSite/_layouts/15/Doc.aspx?sourcedoc=... |
| MIME類型 | STRING |
項目的 MIME 類型。 | application/vnd.ms-excel |
| 由電子郵件建立 | STRING |
是創建該商品的使用者的電子郵件。 | alice@example.onmicrosoft.com |
| 建立者名稱 | STRING |
建立該項目的使用者顯示名稱。 | Alice Example |
| 建立時間戳記 | TIMESTAMP |
物品被創造的時間。 | 2025-12-03 13:33:12 |
| 最後修改者電子郵件 | STRING |
最後修改該項目的使用者的電子郵件。 | alice@example.onmicrosoft.com |
| 上次修改者名稱 | STRING |
最後修改此項目的使用者的顯示名稱。 | Alice Example |
| etag | STRING |
物品的ETag。 當物品或其任何元資料變更時,會改變。 | "{D485E667-FDFB-4810-8E8A-2AFC9EA85279},1" |
| CTAG | STRING |
物件的變更標記。 只有當物品內容改變時才會改變。 | "c:{D485E667-FDFB-4810-8E8A-2AFC9EA85279},1" |
| 描述 | STRING |
物品描述(如果已設定)。 | Q4 financial report |
| additional_metadata | VARIANT |
Microsoft Graph 回傳的其他 driveItem 欄位,但尚未被提取。 | {"shared":{"scope":"users"},...} |
備註
additional_metadata場返回為 VARIANT。 請參閱 VARIANT 類型。
範例
以下範例說明如何在讀取查詢中包含欄位 _sharepoint_metadata 、從欄位中選擇特定欄位,以及從欄位中擷取值 additional_metadataVARIANT 。
Python
df = (spark.read
.format("binaryFile")
.option("databricks.connection", "my_sharepoint_conn")
.load("https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents")
.select("*", "_metadata", "_sharepoint_metadata"))
SQL
SELECT *, _sharepoint_metadata
FROM read_files(
"https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents",
`databricks.connection` => "my_sharepoint_conn",
format => "binaryFile"
);
從 _sharepoint_metadata 結構體中選擇特定欄位:
df = (spark.read
.format("binaryFile")
.option("databricks.connection", "my_sharepoint_conn")
.load("https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents")
.select("_sharepoint_metadata.item_id", "_sharepoint_metadata.etag"))
使用additional_metadata鑄造算符從VARIANT::場中提取數值:
SELECT
*,
_sharepoint_metadata.additional_metadata:shared:scope::STRING AS shared_scope
FROM read_files(
"https://mytenant.sharepoint.com/sites/Marketing/Shared%20Documents",
`databricks.connection` => "my_sharepoint_conn",
format => "binaryFile"
);
匯入 SharePoint 網站頁面
SharePoint 將網站頁面以.aspx檔案形式儲存在網站的網站頁面函式庫中,與文件函式庫分開。 擷取網站頁面需要 Databricks Runtime 19 或更新版本。 欲了解更多關於網站頁面的資訊,請參閱 SharePoint 頁面及頁面模型。
您可以透過以下方式擷取網站頁面:
| 路徑類型 | Description |
|---|---|
| 單頁 | 在 SharePoint 開啟該頁面,並從網址列複製網址。https://mytenant.sharepoint.com/sites/test-site/SitePages/my-news-post.aspx |
| 網站頁面庫 | 從網站內容打開網站頁面,並從地址列複製網址。https://mytenant.sharepoint.com/sites/test-site/SitePages |
| 網站或子網站 | 當您擷取網站或子網站時,Azure Databricks 會連同該網站的文件庫一併擷取網站頁面庫。https://mytenant.sharepoint.com/sites/test-site |
像其他檔案一樣擷取網站頁面:
Python
# Read a single Site Page
df = (spark.read
.format("binaryFile")
.option("databricks.connection", "my_sharepoint_conn")
.load("https://mytenant.sharepoint.com/sites/test-site/SitePages/my-news-post.aspx"))
# Read a site's Site Pages library
df = (spark.read
.format("binaryFile")
.option("databricks.connection", "my_sharepoint_conn")
.option("pathGlobFilter", "*.aspx") # optional. Example: only ingest Site Page files
.load("https://mytenant.sharepoint.com/sites/test-site/SitePages"))
# Read all Site Pages from a site
df = (spark.read
.format("binaryFile")
.option("databricks.connection", "my_sharepoint_conn")
.option("recursiveFileLookup", True)
.option("pathGlobFilter", "*.aspx") # optional. Example: only ingest Site Page files
.load("https://mytenant.sharepoint.com/sites/test-site"))
SQL
-- Read a single Site Page
CREATE TABLE site_page AS
SELECT * FROM read_files(
"https://mytenant.sharepoint.com/sites/test-site/SitePages/my-news-post.aspx",
`databricks.connection` => "my_sharepoint_conn",
format => "binaryFile"
);
-- Read a site's Site Pages library
CREATE TABLE site_pages AS
SELECT * FROM read_files(
"https://mytenant.sharepoint.com/sites/test-site/SitePages",
`databricks.connection` => "my_sharepoint_conn",
format => "binaryFile",
pathGlobFilter => "*.aspx" -- optional. Example: only ingest Site Page files
);
-- Read all Site Pages from a site
CREATE TABLE site_with_pages AS
SELECT * FROM read_files(
"https://mytenant.sharepoint.com/sites/test-site",
`databricks.connection` => "my_sharepoint_conn",
format => "binaryFile",
recursiveFileLookup => true,
pathGlobFilter => "*.aspx" -- optional. Example: only ingest Site Page files
);
自動載入器(Auto Loader) COPY INTO及其他介面與網站頁面的使用方式與其他檔案相同。 請參見「 攝取範例 」。
局限性
SQL 與 Spark API 有以下限制。
- 以下 Microsoft 365 全國雲端部署不被支援:
- 政府社區雲高中 (GCC高) :
.sharepoint.us - 國防部(DoD):
.sharepoint-mil.us - 由世紀互聯營運的中國 Microsoft 365:
.sharepoint.cn
- 政府社區雲高中 (GCC高) :
- 你無法同時輸入多個相同查詢的網站。 要從兩個網站擷取資料,你必須寫兩個獨立的查詢。
- 你可以用
pathGlobFilter篩選檔案名稱的選項。 不支援基於資料夾路徑的過濾。 - SharePoint 清單不被支援。
- 不支援將資料寫回到 SharePoint 伺服器。
- 不支援自動載入器
cleanSource(在擷取後從來源刪除或歸檔檔案)。
後續步驟
- 了解 Auto Loader 以進行進階串流資料獲取模式
- 探索 COPY INTO 冪等增量負載
- 與 雲端物件儲存擷取 模式比較
- 設定 工作排程 來自動化你的資料匯入流程
- 使用 Lakeflow 管線 建立端對端且包含轉換的資料管線