這很重要
這項功能目前處於 公開預覽版。 工作區管理員可以從 「預覽 」頁面控制對此功能的存取。 請參閱 管理 Azure Databricks 預覽。
特徵視圖可讓你在訓練模型時使用符合時間點正確性的特徵計算,並在推論時自動進行特徵查找。 關於定義特徵視圖的資訊,請參見 特徵視圖。
要求
- 功能必須建立成 Feature Views。 請參閱 特色檢視。
- 如需瞭解
CustomUDF和FeatureViewSource的需求,請參閱 Feature Views API 參考文件。
API 方法
create_training_set()
建立特徵檢視後,下一步是建立模型的訓練資料。 為此,將標記資料集傳遞至 create_training_set,自動確保每個特徵值的時間點準確計算。
例如:
FeatureEngineeringClient.create_training_set(
df: DataFrame, # DataFrame with training data
features: Optional[List[Feature]], # List of Feature objects
label: Union[str, List[str], None], # Label column name(s)
exclude_columns: Optional[List[str]] = None, # Optional: columns to exclude
) -> TrainingSet
呼叫 TrainingSet.load_df 將原始訓練資料與按時間點動態計算的特徵結合。
該 df 論證必須符合以下條件:
- 必須包含所有由特徵定義所參考的實體欄位。
- 必須包含特徵定義所參考的時間序列欄位。
- 必須包含任何
RequestSource結構中宣告的所有欄位。 類型會根據宣告的架構進行驗證。 不匹配會產生錯誤(無隱性投擲)。 - 應該包含標籤欄位。
- 實體欄位名稱、時間序列欄位名稱及請求功能欄位名稱必須在所有來源間全球唯一。
時間點正確性(Point-in-time correctness):對於由表格來源支持的聚合功能與特徵,功能僅使用每行的時間戳之前的源數據來計算,以防止未來數據洩漏到模型訓練中。 對於 RequestSource 特徵,值直接取自標註過的資料框架列。
log_model()
使用 MLflow 記錄帶有特徵元資料的模型,以便追蹤血緣關係及推論時自動查找特徵:
FeatureEngineeringClient.log_model(
model, # Trained model object
artifact_path: str, # Path to store model artifact
flavor: ModuleType, # MLflow flavor module (e.g., mlflow.sklearn)
training_set: TrainingSet, # TrainingSet used for training
registered_model_name: Optional[str], # Optional: register model in Unity Catalog
extra_pip_requirements: Optional[List[str]] = None, # Optional: Additional serving dependencies
)
參數 flavor 指定要使用的 MLflow 模型風格 模組,例如 mlflow.sklearn 或 mlflow.xgboost。
使用TrainingSet登錄的模型會自動追蹤訓練中所用特徵的數據譜系。 當訓練集包含 RequestSource 特徵時,欄位 RequestSource 會作為必要的輸入加入 MLflow 模型簽名中。 這確保服務端點的 API 架構反映呼叫者在推論時必須提供的欄位。 詳情請參見 帶有特徵表的火車模型。
對於 FeatureViewSource,在記錄模型前,先註冊導出特徵及其上游特徵。 傳遞相依性所需的請求輸入,在推論時也同樣是必要的。 請參閱 自訂 UDF 相依 性以了解模型套件的需求。
score_batch()
使用自動特徵查詢執行批次推論:
FeatureEngineeringClient.score_batch(
model_uri: str, # URI of logged model
df: DataFrame, # DataFrame with entity keys and timestamps
) -> DataFrame
score_batch 使用隨模型儲存的特徵元資料,自動計算時間點正確的特徵進行推論,確保與訓練一致。 詳情請參見 帶有特徵表的火車模型。
範例工作流程
import mlflow
from databricks.feature_engineering import FeatureEngineeringClient
from sklearn.ensemble import RandomForestClassifier
fe = FeatureEngineeringClient()
# Assume features are registered in UC
# labeled_df should have columns "user_id", "transaction_time", and "is_fraud"
# 1. Create training set using Feature Views
training_set = fe.create_training_set(
df=labeled_df,
features=features,
label="is_fraud",
)
# 2. Load training data with computed features
training_df = training_set.load_df()
X = training_df.drop("is_fraud").toPandas()
y = training_df.select("is_fraud").toPandas().values.ravel()
# 3. Train model
model = RandomForestClassifier().fit(X, y)
# 4. Log model with feature metadata
with mlflow.start_run():
fe.log_model(
model=model,
artifact_path="fraud_model",
flavor=mlflow.sklearn,
training_set=training_set,
registered_model_name="main.ecommerce.fraud_model",
)
# 5. Batch scoring with automatic feature lookup
# inference_df must contain the same entity and timeseries columns
# used during training. Features are automatically computed.
predictions = fe.score_batch(
model_uri="models:/main.ecommerce.fraud_model/1",
df=inference_df,
)
predictions.display()
使用 RequestSource 功能進行訓練
當你的模型需要在推論時提供資料(例如 API 呼叫的交易細節),就要同時使用 RequestSource 功能和資料表支援的功能。 訓練過程中,會從標記的 DataFrame 中擷取 RequestSource 欄位。
from databricks.feature_engineering import FeatureEngineeringClient
from databricks.feature_engineering.entities import (
DeltaTableSource, Feature, FieldDefinition, RequestSource,
ScalarDataType, ColumnSelection,
)
fe = FeatureEngineeringClient()
# RequestSource provides transaction data at inference time
request_source = RequestSource(
schema=[
FieldDefinition(name="transaction_amount", data_type=ScalarDataType.DOUBLE),
FieldDefinition(name="vendor_id", data_type=ScalarDataType.STRING),
FieldDefinition(name="transaction_id", data_type=ScalarDataType.STRING),
FieldDefinition(name="transaction_time", data_type=ScalarDataType.DATE),
]
)
delta_source = DeltaTableSource(
catalog_name="catalog",
schema_name="schema",
table_name="vendor_data",
)
# A column selection feature from the request source (pass-through)
latest_transaction_amount = Feature(
source=request_source,
function=ColumnSelection("transaction_amount"),
name="latest_transaction_amount",
)
# A lookup feature from a delta table
vendor_category = Feature(
source=delta_source,
function=ColumnSelection("vendor_category"),
entity=["vendor_id"],
timeseries_column="transaction_time",
name="vendor_category",
)
# labels_df must contain: transaction_id, transaction_time, vendor_id,
# transaction_amount, and the label column.
ts = fe.create_training_set(
df=labels_df,
features=[latest_transaction_amount, vendor_category],
label="is_fraud",
exclude_columns=["card_id"],
)
import mlflow
from sklearn.ensemble import RandomForestClassifier
with mlflow.start_run():
training_df = ts.load_df().toPandas()
X = training_df.drop(columns=["is_fraud"])
y = training_df["is_fraud"]
model = RandomForestClassifier().fit(X, y)
# log_model() adds RequestSource columns to the MLflow model signature
fe.log_model(
model=model,
artifact_path="fraud_model",
flavor=mlflow.sklearn,
training_set=ts,
registered_model_name="catalog.schema.fraud_model",
)
使用 CustomUDF 轉換請求值
若要轉換請求資料,請使用 CustomUDFColumnSelection。 此範例使用交易記錄對數轉換功能:
log_transaction_amount = fe.get_feature(
full_name="main.ecommerce.log_transaction_amount"
)
request_df = spark.createDataFrame(
[(0.0, 0), (99.0, 1)],
"transaction_amount DOUBLE, label INT",
)
transaction_training_set = fe.create_training_set(
df=request_df,
features=[log_transaction_amount],
label="label",
)
transaction_training_set.load_df().show()
結果包含原始 transaction_amount 和 label 欄位,加上 log_transaction_amount。 UDF 會從每一列讀取 transaction_amount。 計算所需的請求欄位無法在 exclude_columns中列出。
使用 FeatureViewSource 功能來訓練
FeatureViewSource 讓 CustomUDF 使用其他特徵的輸出,包括衍生特徵。 將您想要的輸出傳遞給 create_training_set。 你不需要列出它們的中間相依關係。
對於每個輸入資料列,Azure Databricks 會解析完整的相依性圖:
- 根據其實體鍵、時間戳記和視窗定義,計算以資料表為基礎的上游特徵。 在可用時,它會使用相容的離線實體化結果。
- 從輸入資料框讀取所需的請求欄位,並評估請求支持的功能。
- 依依賴順序評估衍生特徵,使每個 UDF 都能接收其上游結果。
衍生功能不會引入另一個時間窗或時間點查詢。 其上游特徵保留了各自的時間語意。 DataFrame 必須包含這些上游來源所需的實體、時間戳記和請求欄位,即使只要求最終的衍生特徵。
例如,使用已註冊的 保證金功能,其結合了 revenue_sum_7d 和 cost_sum_7d:
from databricks.feature_engineering import FeatureEngineeringClient
fe = FeatureEngineeringClient()
margin = fe.get_feature(full_name="main.ecommerce.margin")
# labeled_df contains customer_id, event_time, and label.
training_set = fe.create_training_set(
df=labeled_df,
features=[margin],
label="label",
exclude_columns=["customer_id", "event_time"],
)
training_df = training_set.load_df()
結果包含 label 和 margin。 收益與成本特徵會被計算,但不會以額外欄位回傳。 若要將收入納入訓練資料,請使用 revenue = fe.get_feature(full_name="main.ecommerce.revenue_sum_7d") 擷取收入,並傳遞 features=[margin, revenue]。 這同樣適用於多層鏈:請求最終功能不會回傳每個中間輸出。
你可以在同一個 features 清單中組合由請求支援、由資料表支援和衍生的功能。 若要在單一 UDF 中結合它們的值,請將請求中的值表示成特徵,並在 FeatureViewSource 中將它們與由資料表支援的特徵一同參照。
為了實驗,請建構局部 Feature 物件,包括其上游圖,但不註冊它們。 使用 create_training_set,也可視需要搭配 label=None,以檢查結果。
compute_features 不支持 RequestSource 或 FeatureViewSource。
Note
每個查詢最多呼叫五次 Unity Catalog UDF 的限制也適用於訓練查詢。 計算整個相依圖所需的 UDF 呼叫,而不僅僅是作為輸出所請求的功能。 此查詢限制與圖的深度限制是分開的。
自訂 UDF 相依性
離線計算時,請在 Unity 目錄 UDF ENVIRONMENT 的子句中宣告 Python 套件。 僅在筆記本安裝套件並不會將其安裝到 UDF 環境中。
對於模型服務,也要明確地將所需套件傳遞給 log_model。 無論是 UDF ENVIRONMENT 還是包含特徵視圖的命名特徵規範,都不會自動提供這些模型需求。 包含上游 UDF 所需的依存項目以及所要求的特徵輸出。
在 transaction_training_set.load_df() 上訓練 scikit-learn 模型後,使用相同的訓練資料集記錄該模型。 包括 NumPy 和相容的查找套件:
import mlflow
fe.log_model(
model=model,
artifact_path="transaction_model",
flavor=mlflow.sklearn,
training_set=transaction_training_set,
registered_model_name="main.ecommerce.transaction_model",
extra_pip_requirements=[
"numpy==1.26.4",
"databricks-feature-lookup>=1.15.0",
],
)
服務端點應自動取得 databricks-feature-lookup 1.15.0 或更新版本,該版本支援按需計算 Unity Catalog UDF 以實現 RequestSource 功能。 保持 UDF 套件版本在離線與服務環境間保持一致,以避免計算值差異。 對於沒有模型的特徵服務端點,請改為在 create_feature_spec 上宣告套件。 請參閱 新增 Python 相依性。
使用串流功能訓練
當你定義 串流時,Databricks 會管理一個將串流資料寫入 Delta 表格的匯入管線。
create_training_set 會從這個擷取資料表讀取,並與已加上標籤的 DataFrame 進行時間點聯結,就像來自 DeltaTableSource 的批次特徵一樣。 如需資料擷取設定、回填和去重複的詳細資訊,請參見 資料擷取與回填。
範例
from databricks.feature_engineering import FeatureEngineeringClient
from databricks.feature_engineering.entities import (
StreamSource,
Feature,
AggregationFunction,
Sum,
RollingWindow,
)
from datetime import timedelta
fe = FeatureEngineeringClient()
# Define a streaming feature
stream_source = StreamSource(full_name="my_catalog.my_schema.my_stream")
streaming_feature = Feature(
name="user_purchase_sum",
source=stream_source,
entity=["value.user_id"],
timeseries_column="value.event_time",
function=AggregationFunction(
operator=Sum(input="value.amount"),
time_window=RollingWindow(window_duration=timedelta(hours=1)),
),
)
# Create training set — reads from the ingestion table
# labeled_df must contain "user_id", "event_time", and label columns.
# Entity and timeseries columns use leaf node names (not value. prefixes).
training_set = fe.create_training_set(
df=labeled_df,
features=[streaming_feature],
label="is_fraud",
)
training_df = training_set.load_df()
混音批次與串流功能
批次與串流功能可在同一訓練集與模型中同時使用。 在提供服務時,批次特徵會從離線或線上儲存庫擷取,而串流特徵則會從線上儲存庫擷取。
training_set = fe.create_training_set(
df=labeled_df,
features=[batch_feature, streaming_feature],
label="is_fraud",
)
被 log_model() 登錄的模型會從線上商店進行特徵查詢,並為兩種來源類型配置模型簽章。
服刑時,原始模型會到達什麼
Feature Store模型包裝器會在將欄位傳遞給原始模型前先過濾。
| 欄類型 | 能達到內在模型嗎? |
|---|---|
明確特徵輸出(ColumnSelection, 聚合) |
是的 |
RequestSource 欄位被宣告為功能特徵 |
是的 |
| 實體欄位(查找鍵) | 沒有(除非明確標示為功能) |
| 時間序列欄位 | 沒有(除非明確標示為功能) |