使用 Kernel SHAP(SHapley 加法解釋)來解釋一個表格分類模型。 Kernel SHAP 是一種與模型無關的方法,用來估算每個特徵對模型預測結果的貢獻。 你會在成人普查收入資料集上訓練邏輯迴歸模型,然後使用 SynapseML TabularSHAP 轉換器計算特徵層級的解釋。
先決條件
取得 Microsoft Fabric 訂用帳戶。 或註冊免費的 Microsoft Fabric 試用版。
登入 Microsoft Fabric。
使用首頁左下角的體驗切換器切換到 Fabric。
- 在你的工作區建立一個新筆記本,並連接到湖邊小屋。 欲了解更多資訊,請參閱 建立筆記本。
SynapseML、PySpark、pandas 和 plotly 都預裝在 Fabric 筆記本環境中。 不需要額外安裝套件。
匯入套件並定義輔助 UDF 函式
在你的 Fabric 筆記本裡,把以下程式碼貼到一個儲存格裡並執行。 此步驟匯入所需的函式庫,並定義兩個使用者自訂函式(UDF)以供後續擷取向量元素。
import pyspark
from synapse.ml.explainers import TabularSHAP
from pyspark.ml import Pipeline
from pyspark.ml.classification import LogisticRegression
from pyspark.ml.feature import StringIndexer, OneHotEncoder, VectorAssembler
from pyspark.sql.types import FloatType, ArrayType
from pyspark.sql.functions import col, lit, rand, broadcast, udf
import pandas as pd
vec_access = udf(lambda v, i: float(v[i]), FloatType())
vec2array = udf(lambda vec: vec.toArray().tolist(), ArrayType(FloatType()))
驗證:在新儲存格執行以下程式碼。 你應該會看到輸出 TabularSHAP imported successfully。
print("TabularSHAP imported successfully")
print(f"PySpark version: {pyspark.__version__}")
載入資料並訓練分類模型
從 Azure Blob 儲存體 載入成人普查收入資料集,索引目標標籤,並訓練邏輯迴歸管線。
df = spark.read.parquet(
"wasbs://publicwasb@mmlspark.blob.core.windows.net/AdultCensusIncome.parquet"
)
labelIndexer = StringIndexer(
inputCol="income", outputCol="label", stringOrderType="alphabetAsc"
).fit(df)
print("Label index assignment: " + str(set(zip(labelIndexer.labels, [0, 1]))))
training = labelIndexer.transform(df).cache()
categorical_features = [
"workclass",
"education",
"marital-status",
"occupation",
"relationship",
"race",
"sex",
"native-country",
]
categorical_features_idx = [feat + "_idx" for feat in categorical_features]
categorical_features_enc = [feat + "_enc" for feat in categorical_features]
numeric_features = [
"age",
"education-num",
"capital-gain",
"capital-loss",
"hours-per-week",
]
strIndexer = StringIndexer(
inputCols=categorical_features, outputCols=categorical_features_idx
)
onehotEnc = OneHotEncoder(
inputCols=categorical_features_idx, outputCols=categorical_features_enc
)
vectAssem = VectorAssembler(
inputCols=categorical_features_enc + numeric_features, outputCol="features"
)
lr = LogisticRegression(featuresCol="features", labelCol="label", weightCol="fnlwgt")
pipeline = Pipeline(stages=[strIndexer, onehotEnc, vectAssem, lr])
model = pipeline.fit(training)
驗證:執行以下儲存格。 你應該會看到訓練資料的列數和流程階段的確認。
print(f"Training rows: {training.count()}")
print(f"Pipeline stages: {[type(s).__name__ for s in model.stages]}")
assert training.count() > 30000, "Dataset should contain over 30,000 rows"
print("Model trained successfully")
# Expected output:
#Training rows: 32561
#Pipeline stages: ['StringIndexerModel', 'OneHotEncoderModel', #'VectorAssembler', 'LogisticRegressionModel']
#Model trained successfully
選擇觀察結果加以說明
從評分訓練資料中隨機選出五個觀察值。 這些觀察就是你用來產生 SHAP 解釋的實例。
explain_instances = (
model.transform(training).orderBy(rand()).limit(5).repartition(200).cache()
)
display(explain_instances)
驗證:確認樣本大小。
count = explain_instances.count()
print(f"Explain instances: {count}")
assert count == 5, f"Expected 5 rows, got {count}"
print("Sample selected successfully")
配置並執行 TabularSHAP
建立 TabularSHAP 一個說明文,並將其應用於所選觀察結果。 關鍵參數包括:
| 參數 | Description |
|---|---|
inputCols |
模型用於預測時所使用的特徵欄位。 |
outputCol |
包含 SHAP 輸出值的欄位名稱。 |
numSamples |
Kernel SHAP 估算用的微擾樣本數。 數值越高越準確,但速度越慢。 |
model |
用訓練好的管線模型來解釋。 |
targetCol |
要說明的模型輸出欄位。 在這個例子中,欄位為 probability。 |
targetClasses |
用階級指數來解釋。
[1] 只解釋了第一類機率。 使用 [0, 1] 說明這兩個類別。 |
backgroundData |
作為整合特徵的參考分布的訓練資料樣本。 |
shap = TabularSHAP(
inputCols=categorical_features + numeric_features,
outputCol="shapValues",
numSamples=5000,
model=model,
targetCol="probability",
targetClasses=[1],
backgroundData=broadcast(training.orderBy(rand()).limit(100).cache()),
)
shap_df = shap.transform(explain_instances)
Note
這個步驟可能需要幾分鐘,視叢集大小而定 numSamples 。 使用 numSamples=5000 和五筆觀測資料時,在預設的 Fabric Spark 叢集上預計需時約 3 到 10 分鐘。
驗證:檢查 SHAP 輸出欄位是否存在。
assert "shapValues" in shap_df.columns, "shapValues column missing"
print(f"SHAP output columns: {shap_df.columns}")
print("TabularSHAP transform completed")
擷取 SHAP 值
從結果資料框中提取類別 1 機率與 SHAP 值。 對於每個觀測值,SHAP 值向量以基礎值(背景資料集的平均輸出)開始,接著是每個特徵的一個值。
shaps = (
shap_df.withColumn("probability", vec_access(col("probability"), lit(1)))
.withColumn("shapValues", vec2array(col("shapValues").getItem(0)))
.select(
["shapValues", "probability", "label"] + categorical_features + numeric_features
)
)
shaps_local = shaps.toPandas()
shaps_local.sort_values("probability", ascending=False, inplace=True, ignore_index=True)
pd.set_option("display.max_colwidth", None)
display(shaps_local)
驗證:確認 pandas DataFrame 的結構。
expected_cols = len(categorical_features) + len(numeric_features) + 3
print(f"DataFrame shape: {shaps_local.shape}")
print(f"Expected columns: {expected_cols}, Actual: {shaps_local.shape[1]}")
assert shaps_local.shape == (5, expected_cols), f"Unexpected shape: {shaps_local.shape}"
print("SHAP values extracted successfully")
視覺化 SHAP 值
為每個觀測建立一條條狀圖,顯示每個特徵如何影響預測機率。
from plotly.subplots import make_subplots
import plotly.graph_objects as go
features = categorical_features + numeric_features
features_with_base = ["Base"] + features
rows = shaps_local.shape[0]
fig = make_subplots(
rows=rows,
cols=1,
subplot_titles="Probability: "
+ shaps_local["probability"].apply("{:.2%}".format)
+ "; Label: "
+ shaps_local["label"].astype(str),
)
for index, row in shaps_local.iterrows():
feature_values = [0] + [row[feature] for feature in features]
shap_values = row["shapValues"]
list_of_tuples = list(zip(features_with_base, feature_values, shap_values))
shap_pdf = pd.DataFrame(list_of_tuples, columns=["name", "value", "shap"])
fig.add_trace(
go.Bar(
x=shap_pdf["name"],
y=shap_pdf["shap"],
hovertext="value: " + shap_pdf["value"].astype(str),
),
row=index + 1,
col=1,
)
fig.update_yaxes(range=[-1, 1], fixedrange=True, zerolinecolor="black")
fig.update_xaxes(type="category", tickangle=45, fixedrange=True)
fig.update_layout(height=400 * rows, title_text="SHAP explanations")
fig.show()
驗證:確認該地塊物件已被創建。
print(f"Figure traces: {len(fig.data)}")
print(f"Figure height: {fig.layout.height}px")
assert len(fig.data) == 5, f"Expected 5 traces, got {len(fig.data)}"
print("Visualization created successfully")
解譯結果
每個子線代表一個觀察點。 這些欄杆顯示:
- 基礎:背景資料集的平均模型輸出(基線機率)。
- 正向 SHAP 值:將預測推向類別 1(收入大於 50K)的特徵。
- 負的SHAP值:使預測趨近於0類(收入小於或等於50K)的特徵。
基值與所有特徵 SHAP 值的總和等於模型對該觀察的預測機率。
Troubleshooting
| Issue | 原因 | Resolution |
|---|---|---|
OutOfMemoryError 在 TabularSHAP 期間 |
numSamples 對可用記憶體而言過大。 |
例如將 減少 numSamples到 1,000,或增加 Spark 執行者記憶體。 |
| SHAP 轉換速度很慢 | 高 numSamples 且多特徵會增加計算時間。 |
將 numSamples 降至 1,000 到 2,000,以更快取得探索結果。 增加用於最終分析的量。 |
FileNotFoundException 適用於拼花地板 |
網路存取 mmlspark.blob.core.windows.net 被封鎖。 |
確認你的 Fabric 工作區是否有外接網路。 或者,也可以將資料集上傳到你的湖屋。 |
shapValues 欄位包含零值 |
若特徵值超出訓練分布,部分觀察可能會失效。 | 檢查輸入特徵是否有空值或意外值。 從結果中過濾空值。 |
display() 顯示無輸出 |
程式碼是在 Fabric 筆記本環境外執行的。 | 在標準Python環境中使用 shaps_local.head() 或 print(shaps_local)。 |
清理
如果你為了這個教學把資料集上傳到 Lakehouse,請移除到免費儲存:
# Remove cached DataFrames from memory
training.unpersist()
explain_instances.unpersist()
print("Cached DataFrames released")