使用 Whisper 進行批次語音轉文字

使用 OpenAI Whisper large-v3-turbo 在已連接的 A10 AI Runtime 上轉錄一批英文語音錄音。 本筆記本示範如何:

  • 使用 Transformers 管線載入 Whisper large-v3-turbo 模型。
  • 從 LibriSpeech 測試集製作一批音訊樣本。
  • 視覺化每個樣本的波形與頻譜圖。
  • 執行批次轉錄,並比較其吞吐量與序列推論。

Note

此範例需要 Databricks AI 環境版本 6 或以上。

連接到無伺服器的 GPU 運算

  1. 從筆記型電腦的運算選擇器中,選擇 無伺服器 GPU。
  2. 在 環境 面板中,選擇 A10 加速器和 AI v6 環境。
  3. 點選 套用,然後確認環境。

Whisper 模型與 LibriSpeech 範例資料集皆為公開,且不需 Hugging Face 驗證。

匯入程式庫

AI 環境包含本筆記本中使用的 PyTorch、Transformers 及 Hugging Face Datasets 套件,因此無需安裝套件。 這個儲存單元會匯入這些資料並確認有 GPU 連接。

import torch
import transformers
import datasets

print(f"Environment")
print(f"   PyTorch:      {torch.__version__}")
print(f"   Transformers: {transformers.__version__}")
print(f"   Datasets:     {datasets.__version__}")
print(f"\nGPU")
print(f"   Available:    {torch.cuda.is_available()}")
if torch.cuda.is_available():
    print(f"   Device:       {torch.cuda.get_device_name(0)}")
    mem_gb = torch.cuda.get_device_properties(0).total_memory / 1e9
    print(f"   Memory:       {mem_gb:.1f} GB")
Environment
   PyTorch:      2.11.0+cu130
   Transformers: 5.8.1
   Datasets:     4.8.5

GPU
   Available:    True
   Device:       NVIDIA A10G
   Memory:       23.7 GB

載入 Whisper 模型

Load openai/whisper-large-v3-turbo是一個精煉的 809M 參數模型,能以更快的推論速度提供接近最先進的準確度。 Transformers 管線可在單一呼叫中處理特徵擷取、分詞化與解碼。

from transformers import pipeline
import torch

whisper_pipe = pipeline(
    "automatic-speech-recognition",
    model="openai/whisper-large-v3-turbo",
    torch_dtype=torch.float16,
    device="cuda",
)
print(f"Model loaded on {whisper_pipe.device}")
Model loaded on cuda

載入並探索音訊樣本

載入 LibriSpeech ASR 測試集,這是一組乾淨的英語語音錄音,並附有參考轉錄,並播放第一個取樣。

from datasets import load_dataset, Audio as AudioFeature
from IPython.display import display, Audio
import numpy as np
import soundfile as sf
import io

# Load the LibriSpeech test samples (decode=False to avoid torchcodec/FFmpeg dependency)
ds = load_dataset(
    "hf-internal-testing/librispeech_asr_dummy", "clean", split="validation"
)
ds = ds.cast_column("audio", AudioFeature(decode=False))
print(f"Loaded {len(ds)} audio samples\n")

def decode_audio(raw):
    """Decode raw audio bytes with soundfile."""
    arr, sr = sf.read(io.BytesIO(raw["bytes"]))
    return {"array": arr, "sampling_rate": sr}

# Show metadata for first few samples
for i in range(5):
    audio = decode_audio(ds[i]["audio"])
    duration = len(audio["array"]) / audio["sampling_rate"]
    text_preview = ds[i]["text"][:80]
    print(f"  Sample {i+1}:  {duration:.2f}s  |  {audio['sampling_rate']} Hz  |  \"{text_preview}...\"")

# Play the first sample inline
print("\n>> Playing Sample 1:")
audio_0 = decode_audio(ds[0]["audio"])
display(Audio(audio_0["array"], rate=audio_0["sampling_rate"]))
Loaded 73 audio samples

  Sample 1:  5.86s  |  16000 Hz  |  "MISTER QUILTER IS THE APOSTLE OF THE MIDDLE CLASSES AND WE ARE GLAD TO WELCOME H..."
  Sample 2:  4.82s  |  16000 Hz  |  "NOR IS MISTER QUILTER'S MANNER LESS INTERESTING THAN HIS MATTER..."
  Sample 3:  12.48s  |  16000 Hz  |  "HE TELLS US THAT AT THIS FESTIVE SEASON OF THE YEAR WITH CHRISTMAS AND ROAST BEE..."
  Sample 4:  9.90s  |  16000 Hz  |  "HE HAS GRAVE DOUBTS WHETHER SIR FREDERICK LEIGHTON'S WORK IS REALLY GREEK AFTER ..."
  Sample 5:  29.40s  |  16000 Hz  |  "LINNELL'S PICTURES ARE A SORT OF UP GUARDS AND AT EM PAINTINGS AND MASON'S EXQUI..."

>> Playing Sample 1:

視覺化音訊波形

將每個樣本的波形與頻譜圖並排繪製。 波形顯示隨時間的振幅變化,頻譜圖則顯示頻率內容,即語音能量集中於100至4,000赫茲的語音頻段。

import matplotlib.pyplot as plt
import numpy as np

NUM_SAMPLES = 4
fig, axes = plt.subplots(NUM_SAMPLES, 2, figsize=(16, 3 * NUM_SAMPLES))
fig.suptitle(
    "Waveform & Spectrogram Profiles", fontsize=16, fontweight="bold", y=1.01
)

for i in range(NUM_SAMPLES):
    audio = decode_audio(ds[i]["audio"])
    samples = audio["array"]
    sr = audio["sampling_rate"]
    t = np.arange(len(samples)) / sr

    # --- Waveform ---
    ax_wave = axes[i, 0]
    ax_wave.plot(t, samples, linewidth=0.4, color="#1f77b4", alpha=0.8)
    ax_wave.fill_between(t, samples, alpha=0.15, color="#1f77b4")
    ax_wave.set_ylabel("Amplitude", fontsize=9)
    ax_wave.set_title(f"Sample {i+1} — Waveform ({len(samples)/sr:.1f}s)", fontsize=10)
    ax_wave.set_xlim(0, t[-1])
    ax_wave.grid(True, alpha=0.3)
    if i == NUM_SAMPLES - 1:
        ax_wave.set_xlabel("Time (seconds)", fontsize=9)

    # --- Spectrogram ---
    ax_spec = axes[i, 1]
    ax_spec.specgram(samples, Fs=sr, NFFT=1024, noverlap=512, cmap="magma")
    ax_spec.set_ylabel("Frequency (Hz)", fontsize=9)
    ax_spec.set_title(f"Sample {i+1} — Spectrogram", fontsize=10)
    ax_spec.set_ylim(0, 8000)  # Focus on speech frequencies
    if i == NUM_SAMPLES - 1:
        ax_spec.set_xlabel("Time (seconds)", fontsize=9)

plt.tight_layout()
plt.show()

轉錄一個樣本

先轉錄一個樣本以驗證管線,然後將預測結果與參考文本進行比較。

import time

audio_input = decode_audio(ds[0]["audio"])

start = time.perf_counter()
result = whisper_pipe(
    audio_input["array"],
    generate_kwargs={"language": "en"},
)
elapsed = time.perf_counter() - start

duration = len(audio_input["array"]) / audio_input["sampling_rate"]

print(f"Inference time:   {elapsed:.2f}s for {duration:.1f}s audio ({duration/elapsed:.1f}x realtime)")
print(f"\nPredicted:  {result['text'].strip()}")
print(f"Reference:  {ds[0]['text']}")
Inference time:   12.23s for 5.9s audio (0.5x realtime)

Predicted:  Mr. Quilter is the apostle of the middle classes, and we are glad to welcome his gospel.
Reference:  MISTER QUILTER IS THE APOSTLE OF THE MIDDLE CLASSES AND WE ARE GLAD TO WELCOME HIS GOSPEL

執行批次推論

以可設定的batch_size轉錄所有樣本。 批次處理讓 GPU 能同時處理多個音訊片段,這比順序推理更能提升吞吐量。 此儲存格也會對循序執行基準進行計時,以供比較。

import time
import pandas as pd

_decoded = [decode_audio(ds[i]["audio"]) for i in range(len(ds))]
audio_inputs = [d["array"] for d in _decoded]
_sr = _decoded[0]["sampling_rate"]

# --- Sequential baseline ---
start = time.perf_counter()
seq_results = [
    whisper_pipe(a, generate_kwargs={"language": "en"}) for a in audio_inputs
]
seq_time = time.perf_counter() - start

# --- Batched inference ---
start = time.perf_counter()
batch_results = whisper_pipe(
    audio_inputs, batch_size=8, generate_kwargs={"language": "en"}
)
batch_time = time.perf_counter() - start

total_audio_sec = sum(
    len(a) / _sr for a in audio_inputs
)

print(f"Performance Comparison ({len(audio_inputs)} samples, {total_audio_sec:.1f}s total audio)")
print(f"   Sequential:  {seq_time:.2f}s  ({total_audio_sec/seq_time:.1f}x realtime)")
print(f"   Batched (8): {batch_time:.2f}s  ({total_audio_sec/batch_time:.1f}x realtime)")
print(f"   Speedup:     {seq_time/batch_time:.2f}x\n")

# --- Results table ---
rows = []
for i, res in enumerate(batch_results):
    duration = len(audio_inputs[i]) / _sr
    rows.append({
        "Sample": i + 1,
        "Duration (s)": round(duration, 1),
        "Transcription": res["text"].strip(),
        "Reference": ds[i]["text"],
    })

df = pd.DataFrame(rows)
display(df)
Performance Comparison (73 samples, 481.0s total audio)
   Sequential:  16.02s  (30.0x realtime)
   Batched (8): 9.56s  (50.3x realtime)
   Speedup:     1.68x

總結

這本筆記本展示了:

  • 免設定 GPU 推論:AI 執行階段環境已預先安裝torch、transformers和datasets;無需%pip install
  • 內嵌音訊播放:使用 IPython.display.Audio 直接在筆記本中聆聽範例音訊
  • 波形與頻譜圖視覺化:以 matplotlib 渲染,且已預先安裝
  • 高效的批次推論:利用 batch_size 參數進行 GPU 平行轉錄
  • Whisper large-v3-turbo:一款快速且精確、適用於正式部署的語音轉文字模型

若要調整此資料以符合你自己的資料,請將 HuggingFace 資料集替換成來自 Unity 目錄 卷或 雲端儲存路徑的音訊檔案。

範例筆記本

使用 Whisper 進行批次語音轉文字

拿筆記本