Whisper Cluster
A K-means codebook for Japanese audio from Lambda Audio Corpus, using the Whisper large-v3-turbo encoder with TensorRT FP16 inference.
Each 1,280-dimensional encoder frame can be assigned to its nearest cluster center to produce a discrete acoustic unit. The codebook has 500 clusters, with IDs from 0 to 499. These units are intended as targets for speech representation learning; they are not text tokens or labeled phonemes.
Representation Format
| Property | Value |
|---|---|
| Encoder | openai/whisper-large-v3-turbo |
| Representation | Final encoder output, before the text decoder |
| Inference | TensorRT FP16 |
| Audio sampling rate | 16,000 Hz, mono |
| Maximum input duration | 30 seconds per chunk |
| Mel channels | 128 |
| Frame interval | 20 ms, or 50 frames per second |
| Dimensions | 1,280 per frame |
| Stored features and centers | Float32, without feature normalization |
| Assignment distance | Squared Euclidean distance |
| Clusters | 500 |
Files
Files are stored under codebook/. Cluster assignment requires only centroids.npy and the matching Whisper encoder. Original audio and transcripts are not included.
| File | Contents |
|---|---|
codebook/centroids.npy |
NumPy float32 array of shape (500, 1280). Row i is the center of cluster ID i. Approximately 2.56 MB. |
codebook/codebook.json |
Feature provenance, K-means settings, completed iterations, total squared-distance inertia, and centroid filename. |
codebook/features.f32 |
Headerless, row-major float32 matrix of shape (sampled_frames, 1280), containing the sampled encoder frames used to fit the codebook. |
codebook/features.json |
Dataset and encoder revisions, preprocessing, accepted audio duration, frame counts, sampling probability, seed, and extraction settings. |
codebook/chunks.npy |
NumPy int64 array of shape (number_of_batches, 2). Each row is a half-open feature row range [start, end). These are feature batch boundaries, not individual audio clip boundaries. |
codebook/audio.jsonl |
One JSON object per accepted audio chunk. Contains source IDs and audio intervals, not audio, features, or cluster IDs. |
Exact frame counts, audio duration, feature file size, and training results are recorded in the JSON metadata.
Audio Manifest
| Field | Description |
|---|---|
id |
Audio record ID in the source dataset. Multiple chunks may share an ID. |
origin |
Source origin as a string. "None" indicates unavailable origin information rather than a known source category. |
start_sample |
Start offset within the source audio after conversion to 16 kHz mono. Divide by 16,000 to obtain seconds. |
num_samples |
Number of accepted audio samples at 16 kHz. Divide by 16,000 to obtain duration in seconds. |
The manifest does not record the random frame selection mask or per-clip feature row offsets. It cannot recover the exact timestamp of each sampled row in features.f32.
Metadata Reference
Both JSON files are UTF-8. Counts and sample offsets are integers; durations, probabilities, and training statistics are JSON numbers. null means that no value was recorded.
Feature Provenance and Representation
These fields appear in features.json and inside the features object of codebook.json.
| Field | Meaning |
|---|---|
dataset, dataset_revision, split |
Source dataset ID, resolved commit, and split. |
model, model_revision |
Encoder model ID and resolved revision. Use these to identify the matching encoder. |
model_sha256 |
Hash of a supplied local model file, or null when no separate file hash was recorded. |
preprocessing |
Padding and Mel configuration: each chunk is padded to 480,000 samples and converted to a 128-channel Whisper log-Mel spectrogram. |
preprocessing_version |
Version of the openai-whisper package used for Mel preprocessing. |
sample_rate |
Audio samples per second: 16,000. |
frame_samples |
Samples between encoder frames: 320, corresponding to 20 ms. This is the temporal stride, not a receptive-field size. |
dimensions |
Feature components per encoder frame and center: 1,280. |
dtype |
Feature storage dtype: float32. FP16 encoder outputs are converted to float32. |
distance |
Assignment metric: squared_euclidean. |
normalized |
false: features are not L2-normalized. |
Audio and Frame Statistics
| Field | Meaning |
|---|---|
audio_samples |
Total accepted 16 kHz samples, excluding padding and rejected chunks. Equals the manifest's summed num_samples. |
audio_hours |
Accepted duration: audio_samples / sample_rate / 3600. Includes silence and non-speech. |
valid_frames |
Frames before random sampling, excluding padding. Each accepted chunk contributes ceil(num_samples / frame_samples), including its final partial frame. |
sampled_frames |
Retained feature rows used for K-means training. |
sample_probability |
Independent retention probability per valid frame. The realized fraction need not equal this probability exactly. |
seed |
Seed for source shuffling and frame retention. |
origins_seconds |
Accepted duration in seconds grouped by the manifest's origin string. |
extraction_settings |
Recorded configuration used to produce features, including requested duration, sampling, batching, shuffling, workers, model and dataset selection, and clustering settings. A recorded output path is historical provenance, not a required download location. |
Frame counts are calculated per chunk, so valid_frames need not equal ceil(audio_samples / 320). These statistics do not describe independent speaker or utterance counts. Requested settings may differ from actual accepted duration or completed training results; use the corresponding result fields.
Codebook and Training Metadata
codebook.json contains three top-level entries:
| Field | Meaning |
|---|---|
features |
Copy of the feature metadata described above. |
training |
Settings and results of the completed K-means fit. |
centroids |
Center filename, relative to codebook.json: centroids.npy. |
Field within training |
Meaning |
|---|---|
algorithm |
Clustering implementation: cuml_kmeans. |
clusters |
Number of centers. |
max_iter, tol |
Iteration limit and convergence tolerance. |
iterations |
Actual completed iterations. |
inertia |
Total squared distance of training frames to their assigned centers. |
batch_size |
Buffering and distance-processing setting; all sampled frames participate in fitting. |
seed |
K-means random seed. |
Binary Layout and Row Ordering
| Artifact | Layout and interpretation |
|---|---|
features.f32 |
Little-endian IEEE 754 float32 in C row-major order. Each row occupies 1280 × 4 = 5120 bytes; row i starts at byte offset i × 5120. File size is sampled_frames × 5120 bytes. There is no NumPy header. |
centroids.npy |
NumPy header records shape and dtype. Row index is the cluster ID. |
chunks.npy |
Consecutive, nonempty ranges start at 0 and end at sampled_frames. A range selects features[start:end]. |
audio.jsonl |
UTF-8 JSON Lines in accepted chunk order. Chunks from different recordings may be interleaved. |
Sampled rows retain chronological order within each chunk and follow accepted chunk order. They are not uniformly spaced: discarded frames and recording boundaries create gaps. Neither chunks.npy nor the manifest supplies a frame-level alignment map.
Read the Training Features
Download the large feature matrix only when needed. Use a memory map to avoid loading it entirely into RAM:
import json
import numpy as np
from huggingface_hub import hf_hub_download
repo_id = "KeisukeMiyamoto/whisper-cluster"
metadata_path = hf_hub_download(repo_id, "codebook/features.json")
with open(metadata_path, encoding="utf-8") as file:
metadata = json.load(file)
features_path = hf_hub_download(repo_id, "codebook/features.f32")
features = np.memmap(
features_path,
dtype=np.float32,
mode="r",
shape=(metadata["sampled_frames"], metadata["dimensions"]),
)
Cluster Assignment
Use the same encoder and chunk preprocessing recorded in the codebook metadata. Assign each valid float32 encoder frame to its nearest center using squared Euclidean distance. Do not L2-normalize features or centers. Cluster assignment can use every valid frame; random frame retention is used for codebook training.
The clustering weights are the centers in codebook/centroids.npy; no additional clustering model is needed. The following CUDA example accepts an encoder feature array of shape (frames, 1280), with padding frames already removed, and returns one cluster ID per frame. The Whisper encoder weights are obtained separately from the model identified in the metadata.
import numpy as np
import torch
from huggingface_hub import hf_hub_download
path = hf_hub_download("KeisukeMiyamoto/whisper-cluster", "codebook/centroids.npy")
centers = torch.as_tensor(np.load(path), dtype=torch.float32, device="cuda")
@torch.inference_mode()
def assign_clusters(features: np.ndarray) -> np.ndarray:
frames = torch.as_tensor(features, dtype=torch.float32, device="cuda")
return torch.cdist(frames, centers).argmin(dim=1).cpu().numpy()
Call assign_clusters(features) for each audio chunk. Euclidean and squared Euclidean distances select the same nearest center; returned IDs are integers from 0 to 499.
Limitations
- Cluster IDs may represent speech, silence, music, noise, or recording conditions; they have no predefined linguistic meaning.
- Downstream speech recognition accuracy and performance on other languages or domains have not been established.
- Frame sampling and rejection of invalid chunks affect the training distribution.
- Cluster IDs from separately trained codebooks are not directly comparable.
- A fixed seed and model revision do not guarantee identical numerical results across software and hardware environments.
Sources and Usage Terms
Consult the Lambda Audio Corpus dataset card and applicable original-source terms before using or redistributing derived representations. See the Whisper large-v3-turbo model card for model details and license information.
- Downloads last month
- -
Model tree for KeisukeMiyamoto/whisper-cluster
Base model
openai/whisper-large-v3