You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Whisper Cluster

A K-means codebook for Japanese audio from Lambda Audio Corpus, using the Whisper large-v3-turbo encoder with TensorRT FP16 inference.

Each 1,280-dimensional encoder frame can be assigned to its nearest cluster center to produce a discrete acoustic unit. The codebook has 500 clusters, with IDs from 0 to 499. These units are intended as targets for speech representation learning; they are not text tokens or labeled phonemes.

Representation Format

Property Value
Encoder openai/whisper-large-v3-turbo
Representation Final encoder output, before the text decoder
Inference TensorRT FP16
Audio sampling rate 16,000 Hz, mono
Maximum input duration 30 seconds per chunk
Mel channels 128
Frame interval 20 ms, or 50 frames per second
Dimensions 1,280 per frame
Stored features and centers Float32, without feature normalization
Assignment distance Squared Euclidean distance
Clusters 500

Files

Files are stored under codebook/. Cluster assignment requires only centroids.npy and the matching Whisper encoder. Original audio and transcripts are not included.

File Contents
codebook/centroids.npy NumPy float32 array of shape (500, 1280). Row i is the center of cluster ID i. Approximately 2.56 MB.
codebook/codebook.json Feature provenance, K-means settings, completed iterations, total squared-distance inertia, and centroid filename.
codebook/features.f32 Headerless, row-major float32 matrix of shape (sampled_frames, 1280), containing the sampled encoder frames used to fit the codebook.
codebook/features.json Dataset and encoder revisions, preprocessing, accepted audio duration, frame counts, sampling probability, seed, and extraction settings.
codebook/chunks.npy NumPy int64 array of shape (number_of_batches, 2). Each row is a half-open feature row range [start, end). These are feature batch boundaries, not individual audio clip boundaries.
codebook/audio.jsonl One JSON object per accepted audio chunk. Contains source IDs and audio intervals, not audio, features, or cluster IDs.

Exact frame counts, audio duration, feature file size, and training results are recorded in the JSON metadata.

Audio Manifest

Field Description
id Audio record ID in the source dataset. Multiple chunks may share an ID.
origin Source origin as a string. "None" indicates unavailable origin information rather than a known source category.
start_sample Start offset within the source audio after conversion to 16 kHz mono. Divide by 16,000 to obtain seconds.
num_samples Number of accepted audio samples at 16 kHz. Divide by 16,000 to obtain duration in seconds.

The manifest does not record the random frame selection mask or per-clip feature row offsets. It cannot recover the exact timestamp of each sampled row in features.f32.

Metadata Reference

Both JSON files are UTF-8. Counts and sample offsets are integers; durations, probabilities, and training statistics are JSON numbers. null means that no value was recorded.

Feature Provenance and Representation

These fields appear in features.json and inside the features object of codebook.json.

Field Meaning
dataset, dataset_revision, split Source dataset ID, resolved commit, and split.
model, model_revision Encoder model ID and resolved revision. Use these to identify the matching encoder.
model_sha256 Hash of a supplied local model file, or null when no separate file hash was recorded.
preprocessing Padding and Mel configuration: each chunk is padded to 480,000 samples and converted to a 128-channel Whisper log-Mel spectrogram.
preprocessing_version Version of the openai-whisper package used for Mel preprocessing.
sample_rate Audio samples per second: 16,000.
frame_samples Samples between encoder frames: 320, corresponding to 20 ms. This is the temporal stride, not a receptive-field size.
dimensions Feature components per encoder frame and center: 1,280.
dtype Feature storage dtype: float32. FP16 encoder outputs are converted to float32.
distance Assignment metric: squared_euclidean.
normalized false: features are not L2-normalized.

Audio and Frame Statistics

Field Meaning
audio_samples Total accepted 16 kHz samples, excluding padding and rejected chunks. Equals the manifest's summed num_samples.
audio_hours Accepted duration: audio_samples / sample_rate / 3600. Includes silence and non-speech.
valid_frames Frames before random sampling, excluding padding. Each accepted chunk contributes ceil(num_samples / frame_samples), including its final partial frame.
sampled_frames Retained feature rows used for K-means training.
sample_probability Independent retention probability per valid frame. The realized fraction need not equal this probability exactly.
seed Seed for source shuffling and frame retention.
origins_seconds Accepted duration in seconds grouped by the manifest's origin string.
extraction_settings Recorded configuration used to produce features, including requested duration, sampling, batching, shuffling, workers, model and dataset selection, and clustering settings. A recorded output path is historical provenance, not a required download location.

Frame counts are calculated per chunk, so valid_frames need not equal ceil(audio_samples / 320). These statistics do not describe independent speaker or utterance counts. Requested settings may differ from actual accepted duration or completed training results; use the corresponding result fields.

Codebook and Training Metadata

codebook.json contains three top-level entries:

Field Meaning
features Copy of the feature metadata described above.
training Settings and results of the completed K-means fit.
centroids Center filename, relative to codebook.json: centroids.npy.
Field within training Meaning
algorithm Clustering implementation: cuml_kmeans.
clusters Number of centers.
max_iter, tol Iteration limit and convergence tolerance.
iterations Actual completed iterations.
inertia Total squared distance of training frames to their assigned centers.
batch_size Buffering and distance-processing setting; all sampled frames participate in fitting.
seed K-means random seed.

Binary Layout and Row Ordering

Artifact Layout and interpretation
features.f32 Little-endian IEEE 754 float32 in C row-major order. Each row occupies 1280 × 4 = 5120 bytes; row i starts at byte offset i × 5120. File size is sampled_frames × 5120 bytes. There is no NumPy header.
centroids.npy NumPy header records shape and dtype. Row index is the cluster ID.
chunks.npy Consecutive, nonempty ranges start at 0 and end at sampled_frames. A range selects features[start:end].
audio.jsonl UTF-8 JSON Lines in accepted chunk order. Chunks from different recordings may be interleaved.

Sampled rows retain chronological order within each chunk and follow accepted chunk order. They are not uniformly spaced: discarded frames and recording boundaries create gaps. Neither chunks.npy nor the manifest supplies a frame-level alignment map.

Read the Training Features

Download the large feature matrix only when needed. Use a memory map to avoid loading it entirely into RAM:

import json

import numpy as np
from huggingface_hub import hf_hub_download

repo_id = "KeisukeMiyamoto/whisper-cluster"
metadata_path = hf_hub_download(repo_id, "codebook/features.json")
with open(metadata_path, encoding="utf-8") as file:
    metadata = json.load(file)

features_path = hf_hub_download(repo_id, "codebook/features.f32")
features = np.memmap(
    features_path,
    dtype=np.float32,
    mode="r",
    shape=(metadata["sampled_frames"], metadata["dimensions"]),
)

Cluster Assignment

Use the same encoder and chunk preprocessing recorded in the codebook metadata. Assign each valid float32 encoder frame to its nearest center using squared Euclidean distance. Do not L2-normalize features or centers. Cluster assignment can use every valid frame; random frame retention is used for codebook training.

The clustering weights are the centers in codebook/centroids.npy; no additional clustering model is needed. The following CUDA example accepts an encoder feature array of shape (frames, 1280), with padding frames already removed, and returns one cluster ID per frame. The Whisper encoder weights are obtained separately from the model identified in the metadata.

import numpy as np
import torch
from huggingface_hub import hf_hub_download

path = hf_hub_download("KeisukeMiyamoto/whisper-cluster", "codebook/centroids.npy")
centers = torch.as_tensor(np.load(path), dtype=torch.float32, device="cuda")

@torch.inference_mode()
def assign_clusters(features: np.ndarray) -> np.ndarray:
    frames = torch.as_tensor(features, dtype=torch.float32, device="cuda")
    return torch.cdist(frames, centers).argmin(dim=1).cpu().numpy()

Call assign_clusters(features) for each audio chunk. Euclidean and squared Euclidean distances select the same nearest center; returned IDs are integers from 0 to 499.

Limitations

  • Cluster IDs may represent speech, silence, music, noise, or recording conditions; they have no predefined linguistic meaning.
  • Downstream speech recognition accuracy and performance on other languages or domains have not been established.
  • Frame sampling and rejection of invalid chunks affect the training distribution.
  • Cluster IDs from separately trained codebooks are not directly comparable.
  • A fixed seed and model revision do not guarantee identical numerical results across software and hardware environments.

Sources and Usage Terms

Consult the Lambda Audio Corpus dataset card and applicable original-source terms before using or redistributing derived representations. See the Whisper large-v3-turbo model card for model details and license information.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KeisukeMiyamoto/whisper-cluster

Finetuned
(638)
this model

Dataset used to train KeisukeMiyamoto/whisper-cluster