Instructions to use TeraSpace/TeraAnaly with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TeraSpace/TeraAnaly with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="TeraSpace/TeraAnaly", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("TeraSpace/TeraAnaly", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
TeraAnaly
Russian speech analysis with one frozen w2v-BERT encoder and independently callable age, gender, emotion, stress and CTC components. Decode audio once, run the encoder once, and reuse its output for any heads you need.
Noncommercial use only for the TeraAnaly code and trained heads, under CC-BY-NC-SA-4.0. The original encoder weights retain their upstream CC-BY-4.0 license. See LICENSE.md.
Files
| File | Contents |
|---|---|
modeling_tera_analy.py |
All model code, audio decoding, frontend and CTC alignment |
example_inference.py |
Example calling each head; edit the variables at the top |
config.json |
Architecture, labels, character vocabulary and source revisions |
encoder.safetensors |
Frozen encoder, about 2.32 GB |
age.safetensors |
Age component |
gender.safetensors |
Gender component |
emotion.safetensors |
Emotion component |
stress.safetensors |
Vowel stress ranker |
ctc.safetensors |
Native 20 ms character CTC component |
weights_manifest.json |
File sizes, SHA-256 checksums and component licenses |
Hugging Face inference
import json
import torch
from transformers import AutoModel
MODEL_ID = "TeraSpace/TeraAnaly"
AUDIO_PATH = "audio.wav"
TEXT = "Конечно, помню! Мы так здорово проводили время!"
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
torch.set_num_threads(4)
model = AutoModel.from_pretrained(MODEL_ID, trust_remote_code=True).to(DEVICE)
waveform = model.read_wav(AUDIO_PATH) # One decode; mono, 16 kHz.
states = model.encoder(waveform) # One frozen encoder pass.
result = {
"age": model.age_head(states),
"gender": model.gender_head(states),
"emotion": model.emotion_head(states),
"stressed_text": model.stress_head(states, TEXT),
}
print(json.dumps(result, ensure_ascii=False, indent=2))
Outputs are an integer age in years, the highest gender label, the highest emotion label, and a transcript with + before stressed vowels. Emotion labels are neutral, angry, positive, sad, other; demographic labels are female, male, child (child is an age category inherited from the teacher).
stress_head(states, text) calls our CTC component internally, pools vowel spans from layer 16, and selects a vowel separately in each Russian word. It requires the matching transcript. It marks single-vowel words too, preserves capitalization and punctuation, and collapses whitespace. It does not require Meta UnitY2, RuAccent, or the teacher models during inference.
Optional character alignment uses the same states:
logits = model.ctc_head(states)
alignment = model.ctc_head.align(states, TEXT)
For a single dictionary directly, use model(waveform, TEXT). This also runs the encoder once.
Batch inference
Pass a list of mono waveforms and a matching list of transcripts. Recordings may have different lengths. The encoder runs once for the entire batch; padding is masked and every result stays in input order.
audio_paths = ["first.wav", "second.wav"]
texts = ["Привет, Маша!", "Как дела?"]
waveforms = [model.read_wav(path) for path in audio_paths]
states = model.encoder(waveforms)
ages = model.age_head(states) # list[int]
genders = model.gender_head(states) # list[str]
emotions = model.emotion_head(states) # list[str]
stressed_texts = model.stress_head(states, texts)
For all four outputs as a list of dictionaries, use results = model(waveforms, texts). For a large collection, process it in smaller batches; example_inference.py does this using BATCH_SIZE.
One waveform still returns a dictionary of encoder states and scalar head outputs. A list (including a one-item list) returns a list of states and one output per recording. Empty lists return empty results. CTC also supports batches: model.ctc_head(states) returns a list of logits trimmed to each recording's frame count, and model.ctc_head.align(states, texts) returns a list of alignments. Stress computes CTC emissions and vowel scores in batches, with forced alignment and word selection performed separately per transcript.
For raw waveforms at different sampling rates, use model.encoder(waveforms, sample_rate=[16000, 22050]), or pass one integer rate for all clips. A dense [batch, samples] tensor is accepted for equally sized waveforms. For different lengths, pass the unpadded waveform list so that normalization and length detection use each recording's actual audio.
Encoder attention, demographic pooling and CTC convolutions mask padding independently. Frontend, pitch statistics and forced CTC decoding are still computed per recording. Batch memory use grows with the number of recordings and the longest clip; reduce the batch size for long audio, and group similar lengths when processing a dataset. Every clip retains the 40-second limit.
Load individual components
Download modeling_tera_analy.py from this repository into your script folder, then import it. Each component has its own loader:
from modeling_tera_analy import (
TeraEncoder, AgeHead, GenderHead, EmotionHead, StressHead, CTCHead, read_wav,
)
MODEL_ID = "TeraSpace/TeraAnaly" # A local model folder also works.
DEVICE = "cuda"
encoder = TeraEncoder.from_pretrained(MODEL_ID, device=DEVICE)
age = AgeHead.from_pretrained(MODEL_ID, device=DEVICE)
gender = GenderHead.from_pretrained(MODEL_ID, device=DEVICE)
emotion = EmotionHead.from_pretrained(MODEL_ID, device=DEVICE)
stress = StressHead.from_pretrained(MODEL_ID, device=DEVICE)
states = encoder(read_wav("audio.wav"))
age_years = age(states)
gender_label = gender(states)
emotion_label = emotion(states)
stressed_text = stress(states, "Привет, Маша!")
Load only the components you need. An age, gender or emotion loader reads its corresponding safetensors and configuration. The encoder loader reads only encoder weights and configuration. The stress loader reads stress and CTC weights; it does not load or rerun the encoder. CTCHead.from_pretrained(...) is available independently too.
states is a dictionary containing hidden_states indexed by layers 4, 8, 16 and 24, recording_features, spectral_features, raw_features, frame_count and frame_seconds. The existing trained heads also use frontend and prosody statistics; these are computed once by encoder(...) alongside the neural encoder. Pass the entire dictionary to a head.
Local example and dependencies
Edit MODEL_ID, AUDIO_PATH, TEXT, DEVICE and BATCH_SIZE at the top of example_inference.py, then run it. For batch inference, set AUDIO_PATH to a list of WAV paths and TEXT to their matching transcript list. It prints one dictionary for a single path, or a list of dictionaries for a path list. There are no command-line arguments. Relative audio paths are resolved relative to that file; absolute paths work too. Supply your own WAV file; recordings are not bundled on Hugging Face.
For offline use, download the repository, set MODEL_ID to that folder, and import TeraAnaly or its individual components directly. Local computation requires only Python's standard library and PyTorch, including a built-in safetensors reader. Hugging Face downloads require huggingface_hub; AutoModel requires transformers. No torchaudio, NumPy, safetensors, or external audio program is used by the model code.
Tested with Python 3.10, PyTorch 2.9.0 and Transformers 5.15.1. This release uses float32 inference. CPU and one CUDA device are supported; clips are limited to 40 seconds. The reader accepts RIFF WAV with PCM 8/16/24/32-bit or IEEE float32 audio, averages channels and resamples to 16 kHz. Alternatively pass a mono PyTorch waveform to encoder(waveform, sample_rate=...). Use Russian transcripts and spell out numbers.
Training and scope
The shared encoder comes from AigizK/w2v-bert-2.0-bashkort-russian-omnivoice, revision e4c4278904c6359e9ac769cf04575be68f99ed4e, based on facebook/w2v-bert-2.0. TeraAnaly training kept this encoder frozen. The unused upstream ASR adapter and ASR output layer are omitted from the export; retained encoder tensor values are unchanged.
Age and gender training used predictions from audeering/wav2vec2-large-robust-24-ft-age-gender, together with available speaker gender labels. Emotion training used compatible recorder labels from Russian emotional dialogues and predictions from xbgoose/hubert-large-speech-emotion-recognition-russian-dusha-finetuned. The training pool also included Russian ewebis, epodcasts and studio recordings. Stress training used RuAccent-derived weak labels and frozen audio embeddings; the trained native CTC component supplies timings at inference.
Age, gender and emotion are model estimates. Stress agreement with automatic RuAccent labels measures agreement with those labels, rather than verified spoken stress accuracy. There is no independent human-labeled spoken-stress accuracy claim for this release.
The export was checked against the existing inference pipeline on a female and a male recording: all four returned outputs matched. Independent encoder loading produced identical hidden states, and separately loaded heads reproduced the bundled heads. Each sample used one encoder pass and stress called CTC internally once. Both local AutoModel(..., trust_remote_code=True) and local execution with optional dependencies disabled were checked. These checks establish export equivalence, not dataset-wide accuracy.
Batch checks compared mixed-length recordings against single inference, including reordered and repeated items, one-item batches, dense waveform tensors, mixed sampling rates and empty inputs. Returned age, gender, emotion and stressed text matched the single-recording results. The encoder, CTC emissions and vowel ranker each ran once per batch. Padding equivalence was also checked with IEEE float32 convolution kernels; CUDA's default TF32 kernels can produce small floating-point differences between single and batch computation.
Attribution and license
Copyright 2026 TeraSpace for the new code and trained heads. These are licensed under CC-BY-NC-SA-4.0. Preserve attribution and distribute adaptations under the same license. The original encoder remains CC-BY-4.0, with attribution to its upstream authors; this exception does not grant commercial use of the TeraAnaly heads or code. Component details and teacher notices are in LICENSE.md, and the full noncommercial license text is in LICENSE.
- Downloads last month
- 41
Model tree for TeraSpace/TeraAnaly
Base model
facebook/w2v-bert-2.0