Instructions to use daltonoscar0/caesura-breaks with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use daltonoscar0/caesura-breaks with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="daltonoscar0/caesura-breaks")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("daltonoscar0/caesura-breaks") model = AutoModelForTokenClassification.from_pretrained("daltonoscar0/caesura-breaks", device_map="auto") - Notebooks
- Google Colab
- Kaggle
caesura-breaks
Predicts prosodic phrase breaks on lowercase, punctuation-free English text, one
label per word boundary: none, b2 (minor break) or b3 (major break). It is
the learned half of caesura, a TTS
front-end stage that has to decide where the voice pauses when the input arrives
without reliable punctuation, which is the normal case for phone transcripts and
LLM output.
Labels
| id | label | meaning |
|---|---|---|
| 0 | none |
no break after this word |
| 1 | b2 |
minor break, roughly a comma-level juncture |
| 2 | b3 |
major break, roughly a sentence-level juncture |
b1 exists in the break inventory of the wider project but is not predicted
here, because it cannot be derived from punctuation.
The training labels are punctuation, not prosody
This is the most important thing to know before using the model. There is no
free ToBI-annotated corpus of usable size, so the labels come from the
punctuation in LibriTTS-R's original transcripts: comma, semicolon, colon, dash
and bracket become b2; full stop, question mark and exclamation mark become
b3; everything else is none. Punctuation and prosody agree often but not
always. Breaks that a reader takes at no punctuation site at all, which is where
the interesting front-end failures live, are systematically labelled none in
training. The model therefore inherits a bias against exactly the breaks that
matter most. The repository quantifies this on a hand-annotated adversarial set.
Training data
LibriTTS-R train-clean-100 transcripts, consecutive same-chapter utterances
stitched into passages of about 45 tokens so that b3 occurs mid-sequence
rather than only on the final token. 10196
passages, roughly 576k boundaries, about 8.7% b2 and 6.4% b3.
Results
Per-boundary F1, from the repository's evaluation:
| test set | b2 F1 | b3 F1 | any-break F1 |
|---|---|---|---|
| libritts_test | 66.0 | 77.5 | 83.1 |
| switchboard | 34.2 | 95.3 | 81.4 |
| adversarial | 28.9 | 96.8 | 75.0 |
switchboard is out-of-domain conversational text; adversarial is 60
hand-annotated sentences built to break syntactic break prediction, so the drop
there is the point of the set rather than a defect of the model.
Usage
from transformers import AutoModelForTokenClassification, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("daltonoscar0/caesura-breaks", add_prefix_space=True)
model = AutoModelForTokenClassification.from_pretrained("daltonoscar0/caesura-breaks")
words = "the man in the grey coat is my uncle".split()
enc = tok(words, is_split_into_words=True, return_tensors="pt")
with torch.no_grad():
logits = model(**enc).logits[0]
# The label for a word sits on that word's LAST subword, not its first.
last = {}
for pos, wid in enumerate(enc.word_ids()):
if wid is not None:
last[wid] = pos
for i, word in enumerate(words):
print(word, model.config.id2label[int(logits[last[i]].argmax())])
Or through the project, which adds the rule system and emphasis marking:
pip install git+https://github.com/daltonoscar0/caesura
python -m caesura --system both "the man in the grey coat is my uncle"
Training details
| base model | distilroberta-base |
| epochs | 1 |
| batch size | 16 |
| learning rate | 5e-05 |
| schedule | one-cycle, 10% warmup |
| seed | 17 |
| selection | dev macro-F1 over b2 and b3 only |
| wall clock | 5594.1 s on cpu |
Labels are attached to the last subword of each word and every other subword
position is masked to -100.
Limitations
- Trained on read speech from public-domain literature. Conversational input is out of domain and scores lower.
- The label set has no
b1, so no minor juncture is available to the caller. - Emphasis is not predicted by this model at all; the repository marks emphasis with rules over a dependency parse.
- On garden-path sentences the model places breaks consistent with the locally favoured reading rather than the globally correct one. The repository documents this case by case.
Licence
MIT.
- Downloads last month
- -
Model tree for daltonoscar0/caesura-breaks
Base model
distilbert/distilroberta-base