Punctuation, True-casing & Sentence Boundary Detection (English) โ ONNX
Mirror of 1-800-BAD-CODE/punctuation_fullstop_truecase_english. Takes lower-cased, unpunctuated English text and, in a single pass, restores punctuation, true-cases words (including acronyms like U.S. and mixed-case words like McDonald's), and detects sentence boundaries.
Mirrored for use with inference4j, an inference-only AI library for Java.
Original Source
- Repository: 1-800-BAD-CODE
- License: apache-2.0
Usage with inference4j
// Java wrapper coming in a future inference4j release.
Model Details
| Property | Value |
|---|---|
| Architecture | Transformer encoder (6 layers, d_model 512) + punctuation, sentence-boundary and true-case heads |
| Task | Punctuation restoration, true-casing, sentence boundary detection |
| Tokenizer | SentencePiece Unigram, 32k vocab, lower-cased (tokenizer.json, converted from spe_32k_lc_en.model); BOS = 1, EOS = 2, PAD = 3, UNK = 0 |
| Max sequence length | 256 subtokens including BOS/EOS |
| Input | input_ids โ [batch, seq] int64, [BOS] + pieces + [EOS] |
| Outputs | pre_preds, post_preds, cap_preds, seg_preds (see below) |
| Punctuation labels | <NULL>, <ACRONYM>, ., ,, ? |
| Training data | WMT News Crawl (~10M lines, 2012 and 2021) |
| Original framework | NeMo (fork), exported to ONNX by the author |
License
This model is licensed under the Apache 2.0 License. Original model and ONNX export by 1-800-BAD-CODE.
Files
| File | Description |
|---|---|
model.onnx |
The model (upstream punct_cap_seg_en.onnx) |
tokenizer.json |
Hugging Face tokenizers Unigram tokenizer, converted from spe_32k_lc_en.model |
spe_32k_lc_en.model |
Original SentencePiece tokenizer, kept for provenance |
config.yaml |
Label sets and max length |
Outputs
All outputs are argmaxed/thresholded inside the graph; no softmax is needed.
| Name | Shape | Type | Meaning |
|---|---|---|---|
pre_preds |
[batch, seq] |
int64 | Pre-token punctuation index into pre_labels. Always <NULL> for English. |
post_preds |
[batch, seq] |
int64 | Post-token punctuation index into post_labels: 0 none, 1 acronym (period after every character), 2 ., 3 ,, 4 ? |
cap_preds |
[batch, seq, 16] |
bool | Per-character upper-case flag for each subtoken; entries past the subtoken's length are ignored |
seg_preds |
[batch, seq] |
bool | Sentence boundary after this subtoken |
Preprocessing
- Lower-case the input, strip punctuation, and collapse runs of whitespace to single
spaces with no leading or trailing space. SentencePiece does this itself;
tokenizer.jsondoes not, and emits extraโtokens for repeated spaces. - Encode with
tokenizer.json(or SentencePiece), then wrap as[1] + ids + [2]. Do not let<s>,</s>,<pad>or<unk>typed in the text match their IDs; SentencePiece encodes them as ordinary characters, with<and>as<unk>. - Inputs longer than 254 pieces must be split into windows; the punctuators package uses overlapping windows and fuses the results
Postprocessing
Ignore the BOS/EOS positions. For each subtoken:
- Upper-case the characters flagged in
cap_preds. Indices cover the raw piece including the SentencePieceโword marker, so forโmarieindex 1 ism. - Append the punctuation from
post_preds; for<ACRONYM>, put a period after every character. - Start a new sentence wherever
seg_predsis true.
Example: marie curie moved from poland to paris and later worked with the us radium institute โ
Marie Curie moved from Poland to Paris and later worked with the U.S. Radium Institute.
- Downloads last month
- 5