Punctuation, True-casing & Sentence Boundary Detection (English) โ€” ONNX

Mirror of 1-800-BAD-CODE/punctuation_fullstop_truecase_english. Takes lower-cased, unpunctuated English text and, in a single pass, restores punctuation, true-cases words (including acronyms like U.S. and mixed-case words like McDonald's), and detects sentence boundaries.

Mirrored for use with inference4j, an inference-only AI library for Java.

Original Source

Usage with inference4j

// Java wrapper coming in a future inference4j release.

Model Details

Property Value
Architecture Transformer encoder (6 layers, d_model 512) + punctuation, sentence-boundary and true-case heads
Task Punctuation restoration, true-casing, sentence boundary detection
Tokenizer SentencePiece Unigram, 32k vocab, lower-cased (tokenizer.json, converted from spe_32k_lc_en.model); BOS = 1, EOS = 2, PAD = 3, UNK = 0
Max sequence length 256 subtokens including BOS/EOS
Input input_ids โ€” [batch, seq] int64, [BOS] + pieces + [EOS]
Outputs pre_preds, post_preds, cap_preds, seg_preds (see below)
Punctuation labels <NULL>, <ACRONYM>, ., ,, ?
Training data WMT News Crawl (~10M lines, 2012 and 2021)
Original framework NeMo (fork), exported to ONNX by the author

License

This model is licensed under the Apache 2.0 License. Original model and ONNX export by 1-800-BAD-CODE.

Files

File Description
model.onnx The model (upstream punct_cap_seg_en.onnx)
tokenizer.json Hugging Face tokenizers Unigram tokenizer, converted from spe_32k_lc_en.model
spe_32k_lc_en.model Original SentencePiece tokenizer, kept for provenance
config.yaml Label sets and max length

Outputs

All outputs are argmaxed/thresholded inside the graph; no softmax is needed.

Name Shape Type Meaning
pre_preds [batch, seq] int64 Pre-token punctuation index into pre_labels. Always <NULL> for English.
post_preds [batch, seq] int64 Post-token punctuation index into post_labels: 0 none, 1 acronym (period after every character), 2 ., 3 ,, 4 ?
cap_preds [batch, seq, 16] bool Per-character upper-case flag for each subtoken; entries past the subtoken's length are ignored
seg_preds [batch, seq] bool Sentence boundary after this subtoken

Preprocessing

  1. Lower-case the input, strip punctuation, and collapse runs of whitespace to single spaces with no leading or trailing space. SentencePiece does this itself; tokenizer.json does not, and emits extra โ– tokens for repeated spaces.
  2. Encode with tokenizer.json (or SentencePiece), then wrap as [1] + ids + [2]. Do not let <s>, </s>, <pad> or <unk> typed in the text match their IDs; SentencePiece encodes them as ordinary characters, with < and > as <unk>.
  3. Inputs longer than 254 pieces must be split into windows; the punctuators package uses overlapping windows and fuses the results

Postprocessing

Ignore the BOS/EOS positions. For each subtoken:

  1. Upper-case the characters flagged in cap_preds. Indices cover the raw piece including the SentencePiece โ– word marker, so for โ–marie index 1 is m.
  2. Append the punctuation from post_preds; for <ACRONYM>, put a period after every character.
  3. Start a new sentence wherever seg_preds is true.

Example: marie curie moved from poland to paris and later worked with the us radium institute โ†’ Marie Curie moved from Poland to Paris and later worked with the U.S. Radium Institute.

Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support