hv-tempo
Reading pace variation predictor. Given a text, output where the reader will slow down and where they'll speed up.
The claim in one sentence
hv-ttu predicts total comprehension time. hv-tempo predicts
variation β the shape of the pace across a text.
What it produces
For each segment of the text (sentence or small group of sentences):
- WPM β estimated reading speed for that span
- duration_s β how long that span will take
- slowdown β multiplicative factor vs the baseline
- reasons β which features caused the slowdown
- features β the raw surface signals
Plus aggregate stats: mean WPM, total duration, pace variance, slowest and fastest spans.
Install
pip install numpy
Actually β no dependencies. Pure stdlib.
## Usage
### Analyze a text
```python
from hv_tempo import HVTempo
m = HVTempo()
report = m.analyze("A black hole is a region of spacetime where gravity...")
print(report.mean_wpm) # ~123
print(report.total_duration_s) # ~54 s
print(report.pace_variance) # 0.181
print(report.slowest.wpm) # 117
print(report.slowest.reasons) # ['long sentences', 'rare vocabulary', ...]
Just the numbers
m.wpm(text) # mean WPM
m.duration(text) # total seconds
m.slowest_spans(text, 5) # top-5 slowest segments
m.fastest_spans(text, 5) # top-5 fastest segments
CLI
python hv_tempo.py # run all demos
python hv_tempo.py --text "..." # analyze a text
python hv_tempo.py --text "..." --wpm # just the mean WPM
python hv_tempo.py --text "..." --duration # just the duration
python hv_tempo.py --text "..." --slowest 5 # top 5 slowest spans
python hv_tempo.py --text "..." --json # JSON output
The model
Per-span features:
| feature | what it captures |
|---|---|
sentence_len_signed |
words per sentence above/below 15 (signed) |
clauses_per_sentence |
commas + subordinators per sentence |
rare_rate |
fraction of words outside a common-word list, length β₯ 7 |
abstract_rate |
fraction of words ending in -tion, -ness, -ity, etc. |
digit_rate |
numbers per word |
negation_rate |
negations per word |
hedge_rate |
hedges per word ("may", "might", "typically") |
conditional_rate |
conditionals per word ("if", "unless") |
passive_rate |
"be + -ed" constructions per sentence |
list_rate |
list markers per sentence (negative weight) |
Slowdown is computed as a multiplicative factor:
slowdown = exp( Ξ£ weight_i Β· feature_i )
wpm = baseline_wpm / slowdown
Weights are signed: positive slows reading, negative speeds it up. Two
features have negative weights: sentence_len_signed (for short
sentences) and list_rate (lists always speed reading up).
Benchmarks
Sample texts
| text | words | mean WPM | duration | pace variance |
|---|---|---|---|---|
| List-heavy | 33 | 304 | 6.5 s | n/a |
| Fiction (Hemingway) | 51 | 270 | 11.3 s | n/a |
| Technical (install) | 83 | 180 | 27.6 s | 0.055 |
| Academic (black hole) | 111 | 123 | 54.0 s | 0.181 |
Reading the table:
- List-heavy is fastest. Short clauses, list bonus.
- Fiction is nearly as fast. Short sentences, common vocabulary.
- Technical is medium. Numerals and rare vocabulary slow it down.
- Academic is slowest. Long sentences, dense clauses, rare vocabulary.
Pace drivers per span
The model reports the top contributors per span. Examples:
- Fiction:
short sentences (8w avg) β faster, dense clauses (0.5/sent) - Academic:
long sentences (24w avg), rare vocabulary (26%), dense clauses (1.0/sent) - Technical:
rare vocabulary (24%), short sentences (12w avg) β faster, dense clauses (1.6/sent) - List-heavy:
short sentences (6w avg) β faster, list structure (4 items) β faster, rare vocabulary (36%)
When to use it
- Typography β where to adjust line length or font weight.
- Editing β where to break long sentences.
- Prose rhythm β testing whether a passage is uniform or varied.
- Audiobook pacing β pre-planning where the narrator slows.
- Ad copy β where the reader's eye drags.
- Technical writing β identifying the load-bearing sentences.
- Text difficulty β comparing two versions of the same content.
When not to use it
- As ground truth. This is a heuristic model. It estimates relative variation, not absolute times. For absolute times, calibrate against real reading data.
- For non-English text. The lexicons and patterns are English-specific.
- For mathematical or code-heavy text. The surface features don't capture symbolic complexity.
Honest limitations
- The weights are hand-tuned. They are plausible, internally consistent, and produce sensible rankings. They are not fit to a real user study.
- The common-word list is ~800 words with a simple stemmer. Handles inflections (walked β walk, friends β friend) but not irregular forms.
- The passive detector is regex-based. Catches
be + -edand a short list of irregular participles, but misses others. - The abstraction detector uses suffixes. Misses abstract words that don't fit suffix patterns (
freedom,justice). - The baseline WPM is fixed. Real readers vary from 150 to 400 WPM. The model's output should be interpreted as a ratio to the baseline, not as an absolute time.
- List detection requires bullets at line starts. A single stray "1." mid-sentence doesn't count as a list.
- No calibration. Extending the model to fit real reading data is future work.
Reference
Part of the reader-model series. Companion to hv-ttu (total
comprehension time) and hv-locality (feature-map locality).
hv-ttu answers how long will this take? hv-tempo answers where
will it be slow?
License
Apache-2.0