FrameRank
An experimental model that scores the text of two video concepts from one creator. Predictive usefulness has not been established. A simple TF-IDF baseline outperformed the embedding ranker on the current small holdout.
FrameRank trains 384 weights on top of frozen BGE small English v1.5 embeddings. The transformer was not fine tuned. Training and inference use local CPU computation. This project used no paid training API or cloud GPU job. The timing in metrics.json measures only the small head fit, excluding downloads, text extraction, embedding, and evaluation.
What goes in and comes out
The input is a JSON request with a schema version, the supported creator, and exactly two candidates. Each candidate must supply an id, on-screen text, caption, and transcript. An unavailable text field can be an empty string, but all three cannot be empty. Views, likes, and unknown fields are rejected. Inputs above 512 encoder tokens are rejected rather than silently cut off.
The output is a JSON object containing two relative scores, their ranking, the score gap, and warnings. Confidence and predicted views are null because neither has been calibrated. Scores do not explain causes or recommend publishing. The current checkpoint supports only the creator whose 36 posts formed the dataset.
See input.schema.json, output.schema.json, example-input.json, and PIPELINE.md.
python -m pip install -r requirements.txt
python predict.py --input example-input.json
config.json pins the encoder revision, feature order, supported creator, and checkpoint hash. ranker.npz is the release head. evaluation_ranker.npz preserves the separate head trained without the nine holdout posts.
Corrected evaluation
The original release misparsed compact calendar dates as epoch nanoseconds, bypassing its 90 day pair filter. Its 61.5 percent result and old checkpoint are superseded. The corrected code has regression tests for date parsing and comparison windows.
After the correction, 27 training posts produce 211 qualifying pairs. Nine later posts produce 26 evaluation pairs. A comparison is eligible only when two posts by the same creator were published within 90 days and their recorded views differ by at least twofold.
| Method | Correct evaluation pairs | Pair accuracy |
|---|---|---|
| Frozen BGE and linear head | 17 of 26 | 65.4 percent |
| TF-IDF and linear head | 18 of 26 | 69.2 percent |
| Chance in expectation | 13 of 26 | 50.0 percent |
The TF-IDF vocabulary and both heads are fit only on training posts. Text-length and age heuristic directions are selected using training pairs. In 200 trials that shuffled training view labels at the post level, 61 fitted heads matched or exceeded the observed BGE accuracy. The corrected one-sided permutation p value is 0.3085. This diagnostic provides no persuasive evidence of useful prediction on this dataset.
The 26 comparisons share nine posts and are not independent examples. Training accuracy is 99.5 percent, which emphasizes the risk of overfitting. The release head was subsequently fit on all 36 posts, yielding 404 pairs. Two old posts have no qualifying comparisons, so only 34 contribute to that final objective. The release head cannot be evaluated on the old holdout as if those posts were unseen.
Features and limitations
The private table combines public captions, speech transcribed with Whisper, and OCR from four sampled frames per Reel. It does not model the video's full visual content, editing, music, or timing. Some OCR captures interface clutter and repeated text. Four of 36 text documents exceed 512 tokens. The corrected historical run explicitly retained the original truncation behavior using --allow-truncation, and reports that choice in the metrics.
Views are rounded public totals captured on September 23, 2026. They are affected by post age, exposure, audience, and distribution. Reach, watch time, and fixed-horizon outcomes were not available in this training table. The private videos, transcripts, captions, and per-post metrics are not published.
The next useful experiment requires reviewing extracted features and scoring new posts before their outcomes are observed, then comparing against the simple baseline at a fixed observation horizon. More transformer training alone would not resolve the current evidence gap.
- Downloads last month
- 11