EmojiFinder
EmojiFinder turns an English word or short phrase into ranked emoji suggestions. It runs locally with ONNX Runtime, supports 1,870 emoji sequences, and requires no network access after download and installation. This release contains the first version of EmojiFinder, in FP16, with 9,553,870 learned parameters and a 20.57 MB graph.
Use it for emoji pickers, keyboard suggestions, and experiments with text-to-emoji search. Short, familiar phrases work best. Unfamiliar wording and broad associations remain unreliable: required-association recall on the 64-phrase test is 19.66%.
Using the weights
This repository distributes model.onnx, the learned FP16 weights and ONNX
inference graph, together with configuration, catalog metadata, fixtures,
evaluation summaries and licenses. Training, inference, preprocessing and
validation source code are not distributed.
The graph runs with a compatible ONNX runtime. Applications must implement the
text preprocessing and output selection described in INTEGRATION.md.
metadata.json supplies the feature-bucket count, label order, complete emoji
strings and acceptance threshold. This is a custom classifier, not a Transformers
pipeline or a text-generating language model. No ready-to-run application is included.
Results must pass the global 0.8 score threshold, so fewer results, including an
empty list, are possible. Scores represent model relevance, not calibrated
probabilities. The graph's inputs are feature tensors rather than raw text.
deployment-fixtures.json contains 46 preprocessing and accepted-set fixtures
for validating an implementation. SHA256SUMS.json provides file integrity hashes.
Model and training
The model has two pre-norm attention layers, width 384, sixteen heads per layer, a tied feature/label embedding table and output bias. It classifies a query in one pass; it does not generate text. Serving uses learned scores without phrase lookups, query rewrites, retrieval, or per-query overrides.
Text becomes hashed word, character n-gram, and adjacent-word features in 16,384 buckets. The contextual path keeps the first 64 words; feature bags retain the first 256 features. There is no external tokenizer download. The ONNX export materializes the label projection from the tied table, so graph storage differs from the training parameter count. Weights use FP16 storage; public floating-point inputs/outputs use FP32. The CPU runtime may promote internal computation to FP32.
EmojiFinder was trained using AdamW, graded raw-logit cross-entropy, judged hard-negative binary cross-entropy, a primary-match margin, and a positive-logit saturation penalty. All parameters train. The final training stage used seed 1919, learning rate 0.00007, FP32 with TF32 disabled, PyTorch 2.6.0+cu124 and one NVIDIA L4. The selected checkpoint is epoch 3 within a 250-second training-stage budget; this is not the total cost or duration of the model's full training.
Training data
The training data combines Unicode Emoji 17.0, CLDR 48 English annotations, and project-authored examples and corrections. emojilib is used only as supplementary reinforcement during training, providing additional English keywords to strengthen word-to-emoji associations. It is not the model's sole training source or an inference-time lookup library.
The catalog caps emoji versions at 15.0 and excludes skin-tone variants. The final correction covers 32 topics and nine specific-item controls. Broad-topic labels are partial judgments: required and forbidden labels do not exhaustively judge every possible suggestion. Full training data, training checkpoints, and optimizer state are not included. Source versions and hashes are in provenance.json.
Evaluation
These are recorded evaluations from October 2, 2026, with assistant-authored labels and familiar concepts. Development checks are not independent accuracy estimates. The suites overlap and must not be summed into one accuracy number. Test phrasing was absent from thirteen available training files, but aggregate development test results were seen before the final ranking pass; this is not a new blind evaluation.
| Check | EmojiFinder result | Interpretation |
|---|---|---|
| Direct development constraints and primary ranking | 148/148 | Reviewed/trained phrases |
| Unchanged user controls | 20/20 | Regression checks |
| Unchanged legacy probes | 191/191 | Four deliberately broadened probes excluded |
| Required associations on 64 test phrases | 206/1,048 (19.66%) | Most required associations are missed |
| Complete required sets on those phrases | 0/64 | No test phrase returns its whole required set |
| Main regression test, exact sets | 3,052/3,111 | Development reference: 3,059/3,111 |
| Additional regression test, exact sets | 2,930/2,944 | Development reference: 2,923/2,944 |
| Phrase validation, exact sets | 1,533/1,629 | Development reference: 1,606/1,629 |
| Phrase test, exact sets | 1,600/1,629 | Development reference: 1,592/1,629 |
| Everyday test, exact sets | 64/112 | Development reference: 61/112 |
Development comparisons showed a tradeoff: the selected model lost 73 exact cases
on phrase validation and seven on the main regression test. Two direct
development examples also contain unjudged associations: library adds ☕ and
movie night adds 📺. Cultural associations, ambiguous intent, exclusions, and
long or unfamiliar wording can produce missing or irrelevant results. Multilingual
quality has not been established. Emoji appearance depends on the device renderer;
this package contains Unicode strings, not emoji artwork.
FP32 ONNX accepted sets matched PyTorch on 9,837 queries. FP16 changed one set,
removing an unjudged 🤧 from bundling up against the winter cold; required,
forbidden and primary labels were preserved. This is bounded set-level parity,
not bitwise equality or universal ranking equivalence. Machine-readable summaries
are in evaluation.json and parity.json.
Resource measurements
On a Modal Linux/gVisor CPU environment, Python 3.11.12, ONNX Runtime 1.23.2, one inference thread and 6,262 queries, EmojiFinder measured 1.316 ms median and 1.614 ms p95 per query. Warm process RSS was 147.2 MiB, rising to 169.8 MiB after the query corpus. The graph is exactly 20,565,199 bytes. No PyTorch import occurred in the serving process. See runtime.json.
These are recorded server CPU measurements, excluding browser and network overhead. They do not establish Mac, phone, native FP16 arithmetic, battery, or browser-only performance. Native mobile and ONNX Runtime Web integrations remain unvalidated.
License
The model weights and original release documentation are released under Apache-2.0. Unicode-derived data retains its Unicode-3.0 notice; emojilib's MIT notice is also retained. See NOTICE. This release license does not apply to the surrounding private Centauri 1 repository or grant a license to unshipped project datasets.
- Downloads last month
- 36