Instructions to use nishparadox/atomizer-gliner-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER
How to use nishparadox/atomizer-gliner-small with GLiNER:
from gliner import GLiNER model = GLiNER.from_pretrained("nishparadox/atomizer-gliner-small") text = "Cristiano Ronaldo dos Santos Aveiro was born on 5 February 1985 in Funchal, Madeira, Portugal." labels = ["person", "date", "location"] entities = model.predict_entities(text, labels) for entity in entities: print(entity["text"], "=>", entity["label"]) - Notebooks
- Google Colab
- Kaggle
atomizer-gliner-small
A single-pass span atomizer with a critic. It splits a text into short, self-contained claims ("atoms") and scores each one:
verifiable: is this a factual claim that can be checked?checkworthy: is it worth fact-checking?
It runs in one forward pass of a 163M-parameter model, taking 6โ46 ms per text on a laptop.
The model does not generate text. It selects words from the input, replaces references ("She" โ "Marie Curie") with earlier spans, and adds closed-class glue words ("is", "the", "'s"). An atom cannot contain a name or number that isn't in the input.
Status: research prototype (v0.1). It works on encyclopedic and scientific prose. It does not handle conversational text, first person or opinions yet (see Limitations).
Usage
The loading code is the atomizer package at NISH1001/atomizer. That repo is
private for now, so access is on request. The weights are a plain PyTorch state dict (model.pt) plus atomizer.json
(backbone, settings, glue and insertion vocabularies).
uv add "atomizer @ git+ssh://git@github.com/NISH1001/atomizer" # needs access to the repo
from atomizer import Atomizer
atomizer = Atomizer.load("nishparadox/atomizer-gliner-small", revision="v0.1")
for atom in atomizer.atomize("Marie Curie was a Polish-born physicist. She won the Nobel Prize in Physics in 1903."):
print(atom.text, round(atom.verifiable, 2), round(atom.checkworthy, 2))
# Marie Curie was a Polish-born physicist. 1.0 0.81
# Marie Curie won the Nobel Prize in Physics in 1903. 1.0 0.88
print(atomizer.explain("...").show()) # every head's predictions
atomizer.critic_scores(["We should invest more in education."]) # [{'verifiable': 0.03, 'checkworthy': 0.0}]
Model
GLiNER small v2.5 (DeBERTa-v3-small encoder, BiLSTM, span layer) plus four heads:
| Head | Predicts |
|---|---|
| anchor | per word: does an atom start here |
| pair | per (atom, word): KEEP / SUB / DROP, and glue after the word |
| antecedent | per reference: which earlier span it means |
| critic | per atom: verifiable, checkworthy |
Assembly from these predictions is deterministic. Details: ARCHITECTURE.md.
Training
- Atomization:
chentong00/propositionizer-wiki-data. 42,857 Wikipedia passages with propositions written by GPT-4, by the dataset's authors (Chen et al., Dense X Retrieval). 3 epochs. - Critic, starting from step 1: ClaimBuster and CLEF CheckThat! 2024 task 1 EN (human-labeled), plus 15k Propositionizer passages. 1 epoch. This is step 1,000 of that run.
Trained on an Apple M3 Max (MPS, bf16 autocast). No API calls were made to create training data.
Evaluation
| Test set | Metric | Score |
|---|---|---|
| Propositionizer test (300 passages) | soft F1 (token overlap, best match) | 0.739 |
| PropSegmEnt test (1,078 sentences, human) | soft F1 | ~0.67 |
| CLEF CheckThat! 2024 task 1 EN, official test (315) | checkworthy F1 / AUC |
0.827 / 0.974 |
| Words not found in the input (Propositionizer test) | rate | 0 |
The share of reference atoms a span model can express at all is about 70% on GPT-4 propositions, so word-exact match is low (exact F1 โ 0.15).
Limitations
- Conversational and first-person text, and opinions, are outside the training data. For example, "I do believe X is a hypocrite" โ "X is hypocrite." (verifiable 0.64). Do not rely on it for chat text yet.
- Over-generates on long scientific passages: about 15 atoms per passage where an LLM atomizer writes about 8.
- One reference resolves to one span. "The study" cannot become a description assembled from several places.
- No reordering and no verb changes. "carrying" cannot become "carried".
- The opening glue word is unstable ("The Marie Curie โฆ" on some devices and precisions).
- The anchor threshold (0.35) is tuned on Wikipedia. Lower it for more atoms.
License
The weights are released under CC-BY-4.0. The atomizer code is MIT.
The training data has its own terms, which you should check for your use:
- Propositionizer labels were generated with GPT-4 by the dataset's authors (OpenAI terms).
- ClaimBuster is CC-BY-SA 4.0.
- CLEF CheckThat! 2024 is under the CLEF lab's terms.
Model tree for nishparadox/atomizer-gliner-small
Base model
gliner-community/gliner_small-v2.5