Ariadne Laya Evidence

Checks whether supplied evidence supports a claim, refutes it or leaves it unresolved. This 4.2 MB specialist interface runs on a shared frozen Laya base. Switching between compatible Ariadne specialists replaces about 1 million parameters, while the 421 million parameter base stays in memory. Each specialist uses its own small interface.

Use

Install the included Python wheel from this downloaded model folder:

pip install ./ariadne_specialists-0.2.0a1-py3-none-any.whl
import ariadne
from ariadne.specialists import Evidence

model = ariadne.load_specialist(Evidence, model=".")
result = model({'claim': 'The event took place in Paris.', 'evidence': 'The event was held in Lyon.'})
print(result.label, result.score)

The base downloads automatically and is cached. Choose a device with device="cpu" or device="cuda:1". A list of inputs returns a list of results. Scores have not been recalibrated for this task.

Load from Hugging Face

After installing the included wheel, you can load this repository directly:

import ariadne
from ariadne.specialists import Evidence

model = ariadne.load_specialist(Evidence, model="GoatHerder/Ariadne-Laya-Evidence")

Use the explicit model= argument with this preview wheel. The interface and pinned base are downloaded automatically and cached. Pass revision="<commit hash>" to pin a particular interface version.

Interface

The interface is a 1,024 × 1,024 linear projection plus a 1,024-element bias: 1,049,600 trainable parameters. It sits after the base's native embeddings and before encoder block 0. It starts as the identity; training updates only this projection. The shared base has 421,293,827 parameters. Compatible specialists share one resident base in the same Python process and on the same device.

Results

Local evaluation uses the same 2,988 inputs for every model. Accuracy is per decision; macro-F1 averages the task's classes (and flag namespaces for Privacy).

Model Accuracy Macro-F1
Base Laya 64.86% 64.07%
Ariadne Evidence 71.99% 70.13%
Yuu-Xie/fever-nli-modernbert-large 78.55% 77.82%
TF-IDF + logistic regression 45.38% 37.69%

Fine-tuned on NLI-FEVER; the model card reports validation on the same official dev source used for our fixed test sample. Common unseen-test status is unverified. This table does not establish a common unseen-test ranking. metrics.json records model revisions, comparison methods, per-class results and existing-task retention.

Training and scope

Trained on pietrolesci/nli_fever, revision 1eddac63112eee1fdf1966e0bca27a5ff248c772. Prepared train/validation/test sizes: 11,927 / 1,496 / 2,988. Overlength exclusions: {'train': 0, 'validation': 0, 'test': 0}.

One seed (0); epoch 1 selected by validation loss, training stopped after epoch 10. LR 1e-4, minimum 10 epochs, patience 3. Only the 1,049,600 exact-identity-initialized interface parameters were trained. The original embeddings, 28 encoder blocks and decision heads stayed frozen and in evaluation mode. Deterministic GPU settings were enabled.

  • NLI-FEVER supplied-evidence task; premise is the claim and hypothesis is retrieved evidence in this export. No retrieval is performed.
  • Official test has no labels, so a stable 3,000-case official dev sample is our test. Train and tuning validation come from official train, grouped by claim ID.
  • Empty evidence is excluded consistently; class counts and every retained ID are recorded. Public comparator used official dev for validation: its test independence is unverified.

English only. Unsupported or ambiguous inputs still receive a prediction. Source datasets retain their own licences. This checkpoint is one training run; it does not establish across-seed variance.

Base revision: 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851. Independent adaptation; no affiliation with the original Laya authors.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GoatHerder/Ariadne-Laya-Evidence

Adapter
(23)
this model

Dataset used to train GoatHerder/Ariadne-Laya-Evidence

Collection including GoatHerder/Ariadne-Laya-Evidence