Locator Healer (v0.1)

When a UI change breaks a test locator, this model finds the element again and returns a robust new locator, or reports that the element is gone instead of guessing.

Most self-healing approaches pick the most similar element on the page, even when the real one has been removed. That turns a clear test failure into a test that silently clicks the wrong thing. This model is trained to heal and to abstain.

Trained and evaluated on the Self-Healing Locators dataset.

Results (test split, 600 examples)

Similarity baseline Locator Healer
Overall 89.8% 99.2%
Easy / medium / hard 99.4% / 83.8% / 86.2% 100.0% / 98.5% / 98.9%
Silent wrong_element failures 78.5% 96.9%
Element removed: correctly answers "gone" 36.4% 93.9%

A prediction counts as correct when it picks the exact target element, or answers "gone" when the target was removed. The test split has 33 removal cases, so read that row with its sample size in mind.

Generalization check

Scores on data from the same generator can flatter a model, so each row below retrains the model with an entire page type or styling convention held out, then tests only on what it never saw.

Held out Examples Baseline overall Model overall Baseline wrong-element Model wrong-element Baseline "gone" Model "gone"
page type: checkout 765 87.5% 99.0% 83.3% 96.3% 23.7% 86.8%
page type: dashboard 732 88.9% 91.7% 87.0% 85.0% 16.7% 83.3%
styling: hashed 1,534 90.4% 98.6% 77.1% 92.6% 56.0% 91.7%
styling: semantic 1,483 86.3% 98.8% 65.1% 93.7% 32.3% 93.5%

What this shows: the model transfers well across styling conventions and to checkout pages. On dashboard pages, which are structurally different (mostly links, no form), overall accuracy drops to 91.7% and it is slightly worse than the baseline on silent wrong-element cases. Detecting removed elements stays far ahead of the baseline in every held-out setting.

Usage

from huggingface_hub import snapshot_download
import sys

path = snapshot_download("Vijayarv07/locator-healer")
sys.path.insert(0, path)
from healer import Healer

healer = Healer(path)
result = healer.heal(
    target_before=what_the_test_knew,       # dict of element features, see below
    original_locator={"strategy": "xpath", "value": "/html/body/main/form/div[1]/input"},
    html_after=new_page_html,
)
# {'status': 'healed', 'confidence': 0.99, 'element': 'html > body > main > form > div:nth-of-type(4) > input',
#  'locator': {'strategy': 'role', 'value': 'textbox[name="Name"]'}}
# or {'status': 'gone', 'confidence': 0.93, 'best_guess': '...'}

target_before holds what a test framework can record when a locator last worked: tag, id, classes, testid, text, aria_label, name, type, placeholder, href, role, accessible_name, css_path, section, label_text, position_in_section. The helper generator.features(element) builds this dict from an lxml element.

How it works

  1. Every interactive element on the new page becomes a candidate.
  2. Each (old element, candidate) pair gets 57 features: exact and fuzzy matches on id, test id, name, visible text, label and accessible name; class overlap; section and position changes; whether the old locator still matches the candidate; and how each candidate compares with the best on the page.
  3. A LightGBM classifier scores each candidate.
  4. If the best score is below a threshold tuned on the validation split (0.19), the answer is "gone".
  5. Otherwise it returns the most robust unique locator for that element (testid > id > role > text > css > css_path > xpath).

The strongest signals are overlap in human-readable names (text, label, aria-label), position within the section, and structural path similarity.

The model file is LightGBM's plain-text format (healer.lgb.txt), not a pickle, so loading it runs no code. It runs on CPU in milliseconds per page.

Training

pip install -r requirements.txt
# download the dataset's data/ folder next to these files, then:
python train_healer.py --data data --out model
python ood_eval.py

Binary objective, learning rate 0.05, 31 leaves, early stopping on validation. Training takes about a minute on a laptop CPU.

Limitations

  • Trained only on synthetic pages from one generator. Real applications have deeper DOMs, shadow DOM, iframes and dynamic content, so expect lower accuracy on real sites until it's validated on them.
  • The dashboard result shows it can struggle with page layouts unlike its training data.
  • It sees element attributes and structure, not rendered visuals.
  • English UI text only.
  • Treat "healed" results as suggestions for a human to confirm in CI rather than silent fixes.

License

Apache 2.0. Dataset: CC BY 4.0.

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train Vijayarv07/locator-healer