Locator Healer (v0.1)
When a UI change breaks a test locator, this model finds the element again and returns a robust new locator, or reports that the element is gone instead of guessing.
Most self-healing approaches pick the most similar element on the page, even when the real one has been removed. That turns a clear test failure into a test that silently clicks the wrong thing. This model is trained to heal and to abstain.
Trained and evaluated on the Self-Healing Locators dataset.
Results (test split, 600 examples)
| Similarity baseline | Locator Healer | |
|---|---|---|
| Overall | 89.8% | 99.2% |
| Easy / medium / hard | 99.4% / 83.8% / 86.2% | 100.0% / 98.5% / 98.9% |
Silent wrong_element failures |
78.5% | 96.9% |
| Element removed: correctly answers "gone" | 36.4% | 93.9% |
A prediction counts as correct when it picks the exact target element, or answers "gone" when the target was removed. The test split has 33 removal cases, so read that row with its sample size in mind.
Generalization check
Scores on data from the same generator can flatter a model, so each row below retrains the model with an entire page type or styling convention held out, then tests only on what it never saw.
| Held out | Examples | Baseline overall | Model overall | Baseline wrong-element | Model wrong-element | Baseline "gone" | Model "gone" |
|---|---|---|---|---|---|---|---|
| page type: checkout | 765 | 87.5% | 99.0% | 83.3% | 96.3% | 23.7% | 86.8% |
| page type: dashboard | 732 | 88.9% | 91.7% | 87.0% | 85.0% | 16.7% | 83.3% |
| styling: hashed | 1,534 | 90.4% | 98.6% | 77.1% | 92.6% | 56.0% | 91.7% |
| styling: semantic | 1,483 | 86.3% | 98.8% | 65.1% | 93.7% | 32.3% | 93.5% |
What this shows: the model transfers well across styling conventions and to checkout pages. On dashboard pages, which are structurally different (mostly links, no form), overall accuracy drops to 91.7% and it is slightly worse than the baseline on silent wrong-element cases. Detecting removed elements stays far ahead of the baseline in every held-out setting.
Usage
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("Vijayarv07/locator-healer")
sys.path.insert(0, path)
from healer import Healer
healer = Healer(path)
result = healer.heal(
target_before=what_the_test_knew, # dict of element features, see below
original_locator={"strategy": "xpath", "value": "/html/body/main/form/div[1]/input"},
html_after=new_page_html,
)
# {'status': 'healed', 'confidence': 0.99, 'element': 'html > body > main > form > div:nth-of-type(4) > input',
# 'locator': {'strategy': 'role', 'value': 'textbox[name="Name"]'}}
# or {'status': 'gone', 'confidence': 0.93, 'best_guess': '...'}
target_before holds what a test framework can record when a locator last worked: tag, id,
classes, testid, text, aria_label, name, type, placeholder, href, role,
accessible_name, css_path, section, label_text, position_in_section. The helper
generator.features(element) builds this dict from an lxml element.
How it works
- Every interactive element on the new page becomes a candidate.
- Each (old element, candidate) pair gets 57 features: exact and fuzzy matches on id, test id, name, visible text, label and accessible name; class overlap; section and position changes; whether the old locator still matches the candidate; and how each candidate compares with the best on the page.
- A LightGBM classifier scores each candidate.
- If the best score is below a threshold tuned on the validation split (0.19), the answer is "gone".
- Otherwise it returns the most robust unique locator for that element
(
testid>id>role>text>css>css_path>xpath).
The strongest signals are overlap in human-readable names (text, label, aria-label), position within the section, and structural path similarity.
The model file is LightGBM's plain-text format (healer.lgb.txt), not a pickle, so loading it runs no code.
It runs on CPU in milliseconds per page.
Training
pip install -r requirements.txt
# download the dataset's data/ folder next to these files, then:
python train_healer.py --data data --out model
python ood_eval.py
Binary objective, learning rate 0.05, 31 leaves, early stopping on validation. Training takes about a minute on a laptop CPU.
Limitations
- Trained only on synthetic pages from one generator. Real applications have deeper DOMs, shadow DOM, iframes and dynamic content, so expect lower accuracy on real sites until it's validated on them.
- The dashboard result shows it can struggle with page layouts unlike its training data.
- It sees element attributes and structure, not rendered visuals.
- English UI text only.
- Treat "healed" results as suggestions for a human to confirm in CI rather than silent fixes.
License
Apache 2.0. Dataset: CC BY 4.0.
- Downloads last month
- 24