ReScraper

ReScraper is a 0.6B refiner that replaces the heuristic HTML-to-text stack of a pretraining data pipeline — a rule-based scraper followed by rule-based cleaning filters — with a single small language model.

Contents

What the model does

The model reads the visible text of a raw web page, rendered one block per line with line ids <lid:n>, and generates a short program:

  1. An <extract> block of line removals (rm 1-4, rm 6, …) that strips headers, menus, boilerplate and other non-content lines.

  2. One operation tag for the page:

    Tag Meaning
    <keep> keep the extracted text as it is
    <edit> remove further noisy lines and spans inside the page
    <delete> drop the page entirely
    <rewrite> rewrite a poorly written but informative page

An executor applies the program to the page, so extraction and cleaning happen in one model instead of a cascade of separate stages.

Files

models/ReScraper/     the refiner (Stage 2 checkpoint), in Hugging Face format

Usage

from vllm import LLM, SamplingParams

llm = LLM(model="cx-cmu/ReScraper", dtype="bfloat16", max_model_len=32768)
params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=3072)

system = open("prompts/student_system_stage2.txt").read()   # from the code repository
page   = "<lid:1> ...\n<lid:2> ...\n"                        # rendered page, one block per line

out = llm.chat(
    [[{"role": "system", "content": system}, {"role": "user", "content": page}]],
    params,
)
print(out[0].outputs[0].text)

The system prompt and the executor that applies the generated program are in the code repository, together with the full pipeline (rendering, inference, post-filtering, deduplication and tokenization).

Training

The model is trained in two supervised stages, distilling three teachers: a main-content extractor, a refinement teacher that decides the operation and edits the text, and a rewriting teacher for pages that are informative but poorly written. Both training sets are released in cx-cmu/ReScraper-Data.

License

Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cx-cmu/ReScraper

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1315)
this model