ReScraper
ReScraper is a 0.6B refiner that replaces the heuristic HTML-to-text stack of a pretraining data pipeline — a rule-based scraper followed by rule-based cleaning filters — with a single small language model.
- Data:
cx-cmu/ReScraper-Data - Code:
cxcscmu/ReScraper
Contents
What the model does
The model reads the visible text of a raw web page, rendered one block per line with line ids
<lid:n>, and generates a short program:
An
<extract>block of line removals (rm 1-4,rm 6, …) that strips headers, menus, boilerplate and other non-content lines.One operation tag for the page:
Tag Meaning <keep>keep the extracted text as it is <edit>remove further noisy lines and spans inside the page <delete>drop the page entirely <rewrite>rewrite a poorly written but informative page
An executor applies the program to the page, so extraction and cleaning happen in one model instead of a cascade of separate stages.
Files
models/ReScraper/ the refiner (Stage 2 checkpoint), in Hugging Face format
Usage
from vllm import LLM, SamplingParams
llm = LLM(model="cx-cmu/ReScraper", dtype="bfloat16", max_model_len=32768)
params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=3072)
system = open("prompts/student_system_stage2.txt").read() # from the code repository
page = "<lid:1> ...\n<lid:2> ...\n" # rendered page, one block per line
out = llm.chat(
[[{"role": "system", "content": system}, {"role": "user", "content": page}]],
params,
)
print(out[0].outputs[0].text)
The system prompt and the executor that applies the generated program are in the code repository, together with the full pipeline (rendering, inference, post-filtering, deduplication and tokenization).
Training
The model is trained in two supervised stages, distilling three teachers: a main-content
extractor, a refinement teacher that decides the operation and edits the text, and a rewriting
teacher for pages that are informative but poorly written. Both training sets are released in
cx-cmu/ReScraper-Data.
License
Apache 2.0.