Verity-4B-RC

Verity-4B-RC is a 4B-parameter privacy analysis model developed for the Veilance ecosystem.

Verity compares a website's privacy policy against browser behavior observed during a specific browsing session. Its purpose is to identify where observed activity appears consistent with, only partially described by, absent from, or potentially inconsistent with the supplied privacy policy.

This release is based on Qwen3-4B and fine-tuned using LoRA on a purpose-built dataset for privacy-policy and browser-telemetry analysis.

Release status: Release Candidate Base model: Qwen/Qwen3-4B Architecture: Causal language model + LoRA fine-tuning Primary task: Privacy policy / observed behavior comparison Maximum training sequence length: 8,192 tokens


Test the model

This repository includes three example inputs under test_records/:

  • test_record_1.json — broadly disclosed analytics/device/storage behavior
  • test_record_2.json — observed cookie activity against an explicit "we do not use cookies" policy statement
  • test_record_3.json — observed fingerprint-related behavior with no usable privacy policy, intended to test indeterminate handling

Install the required packages:

pip install -U torch transformers peft accelerate

Run a test record with the following script:

import json
import sys

import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer


BASE_MODEL = "Qwen/Qwen3-4B"
ADAPTER_MODEL = "<your-hf-org>/Verity-4B-RC"

SYSTEM_PROMPT = """You are Veilance Privacy Analyst. Compare observed browser telemetry against the supplied website privacy policy.

Keep policy claims, observed behavior, and conclusions separate.

Use only the evidence provided in the input. Do not invent policy language, telemetry, trackers, requests, or browser behavior. A behavior not observed during one visit must not be treated as proof that it never occurs. Missing disclosure is not automatically a contradiction.

Return the Verity structured comparison report as valid JSON only."""


def load_record(path):
    with open(path, "r", encoding="utf-8") as f:
        record = json.load(f)

    if "telemetry" not in record or "policy_document" not in record:
        raise ValueError(
            "Test record must contain top-level 'telemetry' and 'policy_document' keys."
        )

    return record


def main():
    path = sys.argv[1] if len(sys.argv) > 1 else "test_records/test_record_1.json"
    record = load_record(path)

    tokenizer = AutoTokenizer.from_pretrained(
        BASE_MODEL,
        trust_remote_code=True,
    )

    base_model = AutoModelForCausalLM.from_pretrained(
        BASE_MODEL,
        torch_dtype="auto",
        device_map="auto",
        trust_remote_code=True,
    )

    model = PeftModel.from_pretrained(base_model, ADAPTER_MODEL)
    model.eval()

    messages = [
        {
            "role": "system",
            "content": SYSTEM_PROMPT,
        },
        {
            "role": "user",
            "content": (
                "Analyze this Veilance observation and privacy policy:\n"
                + json.dumps(
                    record,
                    ensure_ascii=False,
                    separators=(",", ":"),
                )
            ),
        },
    ]

    try:
        prompt = tokenizer.apply_chat_template(
            messages,
            tokenize=False,
            add_generation_prompt=True,
            enable_thinking=False,
        )
    except TypeError:
        prompt = tokenizer.apply_chat_template(
            messages,
            tokenize=False,
            add_generation_prompt=True,
        )

    inputs = tokenizer(
        prompt,
        return_tensors="pt",
    ).to(model.device)

    with torch.inference_mode():
        output = model.generate(
            **inputs,
            max_new_tokens=2600,
            do_sample=False,
            pad_token_id=tokenizer.eos_token_id,
        )

    generated = output[0][inputs["input_ids"].shape[1]:]
    text = tokenizer.decode(
        generated,
        skip_special_tokens=True,
    ).strip()

    print(text)


if __name__ == "__main__":
    main()

Save the example above as test_model.py, then run:

python test_model.py test_records/test_record_1.json
python test_model.py test_records/test_record_2.json
python test_model.py test_records/test_record_3.json

The model should return a JSON comparison report. Verity reports are expected to contain top-level fields such as:

analysis
domain
findings
important_limitations
privacy_policy
visit

analysis.counts may contain:

matched
partially_matched
policy_only
observed_only
possible_contradictions
indeterminate

These test records are designed as behavioral smoke tests, not as benchmark scores. Exact wording may vary, but the classification should remain consistent with the evidence supplied in each record.

Important: Keep the inference prompt aligned with the prompt used during fine-tuning. Changing the system prompt, user wrapper, chat template, or thinking mode can materially change output quality. Verity-4B-RC should be run with Qwen thinking disabled (enable_thinking=False).


Model Description

Verity is designed to analyze two distinct evidence sources:

  1. Privacy policy text
  2. Observed browser telemetry

The model attempts to keep those sources separate and determines how closely documented policy language corresponds with behavior observed during a specific website visit.

Verity is not intended to determine whether a company is legally compliant with a privacy law. It is an evidence-comparison model intended to assist analysts and users in identifying areas that may warrant further review.

Typical observed behavior may include:

  • Third-party network requests
  • Tracker activity
  • Cookies
  • Local and session storage
  • IndexedDB activity
  • Cache usage
  • Canvas access
  • WebGL activity
  • Audio fingerprinting signals
  • Browser and device characteristic access
  • User-Agent access
  • Client hints
  • WebRTC activity
  • Permission requests
  • DOM-related behavior
  • Service worker activity
  • Security-related observations

The exact telemetry supplied to the model depends on the collection system used.


Intended Use

Verity is designed primarily for use with telemetry generated by Veilance, an open-source browser privacy observability platform.

A typical pipeline is:

Website Visit
      |
      v
Veilance Browser Telemetry
      |
      +-------------------+
      |                   |
      v                   v
Observed Behavior    Privacy Policy
      |                   |
      +---------+---------+
                |
                v
              Verity
                |
                v
      Structured Comparison

The model may also be used with independently collected browser telemetry if the input is converted into an appropriate structured format.


Analysis Categories

Verity can classify relationships between policy statements and observed behavior using categories such as:

Matched

The supplied privacy policy reasonably describes the observed behavior.

Partially Matched

The policy broadly or indirectly describes the activity, but the observed implementation contains additional detail or specificity that is not clearly disclosed.

Policy Only

The privacy policy describes a practice that was not observed during the supplied visit.

This does not mean the behavior never occurs.

Observed Only

Behavior was observed for which no sufficiently relevant disclosure could be identified in the supplied policy material.

Possible Contradiction

Observed evidence appears potentially inconsistent with a statement made in the supplied policy.

This classification should be used conservatively.

Indeterminate

The available evidence is insufficient to establish a reliable relationship.


Important Interpretation Principle

Verity separates three concepts:

1. Policy Claims

What the supplied privacy policy states.

2. Observed Behavior

What occurred during the specific browser observation supplied to the model.

3. Analysis

The relationship between those two evidence sources.

This distinction is important.

A single browser visit cannot establish how a website behaves for every:

  • User
  • Account
  • Device
  • Browser
  • Geographic region
  • Consent state
  • Session
  • A/B test
  • Configuration

Likewise, failure to observe behavior during one visit does not establish that the behavior never occurs.


Base Model

Verity-4B-RC is based on:

Qwen/Qwen3-4B

The base model was adapted using parameter-efficient fine-tuning rather than full-model training.


Training

Verity-4B-RC was fine-tuned using LoRA.

Training Configuration

Parameter Value
Base model Qwen/Qwen3-4B
Training examples 69,632
Validation examples 16,384
Epochs 1
Maximum sequence length 8,192
LoRA rank 32
Learning rate 1e-4
Gradient accumulation 16
Distributed world size 4
Evaluation samples 4,096
Evaluation interval 250 steps
Early stopping patience 3

Training was performed using four NVIDIA GPUs.

The development environment used:

4x NVIDIA RTX A6000 48 GB

Training Behavior

During training, loss decreased substantially during the initial portion of the epoch.

Representative training observations included loss declining from approximately:

0.51 -> 0.02

during early training.

Evaluation loss during early checkpoints was approximately:

0.108 - 0.119

Mean evaluation token accuracy was approximately:

0.980 - 0.984

These metrics primarily measure model training behavior and token prediction performance. They should not be interpreted as equivalent to real-world privacy-analysis accuracy.


Task Evaluation

Internal testing estimated Verity-4B-RC at approximately:

~72% task performance

on the project's application-specific evaluation methodology.

Earlier Verity models were estimated at approximately:

~50%

under comparable internal testing.

These numbers are internal development metrics and are not independent benchmark results.

They should not be interpreted as a general 72% accuracy rate for privacy-policy interpretation.

The evaluation emphasizes whether the model:

  • Correctly identifies relevant policy language
  • Correctly identifies observed behavior
  • Maintains separation between evidence types
  • Produces appropriate comparison classifications
  • Avoids unsupported claims
  • Produces structurally valid outputs
  • Avoids overstating evidence from a single visit

Example Input

A simplified Verity request may contain:

{
  "domain_url": "https://example.com",
  "privacy_policy_url": "https://example.com/privacy",
  "visit": {
    "snapshot_id": "example-snapshot",
    "observed_at": "2026-09-20T12:00:00Z",
    "duration_seconds": 300,
    "extension_version": "0.8"
  },
  "seen_behavior": {
    "thirdPartyHosts": [],
    "trackers": [],
    "signals": {},
    "storage": {},
    "security": {}
  },
  "policy": {
    "sections": []
  }
}

Production implementations may provide significantly more detailed telemetry and policy context.


Example Output

A simplified response may resemble:

{
  "analysis": {
    "counts": {
      "matched": 1,
      "partially_matched": 3,
      "policy_only": 1,
      "observed_only": 1,
      "possible_contradictions": 0,
      "indeterminate": 0
    },
    "overall_confidence": 0.79,
    "summary": "The observed visit contained several behaviors broadly addressed by the supplied policy, although some implementation details were not explicitly described."
  },
  "findings": []
}

The production schema may differ depending on the Verity implementation.


Policy Retrieval

Verity itself is intended to analyze supplied evidence.

In the production Veilance architecture, retrieval and validation of privacy policies are handled outside the language model.

The host application is responsible for operations such as:

  • URL validation
  • HTTP retrieval
  • Browser rendering when required
  • Policy identification
  • Document extraction
  • Section selection
  • Text chunking
  • Input-size enforcement
  • Telemetry collection

This separation prevents the model from independently browsing for evidence that was not explicitly provided as part of the analysis.


Why Verity Exists

Privacy policies frequently describe data processing using broad categories such as:

  • Device information
  • Technical information
  • Identifiers
  • Analytics
  • Cookies
  • Fraud prevention
  • Service improvement

Browser instrumentation can reveal significantly more implementation-level detail.

For example, a policy might broadly disclose collection of device information while an observed session contains:

  • Canvas operations
  • WebGL queries
  • Font enumeration
  • Screen characteristic access
  • Browser characteristic reads
  • Client hints

Verity is intended to help determine whether the supplied policy language appears to describe those observed activities and, importantly, how specifically it does so.


Limitations

Verity does not determine legal compliance

The model is not a lawyer and does not independently determine whether observed behavior violates:

  • GDPR
  • CCPA/CPRA
  • COPPA
  • HIPAA
  • State privacy laws
  • Contract law
  • Consumer protection law
  • Other regulatory requirements

Its outputs should be treated as technical analysis rather than legal conclusions.

Observations are session-specific

Telemetry describes behavior observed during a particular session.

Websites can behave differently depending on:

  • Authentication state
  • Geographic location
  • Consent settings
  • Browser configuration
  • Device
  • Cookies
  • Experiments
  • Advertising campaigns
  • Time
  • User behavior

Policy interpretation is probabilistic

Privacy policies frequently contain broad, ambiguous, nested, or legally qualified language.

The model may incorrectly determine whether a disclosure applies to a particular technical behavior.

Missing disclosure is not automatically a contradiction

An observed_only result indicates that the model could not identify an adequate corresponding disclosure in the supplied material.

It does not independently establish deception, illegality, or policy violation.

Policy retrieval matters

Incomplete or incorrectly extracted privacy-policy text can materially affect analysis quality.

Applications using Verity should validate policy retrieval before treating its output as meaningful.


Recommended Use

Verity should be used as part of a larger evidence pipeline rather than as an autonomous privacy authority.

Recommended workflow:

Collect
   |
   v
Validate telemetry
   |
   v
Retrieve policy
   |
   v
Verify policy applicability
   |
   v
Extract relevant sections
   |
   v
Run Verity
   |
   v
Review findings

For consequential investigations, findings should be independently reviewed against the underlying policy and telemetry evidence.


Relationship to Veilance

Verity is part of the broader Veilance privacy observability project.

Veilance collects and structures browser-observable privacy signals.

Verity analyzes those signals against privacy-policy language.

Together, they form a pipeline intended to make the difference between stated privacy practices and observable browser behavior easier to inspect.

Veilance:

https://veilance.org

GitHub:

https://github.com/VeilanceApp


Previous Versions

Verity has undergone several experimental iterations.

Earlier models included approximately 1.7B-parameter variants based on:

Qwen/Qwen3-1.7B

Those releases served as early experiments for:

  • Dataset development
  • Output schema design
  • Policy retrieval
  • Telemetry normalization
  • Finding classification
  • Evaluation methodology

Verity-4B-RC represents a larger and substantially improved generation of the model.


Research and Development Status

This model is a release candidate.

Users should expect:

  • Occasional incorrect classifications
  • Inconsistent interpretation of ambiguous policy language
  • Missed relationships between policy statements and telemetry
  • False positive observed_only findings
  • Output formatting errors under unusual prompts

Further training and evaluation are ongoing.


Responsible Use

Do not represent Verity output as definitive evidence that an organization has violated a law or intentionally misrepresented its practices.

The appropriate interpretation is:

Given the supplied policy material and observed browser telemetry, this is the model's assessment of the relationship between those two evidence sources.

Users should inspect the underlying evidence before making consequential claims.


Citation

If you use Verity in research or analysis, please cite the project and model repository.

@software{verity4brc,
  title = {Verity-4B-RC},
  author = {Revix Technologies, Inc.},
  year = {2026},
  url = {https://huggingface.co/}
}

Update the repository URL after publication.


Disclaimer

Verity is an experimental machine-learning system.

Model outputs may contain errors and should not be considered legal advice, regulatory guidance, or definitive representations regarding a website's privacy practices.

Always verify consequential findings against the original privacy policy, captured telemetry, and other relevant primary evidence.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support