- Verity-4B-RC
- Test the model
- Model Description
- Intended Use
- Analysis Categories
- Important Interpretation Principle
- Base Model
- Training
- Training Behavior
- Task Evaluation
- Example Input
- Example Output
- Policy Retrieval
- Why Verity Exists
- Limitations
- Recommended Use
- Relationship to Veilance
- Previous Versions
- Research and Development Status
- Responsible Use
- Citation
- Disclaimer
- Test the model
Verity-4B-RC
Verity-4B-RC is a 4B-parameter privacy analysis model developed for the Veilance ecosystem.
Verity compares a website's privacy policy against browser behavior observed during a specific browsing session. Its purpose is to identify where observed activity appears consistent with, only partially described by, absent from, or potentially inconsistent with the supplied privacy policy.
This release is based on Qwen3-4B and fine-tuned using LoRA on a purpose-built dataset for privacy-policy and browser-telemetry analysis.
Release status: Release Candidate Base model: Qwen/Qwen3-4B Architecture: Causal language model + LoRA fine-tuning Primary task: Privacy policy / observed behavior comparison Maximum training sequence length: 8,192 tokens
Test the model
This repository includes three example inputs under test_records/:
test_record_1.json— broadly disclosed analytics/device/storage behaviortest_record_2.json— observed cookie activity against an explicit "we do not use cookies" policy statementtest_record_3.json— observed fingerprint-related behavior with no usable privacy policy, intended to test indeterminate handling
Install the required packages:
pip install -U torch transformers peft accelerate
Run a test record with the following script:
import json
import sys
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
BASE_MODEL = "Qwen/Qwen3-4B"
ADAPTER_MODEL = "<your-hf-org>/Verity-4B-RC"
SYSTEM_PROMPT = """You are Veilance Privacy Analyst. Compare observed browser telemetry against the supplied website privacy policy.
Keep policy claims, observed behavior, and conclusions separate.
Use only the evidence provided in the input. Do not invent policy language, telemetry, trackers, requests, or browser behavior. A behavior not observed during one visit must not be treated as proof that it never occurs. Missing disclosure is not automatically a contradiction.
Return the Verity structured comparison report as valid JSON only."""
def load_record(path):
with open(path, "r", encoding="utf-8") as f:
record = json.load(f)
if "telemetry" not in record or "policy_document" not in record:
raise ValueError(
"Test record must contain top-level 'telemetry' and 'policy_document' keys."
)
return record
def main():
path = sys.argv[1] if len(sys.argv) > 1 else "test_records/test_record_1.json"
record = load_record(path)
tokenizer = AutoTokenizer.from_pretrained(
BASE_MODEL,
trust_remote_code=True,
)
base_model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
model = PeftModel.from_pretrained(base_model, ADAPTER_MODEL)
model.eval()
messages = [
{
"role": "system",
"content": SYSTEM_PROMPT,
},
{
"role": "user",
"content": (
"Analyze this Veilance observation and privacy policy:\n"
+ json.dumps(
record,
ensure_ascii=False,
separators=(",", ":"),
)
),
},
]
try:
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
except TypeError:
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(
prompt,
return_tensors="pt",
).to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=2600,
do_sample=False,
pad_token_id=tokenizer.eos_token_id,
)
generated = output[0][inputs["input_ids"].shape[1]:]
text = tokenizer.decode(
generated,
skip_special_tokens=True,
).strip()
print(text)
if __name__ == "__main__":
main()
Save the example above as test_model.py, then run:
python test_model.py test_records/test_record_1.json
python test_model.py test_records/test_record_2.json
python test_model.py test_records/test_record_3.json
The model should return a JSON comparison report. Verity reports are expected to contain top-level fields such as:
analysis
domain
findings
important_limitations
privacy_policy
visit
analysis.counts may contain:
matched
partially_matched
policy_only
observed_only
possible_contradictions
indeterminate
These test records are designed as behavioral smoke tests, not as benchmark scores. Exact wording may vary, but the classification should remain consistent with the evidence supplied in each record.
Important: Keep the inference prompt aligned with the prompt used during fine-tuning. Changing the system prompt, user wrapper, chat template, or thinking mode can materially change output quality. Verity-4B-RC should be run with Qwen thinking disabled (
enable_thinking=False).
Model Description
Verity is designed to analyze two distinct evidence sources:
- Privacy policy text
- Observed browser telemetry
The model attempts to keep those sources separate and determines how closely documented policy language corresponds with behavior observed during a specific website visit.
Verity is not intended to determine whether a company is legally compliant with a privacy law. It is an evidence-comparison model intended to assist analysts and users in identifying areas that may warrant further review.
Typical observed behavior may include:
- Third-party network requests
- Tracker activity
- Cookies
- Local and session storage
- IndexedDB activity
- Cache usage
- Canvas access
- WebGL activity
- Audio fingerprinting signals
- Browser and device characteristic access
- User-Agent access
- Client hints
- WebRTC activity
- Permission requests
- DOM-related behavior
- Service worker activity
- Security-related observations
The exact telemetry supplied to the model depends on the collection system used.
Intended Use
Verity is designed primarily for use with telemetry generated by Veilance, an open-source browser privacy observability platform.
A typical pipeline is:
Website Visit
|
v
Veilance Browser Telemetry
|
+-------------------+
| |
v v
Observed Behavior Privacy Policy
| |
+---------+---------+
|
v
Verity
|
v
Structured Comparison
The model may also be used with independently collected browser telemetry if the input is converted into an appropriate structured format.
Analysis Categories
Verity can classify relationships between policy statements and observed behavior using categories such as:
Matched
The supplied privacy policy reasonably describes the observed behavior.
Partially Matched
The policy broadly or indirectly describes the activity, but the observed implementation contains additional detail or specificity that is not clearly disclosed.
Policy Only
The privacy policy describes a practice that was not observed during the supplied visit.
This does not mean the behavior never occurs.
Observed Only
Behavior was observed for which no sufficiently relevant disclosure could be identified in the supplied policy material.
Possible Contradiction
Observed evidence appears potentially inconsistent with a statement made in the supplied policy.
This classification should be used conservatively.
Indeterminate
The available evidence is insufficient to establish a reliable relationship.
Important Interpretation Principle
Verity separates three concepts:
1. Policy Claims
What the supplied privacy policy states.
2. Observed Behavior
What occurred during the specific browser observation supplied to the model.
3. Analysis
The relationship between those two evidence sources.
This distinction is important.
A single browser visit cannot establish how a website behaves for every:
- User
- Account
- Device
- Browser
- Geographic region
- Consent state
- Session
- A/B test
- Configuration
Likewise, failure to observe behavior during one visit does not establish that the behavior never occurs.
Base Model
Verity-4B-RC is based on:
Qwen/Qwen3-4B
The base model was adapted using parameter-efficient fine-tuning rather than full-model training.
Training
Verity-4B-RC was fine-tuned using LoRA.
Training Configuration
| Parameter | Value |
|---|---|
| Base model | Qwen/Qwen3-4B |
| Training examples | 69,632 |
| Validation examples | 16,384 |
| Epochs | 1 |
| Maximum sequence length | 8,192 |
| LoRA rank | 32 |
| Learning rate | 1e-4 |
| Gradient accumulation | 16 |
| Distributed world size | 4 |
| Evaluation samples | 4,096 |
| Evaluation interval | 250 steps |
| Early stopping patience | 3 |
Training was performed using four NVIDIA GPUs.
The development environment used:
4x NVIDIA RTX A6000 48 GB
Training Behavior
During training, loss decreased substantially during the initial portion of the epoch.
Representative training observations included loss declining from approximately:
0.51 -> 0.02
during early training.
Evaluation loss during early checkpoints was approximately:
0.108 - 0.119
Mean evaluation token accuracy was approximately:
0.980 - 0.984
These metrics primarily measure model training behavior and token prediction performance. They should not be interpreted as equivalent to real-world privacy-analysis accuracy.
Task Evaluation
Internal testing estimated Verity-4B-RC at approximately:
~72% task performance
on the project's application-specific evaluation methodology.
Earlier Verity models were estimated at approximately:
~50%
under comparable internal testing.
These numbers are internal development metrics and are not independent benchmark results.
They should not be interpreted as a general 72% accuracy rate for privacy-policy interpretation.
The evaluation emphasizes whether the model:
- Correctly identifies relevant policy language
- Correctly identifies observed behavior
- Maintains separation between evidence types
- Produces appropriate comparison classifications
- Avoids unsupported claims
- Produces structurally valid outputs
- Avoids overstating evidence from a single visit
Example Input
A simplified Verity request may contain:
{
"domain_url": "https://example.com",
"privacy_policy_url": "https://example.com/privacy",
"visit": {
"snapshot_id": "example-snapshot",
"observed_at": "2026-09-20T12:00:00Z",
"duration_seconds": 300,
"extension_version": "0.8"
},
"seen_behavior": {
"thirdPartyHosts": [],
"trackers": [],
"signals": {},
"storage": {},
"security": {}
},
"policy": {
"sections": []
}
}
Production implementations may provide significantly more detailed telemetry and policy context.
Example Output
A simplified response may resemble:
{
"analysis": {
"counts": {
"matched": 1,
"partially_matched": 3,
"policy_only": 1,
"observed_only": 1,
"possible_contradictions": 0,
"indeterminate": 0
},
"overall_confidence": 0.79,
"summary": "The observed visit contained several behaviors broadly addressed by the supplied policy, although some implementation details were not explicitly described."
},
"findings": []
}
The production schema may differ depending on the Verity implementation.
Policy Retrieval
Verity itself is intended to analyze supplied evidence.
In the production Veilance architecture, retrieval and validation of privacy policies are handled outside the language model.
The host application is responsible for operations such as:
- URL validation
- HTTP retrieval
- Browser rendering when required
- Policy identification
- Document extraction
- Section selection
- Text chunking
- Input-size enforcement
- Telemetry collection
This separation prevents the model from independently browsing for evidence that was not explicitly provided as part of the analysis.
Why Verity Exists
Privacy policies frequently describe data processing using broad categories such as:
- Device information
- Technical information
- Identifiers
- Analytics
- Cookies
- Fraud prevention
- Service improvement
Browser instrumentation can reveal significantly more implementation-level detail.
For example, a policy might broadly disclose collection of device information while an observed session contains:
- Canvas operations
- WebGL queries
- Font enumeration
- Screen characteristic access
- Browser characteristic reads
- Client hints
Verity is intended to help determine whether the supplied policy language appears to describe those observed activities and, importantly, how specifically it does so.
Limitations
Verity does not determine legal compliance
The model is not a lawyer and does not independently determine whether observed behavior violates:
- GDPR
- CCPA/CPRA
- COPPA
- HIPAA
- State privacy laws
- Contract law
- Consumer protection law
- Other regulatory requirements
Its outputs should be treated as technical analysis rather than legal conclusions.
Observations are session-specific
Telemetry describes behavior observed during a particular session.
Websites can behave differently depending on:
- Authentication state
- Geographic location
- Consent settings
- Browser configuration
- Device
- Cookies
- Experiments
- Advertising campaigns
- Time
- User behavior
Policy interpretation is probabilistic
Privacy policies frequently contain broad, ambiguous, nested, or legally qualified language.
The model may incorrectly determine whether a disclosure applies to a particular technical behavior.
Missing disclosure is not automatically a contradiction
An observed_only result indicates that the model could not identify an adequate corresponding disclosure in the supplied material.
It does not independently establish deception, illegality, or policy violation.
Policy retrieval matters
Incomplete or incorrectly extracted privacy-policy text can materially affect analysis quality.
Applications using Verity should validate policy retrieval before treating its output as meaningful.
Recommended Use
Verity should be used as part of a larger evidence pipeline rather than as an autonomous privacy authority.
Recommended workflow:
Collect
|
v
Validate telemetry
|
v
Retrieve policy
|
v
Verify policy applicability
|
v
Extract relevant sections
|
v
Run Verity
|
v
Review findings
For consequential investigations, findings should be independently reviewed against the underlying policy and telemetry evidence.
Relationship to Veilance
Verity is part of the broader Veilance privacy observability project.
Veilance collects and structures browser-observable privacy signals.
Verity analyzes those signals against privacy-policy language.
Together, they form a pipeline intended to make the difference between stated privacy practices and observable browser behavior easier to inspect.
Veilance:
GitHub:
https://github.com/VeilanceApp
Previous Versions
Verity has undergone several experimental iterations.
Earlier models included approximately 1.7B-parameter variants based on:
Qwen/Qwen3-1.7B
Those releases served as early experiments for:
- Dataset development
- Output schema design
- Policy retrieval
- Telemetry normalization
- Finding classification
- Evaluation methodology
Verity-4B-RC represents a larger and substantially improved generation of the model.
Research and Development Status
This model is a release candidate.
Users should expect:
- Occasional incorrect classifications
- Inconsistent interpretation of ambiguous policy language
- Missed relationships between policy statements and telemetry
- False positive
observed_onlyfindings - Output formatting errors under unusual prompts
Further training and evaluation are ongoing.
Responsible Use
Do not represent Verity output as definitive evidence that an organization has violated a law or intentionally misrepresented its practices.
The appropriate interpretation is:
Given the supplied policy material and observed browser telemetry, this is the model's assessment of the relationship between those two evidence sources.
Users should inspect the underlying evidence before making consequential claims.
Citation
If you use Verity in research or analysis, please cite the project and model repository.
@software{verity4brc,
title = {Verity-4B-RC},
author = {Revix Technologies, Inc.},
year = {2026},
url = {https://huggingface.co/}
}
Update the repository URL after publication.
Disclaimer
Verity is an experimental machine-learning system.
Model outputs may contain errors and should not be considered legal advice, regulatory guidance, or definitive representations regarding a website's privacy practices.
Always verify consequential findings against the original privacy policy, captured telemetry, and other relevant primary evidence.