Instructions to use itspublu/EgoSieve-S with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use itspublu/EgoSieve-S with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("video-classification", model="itspublu/EgoSieve-S", trust_remote_code=True)# Load model directly from transformers import AutoModelForVideoClassification model = AutoModelForVideoClassification.from_pretrained("itspublu/EgoSieve-S", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
EgoSieve-S
EgoSieve-S ranks manipulation-ready spans in first-person video. It produces
three readiness logits (KEEP, REVIEW, REJECT), seven observable issue
scores, diagnostic start/end boundary proposals, and a normalized retrieval
embedding. It is a dataset-curation model, not a robot policy.
Usage
from transformers import AutoModelForVideoClassification, AutoProcessor
processor = AutoProcessor.from_pretrained(
"itspublu/EgoSieve-S", revision="v0.1.0", trust_remote_code=True
)
model = AutoModelForVideoClassification.from_pretrained(
"itspublu/EgoSieve-S", revision="v0.1.0", trust_remote_code=True
).eval()
outputs = model(**processor(frames, return_tensors="pt"))
The timestamp-aware scanner and JSONL compiler are provided by the egosieve
package. The checkpoint expects 12 center-sampled
RGB frames per window; use its bundled processor.
Training and evaluation
Data represented in the held-out evaluation: HoloAssist (CDLA-Permissive-2.0), HoloAssist controlled corruptions (CDLA-Permissive-2.0). Splits are grouped by original capture unit. Readiness, calibration, and boundary results use 142 human-grounded readiness rows: 0 direct human and 142 human-derived. Boundary results use 97 human-grounded boundary rows. Issue results use 145 labeled rows: 0 human, 109 human-derived, and 36 programmatic controlled corruptions. Unlabeled task targets are masked. The release bundle includes raw held-out predictions with task-level provenance, split assignments, exact run configuration, and metric provenance.
Human-derived rows are not direct, independent EgoSieve rubric judgments. Treat them as proxy evidence and consult the source dataset cards and their dataset-specific proxy details before comparing or interpreting these metrics.
For v0.1, readiness and boundaries come from a fixed-grid occupancy rule over
reviewed HoloAssist fine-action intervals. low_hand_activity is an occupancy
proxy, and acting_hand_not_visible follows HoloAssist's acting-hand modifier;
it does not assert that every hand is absent. The other five issue metrics
measure injected-corruption versus unmodified-reference discrimination. Those
references were not independently audited as natural issue negatives.
- Readiness macro F1: 0.6739
- Issue macro AUROC: 0.7407
- Issue macro average precision: 0.7222
- Boundary micro F1 at 0.30s: 0.1030
- Readiness ECE: 0.1018
Intended use and limitations
Use the model to rank raw egocentric windows, route uncertain spans for review, and create embeddings for near-duplicate search. Readiness remains dependent on the published rubric and capture domain. RGB cannot establish force, physical success, consent, safety, metric depth, or legal publishability. Boundary scores are proposals and are diagnostic-only in the v0.1 compiler.
First-person recordings can contain faces, screens, homes, and bystanders. Apply a separate privacy and consent review before sharing any media.
- Downloads last month
- 20
Model tree for itspublu/EgoSieve-S
Base model
facebook/dinov2-smallDataset used to train itspublu/EgoSieve-S
Space using itspublu/EgoSieve-S 1
Evaluation results
- Readiness macro F1 on EgoSieve-Evaltest set self-reported0.674
- Issue macro AUROC on EgoSieve-Evaltest set self-reported0.741
- Issue macro average precision on EgoSieve-Evaltest set self-reported0.722
- Boundary micro F1 on EgoSieve-Evaltest set self-reported0.103