⚑ At a Glance

Stage Model Role
1 french-semantic-stage1 Fast semantic retrieval, candidate generation
2 french-semantic-stage2 Meaning verification: SAME / DIFFERENT
+ rule guards Deterministic protection for critical changes

πŸ”΅ Quick summary: Stage 1 finds what is worth comparing. Stage 2 determines whether meaning is preserved. Rule guards add deterministic protection. Together they form a practical two-stage French semantic quality pipeline for dataset cleaning and semantic deduplication.


πŸ—οΈ Architecture

French dataset
      ↓
Stage 1 β€” semantic embeddings
      ↓
FAISS / vector search
      ↓
high-similarity candidate pairs
      ↓
Stage 2 β€” semantic verifier
      ↓
SAME / DIFFERENT
      ↓
rule guards
      ↓
KEEP / DEDUP / CONFLICT / REVIEW

πŸ” Stage 1 β€” Semantic Retrieval

France-ML/french-semantic-stage1

A French SentenceTransformer-based embedding model used for fast semantic retrieval.

It converts French sentences into vectors so similar sentences can be found efficiently using FAISS or another vector-search system. Stage 1 is intentionally used as a candidate generator, not the final decision-maker.

Main uses

  • Semantic search
  • Duplicate / paraphrase candidate retrieval
  • Dataset clustering
  • Similarity filtering
  • Large-scale candidate generation

🧠 Stage 2 β€” Meaning Verification

France-ML/french-semantic-stage2

A French pair-classification model that performs deeper semantic verification and predicts:

  • SAME: meaning is sufficiently preserved
  • DIFFERENT: meaning has changed

Useful for detecting meaning changes involving negation, numbers, dates, quantities, units, tense, roles, directions, locations, cause/effect, and conditions.

Examples where lexical similarity hides a meaning change:

Il a vendu 5 voitures.
Il a vendu 50 voitures.
Il n'est pas venu.
Il est venu.
Marie a donnΓ© le livre Γ  Paul.
Paul a donnΓ© le livre Γ  Marie.

These sentences can be lexically very similar while having different meanings.


πŸ›‘οΈ Rule Guards

Model decisions can be combined with deterministic checks for important semantic features:

  • Numbers
  • Dates
  • Units
  • Negation words
  • Quantities
  • Named entities
  • Directional terms
  • Other task-specific constraints

The final decision can be more conservative than relying on neural similarity alone.

Possible outputs

Label Meaning
KEEP No problematic match detected
DEDUP Sufficiently equivalent duplicate / paraphrase
CONFLICT Similar text but meaning-changing information detected
REVIEW Uncertain or requires human inspection

🎯 Applications

  • French dataset deduplication
  • Training-data cleaning
  • Paraphrase detection
  • Semantic search preprocessing
  • Near-duplicate detection
  • Dataset quality control
  • Conflict detection
  • Retrieval candidate generation
  • Filtering noisy training data
  • Preparing French NLP datasets for model training

πŸ—‚οΈ Built for Datasets

The engine is made for French dataset work β€” cleaning, deduplicating, and quality-checking text corpora before they are used to train or evaluate models.

Most dataset problems are not exact duplicates. They are near-duplicates, paraphrases, and near-identical sentences that quietly disagree on one important detail. String matching misses these. Pure embedding similarity also misses them, because the sentences look the same to a vector model even when the meaning changed.

This engine is built to catch that gap.

Dataset Problem How the Engine Handles It
Exact and near-duplicate rows Stage 1 retrieval surfaces candidates, Stage 2 confirms and labels them DEDUP
Paraphrases that mean the same thing Stage 2 predicts SAME even when wording differs
Similar sentences with a changed detail (numbers, negation, roles) Stage 2 flags them and rule guards raise CONFLICT
Ambiguous or borderline cases Marked REVIEW instead of forcing a wrong decision

Why it is built for scale

Running a deep verifier across every possible sentence pair in a large corpus would be far too expensive. The two-stage design avoids that:

Millions of French sentences
        ↓
Stage 1 β€” vector search  (fast, cheap)
        ↓
small set of likely matches
        ↓
Stage 2 β€” meaning verification  (slow, careful)
        ↓
rule guards
        ↓
KEEP / DEDUP / CONFLICT / REVIEW

βœ… What It Does Well

The system is built around semantic similarity followed by meaning verification, not simple string matching. This allows it to identify pairs that:

  • Use different wording but preserve meaning
  • Are highly similar but contain important semantic changes
  • Should be considered potential duplicates
  • Should be separated because a critical detail changed

Stage 1 is optimized for efficient retrieval. Stage 2 provides a focused semantic decision.


⚠️ Limitations

This is not a perfect semantic truth detector. Neural models can still make mistakes, especially with:

  • Ambiguous sentences
  • Very long or complicated text
  • Rare vocabulary
  • Unusual French constructions
  • World knowledge
  • Sarcasm or implicit meaning
  • Complex multi-step reasoning
  • Subtle factual differences
  • Domain-specific terminology

Stage 2 is a semantic verifier, not an absolute authority. For high-value datasets, use deterministic rule guards and/or human review for uncertain cases.

Architecture limitation. Stage 1 and Stage 2 are designed to work together. Stage 2 should not be treated as a general unrelated-text classifier. Its purpose is to verify candidate pairs already flagged as semantically related by Stage 1.


πŸ“ˆ Evaluation

  • Stage 1 was evaluated using validation data, paraphrases, and hard-negative examples. The selected checkpoint prioritized generalization and separation of unrelated text, not just maximal validation similarity.
  • Stage 2 achieved strong results on the project's validation and targeted evaluation sets, including difficult negative examples.

These are project-specific evaluation results, not a guarantee of performance on every French dataset or domain.


πŸ“ Contents

French-semantic-engine/
β”œβ”€β”€ stage1/
β”‚   β”œβ”€β”€ 1_Pooling/
β”‚   β”œβ”€β”€ config.json
β”‚   β”œβ”€β”€ config_sentence_transformers.json
β”‚   β”œβ”€β”€ model.safetensors
β”‚   β”œβ”€β”€ modules.json
β”‚   β”œβ”€β”€ sentence_bert_config.json
β”‚   β”œβ”€β”€ tokenizer.json
β”‚   └── tokenizer_config.json
β”‚
└── stage2/
    β”œβ”€β”€ config.json
    β”œβ”€β”€ model.safetensors
    β”œβ”€β”€ tokenizer.json
    └── tokenizer_config.json

🧾 Run Summary

Model France-ML/French-semantic-engine
Pipeline Stage 1 β†’ vector search β†’ Stage 2 β†’ rule guards
Language French
Task sentence-similarity Β· dedup Β· paraphrase Β· conflict detection
Outputs KEEP Β· DEDUP Β· CONFLICT Β· REVIEW
Library transformers Β· sentence-transformers
License Apache 2.0

⭐ If this engine helped you, consider leaving a like on the Hub.



Hugging Face    Hugging Face model



France Made For French ML Β· 2026

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support