Instructions to use France-ML/French-semantic-engine with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use France-ML/French-semantic-engine with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("France-ML/French-semantic-engine") sentences = [ "C'est une personne heureuse", "C'est un chien heureux", "C'est une personne très heureuse", "Aujourd'hui est une journée ensoleillée" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
β‘ At a Glance
| Stage | Model | Role |
|---|---|---|
| 1 | french-semantic-stage1 |
Fast semantic retrieval, candidate generation |
| 2 | french-semantic-stage2 |
Meaning verification: SAME / DIFFERENT |
| + | rule guards | Deterministic protection for critical changes |
π΅ Quick summary: Stage 1 finds what is worth comparing. Stage 2 determines whether meaning is preserved. Rule guards add deterministic protection. Together they form a practical two-stage French semantic quality pipeline for dataset cleaning and semantic deduplication.
ποΈ Architecture
French dataset
β
Stage 1 β semantic embeddings
β
FAISS / vector search
β
high-similarity candidate pairs
β
Stage 2 β semantic verifier
β
SAME / DIFFERENT
β
rule guards
β
KEEP / DEDUP / CONFLICT / REVIEW
π Stage 1 β Semantic Retrieval
France-ML/french-semantic-stage1
A French SentenceTransformer-based embedding model used for fast semantic retrieval.
It converts French sentences into vectors so similar sentences can be found efficiently using FAISS or another vector-search system. Stage 1 is intentionally used as a candidate generator, not the final decision-maker.
Main uses
- Semantic search
- Duplicate / paraphrase candidate retrieval
- Dataset clustering
- Similarity filtering
- Large-scale candidate generation
π§ Stage 2 β Meaning Verification
France-ML/french-semantic-stage2
A French pair-classification model that performs deeper semantic verification and predicts:
SAME: meaning is sufficiently preservedDIFFERENT: meaning has changed
Useful for detecting meaning changes involving negation, numbers, dates, quantities, units, tense, roles, directions, locations, cause/effect, and conditions.
Examples where lexical similarity hides a meaning change:
Il a vendu 5 voitures.
Il a vendu 50 voitures.
Il n'est pas venu.
Il est venu.
Marie a donnΓ© le livre Γ Paul.
Paul a donnΓ© le livre Γ Marie.
These sentences can be lexically very similar while having different meanings.
π‘οΈ Rule Guards
Model decisions can be combined with deterministic checks for important semantic features:
- Numbers
- Dates
- Units
- Negation words
- Quantities
- Named entities
- Directional terms
- Other task-specific constraints
The final decision can be more conservative than relying on neural similarity alone.
Possible outputs
| Label | Meaning |
|---|---|
KEEP |
No problematic match detected |
DEDUP |
Sufficiently equivalent duplicate / paraphrase |
CONFLICT |
Similar text but meaning-changing information detected |
REVIEW |
Uncertain or requires human inspection |
π― Applications
- French dataset deduplication
- Training-data cleaning
- Paraphrase detection
- Semantic search preprocessing
- Near-duplicate detection
- Dataset quality control
- Conflict detection
- Retrieval candidate generation
- Filtering noisy training data
- Preparing French NLP datasets for model training
The engine is made for French dataset work β cleaning, deduplicating, and quality-checking text corpora before they are used to train or evaluate models.
Most dataset problems are not exact duplicates. They are near-duplicates, paraphrases, and near-identical sentences that quietly disagree on one important detail. String matching misses these. Pure embedding similarity also misses them, because the sentences look the same to a vector model even when the meaning changed.
This engine is built to catch that gap.
| Dataset Problem | How the Engine Handles It |
|---|---|
| Exact and near-duplicate rows | Stage 1 retrieval surfaces candidates, Stage 2 confirms and labels them DEDUP |
| Paraphrases that mean the same thing | Stage 2 predicts SAME even when wording differs |
| Similar sentences with a changed detail (numbers, negation, roles) | Stage 2 flags them and rule guards raise CONFLICT |
| Ambiguous or borderline cases | Marked REVIEW instead of forcing a wrong decision |
Why it is built for scale
Running a deep verifier across every possible sentence pair in a large corpus would be far too expensive. The two-stage design avoids that:
Millions of French sentences
β
Stage 1 β vector search (fast, cheap)
β
small set of likely matches
β
Stage 2 β meaning verification (slow, careful)
β
rule guards
β
KEEP / DEDUP / CONFLICT / REVIEW
β What It Does Well
The system is built around semantic similarity followed by meaning verification, not simple string matching. This allows it to identify pairs that:
- Use different wording but preserve meaning
- Are highly similar but contain important semantic changes
- Should be considered potential duplicates
- Should be separated because a critical detail changed
Stage 1 is optimized for efficient retrieval. Stage 2 provides a focused semantic decision.
β οΈ Limitations
This is not a perfect semantic truth detector. Neural models can still make mistakes, especially with:
- Ambiguous sentences
- Very long or complicated text
- Rare vocabulary
- Unusual French constructions
- World knowledge
- Sarcasm or implicit meaning
- Complex multi-step reasoning
- Subtle factual differences
- Domain-specific terminology
Stage 2 is a semantic verifier, not an absolute authority. For high-value datasets, use deterministic rule guards and/or human review for uncertain cases.
Architecture limitation. Stage 1 and Stage 2 are designed to work together. Stage 2 should not be treated as a general unrelated-text classifier. Its purpose is to verify candidate pairs already flagged as semantically related by Stage 1.
π Evaluation
- Stage 1 was evaluated using validation data, paraphrases, and hard-negative examples. The selected checkpoint prioritized generalization and separation of unrelated text, not just maximal validation similarity.
- Stage 2 achieved strong results on the project's validation and targeted evaluation sets, including difficult negative examples.
These are project-specific evaluation results, not a guarantee of performance on every French dataset or domain.
π Contents
French-semantic-engine/
βββ stage1/
β βββ 1_Pooling/
β βββ config.json
β βββ config_sentence_transformers.json
β βββ model.safetensors
β βββ modules.json
β βββ sentence_bert_config.json
β βββ tokenizer.json
β βββ tokenizer_config.json
β
βββ stage2/
βββ config.json
βββ model.safetensors
βββ tokenizer.json
βββ tokenizer_config.json
π§Ύ Run Summary
| Model | France-ML/French-semantic-engine |
| Pipeline | Stage 1 β vector search β Stage 2 β rule guards |
| Language | French |
| Task | sentence-similarity Β· dedup Β· paraphrase Β· conflict detection |
| Outputs | KEEP Β· DEDUP Β· CONFLICT Β· REVIEW |
| Library | transformers Β· sentence-transformers |
| License | Apache 2.0 |