Instructions to use Nathalie00/code-search-net-tokenizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Nathalie00/code-search-net-tokenizer with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Nathalie00/code-search-net-tokenizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Code Search Net Tokenizer
Description
This repository contains a custom tokenizer trained for Python source code using the Hugging Face transformers library.
The tokenizer was created by training a new vocabulary from an existing tokenizer and a large corpus of Python functions from the CodeSearchNet dataset.
The objective was to obtain a tokenizer that is better adapted to the structure and vocabulary of Python source code.
Important: This repository contains a tokenizer, not a trained language model.
Project Overview
The tokenizer was trained following these main steps:
- Load the Python portion of the CodeSearchNet dataset.
- Extract the complete Python functions from the
whole_func_stringfield. - Build a memory-efficient training corpus using a generator.
- Load the existing
asi/gpt-fr-cased-basetokenizer. - Train a new tokenizer vocabulary from the Python corpus.
- Compare the tokenization produced by the original and new tokenizers.
- Save the new tokenizer locally.
- Upload the tokenizer to the Hugging Face Hub.
Training Dataset
The tokenizer was trained using the Python subset of the CodeSearchNet dataset:
from datasets import load_dataset
raw_datasets = load_dataset(
"code-search-net/code_search_net",
"python"
)
The training corpus is built from the:
whole_func_string
field, which contains the complete source code of Python functions.
Because the dataset is large, the training corpus is provided to the tokenizer progressively using a generator rather than loading the entire corpus into memory at once.
Base Tokenizer
The new tokenizer was initialized from:
asi/gpt-fr-cased-base
The original tokenizer provides the starting tokenizer configuration, while its vocabulary is adapted to the Python code corpus.
from transformers import AutoTokenizer
old_tokenizer = AutoTokenizer.from_pretrained(
"asi/gpt-fr-cased-base"
)
Training the New Tokenizer
The tokenizer was trained using:
tokenizer = old_tokenizer.train_new_from_iterator(
training_corpus,
52000
)
The target vocabulary size was set to 52,000 tokens.
The training corpus consists of Python source code extracted from CodeSearchNet.
Why Train a New Tokenizer?
A general-purpose tokenizer may not represent programming code as efficiently as a tokenizer specifically exposed to source code.
Python contains many recurring structures such as:
defclassselfreturn_.:- parentheses and brackets
- indentation
- function and variable names
- common programming patterns
Training the tokenizer on Python code allows it to learn tokenization patterns that are more representative of this type of data.
Comparing the Tokenizers
The original and new tokenizers were compared using Python code examples.
For example:
tokens = old_tokenizer.tokenize(example)
and:
tokens = tokenizer.tokenize(example)
The number of generated tokens was also compared:
len(tokens)
This makes it possible to observe how the new vocabulary represents Python source code compared with the original tokenizer.
How to Use
Load the tokenizer directly from the Hugging Face Hub:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(
"Nathalie00/code-search-net-tokenizer"
)
Then tokenize Python code:
code = """
def hello(name):
return f"Hello {name}"
"""
tokens = tokenizer.tokenize(code)
print(tokens)
Repository Contents
This repository contains the files required to load and use the trained tokenizer, including its tokenizer configuration and vocabulary files.
The tokenizer can therefore be loaded directly with:
AutoTokenizer.from_pretrained(
"Nathalie00/code-search-net-tokenizer"
)
Limitations
This project has several limitations:
- The tokenizer was trained specifically on Python source code.
- Its performance may differ on other programming languages.
- The tokenizer itself does not generate text or perform code completion.
- No language model was trained as part of this project.
- The tokenizer was evaluated through qualitative tokenization comparisons rather than through a complete downstream model benchmark.
Intended Use
This tokenizer can be used as a preprocessing component for NLP or machine-learning projects involving Python source code, for example:
- source-code analysis
- code classification
- code representation
- code search
- machine-learning experiments on Python code
- preparation of Python code for a language model
It should be used as a tokenizer component rather than as a standalone generative model.
Training Resources
The tokenizer was trained using the Hugging Face:
datasetslibrarytransformerslibrary- CodeSearchNet Python dataset
Acknowledgements
This project follows the Hugging Face course section on training a new tokenizer from an existing tokenizer.
Dataset:
CodeSearchNet
The tokenizer training methodology is based on the Hugging Face train_new_from_iterator() approach.
Author
Nathalie Ramanampamonjy
Master 1 — Gouvernance et Ingénierie des Données ENI Fianarantsoa, Madagascar
GitHub: A-Thecle
Hugging Face: Nathalie00
Citation
If you use this tokenizer in a project, please refer to this repository:
Nathalie Ramanampamonjy,
Code Search Net Tokenizer,
Hugging Face Hub.