Code Search Net Tokenizer

Description

This repository contains a custom tokenizer trained for Python source code using the Hugging Face transformers library.

The tokenizer was created by training a new vocabulary from an existing tokenizer and a large corpus of Python functions from the CodeSearchNet dataset.

The objective was to obtain a tokenizer that is better adapted to the structure and vocabulary of Python source code.

Important: This repository contains a tokenizer, not a trained language model.


Project Overview

The tokenizer was trained following these main steps:

  1. Load the Python portion of the CodeSearchNet dataset.
  2. Extract the complete Python functions from the whole_func_string field.
  3. Build a memory-efficient training corpus using a generator.
  4. Load the existing asi/gpt-fr-cased-base tokenizer.
  5. Train a new tokenizer vocabulary from the Python corpus.
  6. Compare the tokenization produced by the original and new tokenizers.
  7. Save the new tokenizer locally.
  8. Upload the tokenizer to the Hugging Face Hub.

Training Dataset

The tokenizer was trained using the Python subset of the CodeSearchNet dataset:

from datasets import load_dataset

raw_datasets = load_dataset(
    "code-search-net/code_search_net",
    "python"
)

The training corpus is built from the:

whole_func_string

field, which contains the complete source code of Python functions.

Because the dataset is large, the training corpus is provided to the tokenizer progressively using a generator rather than loading the entire corpus into memory at once.


Base Tokenizer

The new tokenizer was initialized from:

asi/gpt-fr-cased-base

The original tokenizer provides the starting tokenizer configuration, while its vocabulary is adapted to the Python code corpus.

from transformers import AutoTokenizer

old_tokenizer = AutoTokenizer.from_pretrained(
    "asi/gpt-fr-cased-base"
)

Training the New Tokenizer

The tokenizer was trained using:

tokenizer = old_tokenizer.train_new_from_iterator(
    training_corpus,
    52000
)

The target vocabulary size was set to 52,000 tokens.

The training corpus consists of Python source code extracted from CodeSearchNet.


Why Train a New Tokenizer?

A general-purpose tokenizer may not represent programming code as efficiently as a tokenizer specifically exposed to source code.

Python contains many recurring structures such as:

  • def
  • class
  • self
  • return
  • _
  • .
  • :
  • parentheses and brackets
  • indentation
  • function and variable names
  • common programming patterns

Training the tokenizer on Python code allows it to learn tokenization patterns that are more representative of this type of data.


Comparing the Tokenizers

The original and new tokenizers were compared using Python code examples.

For example:

tokens = old_tokenizer.tokenize(example)

and:

tokens = tokenizer.tokenize(example)

The number of generated tokens was also compared:

len(tokens)

This makes it possible to observe how the new vocabulary represents Python source code compared with the original tokenizer.


How to Use

Load the tokenizer directly from the Hugging Face Hub:

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained(
    "Nathalie00/code-search-net-tokenizer"
)

Then tokenize Python code:

code = """
def hello(name):
    return f"Hello {name}"
"""

tokens = tokenizer.tokenize(code)

print(tokens)

Repository Contents

This repository contains the files required to load and use the trained tokenizer, including its tokenizer configuration and vocabulary files.

The tokenizer can therefore be loaded directly with:

AutoTokenizer.from_pretrained(
    "Nathalie00/code-search-net-tokenizer"
)

Limitations

This project has several limitations:

  • The tokenizer was trained specifically on Python source code.
  • Its performance may differ on other programming languages.
  • The tokenizer itself does not generate text or perform code completion.
  • No language model was trained as part of this project.
  • The tokenizer was evaluated through qualitative tokenization comparisons rather than through a complete downstream model benchmark.

Intended Use

This tokenizer can be used as a preprocessing component for NLP or machine-learning projects involving Python source code, for example:

  • source-code analysis
  • code classification
  • code representation
  • code search
  • machine-learning experiments on Python code
  • preparation of Python code for a language model

It should be used as a tokenizer component rather than as a standalone generative model.


Training Resources

The tokenizer was trained using the Hugging Face:

  • datasets library
  • transformers library
  • CodeSearchNet Python dataset

Acknowledgements

This project follows the Hugging Face course section on training a new tokenizer from an existing tokenizer.

Dataset:

CodeSearchNet

The tokenizer training methodology is based on the Hugging Face train_new_from_iterator() approach.


Author

Nathalie Ramanampamonjy

Master 1 — Gouvernance et Ingénierie des Données ENI Fianarantsoa, Madagascar

GitHub: A-Thecle

Hugging Face: Nathalie00


Citation

If you use this tokenizer in a project, please refer to this repository:

Nathalie Ramanampamonjy,
Code Search Net Tokenizer,
Hugging Face Hub.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support