ALBERT

class transformers.AlbertConfig

( vocab_size = 30000 embedding_size = 128 hidden_size = 4096 num_hidden_layers = 12 num_hidden_groups = 1 num_attention_heads = 64 intermediate_size = 16384 inner_group_num = 1 hidden_act = 'gelu_new' hidden_dropout_prob = 0 attention_probs_dropout_prob = 0 max_position_embeddings = 512 type_vocab_size = 2 initializer_range = 0.02 layer_norm_eps = 1e-12 classifier_dropout_prob = 0.1 position_embedding_type = 'absolute' pad_token_id = 0 bos_token_id = 2 eos_token_id = 3 **kwargs )

Parameters

vocab_size (int, optional, defaults to 30000) — Vocabulary size of the ALBERT model. Defines the number of different tokens that can be represented by the inputs_ids passed when calling AlbertModel or TFAlbertModel.
embedding_size (int, optional, defaults to 128) — Dimensionality of vocabulary embeddings.
hidden_size (int, optional, defaults to 4096) — Dimensionality of the encoder layers and the pooler layer.
num_hidden_layers (int, optional, defaults to 12) — Number of hidden layers in the Transformer encoder.
num_hidden_groups (int, optional, defaults to 1) — Number of groups for the hidden layers, parameters in the same group are shared.
num_attention_heads (int, optional, defaults to 64) — Number of attention heads for each attention layer in the Transformer encoder.
intermediate_size (int, optional, defaults to 16384) — The dimensionality of the “intermediate” (often named feed-forward) layer in the Transformer encoder.
inner_group_num (int, optional, defaults to 1) — The number of inner repetition of attention and ffn.
hidden_act (str or Callable, optional, defaults to "gelu_new") — The non-linear activation function (function or string) in the encoder and pooler. If string, "gelu", "relu", "silu" and "gelu_new" are supported.
hidden_dropout_prob (float, optional, defaults to 0) — The dropout probability for all fully connected layers in the embeddings, encoder, and pooler.
attention_probs_dropout_prob (float, optional, defaults to 0) — The dropout ratio for the attention probabilities.
max_position_embeddings (int, optional, defaults to 512) — The maximum sequence length that this model might ever be used with. Typically set this to something large (e.g., 512 or 1024 or 2048).
type_vocab_size (int, optional, defaults to 2) — The vocabulary size of the token_type_ids passed when calling AlbertModel or TFAlbertModel.
initializer_range (float, optional, defaults to 0.02) — The standard deviation of the truncated_normal_initializer for initializing all weight matrices.
layer_norm_eps (float, optional, defaults to 1e-12) — The epsilon used by the layer normalization layers.
classifier_dropout_prob (float, optional, defaults to 0.1) — The dropout ratio for attached classifiers.
position_embedding_type (str, optional, defaults to "absolute") — Type of position embedding. Choose one of "absolute", "relative_key", "relative_key_query". For positional embeddings use "absolute". For more information on "relative_key", please refer to Self-Attention with Relative Position Representations (Shaw et al.). For more information on "relative_key_query", please refer to Method 4 in Improve Transformer Models with Better Relative Position Embeddings (Huang et al.).

This is the configuration class to store the configuration of a AlbertModel or a TFAlbertModel. It is used to instantiate an ALBERT model according to the specified arguments, defining the model architecture. Instantiating a configuration with the defaults will yield a similar configuration to that of the ALBERT albert-xxlarge-v2 architecture.

Configuration objects inherit from PretrainedConfig and can be used to control the model outputs. Read the documentation from PretrainedConfig for more information.

Examples:

>>> from transformers import AlbertConfig, AlbertModel

>>> # Initializing an ALBERT-xxlarge style configuration
>>> albert_xxlarge_configuration = AlbertConfig()

>>> # Initializing an ALBERT-base style configuration
>>> albert_base_configuration = AlbertConfig(
...     hidden_size=768,
...     num_attention_heads=12,
...     intermediate_size=3072,
... )

>>> # Initializing a model from the ALBERT-base style configuration
>>> model = AlbertModel(albert_xxlarge_configuration)

>>> # Accessing the model configuration
>>> configuration = model.config

Transformers

ALBERT

Overview

AlbertConfig

class transformers.AlbertConfig

AlbertTokenizer

class transformers.AlbertTokenizer

build_inputs_with_special_tokens

get_special_tokens_mask

create_token_type_ids_from_sequences

save_vocabulary

AlbertTokenizerFast

class transformers.AlbertTokenizerFast

build_inputs_with_special_tokens

create_token_type_ids_from_sequences

Albert specific outputs

class transformers.models.albert.modeling_albert.AlbertForPreTrainingOutput

class transformers.models.albert.modeling_tf_albert.TFAlbertForPreTrainingOutput

AlbertModel

class transformers.AlbertModel

forward

AlbertForPreTraining

class transformers.AlbertForPreTraining

forward

AlbertForMaskedLM

class transformers.AlbertForMaskedLM

forward

AlbertForSequenceClassification

class transformers.AlbertForSequenceClassification

forward

AlbertForMultipleChoice

class transformers.AlbertForMultipleChoice

forward

AlbertForTokenClassification

class transformers.AlbertForTokenClassification

forward

AlbertForQuestionAnswering

class transformers.AlbertForQuestionAnswering

forward

TFAlbertModel

class transformers.TFAlbertModel

call

TFAlbertForPreTraining

class transformers.TFAlbertForPreTraining

call

TFAlbertForMaskedLM

class transformers.TFAlbertForMaskedLM

call

TFAlbertForSequenceClassification

class transformers.TFAlbertForSequenceClassification

call

TFAlbertForMultipleChoice

class transformers.TFAlbertForMultipleChoice

call

TFAlbertForTokenClassification

class transformers.TFAlbertForTokenClassification

call

TFAlbertForQuestionAnswering

class transformers.TFAlbertForQuestionAnswering

call

FlaxAlbertModel

class transformers.FlaxAlbertModel

__call__

FlaxAlbertForPreTraining

class transformers.FlaxAlbertForPreTraining

__call__

FlaxAlbertForMaskedLM

class transformers.FlaxAlbertForMaskedLM

__call__

FlaxAlbertForSequenceClassification

class transformers.FlaxAlbertForSequenceClassification

__call__

FlaxAlbertForMultipleChoice

class transformers.FlaxAlbertForMultipleChoice

__call__

FlaxAlbertForTokenClassification

class transformers.FlaxAlbertForTokenClassification

__call__

FlaxAlbertForQuestionAnswering

class transformers.FlaxAlbertForQuestionAnswering

call

call

call

call

call

call

call