Config 1: Standard Subword Transformer (Post-LN + Absolute PE + MHA)

This model is part of the ANLP Assignment 1: Custom Transformers & Byte Latent Transformers at IIIT Hyderabad.

Architecture

  • Description: Standard Encoder-Decoder Seq2Seq Transformer, Post-LayerNorm, Sinusoidal Positional Encoding, Multi-Head Attention (MHA).
  • Parameters: 11.67M

Test Evaluation Results

  • Bit-Level Accuracy: 93.60%
  • Sequence Accuracy: 68.07%
  • BLEU Score: 98.84%
  • ROUGE-1 / ROUGE-2 / ROUGE-L: 99.38% / 98.15% / 99.38%
  • Average Levenshtein Distance: 0.41

Repository & Links

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support