YAML Metadata Warning:The pipeline tag "text2text-generation" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other

ASTRA 1.4B Model Description

The ASTRA 1.4B is a highly efficient large language model designed for diverse natural language processing (NLP) tasks. Below are the specifications and architectural features of the model:

Model Specifications

  • Vocabulary Size: The model can recognize and utilize a vocabulary of 100,277 tokens, ensuring broad linguistic and contextual coverage.
  • Context Length: The maximum sequence length that the model can process is 2,048 tokens, making it suitable for long-form text generation and comprehension.
  • Embedding Dimension: The model uses an embedding size of 2,048 dimensions, providing a rich representation of input tokens.
  • Number of Attention Heads: With 32 attention heads, the model achieves high parallelization and fine-grained attention distribution.
  • Number of Transformer Layers: The architecture comprises 24 transformer layers, enabling the model to capture deep hierarchical language features.
  • Dropout Rate: A 10% dropout rate is applied during training to prevent overfitting and improve generalization.
  • Query-Key-Value Bias: The model incorporates bias terms in its attention mechanisms, allowing for better flexibility in capturing relationships.

Key Features

  1. Scalability: With 1.4 billion parameters, ASTRA 1.4B strikes a balance between computational efficiency and performance, making it suitable for deployment in resource-constrained environments.
  2. Adaptability: Its architectural design ensures high adaptability for various NLP tasks, including:
    • Text generation
    • Machine translation
    • Sentiment analysis
    • Question answering
  3. Optimized Training: The inclusion of bias terms (qkv_bias) and a tuned dropout rate enhances both training stability and model performance on unseen data.
  4. Rich Representations: The 2,048-dimensional embeddings and 32 attention heads contribute to capturing nuanced patterns and relationships in language.

Use Cases

  • Content Generation: Generate coherent and contextually relevant text for creative and informational purposes.
  • Conversational AI: Build advanced chatbots and virtual assistants capable of maintaining meaningful and context-aware conversations.
  • Document Summarization: Extract concise summaries from lengthy documents.
  • Language Understanding: Solve complex tasks like named entity recognition and part-of-speech tagging.

Model Summary

The ASTRA 1.4B model combines modern transformer-based architecture with a well-balanced parameterization, offering robust performance across a wide range of NLP tasks. Its design allows for efficient training and inference, making it an ideal choice for applications requiring both scalability and precision.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support