Amharic Word Embeddings (from Scratch)

A Skip-Gram word embedding model trained from first principles on an Amharic corpus without neural network libraries, autograd, or NumPy.

Model Summary

  • Architecture: Skip-Gram with Full Categorical Softmax
  • Language: Amharic (αŠ αˆ›αˆ­αŠ›), Ge'ez script
  • Embedding Dimension ($d$): 30
  • Vocabulary Size ($V$): 157 (including <UNK>)
  • Total Parameters: $2 \times V \times d = 9,420$ floats
  • Context Window ($w$): 2 (sentence-bounded)
  • Optimizer: Vanilla Stochastic Gradient Descent (SGD), $\text{lr} = 0.05$

Evaluation Metrics & Technical Description

Metric Value Description
Final Dataset Loss ($J$) 2.9256 Negative log-likelihood evaluated over fixed weights across all $N = 7,950$ training pairs. Started at $5.0563$ at epoch 0 ($42.13%$ total loss reduction).
Perplexity ($\text{PPL} = e^J$) 18.65 Measures uncertainty when predicting context tokens. A random baseline over $V=157$ is $157.0$. The trained model demonstrates an $88.1%$ reduction in entropy.
Semantic Purity@3 100% (1.00) Percentage of the 3 closest vectors matching the semantic category of the query word across test categories (beverages, kinship, crops, cities, colors).
Out-of-Vocabulary (OOV) Rate 2.3% Only $75$ out of $3,264$ tokens fell below the minimum occurrence threshold (min_count = 2) and were mapped to <UNK>.
Numerical Gradient Precision $< 10^{-10}$ Maximum relative error verified using central differences $\frac{L(\theta+\epsilon) - L(\theta-\epsilon)}{2\epsilon}$ with $\epsilon = 10^{-5}$.

Semantic Clustering Benchmarks (Top-3 Nearest Neighbors)

Query Word English Translation Nearest Neighbors (Cosine Similarity) Semantic Category
αŠ₯αŠ“α‰΄ My mother 1. αŠ αŠ­αˆ΅α‰΄ (0.9140) Β· 2. αŠ₯αˆ…α‰΄ (0.8522) Β· 3. αˆαŒ…α‰· (0.8266) Kinship / Female Relations
α‰‘αŠ“ Coffee 1. α‹αˆƒ (0.9280) Β· 2. αˆ»α‹­ (0.9266) Β· 3. α‹ˆα‰°α‰΅ (0.9119) Beverages / Consumables
ጀፍ Teff 1. α‰ α‰†αˆŽ (0.9671) Β· 2. ገα‰₯ሡ (0.9495) Β· 3. αˆ›αˆ½αˆ‹ (0.9322) Ethiopian Crops & Grains
αŒŽαŠ•α‹°αˆ­ Gondar 1. αˆ€α‹‹αˆ³ (0.9548) Β· 2. αŒ…αˆ› (0.9512) Β· 3. αˆ€αˆ¨αˆ­ (0.9471) Major Ethiopian Cities
ቀይ Red 1. αŠ αˆ¨αŠ•αŒ“α‹΄ (0.9714) Β· 2. αŒ₯α‰αˆ­ (0.9619) Β· 3. ነጭ (0.9554) Primary Colors

Preprocessing Pipeline

  1. Sentence Splitting: Strictly segmented at Ethiopic full stops (ፒ), question marks (፧), !, ?, and line breaks. Context windows never cross sentence boundaries.
  2. Homophone Normalization: Character-level mapping across all 7 vowel orders for Amharic homophones:
    • ሐ (hha) / αŠ€ (xa) $\rightarrow$ αˆ€ (ha)
    • ሠ (sza) $\rightarrow$ ሰ (sa)
    • ዐ (pharyngeal a) $\rightarrow$ አ (glottal a)
    • ፀ (tza) $\rightarrow$ ጸ (tsa)
  3. Filtering: Strips punctuation and non-Ethiopic script to prevent token agglutination.

How to Use

Load and query the model directly using pure Python:

import json
import math

# Load model checkpoint
with open("model.json", "r", encoding="utf-8") as f:
    model = json.load(f)

E = model["E"]
vocab = model["vocab"]
word_to_id = {w: idx for idx, w in enumerate(vocab)}

def cosine(a, b):
    dot = sum(ak * bk for ak, bk in zip(a, b))
    norm_a = math.sqrt(sum(ak * ak for ak in a))
    norm_b = math.sqrt(sum(bk * bk for bk in b))
    return dot / (norm_a * norm_b) if norm_a and norm_b else 0.0

def find_neighbors(word, k=3):
    if word not in word_to_id:
        return f"Word '{word}' not in vocabulary."
    q_idx = word_to_id[word]
    scores = [(w, cosine(E[q_idx], E[i])) for i, w in enumerate(vocab) if i != q_idx and w != "<UNK>"]
    scores.sort(key=lambda x: x[1], reverse=True)
    return scores[:k]

# Example Query
print(find_neighbors("α‰‘αŠ“", k=3))
# Output: [('α‹αˆƒ', 0.9280), ('αˆ»α‹­', 0.9266), ('α‹ˆα‰°α‰΅', 0.9119)]

---
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results

  • Dataset Cross-Entropy Loss (J) on Custom Amharic Corpus
    self-reported
    2.926
  • Perplexity (PPL) on Custom Amharic Corpus
    self-reported
    18.650
  • Semantic Purity@3 on Custom Amharic Corpus
    self-reported
    1.000
  • Out-of-Vocabulary (OOV) Rate on Custom Amharic Corpus
    self-reported
    0.023
  • Numerical Gradient Relative Error on Custom Amharic Corpus
    self-reported
    0.000