Amharic Word Embeddings (from Scratch)
A Skip-Gram word embedding model trained from first principles on an Amharic corpus without neural network libraries, autograd, or NumPy.
Model Summary
- Architecture: Skip-Gram with Full Categorical Softmax
- Language: Amharic (α ααα), Ge'ez script
- Embedding Dimension ($d$): 30
- Vocabulary Size ($V$): 157 (including
<UNK>) - Total Parameters: $2 \times V \times d = 9,420$ floats
- Context Window ($w$): 2 (sentence-bounded)
- Optimizer: Vanilla Stochastic Gradient Descent (SGD), $\text{lr} = 0.05$
Evaluation Metrics & Technical Description
| Metric | Value | Description |
|---|---|---|
| Final Dataset Loss ($J$) | 2.9256 | Negative log-likelihood evaluated over fixed weights across all $N = 7,950$ training pairs. Started at $5.0563$ at epoch 0 ($42.13%$ total loss reduction). |
| Perplexity ($\text{PPL} = e^J$) | 18.65 | Measures uncertainty when predicting context tokens. A random baseline over $V=157$ is $157.0$. The trained model demonstrates an $88.1%$ reduction in entropy. |
| Semantic Purity@3 | 100% (1.00) | Percentage of the 3 closest vectors matching the semantic category of the query word across test categories (beverages, kinship, crops, cities, colors). |
| Out-of-Vocabulary (OOV) Rate | 2.3% | Only $75$ out of $3,264$ tokens fell below the minimum occurrence threshold (min_count = 2) and were mapped to <UNK>. |
| Numerical Gradient Precision | $< 10^{-10}$ | Maximum relative error verified using central differences $\frac{L(\theta+\epsilon) - L(\theta-\epsilon)}{2\epsilon}$ with $\epsilon = 10^{-5}$. |
Semantic Clustering Benchmarks (Top-3 Nearest Neighbors)
| Query Word | English Translation | Nearest Neighbors (Cosine Similarity) | Semantic Category |
|---|---|---|---|
| α₯αα΄ | My mother | 1. α αα΅α΄ (0.9140) Β· 2. α₯α α΄ (0.8522) Β· 3. αα α· (0.8266) | Kinship / Female Relations |
| α‘α | Coffee | 1. αα (0.9280) Β· 2. α»α (0.9266) Β· 3. αα°α΅ (0.9119) | Beverages / Consumables |
| α€α | Teff | 1. α αα (0.9671) Β· 2. αα₯α΅ (0.9495) Β· 3. αα½α (0.9322) | Ethiopian Crops & Grains |
| ααα°α | Gondar | 1. ααα³ (0.9548) Β· 2. α α (0.9512) Β· 3. αα¨α (0.9471) | Major Ethiopian Cities |
| αα | Red | 1. α α¨ααα΄ (0.9714) Β· 2. α₯αα (0.9619) Β· 3. αα (0.9554) | Primary Colors |
Preprocessing Pipeline
- Sentence Splitting: Strictly segmented at Ethiopic full stops (
α’), question marks (α§),!,?, and line breaks. Context windows never cross sentence boundaries. - Homophone Normalization: Character-level mapping across all 7 vowel orders for Amharic homophones:
- α (hha) / α (xa) $\rightarrow$ α (ha)
- α (sza) $\rightarrow$ α° (sa)
- α (pharyngeal a) $\rightarrow$ α (glottal a)
- α (tza) $\rightarrow$ αΈ (tsa)
- Filtering: Strips punctuation and non-Ethiopic script to prevent token agglutination.
How to Use
Load and query the model directly using pure Python:
import json
import math
# Load model checkpoint
with open("model.json", "r", encoding="utf-8") as f:
model = json.load(f)
E = model["E"]
vocab = model["vocab"]
word_to_id = {w: idx for idx, w in enumerate(vocab)}
def cosine(a, b):
dot = sum(ak * bk for ak, bk in zip(a, b))
norm_a = math.sqrt(sum(ak * ak for ak in a))
norm_b = math.sqrt(sum(bk * bk for bk in b))
return dot / (norm_a * norm_b) if norm_a and norm_b else 0.0
def find_neighbors(word, k=3):
if word not in word_to_id:
return f"Word '{word}' not in vocabulary."
q_idx = word_to_id[word]
scores = [(w, cosine(E[q_idx], E[i])) for i, w in enumerate(vocab) if i != q_idx and w != "<UNK>"]
scores.sort(key=lambda x: x[1], reverse=True)
return scores[:k]
# Example Query
print(find_neighbors("α‘α", k=3))
# Output: [('αα', 0.9280), ('α»α', 0.9266), ('αα°α΅', 0.9119)]
---
Evaluation results
- Dataset Cross-Entropy Loss (J) on Custom Amharic Corpusself-reported2.926
- Perplexity (PPL) on Custom Amharic Corpusself-reported18.650
- Semantic Purity@3 on Custom Amharic Corpusself-reported1.000
- Out-of-Vocabulary (OOV) Rate on Custom Amharic Corpusself-reported0.023
- Numerical Gradient Relative Error on Custom Amharic Corpusself-reported0.000