Glyph — Multi-Task Byte-Level Text Classifier

4M parameters · 5 tasks · No tokenizer needed · Runs on CPU

Glyph is an ultra-compact multi-task text classification model that operates directly on raw UTF-8 bytes. A single shared backbone serves 5 classification heads simultaneously.

Tasks & Performance

Task Labels Val Accuracy Description
lang_id 351 94.3% Language identification
prog_lang 30 94.4% Programming language detection
spam 2 96.1% Spam vs ham classification
toxic 2 67.3% Toxicity detection (multilingual)

Average validation accuracy: 88.0%

Architecture

Glyph uses a byte-level hybrid architecture inspired by CommonLingua:

  • Input: Raw UTF-8 bytes (no tokenizer), padded to 512 bytes
  • Trigram hash embedding: Polynomial rolling hash of byte 3-grams → 8192-bucket embedding table
  • Byte unigram embedding: Standard embedding for individual bytes
  • 4× Conv1D blocks: Causal convolutions + BatchNorm + GELU + residual
  • 2× Bidirectional attention: Multi-head self-attention with RoPE
  • Global average pooling → per-task classification heads

Total: ~4M shared parameters + small per-task linear heads.

Usage

import torch, importlib, sys
from huggingface_hub import hf_hub_download

# Download model + weights
model_py_path = hf_hub_download("ThingAI/Glyph", "model.py")
ckpt_path = hf_hub_download("ThingAI/Glyph", "model.pt")

# Load model definition
import importlib.util
spec = importlib.util.spec_from_file_location("model", model_py_path)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)

# Load weights
ckpt = torch.load(ckpt_path, map_location="cpu", weights_only=False)
model = mod.MultiTaskLID(ckpt["task_configs"]).eval()
model.load_state_dict(ckpt["model"])

# Predict
def predict(text, task="lang_id", top_k=3):
    raw = text.encode("utf-8")[:512]
    byte_ids = list(raw) + [0] * (512 - len(raw))
    inp = torch.tensor([byte_ids], dtype=torch.long)
    with torch.no_grad():
        logits = model(inp, task)["logits"][0]
    probs = torch.softmax(logits, dim=-1)
    idx2label = {v: k for k, v in ckpt["label_maps"][task].items()}
    k = min(top_k, len(probs))
    topk = probs.topk(k)
    return [(idx2label[topk.indices[i].item()], topk.values[i].item()) for i in range(k)]

# Examples
print(predict("La pizza napoletana è patrimonio UNESCO", task="lang_id"))
print(predict("def foo(x): return x + 1", task="prog_lang"))
print(predict("You won a FREE iPhone!!!", task="spam"))

Quick predict script

python predict.py --text "Ciao, come stai?"
# lang_id: ita (98.2%)

python predict.py --task prog_lang --text "fn main() { println!("hello"); }"
# prog_lang: Rust (97.1%)

python predict.py --task spam --text "URGENT: Click here to win!"
# spam: spam (99.5%)

Training Data

Task Dataset Samples
Language ID PleIAs/CommonLingua-Train 200K (capped)
Programming Language cakiki/rosetta-code ~26K
Spam Detection ucirvine/sms_spam 5.5K
Toxicity textdetox/multilingual_toxicity_dataset ~45K

Key Design Decisions

  • Byte-level input: No tokenizer means it works on any language, script, or encoding without preprocessing
  • Trigram hashing: +1.2 F1 over unigram-only baseline, acts as regularization via hash collisions
  • Multi-task learning: Shared backbone learns universal text representations; per-task heads are tiny (~130K params each)
  • No attention masking: Bidirectional attention for classification (not causal)
  • OneCycleLR: Fast convergence in few epochs

Limitations

  • Designed for paragraph-level classification (50+ bytes). Short texts (<20 bytes) may be unreliable
  • Toxicity detection accuracy is lower than specialized models (trained on limited multilingual data)
  • Programming language detection trained on Rosetta Code samples which have a specific style

Citation

@misc{glyph2026,
  author = {ThingAI},
  title  = {Glyph: Multi-Task Byte-Level Text Classifier},
  year   = {2026},
  url    = {https://huggingface.co/ThingAI/Glyph}
}

License

Apache 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support