Language: Python

This is an experiment with building a language model from scratch. Instead of using subword tokens, my experiments operate directly on bytes (UTF-8), allowing the model to learn from the raw byte stream.

Inspired by how the human brain continuously predicts future sensory inputs, I wanted to explore whether a compact byte-level architecture could learn meaningful structural representations of programming languages while combining autoregressive prediction with latent representation learning.

The current prototype consists of:

A byte embedding layer (384 dimensions) A shared encoder built from six Selective SSM blocks with RMSNorm and SwiGLU feed-forward layers An autoregressive head for next-byte prediction A JEPA-inspired predictive branch with an EMA target encoder for learning latent representations

(Architecture diagram attached.)

The project is still in its early stages, but several observations have been encouraging.

The model has progressed from generating random byte sequences to producing outputs with recognizable Python syntax and indentation. It has started learning structural patterns such as keywords, punctuation, and code layout, although longer-range semantic coherence is still limited. The current implementation is constrained by my available hardware and a relatively small training corpus, so it should be viewed as an exploratory research prototype rather than a competitive language model.

There are also several improvements planned for the next iterations:

Expand training to substantially larger programming-language datasets. Replace the current placeholder Selective SSM with a more faithful state-space implementation. Add EoPE (Enhanced/Extended Positional Encoding) to improve long-range sequence modelling. Continue evaluating how JEPA-style latent prediction complements autoregressive byte prediction.

This series is an experiment in understanding language-model architecture from first principles rather than treating existing models as black boxes. In future articles, I'll share what works, what doesn't, and how the architecture evolves as additional components are introduced.

image

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support