🚀 Eblan 2.0 Flash Lite (EblanForCausalLM)
Eblan 2.0 Flash Lite is a next-generation lightweight language model built on the custom EblanForCausalLM architecture.
By leveraging an innovative O(1) polynomial recurrent state kernel, the model delivers ultra-fast inference with virtually zero memory overhead, running natively via NumPy and Tiktoken.
🏗️ Architecture (EblanForCausalLM)
Unlike traditional Transformer-based models that rely on heavy Attention blocks (Q, K, V), EblanForCausalLM utilizes a single scalar weight vector W (in R¹) with a direct scalar projection layer.
Forward Pass Formulation:
Where:
- x =
current_id/vocab_size— normalized input token index. - W — trained scalar weight (
eblan-2.0-flash-lite.npy). - Logits formula for distance-based logit calculation prior to Softmax sampling:
📜 License
This project is licensed under the MIT License.
Evaluation results
- Accuracy on Synthetic Benchmarkself-reported314.000