AI & ML interests

None defined yet.

Recent Activity

Organization Card

🐻🔥 Model Rampage

Building Hardware-Portable, Pure-GEMM Sub-Quadratic LLM Architectures

GitHub Repo Paper PDF License


🌟 About Model Rampage

Model Rampage is an independent AI research organization focused on breaking the custom kernel lock-in and long-context memory wall in modern language models.

Our flagship framework, BareTorch, formulates next-generation sub-quadratic sequence mixers (like CS-LRAD) using strictly pure, high-level matrix multiplication (GEMM) equations. By eliminating hardware-specific CUDA/Triton kernels, our models run with linear $O(N)$ execution scaling and constant $O(1)$ memory state updates natively across NVIDIA CUDA, Apple Silicon MLX, WebGPU, and TPUs.


🚀 Featured Models


⚡ Key Performance Breakthroughs (32,768 Context)

  • 44.98x Faster Edge Decoding (Apple Silicon MLX): Streams at 29.69 tok/s on an M1 MacBook Pro at 32k context where standard MLX Transformer baselines collapse to 0.66 tok/s.
  • 12.51x Faster CUDA Decoding (NVIDIA RTX 4090): Reaches 164.49 tok/s at 32k context vs 13.15 tok/s for standard attention.
  • 80.6% VRAM Footprint Savings: Replaces linear KV-caches with constant-sized recurrent state updates, running 32k prompt generation seamlessly where standard baselines crash with Out-Of-Memory (OOM) errors.

📊 Benchmark Summary (BareTorch-500M-SFT)

Benchmark Task Metric Pre-Trained Base SFT Instruction Aligned
HellaSwag Acc-Norm 43.69% 52.52%
ARC-Easy Acc-Norm 53.58% 59.18%
ARC-Challenge Acc-Norm 28.92% 35.41%
WinoGrande Accuracy 51.30% 55.80%
MMLU (Overall) Accuracy 24.70% 25.57%

🔗 Resources & Citation

@article{kovacevic2026baretorch,
  title={BareTorch: Challenging State-of-The-Art Sequence Mixing Topologies via Kernel-Free, Pure GEMM-Compliant Architectures},
  author={Kovacevic Buvinic, Martin Ignacio},
  journal={BareTorch Framework Laboratory Technical Report},
  year={2026}
}

datasets 0

None public yet