🦙 Meta Llama-3.2-3B-PolySurgery
Subtitle: Closed-Form In-Situ Surgery on Meta Llama-3.2-3B: Eliminating 1.06 Billion Parameters (-50.00% SwiGLU Weights) in ~40 Seconds.
Author: Dr. A. Emre ÇETİN (Computational Systems and Cognitive Architectures, Izmir, Turkey)
Official Paper: Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks
Patent Base: U.S. Patent Application No.64/149,540&64/148,668
GitHub Repository: github.com/aemre-cetin/idempotent-poly
💡 What is Llama-3.2-3B-PolySurgery?
Meta's Llama-3.2-3B utilizes SwiGLU feed-forward blocks with hidden size $d=3072$ and intermediate size $d_{ff}=8192$. Across 28 layers, this requires $2.11$ Billion parameters in MLP weights alone.
By computing orthogonal Chebyshev polynomial tensor projections ($K=3$) using closed-form normal equations, intermediate projections are eliminated, cutting per-layer FFN parameters from 75.50M down to 37.75M (-50.00%).
📊 Benchmark Results
| Metric | Original Meta Llama-3.2-3B | Llama-3.2-3B-PolySurgery (Ours) | Savings / Gain |
|---|---|---|---|
| FFN Params Per Layer | 75,497,472 (75.50M) | 37,748,736 (37.75M) | -50.00% FFN Reduction |
| Total Model FFN Params | 2,113,929,216 (2.11B) | 1,056,964,608 (1.06B) | -1.057 Billion Parameters Removed! |
| Total Model Parameters | 3,212,749,824 (3.21B) | 2,155,785,216 (2.16B) | -32.89% Total Model Shrinkage |
| Full Surgery Time (28 Layers) | Days of gradient tuning | 11.13 Seconds (397.4 ms/layer) | Zero-Backprop Closed-Form |
| Inference Acceleration | 1.00x (Baseline) | ~2.2x Faster | Reduced Projection Bottleneck |
| VRAM Footprint (FP16) | 6.42 GB | 4.31 GB | Fits in 4GB-6GB Edge GPUs! |
💻 Quickstart & Verification
# pip install idempotent-poly torch
from idempotent_poly.chebyshev import ChebyshevPolyFFN
# Degree 3 Chebyshev tensor replaces SwiGLU:
cheb_ffn = ChebyshevPolyFFN(d_model=3072, degree=3)
print("PolyFFN Params:", sum(p.numel() for p in cheb_ffn.parameters()))
# Output: 37,748,736 (-50.00% vs 75,497,472)
📜 Citation
@article{cetin2026orthogonalpoly,
title={Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks},
author={Çetin, A. Emre},
journal={arXiv preprint arXiv:2609.xxxxx},
year={2026}
}
- Downloads last month
- -
Model tree for aecetin/Llama-3.2-3B-PolySurgery
Base model
meta-llama/Llama-3.2-3B