🦙 Meta Llama-3.2-3B-PolySurgery

Subtitle: Closed-Form In-Situ Surgery on Meta Llama-3.2-3B: Eliminating 1.06 Billion Parameters (-50.00% SwiGLU Weights) in ~40 Seconds.
Author: Dr. A. Emre ÇETİN (Computational Systems and Cognitive Architectures, Izmir, Turkey)
Official Paper: Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks
Patent Base: U.S. Patent Application No. 64/149,540 & 64/148,668
GitHub Repository: github.com/aemre-cetin/idempotent-poly


💡 What is Llama-3.2-3B-PolySurgery?

Meta's Llama-3.2-3B utilizes SwiGLU feed-forward blocks with hidden size $d=3072$ and intermediate size $d_{ff}=8192$. Across 28 layers, this requires $2.11$ Billion parameters in MLP weights alone.

By computing orthogonal Chebyshev polynomial tensor projections ($K=3$) using closed-form normal equations, intermediate projections are eliminated, cutting per-layer FFN parameters from 75.50M down to 37.75M (-50.00%).

📊 Benchmark Results

Metric Original Meta Llama-3.2-3B Llama-3.2-3B-PolySurgery (Ours) Savings / Gain
FFN Params Per Layer 75,497,472 (75.50M) 37,748,736 (37.75M) -50.00% FFN Reduction
Total Model FFN Params 2,113,929,216 (2.11B) 1,056,964,608 (1.06B) -1.057 Billion Parameters Removed!
Total Model Parameters 3,212,749,824 (3.21B) 2,155,785,216 (2.16B) -32.89% Total Model Shrinkage
Full Surgery Time (28 Layers) Days of gradient tuning 11.13 Seconds (397.4 ms/layer) Zero-Backprop Closed-Form
Inference Acceleration 1.00x (Baseline) ~2.2x Faster Reduced Projection Bottleneck
VRAM Footprint (FP16) 6.42 GB 4.31 GB Fits in 4GB-6GB Edge GPUs!

💻 Quickstart & Verification

# pip install idempotent-poly torch
from idempotent_poly.chebyshev import ChebyshevPolyFFN

# Degree 3 Chebyshev tensor replaces SwiGLU:
cheb_ffn = ChebyshevPolyFFN(d_model=3072, degree=3)
print("PolyFFN Params:", sum(p.numel() for p in cheb_ffn.parameters()))
# Output: 37,748,736 (-50.00% vs 75,497,472)

📜 Citation

@article{cetin2026orthogonalpoly,
  title={Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks},
  author={Çetin, A. Emre},
  journal={arXiv preprint arXiv:2609.xxxxx},
  year={2026}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aecetin/Llama-3.2-3B-PolySurgery

Finetuned
(524)
this model