🥋 Alibaba Qwen-2.5-7B-PolySurgery

Subtitle: Eliminating 4.26 Billion Parameters (-74.78% SwiGLU Weights) from Qwen-2.5-7B via Closed-Form Chebyshev Polynomial Tensor Surgery.
Author: Dr. A. Emre ÇETİN (Computational Systems and Cognitive Architectures, Izmir, Turkey)
Official Paper: Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks
Patent Base: U.S. Patent Application No. 64/149,540 & 64/148,668
GitHub Repository: github.com/aemre-cetin/idempotent-poly


💡 What is Qwen-2.5-7B-PolySurgery?

Alibaba's Qwen-2.5-7B has an extraordinarily wide intermediate feed-forward dimension: $d=3584$ with $d_{ff}=18944$. This results in 203.69 Million parameters per layer in standard SwiGLU across 28 layers, consuming $5.70$ Billion parameters (75% of the total model).

Through Orthogonal Polynomial Surgery, this massive bottleneck is converted into a 4-term orthogonal tensor ($K=3$), reducing per-layer FFN parameters to 51.38 Million (-74.78% reduction).

📊 Benchmark Results

Metric Original Alibaba Qwen-2.5-7B Qwen-2.5-7B-PolySurgery (Ours) Savings / Gain
FFN Params Per Layer 203,685,888 (203.69M) 51,380,224 (51.38M) -74.78% FFN Reduction
Total Model FFN Params 5,703,204,864 (5.70B) 1,438,646,272 (1.44B) -4.265 Billion Parameters Removed!
Total Model Parameters 7.61 Billion 3.34 Billion -56.07% Total Model Shrinkage!
Full Surgery Time (28 Layers) Weeks of compute cluster 28.25 Seconds (1009 ms/layer) Closed-Form Algebraic SVD
VRAM Footprint (FP16) 15.22 GB 6.68 GB Shrinks a 7.6B model into 3.3B!

📜 Citation

@article{cetin2026orthogonalpoly,
  title={Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks},
  author={Çetin, A. Emre},
  journal={arXiv preprint arXiv:2609.xxxxx},
  year={2026}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aecetin/Qwen-2.5-7B-PolySurgery

Base model

Qwen/Qwen2.5-7B
Finetuned
(950)
this model