🥋 Alibaba Qwen-2.5-7B-PolySurgery
Subtitle: Eliminating 4.26 Billion Parameters (-74.78% SwiGLU Weights) from Qwen-2.5-7B via Closed-Form Chebyshev Polynomial Tensor Surgery.
Author: Dr. A. Emre ÇETİN (Computational Systems and Cognitive Architectures, Izmir, Turkey)
Official Paper: Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks
Patent Base: U.S. Patent Application No.64/149,540&64/148,668
GitHub Repository: github.com/aemre-cetin/idempotent-poly
💡 What is Qwen-2.5-7B-PolySurgery?
Alibaba's Qwen-2.5-7B has an extraordinarily wide intermediate feed-forward dimension: $d=3584$ with $d_{ff}=18944$. This results in 203.69 Million parameters per layer in standard SwiGLU across 28 layers, consuming $5.70$ Billion parameters (75% of the total model).
Through Orthogonal Polynomial Surgery, this massive bottleneck is converted into a 4-term orthogonal tensor ($K=3$), reducing per-layer FFN parameters to 51.38 Million (-74.78% reduction).
📊 Benchmark Results
| Metric | Original Alibaba Qwen-2.5-7B | Qwen-2.5-7B-PolySurgery (Ours) | Savings / Gain |
|---|---|---|---|
| FFN Params Per Layer | 203,685,888 (203.69M) | 51,380,224 (51.38M) | -74.78% FFN Reduction |
| Total Model FFN Params | 5,703,204,864 (5.70B) | 1,438,646,272 (1.44B) | -4.265 Billion Parameters Removed! |
| Total Model Parameters | 7.61 Billion | 3.34 Billion | -56.07% Total Model Shrinkage! |
| Full Surgery Time (28 Layers) | Weeks of compute cluster | 28.25 Seconds (1009 ms/layer) | Closed-Form Algebraic SVD |
| VRAM Footprint (FP16) | 15.22 GB | 6.68 GB | Shrinks a 7.6B model into 3.3B! |
📜 Citation
@article{cetin2026orthogonalpoly,
title={Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks},
author={Çetin, A. Emre},
journal={arXiv preprint arXiv:2609.xxxxx},
year={2026}
}
- Downloads last month
- -
Model tree for aecetin/Qwen-2.5-7B-PolySurgery
Base model
Qwen/Qwen2.5-7B