ai21labs/Jamba-v0.1 · Would there a chance Jamba to be train in 1.58bit weight?

No one has tested SSMs with the 1.58Bit strategy. It is likely that research would need to be done on Mamba before it is done on Jamba as the BitNet only creates BitLinear and not BitConv1D (yet) and we don't know how something like this will perform. Then when you consider the fact that Mamba already provides efficient inference and it is more likely that other methods (speculative decoding for example) are used for more efficient inference before this kind of quantization aware training. All that being said it could still prove to be an interesting line of research, but it is not just low-hanging fruit.