Instructions to use poolside/Laguna-S-2.1-NVFP4-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use poolside/Laguna-S-2.1-NVFP4-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Laguna-S-2.1-NVFP4-mlx poolside/Laguna-S-2.1-NVFP4-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
Endless Looping
Thank you for the model. This model loops endlessly on coding tasks and with reasoning on it keeps on thinking in a loop and never converges to an answer, I have also seen it decide on an answer and then restart the whole thinking process again. I am running with poolside's llama.cpp branch.
I also tried pipenetworks MLX 6 bit all of these seem to be exhibiting the same issue.
Hi there! We've updated the NVFP4 on poolside/Laguna-S-2.1-NVFP4. I'm planning on rolling out the new NVFP4 here too in the next few hours. In our testing it exhibits far less looping.
Please note that we have another variant planned in the next day or so that should resolve all such issues, so feel free to wait for that too.
Hi there! We've updated the NVFP4 on
poolside/Laguna-S-2.1-NVFP4. I'm planning on rolling out the new NVFP4 here too in the next few hours. In our testing it exhibits far less looping.Please note that we have another variant planned in the next day or so that should resolve all such issues, so feel free to wait for that too.
Thank you, looking forward to running it, will test when the new version drops and report back here.
Hi @abhiksngh , the first of the two updates has landed. I'll update again when we have the follow-up :)
Hi @abhiksngh , the first of the two updates has landed. I'll update again when we have the follow-up :)
Thanks, what exactly is in the first update? I am only seeing the generation_config.json file being updated and not the model files itself. Am i missing something?
Cuando le pregunto que modelo LLM eres?. El modelo responde :"¡Hola! Soy GPT-3 (Generative Pre-trained Transformer 3), un modelo de lenguaje de gran tamaño desarrollado por OpenAI. Mi tecnología se basa en la arquitectura transformers, que utiliza mecanismos de atención (attention) para procesar y generar texto de manera coherente y contextual.
Fui entrenado con una gran cantidad de texto en múltiples idiomas, incluido el español, lo que me permite comprender y generar respuestas en este idioma. Mi versión específica está optimizada para tareas como responder preguntas, ayudar con tareas creativas o técnicas, y mantener conversaciones naturales." es normal o algún problema?
Has anyone tested the performance of this model on a MacBook Pro with an M Max chip and 128 GB of RAM?