Instructions to use alexpro100/sci-rus-tiny4-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use alexpro100/sci-rus-tiny4-GGUF with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("alexpro100/sci-rus-tiny4-GGUF") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use alexpro100/sci-rus-tiny4-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf alexpro100/sci-rus-tiny4-GGUF:F16 # Run inference directly in the terminal: llama cli -hf alexpro100/sci-rus-tiny4-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf alexpro100/sci-rus-tiny4-GGUF:F16 # Run inference directly in the terminal: llama cli -hf alexpro100/sci-rus-tiny4-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf alexpro100/sci-rus-tiny4-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf alexpro100/sci-rus-tiny4-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf alexpro100/sci-rus-tiny4-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf alexpro100/sci-rus-tiny4-GGUF:F16
Use Docker
docker model run hf.co/alexpro100/sci-rus-tiny4-GGUF:F16
- LM Studio
- Jan
- Ollama
How to use alexpro100/sci-rus-tiny4-GGUF with Ollama:
ollama run hf.co/alexpro100/sci-rus-tiny4-GGUF:F16
- Unsloth Desktop
- Docker Model Runner
How to use alexpro100/sci-rus-tiny4-GGUF with Docker Model Runner:
docker model run hf.co/alexpro100/sci-rus-tiny4-GGUF:F16
- Lemonade
How to use alexpro100/sci-rus-tiny4-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull alexpro100/sci-rus-tiny4-GGUF:F16
Run and chat with the model
lemonade run user.sci-rus-tiny4-GGUF-F16
List all available models
lemonade list
- Atomic Chat
Description
This model was converted to GGUF format from mlsa-iai-msu-lab/sci-rus-tiny4 using llama.cpp.
For more information go to here.
Test:
import numpy as np
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
import openai
# ./llama.cpp/build/bin/llama-server --models-dir sci-rus-tiny4-GGUF/ --port 8081 --embeddings
openai_client = openai.OpenAI(
base_url="http://127.0.0.1:8081/v1",
api_key="sk-no-key-required",
)
# embedding get
def get_embedding(text: str, limit_tokens: int=2048, model="embedding") -> list[float]:
response = openai_client.embeddings.create(
input=text[:limit_tokens],
model=model,
)
return response.data[0].embedding
model = SentenceTransformer(
"mlsa-iai-msu-lab/sci-rus-tiny4",
)
text = """Текст для математики. Пусть у нас есть функция f(x) = x^2 + 3x + 2. Найдите производную этой функции и определите ее критические точки."""
embed1 = model.encode(text)
for quant in ["Q8_0", "F16", "F32"]:
embed2 = np.array(get_embedding(text, model=f"sci-rus-tiny4-{quant}"), dtype=np.float32)
print(f"Cosine Similarity with {quant}: {cos_sim(embed1, embed2).item()}")
Output:
Cosine Similarity with Q8_0: 0.9999872446060181
Cosine Similarity with F16: 0.9999996423721313
Cosine Similarity with F32: 0.9999997615814209
Converting
To get the GGUF file, you have to:
- Patch
llama.cpp/conversion/base.pyto add the new model:
# (after res = "modern-bert")
# To get this hash, just run `./llama.cpp/convert_hf_to_gguf.py sci-rus-tiny4` to show the hash in the output.
if chkhsh == "762dda4b8f4cebbb9a1e702c6994236675d3acd5cf50cd3d27dcd47bb7b3f599":
# ref: https://huggingface.co/mlsa-iai-msu-lab/sci-rus-tiny4
res = "modern-bert"
- And run
./llama.cpp/convert_hf_to_gguf.py sci-rus-tiny4 --outtype q8_0 --outfile sci-rus-tiny4-GGUF/sci-rus-tiny4-Q8_0.gguf
./llama.cpp/convert_hf_to_gguf.py sci-rus-tiny4 --outtype f16 --outfile sci-rus-tiny4-GGUF/sci-rus-tiny4-F16.gguf
./llama.cpp/convert_hf_to_gguf.py sci-rus-tiny4 --outtype f32 --outfile sci-rus-tiny4-GGUF/sci-rus-tiny4-F32.gguf
- Downloads last month
- 21
Hardware compatibility
Log In to add your hardware
8-bit
16-bit
32-bit
Model tree for alexpro100/sci-rus-tiny4-GGUF
Base model
mlsa-iai-msu-lab/sci-rus-tiny4