Description

This model was converted to GGUF format from tbs17/MathBERT using llama.cpp.

For more information go to here.

Test:

import numpy as np
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
import openai

# ./llama.cpp/build/bin/llama-server --models-dir MathBERT-GGUF/ --embeddings
openai_client = openai.OpenAI(
    base_url="http://127.0.0.1:8080/v1",
    api_key="sk-no-key-required",
)

# embedding get
def get_embedding(text: str, limit_tokens: int=2048, model="embedding") -> list[float]:
    response = openai_client.embeddings.create(
        input=text[:limit_tokens],
        model=model,
    )
    return response.data[0].embedding

model = SentenceTransformer(
    "tbs17/MathBERT",
)

text = """This is text for math. It contains some math formulas, such as $E=mc^2$ and $\\int_0^\\infty e^{-x} dx = 1$."""

embed1 = model.encode(text)

for quant in ["Q8_0", "F16", "F32"]:
    embed2 = np.array(get_embedding(text, model=f"MathBERT-0.1B-{quant}"), dtype=np.float32)
    print(f"Cosine Similarity with {quant}: {cos_sim(embed1, embed2).item()}")

Output:

Cosine Similarity with Q8_0: 0.9999127984046936
Cosine Similarity with F16: 0.9999980926513672
Cosine Similarity with F32: 0.9999997019767761

Converting

To get the GGUF file, you have to:

  1. Add file MathBERT/1_Pooling/config.json:
{"word_embedding_dimension":768,"pooling_mode":"mean","pooling_mode_cls_token":false,"pooling_mode_mean_tokens":true,"pooling_mode_max_tokens":false,"pooling_mode_mean_sqrt_len_tokens":false}
  1. Add file 'MathBERT/modules.json':
[
  {
    "idx": 0,
    "name": "0",
    "path": "",
    "type": "sentence_transformers.models.Transformer"
  },
  {
    "idx": 1,
    "name": "1",
    "path": "1_Pooling",
    "type": "sentence_transformers.models.Pooling"
  }
]
  1. And run
./llama.cpp/convert_hf_to_gguf.py MathBERT --outtype q8_0 --outfile MathBERT-GGUF/MathBERT-0.1B-Q8_0.gguf
./llama.cpp/convert_hf_to_gguf.py MathBERT --outtype f16 --outfile MathBERT-GGUF/MathBERT-0.1B-F16.gguf
./llama.cpp/convert_hf_to_gguf.py MathBERT --outtype f32 --outfile MathBERT-GGUF/MathBERT-0.1B-F32.gguf
Downloads last month
16
GGUF
Model size
0.1B params
Architecture
bert
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for alexpro100/MathBERT-GGUF

Base model

tbs17/MathBERT
Quantized
(1)
this model