PsarAI-2B

PsarAI-2B is a PsarAI chat model exported in Hugging Face format.

The model uses a Gemma4-style architecture and a PsarAI chat template. The assistant identity in the template is:

You are PsarAI, created by the PsarAI team under the leadership of an ITC lecturer.

Files

This repository contains the standard Hugging Face model export:

File Purpose
model.safetensors model weights
config.json model architecture/config
tokenizer.json tokenizer
tokenizer_config.json tokenizer metadata and special tokens
processor_config.json multimodal processor config
chat_template.jinja chat formatting template
generation_config.json generation defaults

Quick Start

import torch
from transformers import AutoProcessor, AutoModelForCausalLM

repo_id = "nphearum/PsarAI-2B"

processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)

messages = [
    {"role": "user", "content": "Who created you?"}
]

prompt = processor.tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,
)

inputs = processor.tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
    **inputs,
    max_new_tokens=256,
    temperature=0.7,
    top_p=0.9,
)

print(processor.tokenizer.decode(outputs[0], skip_special_tokens=False))

Chat Template

The template uses Gemma-style tokens:

  • <|turn>system
  • <|turn>user
  • <|turn>model
  • <turn|>
  • <|channel>thought
  • <|tool_call>
  • <|tool_response>

For normal chatbot use, disable visible thinking when your runtime supports template kwargs:

enable_thinking=False

Suggested Generation Settings

temperature = 0.7
top_p = 0.9
max_new_tokens = 512

Use lower temperature, such as 0.2, for factual or deterministic answers.

Multimodal Notes

The config includes image, audio, and video processor metadata. Runtime support depends on the installed transformers version and model implementation availability.

For GGUF/llama.cpp usage, use the sibling GGUF export repo instead:

nphearum/PsarAI-2B-GGUF

Attribution

Base model metadata in this export is:

phearum/psarai-2b

Keep this metadata for traceability when publishing derived formats.

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support