Instructions to use patriotmemory-ai/PMA-1.3-Rosa-256M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use patriotmemory-ai/PMA-1.3-Rosa-256M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="patriotmemory-ai/PMA-1.3-Rosa-256M", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("patriotmemory-ai/PMA-1.3-Rosa-256M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use patriotmemory-ai/PMA-1.3-Rosa-256M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "patriotmemory-ai/PMA-1.3-Rosa-256M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "patriotmemory-ai/PMA-1.3-Rosa-256M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/patriotmemory-ai/PMA-1.3-Rosa-256M
- SGLang
How to use patriotmemory-ai/PMA-1.3-Rosa-256M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "patriotmemory-ai/PMA-1.3-Rosa-256M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "patriotmemory-ai/PMA-1.3-Rosa-256M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "patriotmemory-ai/PMA-1.3-Rosa-256M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "patriotmemory-ai/PMA-1.3-Rosa-256M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use patriotmemory-ai/PMA-1.3-Rosa-256M with Docker Model Runner:
docker model run hf.co/patriotmemory-ai/PMA-1.3-Rosa-256M
Model Overview
Rosa is Patriot Memory's 256M-parameter edge assistant for English and Traditional Chinese — built to fit the edge devices you actually ship, with a vocabulary trained natively on Traditional Chinese. It provides accurate information regarding:
- DDR4 & DDR5 RAM: Specifications, XMP 3.0 / EXPO profile support, dual-channel setups, and overclocking guidance.
- PCIe & SATA SSDs: Gen3/Gen4/Gen5 compatibility, read/write performance specifications, and installation troubleshooting.
- Gaming Peripherals & Storage: USB drives, flash cards, and Viper Gaming gear.
- Tool / Function Calling: Seamless integration with backend APIs (e.g., checking warranty status, looking up technical specs via S/N).
Architecture
| Parameters | 253,283,329 (253M-class; fits 256M edge budget) |
| Memory | ~1.0 GB fp32 master weights (cast to fp16 at load for ~500 MB) |
| Architecture | PMA spine — 19 layers, hidden 1024, GQA 8q/2kv, gated attention output, value residuals, Norm-Head output, tied embeddings |
| Tokenizer | OWN SentencePiece 8k unigram, trained on Traditional-Chinese + English corpus — Traditional-exclusive glyphs encode as real pieces |
| Context | 4096 positions (2048-token training windows) |
| Reasoning | English rationale format: Reasoning: … Answer: … |
| Training | 2.3B-token adapted pretrain on the TW tokenizer (warm-started), SFT with identity/format/rationale pools + targeted repair pass |
Quickstart
pip install -U torch transformers==4.51.0 accelerate sentencepiece protobuf
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
name = "patriotmemory-ai/PMA-1.3-Rosa-256M"
tok = AutoTokenizer.from_pretrained(name,trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
name, torch_dtype=torch.float16, device_map="auto",
trust_remote_code=True)
msgs = [{"role": "user", "content": "博帝的 Viper DDR5 支援 XMP 3.0 嗎?"}]
ids = tok.apply_chat_template(msgs, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=200)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
Decode: temperature 0.8, top_p 0.9, repetition_penalty 1.1
If you run into multi-GPU tensor device mismatch errors: RuntimeError: Expected all tensors to be on the same device... Run the script with CUDA_VISIBLE_DEVICES=0 to isolate execution to GPU 0.
Acceptance gates (measured, 3-sample pooled)
| Gate | Result |
|---|---|
| A — story shape (EN + zh-TW) | 29/30 |
| B — number format | 27/30 |
| D — identity + injection defense | 18/18 |
F — Reasoning: … Answer: rationale format |
27/30 |
| T — Traditional-Chinese purity (all zh generations) | 36/36 |
What she's good at / not good at
Good: introducing herself in both languages; answering arithmetic word problems with visible English reasoning steps; clean Traditional Chinese — glyphs, vocabulary and register; running fully offline on edge hardware.
Not: her reasoning traces are formatted correctly ~90% of the time but at 253M the arithmetic inside them is occasionally wrong-but-confident — treat the steps as display, verify the numbers. Long open-ended creative Chinese can drift into polite deflection, and unusual open-ended generation tasks (draw-me-this, write-me-that in novel domains) can loop in paraphrase instead of producing the artifact — not for long-document streaming.
Research preview: draw
One extra SVG-taught pass — generates valid SVG directly; these are its unretouched outputs, source files included in images/:
Source: images/pelican_bike.svg, images/pelican.svg, images/bicycle.svg, images/cat.svg.
Limitations & Responsible Use
PMA-1.3-Rosa-256M is a probabilistic language model trained on statistical patterns. Please keep the following in mind when deploying or evaluating this model:
- Generation Risks: The model may generate inaccurate, hallucinated, biased, or objectionable content. Outputs should always be independently verified—especially in high-stakes domain applications (e.g., medical, legal, or financial).
- Preview Release: As an experimental preview, model behavior, outputs, and performance metrics may vary between updates and versions.
- User Responsibility: Users and developers are responsible for implementing appropriate safety guardrails, evaluating outputs for their specific use cases, and ensuring compliance with applicable laws, regulations, and platform safety guidelines.
Official Links
Official Website: patriotmemory.com
Viper Gaming: viper.patriotmemory.com
ACPI Technology: acpitechnology.com
Model Inquiries & Feedback: danton.chu hunter.wang oda.chang york.lin@acpitechnology.com
License & attribution
Apache-2.0. Built by Patriot Memory (patriotmemory.com).
- Downloads last month
- 719
