Instructions to use APMIC/ACE-3-E4B-Preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use APMIC/ACE-3-E4B-Preview with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("APMIC/ACE-3-E4B-Preview") model = AutoModelForMultimodalLM.from_pretrained("APMIC/ACE-3-E4B-Preview", device_map="auto") - Notebooks
- Google Colab
- Kaggle
ACE-3-E4B-Preview
Model Description
ACE-3-E4B-Preview is a preview release of the edge-sized member of APMIC's ACE-3 model family, intended for Traditional Chinese (Taiwan) enterprise scenarios and agentic workflows on resource-constrained hardware.
The model is based on google/gemma-4-e4b-it, a dense multimodal model with about 4.5B effective parameters (8B including Per-Layer Embeddings — the "E" in E4B stands for "effective"). It accepts text, image and audio input, generates text, and supports a 128K-token context window.
Model Details
- Developed by: APMIC
- Model type: Gemma4ForConditionalGeneration (Transformers), dense with Per-Layer Embeddings (~4.5B effective / 8B total)
- Base model: google/gemma-4-e4b-it
- Modalities: text, image and audio in; text out
- Context length: 128K tokens
- Language(s) (NLP): Traditional Chinese & English
- Weights precision: bfloat16 (not quantized)
- License: Apache 2.0 (inherited from Gemma 4)
Usage
messages = [{"role": "user", "content": "請用繁體中文簡單說明什麼是邊緣運算。"}]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False, # True to let the model produce a reasoning block first
)
Transformers
import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "APMIC/ACE-3-E4B-Preview"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto"
)
messages = [{"role": "user", "content": "請用繁體中文簡單說明什麼是邊緣運算。"}]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, enable_thinking=False,
tokenize=True, return_dict=True, return_tensors="pt",
).to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(processor.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Recommended sampling
The bundled generation_config.json uses temperature=1.0, top_p=0.95, top_k=64. These follow the base model's
recommended settings.
Serving
The model can be served with any inference engine that supports Gemma 4 (for example vLLM). Because the weights are bfloat16, plan for roughly 16 GB of GPU memory for the weights alone, plus KV cache.
Intended Use and Limitations
Intended for
- Traditional Chinese (Taiwan) assistants and enterprise applications on edge or single-GPU deployments
- Function-calling and agent frameworks that need a small, locally deployable model
- Research and evaluation of the ACE-3 model family
Limitations
- This is a preview release; behavior may change in later ACE-3 versions.
- Like all LLMs, the model can produce incorrect or fabricated content. For high-risk use (financial, legal, medical), keep a human review or enterprise control layer in place.
- Tool-call outputs should be validated (schema and argument checks) before being executed.
- Performance on languages other than Traditional Chinese and English has not been specifically evaluated.
License
This model is a derivative of Gemma 4 and is distributed under the Apache License 2.0, following the base model's Gemma 4 license.
- Downloads last month
- -

