You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

ACE-3-E4B-Preview

APMIC-logo-橫-黑 NVIDIA-NeMo

Model Description

ACE-3-E4B-Preview is a preview release of the edge-sized member of APMIC's ACE-3 model family, intended for Traditional Chinese (Taiwan) enterprise scenarios and agentic workflows on resource-constrained hardware.

The model is based on google/gemma-4-e4b-it, a dense multimodal model with about 4.5B effective parameters (8B including Per-Layer Embeddings — the "E" in E4B stands for "effective"). It accepts text, image and audio input, generates text, and supports a 128K-token context window.


Model Details

  • Developed by: APMIC
  • Model type: Gemma4ForConditionalGeneration (Transformers), dense with Per-Layer Embeddings (~4.5B effective / 8B total)
  • Base model: google/gemma-4-e4b-it
  • Modalities: text, image and audio in; text out
  • Context length: 128K tokens
  • Language(s) (NLP): Traditional Chinese & English
  • Weights precision: bfloat16 (not quantized)
  • License: Apache 2.0 (inherited from Gemma 4)

Usage

messages = [{"role": "user", "content": "請用繁體中文簡單說明什麼是邊緣運算。"}]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,   # True to let the model produce a reasoning block first
)

Transformers

import torch
from transformers import AutoProcessor, AutoModelForImageTextToText

model_id = "APMIC/ACE-3-E4B-Preview"

processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

messages = [{"role": "user", "content": "請用繁體中文簡單說明什麼是邊緣運算。"}]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, enable_thinking=False,
    tokenize=True, return_dict=True, return_tensors="pt",
).to(model.device)

out = model.generate(**inputs, max_new_tokens=512)
print(processor.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Recommended sampling

The bundled generation_config.json uses temperature=1.0, top_p=0.95, top_k=64. These follow the base model's recommended settings.

Serving

The model can be served with any inference engine that supports Gemma 4 (for example vLLM). Because the weights are bfloat16, plan for roughly 16 GB of GPU memory for the weights alone, plus KV cache.


Intended Use and Limitations

Intended for

  • Traditional Chinese (Taiwan) assistants and enterprise applications on edge or single-GPU deployments
  • Function-calling and agent frameworks that need a small, locally deployable model
  • Research and evaluation of the ACE-3 model family

Limitations

  • This is a preview release; behavior may change in later ACE-3 versions.
  • Like all LLMs, the model can produce incorrect or fabricated content. For high-risk use (financial, legal, medical), keep a human review or enterprise control layer in place.
  • Tool-call outputs should be validated (schema and argument checks) before being executed.
  • Performance on languages other than Traditional Chinese and English has not been specifically evaluated.

License

This model is a derivative of Gemma 4 and is distributed under the Apache License 2.0, following the base model's Gemma 4 license.

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support