ATTENTION

Due to underperformance, this model is being put through background changes to improve performance. You can use this checkpoint, but know that it performs very badly compared to same class competitors.

UPDATE

The RMSNorm layers in the model have never updated due to a software bug causing underperformance. The fixed model will be published today.

Ivme-Conversate-XL-v1-Base

Conversate-XL-v1 Logo

Dense decoder-only transformer, 125.6M parameters, trained from scratch by IvmeLabs. Part of the Conversate family — see the IvmeLabs organization page for related models (Conversate-S, mainline Conversate, and this XL tier).

Architecture

  • 12 layers, hidden size 768, 12 attention heads (head_dim 64)
  • SwiGLU feed-forward, ffn_dim 3072
  • RoPE positional encoding (theta=10000.0)
  • RMSNorm (pre-norm), tied input/output embeddings, no bias terms
  • Vocabulary: 16000 tokens (BPE)
  • Max sequence length: 1024

Training

Trained on a 5.0B-token mix (backbone: DCLM-baseline, FineWeb-Edu, FineMath; supplement: Wikipedia-en, Project Gutenberg-en) using Muon (body weights) + AdamW (embeddings/norms), on a single AMD Instinct MI300X (ROCm 7.14.0, PyTorch 2.12.0).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "IvmeLabs/Ivme-Conversate-XL-v1-Base", trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("IvmeLabs/Ivme-Conversate-XL-v1-Base")

inputs = tokenizer("Hello, my name is", return_tensors="pt")
outputs = model.generate(inputs["input_ids"], max_new_tokens=50)
print(tokenizer.decode(outputs[0]))

Note: requires trust_remote_code=True since this uses a custom architecture (modeling_ivme.py in this repo), not a built-in transformers model class. Review that file before trusting it, as with any trust_remote_code=True model.

Checkpoint

This repo contains checkpoint(s) from step(s): 160, 320, 480, 640, 800, 960, 1120, 1280, 1440, 1600, 1760, 1920, 2080, 2240, 2400, 2560, 2720, 2880, 3040, 3200, 3318

Downloads last month
355
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support