RT-J β€” int8 (Q8_0-style) quantized checkpoints

Int8 quantizations of stanford-star/rt-j, the Relational Transformer foundation model for in-context learning over relational databases. Produced and consumed by the RelativeDB native inference engine (cpp/rt_quantize --type q8). Weights stay quantized-resident at inference: the engine's CPU (Accelerate / portable SIMD) and Metal/MPS kernels dequantize inside the GEMM, so DRAM weight traffic is the int8 payload.

Sibling repos: rt-j-int4 Β· rt-j-fp16

File Task head Size
classification/model.q8.safetensors classification / ranking (logits β€” apply sigmoid) 86 MB
regression/model.q8.safetensors regression / forecasting (normalized values) 86 MB

(vs. 171 MB bf16 upstream, 342 MB fp32 in memory.)

Quantization format

Q8_0-style, plain safetensors β€” no custom container:

  • Every transformer-block projection (wq/wk/wv/wg, wo, ffn.w1/w2/w3 across all 12 blocks β€” ~99% of parameters) is stored as an I8 tensor with a per-output-row symmetric fp32 scale in a <name>.q_scale companion tensor: W[o,i] β‰ˆ q[o,i] * scale[o], scale[o] = max|W[o,:]| / 127.
  • The value/col-name encoders, decoder head, norms, biases and mask embeddings stay fp32 (like llama.cpp's embedding/output layers): input-side quantization error would propagate through all 12 blocks for a negligible size win.

Any safetensors reader can load these files; dequantization is one line per row. The RelativeDB engine keeps the int8 payload resident and dequantizes 64-row tiles (CPU) or in-register while staging GEMM tiles (Metal); the CUDA backend is fp32-only for now.

Accuracy

Measured on the RelativeDB golden batch (B=5, S=16 churn example) against the PyTorch reference, identical on CPU and Metal/MPS:

Metric fp32 int8
yhat max abs. error vs. torch 3.9e-3 1.1e-2
yhat mean abs. error 5.1e-4 3.5e-3
target-score sign / ranking preserved preserved

Worst per-row mean relative weight error: β‰ˆ5%.

Usage (RelativeDB native engine)

# direct path
./build/rt_test testdata classification/model.q8.safetensors --quantized --device mps

# via the Java / Python / Rust bindings: place the .q8 file next to the fp32
# checkpoint (or point at a directory containing it) and opt in with
export RELATIVEDB_RT_QUANTIZED=q8   # (or 1/true; q4 and f16 select siblings)

The C ABI (rt_model_load) accepts these files directly β€” the format is auto-detected from the tensor dtypes.

Reproduce

cmake -B build -S cpp && cmake --build build -j
./build/rt_quantize <rt-j>/classification/model.safetensors classification/model.q8.safetensors
./build/rt_quantize <rt-j>/regression/model.safetensors regression/model.q8.safetensors

License & attribution

Derivative of stanford-star/rt-j (Stanford STAR lab), redistributed under the same CC-BY-NC-SA-4.0 license. Architecture and training details are described in the upstream model card; only the weight storage format differs here.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for RelativeDB/rt-j-int8

Finetuned
(3)
this model

Datasets used to train RelativeDB/rt-j-int8