YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Companion Forge L4 ONNX/TensorRT Runtime
Precompiled ONNX and TensorRT runtime for patdev/Companion-Forge, targeted at Hugging Face Jobs l4x1 (NVIDIA L4 / SM89 / 24 GB class VRAM).
Production coverage
| Component | ONNX | TensorRT L4 | Notes |
|---|---|---|---|
| FLUX.2 Klein 4B transformer | opset 23 | β | static 512 profile |
| DINOv2 ViT-L/14-reg | β | β | 4.83Γ vs ORT CUDA in validation |
| DSINE | β | β | 3.62Γ vs native PyTorch |
AniGen ss_flow_solo |
opset 23 | β | complete model; 1.28Γ per forward |
AniGen ss_dae decoder |
opset 23 | β | complete model; 2.02Γ |
AniGen slat_flow_auto transformer core |
opset 18 dynamic | β | 7.30Γ vs sparse native core |
| AniGen SLat sparse shell | β | native spconv | only ~188 MB shell weights; core lives in TRT |
| AniGen skin decoder | opset 18 dynamic | β | 6.80Γ; profile up to 300k vertices / 64 joints |
| AniGen SLat decoder sparse backbone | custom sparse ops | native spconv + FlashAttention | windowed sparse attention/topology changes are not standard dense ONNX ops |
| BiRefNet/rembg fallback | upstream ONNX | ORT GPU | used only for complex backgrounds |
The remaining native path is not another dense model: it is the sparse/topology operator layer (SparseConv3D, window-partitioned sparse attention, sparse up/downsample and subdivide). TensorRT Plugin V3 schemas are tracked under plugins/ for future replacement.
Final v6.4 L4 validation
Fast rigged Code-Pet generation completed end-to-end with:
- runtime:
hybrid-staged-dino-trt+dsine-trt+slat-shell-trt+skin-trt+ss-trt - worker time: 90.443 s including cold model/runtime preparation
- conditioning: 13.939 s
- runtime/model load: 12.997 s
- SS + SLat sampling: 7.895 s
- mesh post-process: 3.169 s
- PyTorch-tracked peak: 6.972 GiB
- real CUDA device peak: 12.734 GiB
- output: rigged
companion.glb+ separateskeleton.glb - no PyTorch SS/SLat flow-model fallback in the validated run
The earlier hybrid implementation took about 115.5 s and loaded full native AniGen flow checkpoints. v6.4 removes those flow checkpoints from the normal path.
Important artifacts
onnx/flux2-klein-4b/512/onnx/dinov2/onnx/dsine/onnx/anigen/ss-flow-solo/onnx/anigen/ss-decoder/onnx/anigen/slat-flow-core/onnx/anigen/skin-decoder/engines/l4-sm89/weights/anigen/slat-flow-shell/runtime/bench/plugins/
SLat split
The original slat_flow_auto checkpoint is about 2.46 GB. The production split keeps only the learned sparse IO shell (188 MB) and moves the 24 heavy transformer blocks into the SM89 TensorRT plan (1.71 GB). The full PyTorch checkpoint is downloaded only by the lazy fallback when an input exceeds the validated dynamic engine profile.
See runtime_manifest.json for the machine-readable runtime order, profiles and validation records.