Gestalt: Large Multimodal Interplay Model
Zequn Yang†
Yu Miao†
Haotian Ni†
Ziheng Chen†
Chengxiang Huang†
Dongzhan Zhou
Kai Chen
Qi Zhang
Ji-Rong Wen
Yake Wei‡
Di Hu‡,✉
† Equal contribution ‡ Team leader ✉ Corresponding author
Gestalt is a new paradigm of large multimodal model built around multimodal interplay. Guided by a multimodal interplay pyramid — from modality-specific modeling, through cross-modal alignment, to multimodal synergy — Gestalt adopts a unified discrete diffusion framework with an interplay-partitioned architecture, where learnable interplay tokens mediate cross-modal exchange and integration. The name is inspired by Gestalt psychology: the whole is greater than the sum of its parts.
Model Description
Gestalt is an interplay-centric large multimodal model built on a unified discrete diffusion framework. Through multimodal pretraining, continual pretraining, and interplay-oriented supervised fine-tuning,Gestalt supports both understanding and generation.
| Property | Value |
|---|---|
| Architecture | GestaltModelLM (discrete diffusion with interplay tokens) |
| Parameters | ~8B |
| Hidden size | 4096 |
| Layers | 32 |
| Attention heads | 32 |
| Vocab size | 142,848 |
| Max sequence length | 4,096 |
| Precision | bfloat16 |
Download
# Gestalt model
huggingface-cli download GeWuLab/Gestalt --local-dir /path/to/gestalt
# IBQ vision tokenizer (required for encoding/decoding images)
huggingface-cli download TencentARC/IBQ-Tokenizer-16384 --local-dir /path/to/ibq
The IBQ tokenizer is from TencentARC/IBQ-Tokenizer-16384, used to convert between images and discrete visual tokens.
Training Pipeline
| Stage | Data | Objective |
|---|---|---|
| Multimodal Pretraining | 70M | Joint masked prediction |
| Continual Pretraining ← this checkpoint | 8M | Conditional masked prediction |
| Supervised Fine-Tuning | 13.7M (+≈2.5M T2I) | Four interplay categories, three-phase curriculum |
Citation
If you find Gestalt useful for your research, please cite:
@article{gestalt2026,
title = {Gestalt: Large Multimodal Interplay Model},
author = {Yang, Zequn and Miao, Yu and Ni, Haotian and Chen, Ziheng and
Huang, Chengxiang and Zhou, Dongzhan and Chen, Kai and Zhang, Qi and
Wen, Ji-Rong and Wei, Yake and Hu, Di},
year = {2026},
url = {https://github.com/GeWu-Lab/Gestalt}
}
License
This project is released under the Apache 2.0 license.
Author Contributions
Zequn Yang, Yake Wei, and Di Hu drove the overall advancement of the project. Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Chengxiang Huang contributed equally to this work. Zequn Yang conducted model pretraining and supervised fine-tuning. Zequn Yang, Yu Miao, and Haotian Ni developed the model architecture and conducted the core experiments. Yu Miao, Ziheng Chen, and Chengxiang Huang contributed to data processing and organization. Yu Miao conducted the evaluation of image generation capabilities, Haotian Ni conducted the multimodal understanding evaluation, and Ziheng Chen conducted the text-only evaluation. Dongzhan Zhou, Kai Chen, Qi Zhang, and Ji-Rong Wen contributed to discussions on the technical design and methodology. Yake Wei, Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, and Di Hu contributed to writing and revising the manuscript. Di Hu initiated the project. Yake Wei and Di Hu supervised and advised the project.
- Downloads last month
- 41