AI & ML interests
None defined yet.
Recent Activity
Team Phoenix is the LLM AI R&D team at HTX, the Home Team Science and Technology Agency, Singapore. We build Phoenix, Singapore's sovereign large language model family for the Home Team and the wider public service.
This organisation is where we publish model cards, technical reports and selected datasets. Most of our work runs inside government, so what appears here is a subset of what we build.
Why a sovereign model family
High-impact sovereign and enterprise deployments require understanding of specific local legislation, policy, institutional terminology and operational context. Phoenix holds that knowledge in its parameters rather than retrieving it by web search, which is what makes secure, air-gapped deployment possible. Localised model efforts so far have concentrated on smaller scales and text-only modalities; Phoenix tests whether a frontier multimodal model can be deeply adapted to local context without giving up broad general capability.
Building Phoenix also develops sustained local expertise in multilingual and multimodal data curation, continued pre-training, and safety.
The Phoenix family
- Phoenix Medium โ 123B, continued pre-training. Our flagship sovereign multimodal model, adapted on 1 trillion tokens of Singapore and regional content while remaining globally competitive as a general-purpose model. Technical report
- Phoenix Small โ 24B, continued pre-training. The first Phoenix model, built for Singapore and the Home Team. Technical report
- Phoenix Nano โ 3B, pre-trained from scratch on a training dataset we built in house. How we built the training dataset
Datasets
Our domain corpus spans Singapore law, government, history and culture across English, Chinese, Malay and Tamil, with annotation done by native speakers.
Evaluation
We run open source benchmarks and build our own on top of them. Four so far:
- HT-Lexicon โ Home Team terminology and operational language
- SG-Gov โ Singapore government policy and procedure
- SG-Legal โ Singapore statutes and legal reasoning
- SG-Multimodal โ Singapore imagery, human annotated
Research
Alongside the product lines, we work on reinforcement learning for knowledge recall, multilingual transfer, multimodal continued pre-training, and high-quality human annotation.
