A family of large language models optimized for on-device inference, available in 1B, 270M and 1M parameter sizes