YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
OmniModalLLM
OmniModalLLM is a versatile and powerful multimodal language model designed to handle both text and image inputs, enabling sophisticated conversational AI applications similar to ChatGPT. Leveraging advanced architectures like Mixture of Experts (MoE) and Vector Quantized Variational Autoencoders (VQVAE), OmniModalLLM offers robust performance and adaptability across various tasks.
Features
- Multimodal Support: Handles both text and image data seamlessly.
- Mixture of Experts (MoE): Enhances model performance by leveraging multiple specialized experts.
- Vector Quantized VAE (VQVAE): Efficiently tokenizes and reconstructs images.
- Dynamic Configuration: Adjusts model components dynamically based on input data.
- Chat API: Provides a ChatGPT-like conversational interface using FastAPI.
- Memory Optimizations: Implements techniques like gradient checkpointing and mixed precision training to prevent CUDA Out-Of-Memory (OOM) errors.
- Rate Limiting: Protects the API from abuse using
slowapi. - Dynamic Response Generation Loop: Enables the assistant to generate responses of arbitrary length by iteratively predicting tokens until an end-of-sequence token is encountered or a maximum token limit is reached.
Table of Contents
Installation
Prerequisites
- Python 3.8 or higher
- Git
- CUDA-enabled GPU (optional, for training and inference acceleration)
Clone the Repository
git clone https://github.com/kirill670/OmniModalLLM.git
cd OmniModalLLM
Create a Virtual Environment
It's recommended to use a virtual environment to manage dependencies.
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
Install Dependencies
pip install --upgrade pip
pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu116
pip install transformers datasets pillow fastapi uvicorn tiktoken einops tensorboard faiss-cpu slowapi
Note: If you're using a TPU or CPU, adjust the PyTorch installation accordingly.
Usage
Training
OmniModalLLM is pre-configured to train on the Flickr30k and DailyDialog datasets. Ensure you have sufficient computational resources before initiating training.
python training_script.py
This command will:
- Load and preprocess the Flickr30k and DailyDialog datasets.
- Initialize the OmniModalLLM model and tokenizer.
- Start the training loop with mixed precision and gradient checkpointing.
- Save model checkpoints upon improvement.
- Launch the FastAPI server concurrently.
Training Parameters:
- Epochs: 5
- Batch Size: Adjusted based on available device (GPU, TPU, CPU)
- Learning Rate: 1e-4
- Early Stopping: Triggered after 3 epochs without improvement
API Deployment
The FastAPI server provides a /chat/ endpoint for interactive conversations. Once the training completes, the API server will be accessible at http://0.0.0.0:8000/chat/.
Running the API Server Independently
If you wish to run the API server without training, ensure the model is trained and load the saved checkpoint.
python api_server.py
This command will:
- Load the trained OmniModalLLM model and tokenizer.
- Start the FastAPI server listening on
http://0.0.0.0:8000. - Expose the
/chat/endpoint for interactive chat.
API Endpoints
POST /chat/
Generates a response based on the provided chat messages.
Request Body:
{
"session_id": "optional-session-id",
"messages": [
{
"role": "user",
"content": "Hello, how are you?"
}
]
}
session_id(optional): Unique identifier to maintain conversation context across multiple requests. If not provided, a new session will be created.messages: List of message objects containing the role (user,assistant, orsystem) and the content.
Response:
{
"session_id": "unique-session-id",
"message": {
"role": "assistant",
"content": "I'm a model designed to assist you. How can I help today?"
}
}
Examples
Using curl
Initial Request (Start a New Session):
curl -X POST "http://localhost:8000/chat/" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "user", "content": "Hello, how are you?"}
]
}'
Response:
{
"session_id": "generated-session-id",
"message": {
"role": "assistant",
"content": "I'm doing well, thank you! How can I assist you today?"
}
}
Subsequent Request (Continue the Conversation):
curl -X POST "http://localhost:8000/chat/" \
-H "Content-Type: application/json" \
-d '{
"session_id": "existing-session-id",
"messages": [
{"role": "user", "content": "Can you tell me a joke?"}
]
}'
Response:
{
"session_id": "existing-session-id",
"message": {
"role": "assistant",
"content": "Sure! Why did the computer show up at work late? It had a hard drive!"
}
}
Using Postman
Create a New POST Request:
- URL:
http://localhost:8000/chat/
- URL:
Set Headers:
Content-Type:application/json
Set Body:
- Choose
rawandJSONformat. - Input the JSON payload as shown in the
curlexamples.
- Choose
Send the Request:
- Observe the assistant's reply in the response section.
Creating a Simple Frontend
For a more interactive experience, consider creating a simple frontend using frameworks like React, Vue, or even plain HTML/CSS/JavaScript. This frontend can interact with the FastAPI backend via the /chat/ endpoint, allowing users to engage in conversations with the assistant through a web interface.
Contributing
Contributions are welcome! Please follow these steps:
- Fork the repository.
- Create a new branch (
git checkout -b feature/YourFeature). - Commit your changes (
git commit -m 'Add some feature'). - Push to the branch (
git push origin feature/YourFeature). - Open a Pull Request.
Please ensure your code adheres to the project's coding standards and includes appropriate tests.
License
This project is licensed under the MIT License.
Russian Version
OmniModalLLM
OmniModalLLM — это универсальная и мощная мультимодальная языковая модель, разработанная для обработки как текстовых, так и изображенческих данных. Она позволяет создавать сложные приложения для разговорного ИИ, аналогичные ChatGPT. Используя передовые архитектуры, такие как Mixture of Experts (MoE) и Vector Quantized Variational Autoencoders (VQVAE), OmniModalLLM обеспечивает высокую производительность и адаптивность для различных задач.
Особенности
- Мультимодальная поддержка: Бесшовно обрабатывает как текстовые, так и изображенческие данные.
- Mixture of Experts (MoE): Улучшает производительность модели за счет использования нескольких специализированных экспертов.
- Vector Quantized VAE (VQVAE): Эффективно токенизирует и восстанавливает изображения.
- Динамическая конфигурация: Настраивает компоненты модели динамически на основе входных данных.
- Чат API: Предоставляет интерфейс для разговоров, похожий на ChatGPT, используя FastAPI.
- Оптимизация памяти: Реализует такие методы, как градиентный чекпоинтинг и обучение с смешанной точностью, чтобы избежать ошибок Out-Of-Memory (OOM) на CUDA.
- Ограничение скорости запросов: Защищает API от злоупотреблений с помощью
slowapi. - Динамический цикл генерации ответов: Позволяет ассистенту генерировать ответы произвольной длины, итеративно предсказывая токены до появления токена конца последовательности или достижения максимального лимита токенов.
Содержание
Установка
Предварительные требования
- Python 3.8 или выше
- Git
- CUDA-совместимый GPU (опционально, для ускорения обучения и инференса)
Клонирование репозитория
git clone https://github.com/kirill670/OmniModalLLM.git
cd OmniModalLLM
Создание виртуального окружения
Рекомендуется использовать виртуальное окружение для управления зависимостями.
python -m venv venv
source venv/bin/activate # В Windows: venv\Scripts\activate
Установка зависимостей
pip install --upgrade pip
pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu116
pip install transformers datasets pillow fastapi uvicorn tiktoken einops tensorboard faiss-cpu slowapi
Примечание: Если вы используете TPU или CPU, настройте установку PyTorch соответствующим образом.
Использование
Обучение
OmniModalLLM преднастроена для обучения на датасетах Flickr30k и DailyDialog. Убедитесь, что у вас есть достаточные вычислительные ресурсы перед началом обучения.
python training_script.py
Эта команда выполнит следующие действия:
- Загрузит и предварительно обработает датасеты Flickr30k и DailyDialog.
- Инициализирует модель OmniModalLLM и токенизатор.
- Запустит цикл обучения с использованием смешанной точности и градиентного чекпоинтинга.
- Сохранит чекпоинты модели при улучшении результатов.
- Одновременно запустит сервер FastAPI.
Параметры обучения:
- Эпохи: 5
- Размер батча: Настраивается в зависимости от доступного устройства (GPU, TPU, CPU)
- Скорость обучения: 1e-4
- Ранняя остановка: Активируется после 3 эпох без улучшения
Развертывание API
Сервер FastAPI предоставляет эндпоинт /chat/ для интерактивных разговоров. После завершения обучения API-сервер будет доступен по адресу http://0.0.0.0:8000/chat/.
Запуск API-сервера отдельно
Если вы хотите запустить API-сервер без обучения, убедитесь, что модель обучена и загрузите сохраненный чекпоинт.
python api_server.py
Эта команда выполнит следующие действия:
- Загрузит обученную модель OmniModalLLM и токенизатор.
- Запустит сервер FastAPI, слушающий на
http://0.0.0.0:8000. - Предоставит эндпоинт
/chat/для интерактивного чата.
API Эндпоинты
POST /chat/
Генерирует ответ на основе предоставленных сообщений чата.
Тело запроса:
{
"session_id": "optional-session-id",
"messages": [
{
"role": "user",
"content": "Привет, как дела?"
}
]
}
session_id(опционально): Уникальный идентификатор для поддержания контекста разговора между несколькими запросами. Если не предоставлен, будет создана новая сессия.messages: Список объектов сообщений, содержащих роль (user,assistantилиsystem) и содержимое.
Ответ:
{
"session_id": "unique-session-id",
"message": {
"role": "assistant",
"content": "Я модель, созданная для помощи вам. Чем могу помочь сегодня?"
}
}
Примеры
Использование curl
Начальный запрос (Создание новой сессии):
curl -X POST "http://localhost:8000/chat/" \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "user", "content": "Привет, как дела?"}
]
}'
Ответ:
{
"session_id": "generated-session-id",
"message": {
"role": "assistant",
"content": "Я делаю хорошо, спасибо! Чем могу помочь вам сегодня?"
}
}
Последующий запрос (Продолжение разговора):
curl -X POST "http://localhost:8000/chat/" \
-H "Content-Type: application/json" \
-d '{
"session_id": "existing-session-id",
"messages": [
{"role": "user", "content": "Расскажи анекдот."}
]
}'
Ответ:
{
"session_id": "existing-session-id",
"message": {
"role": "assistant",
"content": "Конечно! Почему компьютер опоздал на работу? Потому что у него был жесткий диск!"
}
}
Использование Postman
Создайте новый POST-запрос:
- URL:
http://localhost:8000/chat/
- URL:
Установите заголовки:
Content-Type:application/json
Установите тело запроса:
- Выберите
rawи форматJSON. - Введите JSON-пейлоад, как показано в примерах
curl.
- Выберите
Отправьте запрос:
- Наблюдайте ответ ассистента в разделе ответа.
Создание простого фронтенда
Для более интерактивного опыта рассмотрите возможность создания простого фронтенда с использованием таких фреймворков, как React, Vue или даже простого HTML/CSS/JavaScript. Этот фронтенд может взаимодействовать с бэкендом FastAPI через эндпоинт /chat/, позволяя пользователям вести диалог с ассистентом через веб-интерфейс.
Вклад
Вклады приветствуются! Пожалуйста, следуйте этим шагам:
- Форкните репозиторий.
- Создайте новую ветку (
git checkout -b feature/YourFeature). - Закоммитьте свои изменения (
git commit -m 'Добавить некоторую функцию'). - Запушьте в ветку (
git push origin feature/YourFeature). - Откройте Pull Request.
Пожалуйста, убедитесь, что ваш код соответствует стандартам кодирования проекта и включает соответствующие тесты.
Лицензия
Этот проект лицензирован под MIT License.
Additional Information
For further assistance, questions, or suggestions, please feel free to open an issue on the GitHub repository.
Summary of Critical Additions and Corrections:
Dynamic Response Generation Loop in API Server:
- Purpose: Allows the assistant to generate responses of arbitrary length by iteratively predicting the next token until an end-of-sequence token is generated or a maximum number of tokens is reached.
- Implementation: Added a loop in the
generate_response_apifunction within theapi_server.pyscript that handles token generation, temperature scaling, top-k and top-p sampling, and termination conditions.
Updated Features Section:
- Included the Dynamic Response Generation Loop to highlight the model's capability to generate extensive and coherent responses.
Usage Instructions:
- Enhanced the API Deployment section to explain the independent running of the API server and its functionalities.
- Updated the Examples section to reflect the ability to handle extended conversations and generate longer responses.
Concurrency and Thread Safety:
- Ensured that the conversation history management in the API server is thread-safe using
threading.Lockto prevent race conditions during concurrent access.
- Ensured that the conversation history management in the API server is thread-safe using
Rate Limiting:
- Implemented rate limiting using
slowapito protect the API from abuse, with configurable request limits.
- Implemented rate limiting using
Error Handling Enhancements:
- Added checks to handle scenarios where user messages are not provided, ensuring the assistant responds appropriately.
Device Compatibility:
- Ensured that all tensors are correctly moved to the designated device (
CPU,GPU, orTPU) to prevent device mismatch errors during training and inference.
- Ensured that all tensors are correctly moved to the designated device (
Documentation Improvements:
- Provided detailed instructions for setting up, training, and deploying the model.
- Included examples for using both
curland Postman to interact with the API. - Suggested creating a frontend for enhanced user interaction.
By incorporating these updates, the OmniModalLLM project now offers a more robust and flexible framework for developing advanced multimodal conversational AI applications.