YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Vortex AI - Optimized AI Assistant for Low-Resource Hardware
Vortex AI is a lightweight, efficient AI assistant specifically designed to run on systems with limited resources (like 2GB GPU, 4-core CPU, 12GB RAM). It uses state-of-the-art optimization techniques to deliver powerful AI capabilities while maintaining efficiency.
Features
- Lightweight Design: Optimized for systems with limited resources
- Smart Model Selection: Automatically selects the best model for your hardware
- 4-bit Quantization: Reduces memory usage by up to 75%
- Web Search Integration: Can search the web for current information
- Memory Management: Intelligent memory management for 2GB GPUs
- Command-line Interface: Easy-to-use CLI for various tasks
- Interactive Mode: Chat-like experience for extended conversations
Installation
Prerequisites
- Python 3.8 or higher
- PyTorch compatible with your CUDA version (if using GPU)
- At least 8GB free disk space for models
Install Dependencies
pip install torch transformers accelerate bitsandbytes sentence-transformers
pip install GPUtil psutil requests
pip install peft # For LoRA adapters (optional)
Usage
Command Line Interface
Basic Usage
python vortex_cli.py -q "What is the capital of France?"
With Custom Parameters
python vortex_cli.py --model microsoft/Phi-3-mini-4k-instruct --temp 0.7 --max-tokens 512 "Explain quantum computing"
Interactive Mode
python vortex_cli.py --interactive
Get System Info
python vortex_cli.py --hardware-info
Python API Usage
from vortex_engine import VortexEngine
# Initialize Vortex
vortex = VortexEngine(
model_name="microsoft/Phi-3-mini-4k-instruct",
quantization="4bit" # Use 4-bit quantization to save memory
)
# Generate text
response = vortex.generate("Hello, how are you?")
print(response)
# Perform search and respond with current information
response = vortex.search_and_respond("What is the latest version of Python?")
print(response)
# Cleanup when done
vortex.cleanup()
Hardware Optimization
Vortex intelligently adapts to your hardware:
- For 2GB+ GPU: Uses Phi-3 Mini with 4-bit quantization
- For 4GB+ GPU: Uses Phi-3 Mini with 4-bit or 8-bit quantization
- For 8GB+ GPU: Can use larger models like Phi-3 Small
- For CPU-only: Falls back to efficient CPU-optimized models
Architecture
┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
│ User Input │───▶│ Vortex Engine │───▶│ Output/Result │
└─────────────────┘ │ │ └─────────────────┘
│ • Model Loading │
│ • Quantization │
│ • Generation │
│ • Memory Mgmt │
└─────────────────┘
│
┌──────────────────┐
│ Utilities │
│ • Model Selector │
│ • Memory Manager │
│ • Quantization │
└──────────────────┘
│
┌──────────────────┐
│ Search Engine │
│ • Web Search │
│ • Document Search│
└──────────────────┘
Performance Tips
- For Maximum Speed: Use lower max_length values
- For Better Quality: Use higher temperature (0.7-0.9) and top_p (0.9)
- For Memory Conservation: Use 4-bit quantization and lower max_length
- For Current Information: Use the search_and_respond method
Model Recommendations
Vortex AI uses these models based on hardware:
- microsoft/Phi-3-mini-4k-instruct (Recommended for 2GB+ GPUs)
- microsoft/Phi-3-small-8k-instruct (For 4GB+ GPUs)
- microsoft/Phi-3-medium-4k-instruct (For 8GB+ GPUs)
All models are optimized using 4-bit or 8-bit quantization to reduce memory usage.
Troubleshooting
Common Issues
Out of Memory Error:
- Reduce max_length parameter
- Ensure 4-bit quantization is enabled
- Close other GPU-intensive applications
Slow Performance:
- Check if model loaded properly
- Verify GPU utilization
- Consider using lower max_length for generation
Model Loading Issues:
- Ensure sufficient disk space
- Check internet connection for model download
- Verify PyTorch/CUDA compatibility
Memory Management
Vortex includes sophisticated memory management:
- Automatic GPU cache clearing
- CPU memory monitoring
- Efficient tensor handling
- Garbage collection optimization
Ollama Integration
Vortex AI can also be integrated with Ollama for an even more efficient experience on low-resource hardware!
Installing Vortex in Ollama
Make sure Ollama is installed and running:
# Install Ollama from https://ollama.com/ ollama serve # Run this in a separate terminalNavigate to the Vortex directory and run the integration script:
cd /home/astracat/vortex-ai ./ollama/install_vortex_ollama.shOr use the Python integration script:
cd /home/astracat/vortex-ai python ollama/ollama_integration.py
Using Vortex with Ollama
Once installed, you can use Vortex through Ollama:
# Interactive chat
ollama run vortex:latest
# Single query
ollama generate vortex:latest "What is the capital of France?"
# Get model information
ollama show vortex:latest
# Use with API
curl http://localhost:11434/api/generate -d '{
"model": "vortex:latest",
"prompt": "Hello!"
}'
License
This project is licensed under the MIT License - see the LICENSE file for details.
Contributing
We welcome contributions to Vortex AI! Feel free to submit issues or pull requests to improve:
- Model optimization
- Memory management
- Search capabilities
- Documentation
- Performance improvements
Acknowledgments
- Microsoft for Phi-3 models
- Hugging Face for Transformers library
- Ollama for the amazing local LLM platform
- The open-source AI community