YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Vortex AI - Optimized AI Assistant for Low-Resource Hardware

Vortex AI is a lightweight, efficient AI assistant specifically designed to run on systems with limited resources (like 2GB GPU, 4-core CPU, 12GB RAM). It uses state-of-the-art optimization techniques to deliver powerful AI capabilities while maintaining efficiency.

Features

  • Lightweight Design: Optimized for systems with limited resources
  • Smart Model Selection: Automatically selects the best model for your hardware
  • 4-bit Quantization: Reduces memory usage by up to 75%
  • Web Search Integration: Can search the web for current information
  • Memory Management: Intelligent memory management for 2GB GPUs
  • Command-line Interface: Easy-to-use CLI for various tasks
  • Interactive Mode: Chat-like experience for extended conversations

Installation

Prerequisites

  • Python 3.8 or higher
  • PyTorch compatible with your CUDA version (if using GPU)
  • At least 8GB free disk space for models

Install Dependencies

pip install torch transformers accelerate bitsandbytes sentence-transformers
pip install GPUtil psutil requests
pip install peft  # For LoRA adapters (optional)

Usage

Command Line Interface

Basic Usage

python vortex_cli.py -q "What is the capital of France?"

With Custom Parameters

python vortex_cli.py --model microsoft/Phi-3-mini-4k-instruct --temp 0.7 --max-tokens 512 "Explain quantum computing"

Interactive Mode

python vortex_cli.py --interactive

Get System Info

python vortex_cli.py --hardware-info

Python API Usage

from vortex_engine import VortexEngine

# Initialize Vortex
vortex = VortexEngine(
    model_name="microsoft/Phi-3-mini-4k-instruct",
    quantization="4bit"  # Use 4-bit quantization to save memory
)

# Generate text
response = vortex.generate("Hello, how are you?")
print(response)

# Perform search and respond with current information
response = vortex.search_and_respond("What is the latest version of Python?")
print(response)

# Cleanup when done
vortex.cleanup()

Hardware Optimization

Vortex intelligently adapts to your hardware:

  • For 2GB+ GPU: Uses Phi-3 Mini with 4-bit quantization
  • For 4GB+ GPU: Uses Phi-3 Mini with 4-bit or 8-bit quantization
  • For 8GB+ GPU: Can use larger models like Phi-3 Small
  • For CPU-only: Falls back to efficient CPU-optimized models

Architecture

┌─────────────────┐    ┌──────────────────┐    ┌─────────────────┐
│   User Input    │───▶│ Vortex Engine    │───▶│  Output/Result  │
└─────────────────┘    │                 │    └─────────────────┘
                       │ • Model Loading │
                       │ • Quantization  │
                       │ • Generation    │
                       │ • Memory Mgmt   │
                       └─────────────────┘
                                │
                       ┌──────────────────┐
                       │  Utilities       │
                       │ • Model Selector │
                       │ • Memory Manager │
                       │ • Quantization   │
                       └──────────────────┘
                                │
                       ┌──────────────────┐
                       │  Search Engine   │
                       │ • Web Search     │
                       │ • Document Search│
                       └──────────────────┘

Performance Tips

  1. For Maximum Speed: Use lower max_length values
  2. For Better Quality: Use higher temperature (0.7-0.9) and top_p (0.9)
  3. For Memory Conservation: Use 4-bit quantization and lower max_length
  4. For Current Information: Use the search_and_respond method

Model Recommendations

Vortex AI uses these models based on hardware:

  • microsoft/Phi-3-mini-4k-instruct (Recommended for 2GB+ GPUs)
  • microsoft/Phi-3-small-8k-instruct (For 4GB+ GPUs)
  • microsoft/Phi-3-medium-4k-instruct (For 8GB+ GPUs)

All models are optimized using 4-bit or 8-bit quantization to reduce memory usage.

Troubleshooting

Common Issues

  1. Out of Memory Error:

    • Reduce max_length parameter
    • Ensure 4-bit quantization is enabled
    • Close other GPU-intensive applications
  2. Slow Performance:

    • Check if model loaded properly
    • Verify GPU utilization
    • Consider using lower max_length for generation
  3. Model Loading Issues:

    • Ensure sufficient disk space
    • Check internet connection for model download
    • Verify PyTorch/CUDA compatibility

Memory Management

Vortex includes sophisticated memory management:

  • Automatic GPU cache clearing
  • CPU memory monitoring
  • Efficient tensor handling
  • Garbage collection optimization

Ollama Integration

Vortex AI can also be integrated with Ollama for an even more efficient experience on low-resource hardware!

Installing Vortex in Ollama

  1. Make sure Ollama is installed and running:

    # Install Ollama from https://ollama.com/
    ollama serve  # Run this in a separate terminal
    
  2. Navigate to the Vortex directory and run the integration script:

    cd /home/astracat/vortex-ai
    ./ollama/install_vortex_ollama.sh
    
  3. Or use the Python integration script:

    cd /home/astracat/vortex-ai
    python ollama/ollama_integration.py
    

Using Vortex with Ollama

Once installed, you can use Vortex through Ollama:

# Interactive chat
ollama run vortex:latest

# Single query
ollama generate vortex:latest "What is the capital of France?"

# Get model information
ollama show vortex:latest

# Use with API
curl http://localhost:11434/api/generate -d '{
  "model": "vortex:latest",
  "prompt": "Hello!"
}'

License

This project is licensed under the MIT License - see the LICENSE file for details.

Contributing

We welcome contributions to Vortex AI! Feel free to submit issues or pull requests to improve:

  • Model optimization
  • Memory management
  • Search capabilities
  • Documentation
  • Performance improvements

Acknowledgments

  • Microsoft for Phi-3 models
  • Hugging Face for Transformers library
  • Ollama for the amazing local LLM platform
  • The open-source AI community
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support