Ollama_VLM πŸ€–πŸ§ πŸ“·

This notebook demonstrates a multimodal interaction system using:

  • 🎀 Voice input (SpeechRecognition)
  • πŸ“· Live image input (OpenCV)
  • 🧠 Vision-Language reasoning (LLaVA via Ollama)
  • πŸ”Š Spoken responses (pyttsx3)

How it works:

  1. Captures webcam image
  2. Listens for a question
  3. Sends both to LLaVA (Ollama) for processing
  4. Speaks the result

Perfect for building AI-powered robotic assistants!


Run locally with: pip install cv2 pyttsx3 speechrecognition ollama numpy

Ensure ollama and llava are properly set up on your machine.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support