Ollama_VLM π€π§ π·
This notebook demonstrates a multimodal interaction system using:
- π€ Voice input (SpeechRecognition)
- π· Live image input (OpenCV)
- π§ Vision-Language reasoning (LLaVA via Ollama)
- π Spoken responses (pyttsx3)
How it works:
- Captures webcam image
- Listens for a question
- Sends both to LLaVA (Ollama) for processing
- Speaks the result
Perfect for building AI-powered robotic assistants!
Run locally with: pip install cv2 pyttsx3 speechrecognition ollama numpy
Ensure ollama and llava are properly set up on your machine.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support