YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
π©Ί MediSight AI
Vision & Voice Assisted Medical AI (Educational Project)
MediSight AI is a multimodal AI medical assistant that analyzes user-uploaded images and spoken health concerns to generate concise, doctor-style responses.
The system integrates speech recognition, vision-language models, and text-to-speech to simulate a real-world medical consultation experience.
β οΈ This project is for educational and research purposes only and does not replace professional medical advice.
π Features
- π€ Voice-based medical query input
- πΌοΈ Image-based health analysis
- π§ Multimodal LLM reasoning
- π Spoken doctor-style responses
- π Gradio web interface
- βοΈ Deployable on Hugging Face Spaces
π§ Tech Stack
- Frontend: Gradio
- Speech-to-Text: Whisper (via Groq)
- Multimodal LLM: Meta LLaMA 4 (Groq)
- Text-to-Speech: gTTS, ElevenLabs
- Backend: Python
- Deployment: Hugging Face Spaces
π Project Structure
artificial-doctor/ β βββ app.py # Gradio application βββ Brain_of_doctor.py # Multimodal LLM (image + text) βββ voice_of_patient.py # Speech-to-text pipeline βββ voice_of_doctor.py # Text-to-speech pipeline β βββ requirements.txt βββ README.md βββ .gitignore