YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

🩺 MediSight AI

Vision & Voice Assisted Medical AI (Educational Project)

MediSight AI is a multimodal AI medical assistant that analyzes user-uploaded images and spoken health concerns to generate concise, doctor-style responses.
The system integrates speech recognition, vision-language models, and text-to-speech to simulate a real-world medical consultation experience.

⚠️ This project is for educational and research purposes only and does not replace professional medical advice.


πŸš€ Features

  • 🎀 Voice-based medical query input
  • πŸ–ΌοΈ Image-based health analysis
  • 🧠 Multimodal LLM reasoning
  • πŸ”Š Spoken doctor-style responses
  • 🌐 Gradio web interface
  • ☁️ Deployable on Hugging Face Spaces

🧠 Tech Stack

  • Frontend: Gradio
  • Speech-to-Text: Whisper (via Groq)
  • Multimodal LLM: Meta LLaMA 4 (Groq)
  • Text-to-Speech: gTTS, ElevenLabs
  • Backend: Python
  • Deployment: Hugging Face Spaces

πŸ“ Project Structure

artificial-doctor/ β”‚ β”œβ”€β”€ app.py # Gradio application β”œβ”€β”€ Brain_of_doctor.py # Multimodal LLM (image + text) β”œβ”€β”€ voice_of_patient.py # Speech-to-text pipeline β”œβ”€β”€ voice_of_doctor.py # Text-to-speech pipeline β”‚ β”œβ”€β”€ requirements.txt β”œβ”€β”€ README.md └── .gitignore

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support