YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
yc-startup-rag-chatbot
π YC Startup RAG Chatbot
A Retrieval-Augmented Generation (RAG) chatbot trained on Y Combinatorβs βHow to Start a Startupβ lectures.
This project is an end-to-end RAG system that allows users to ask questions about YCβs Startup School lectures and get accurate, grounded answers with citations from the original transcript. It uses:
Python
Ollama (local LLM inference)
BGE-M3 embeddings
Streamlit (interactive chat UI)
Custom chunking + vector search
Local Retrieval-Augmented Generation pipeline
π§ Features
β Chunking & Embeddings
Lecture transcripts are chunked using a Recursive Character Text Splitter.
Embeddings generated using BGE-M3 via Ollama.
Stored efficiently in embeddings.joblib.
β RAG Pipeline
Retrieve top-K most relevant chunks using cosine similarity.
Construct structured prompts using retrieved YC lecture content.
Generate grounded answers using lightweight local LLMs:
llama3.2:1b
β Interactive Chat UI
Built with Streamlit, showing:
Chat messages
Retrieved lecture chunks (sources)
Clean and readable answers
Session chat history
β Local, Privacy-Friendly & Fast
Everything runs fully offline using Ollama on your machine.
π Project Structure
βββ app.py # Streamlit chat application βββ chunking.py # Splits transcripts into chunks βββ read_chunk.py # Generates embeddings + saves joblib βββ process_incoming.py # CLI-based RAG pipeline βββ requirements.txt βββ .gitignore βββ transcript/ # raw lecture transcripts (ignored) βββ json/ # chunk JSON files (ignored) βββ embeddings.joblib # embeddings (ignored) βββ videos/, audios/ # raw data (ignored)
π Getting Started
1οΈβ£ Clone the repository
git clone https://github.com//yc-startup-rag-chatbot.git cd yc-startup-rag-chatbot
2οΈβ£ Install dependencies
Create environment (optional):
conda create -n yc-rag python=3.10 -y conda activate yc-rag
Install packages:
pip install -r requirements.txt
3οΈβ£ Install and start Ollama
Download Ollama: https://ollama.com/download
Serve models:
ollama pull bge-m3 ollama pull llama3.2:1b # fastest
or ollama pull phi3 # best speed + quality
Start Ollama:
ollama serve
4οΈβ£ Prepare Data
Chunk transcripts
python chunking.py
Generate embeddings
python read_chunk.py
5οΈβ£ Run the Streamlit App
streamlit run app.py
Your browser will open the chatbot UI at:
π§ͺ Example Questions
Try asking:
"What does Paul Graham say about generating startup ideas?"
"How should founders think about growth?"
"What is the most important quality in a co-founder?"
The app will show:
The answer
The exact lecture chunks used as context
ποΈ RAG Architecture
User Query β Create Embedding (BGE-M3) β Vector Search (Cosine Similarity) β Retrieve Top-K Lecture Chunks β Build Structured Prompt β Local LLM (Llama3.2 / Phi3) β Grounded Answer + Sources
π§© Technologies Used
Python
Streamlit
Ollama
BGE-M3 embeddings
Numpy / Pandas
Scikit-Learn
Joblib
π§βπ» Author
Ashwani Jha RAG Developer | Machine Learning | LLMs