YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
TRIDENT β 3-Head Tiny LM + P2P RAG + MCP
~1M parameter transformer with three specialist heads: Code, Math, Research. Runs on CPU (Termux/iPad/laptop). Knowledge retrieval via P2P WebRTC β no cloud needed.
Files
model.pyβ Trident architecture (backbone + 3 heads + RAG gate)rag.pyβ Python chunk store + P2P RAG node (aiortc)rag_client.jsβ Browser P2P RAG client (WebRTC DataChannel)signal_server.pyβ Minimal WebSocket signaling (handshake only, ~50 lines)mcp_server.pyβ FastMCP server exposing all tools to Claude/any LLMtrain.pyβ Training loop (toy data β real data)
Quick Start
# Install
pip install -r requirements.txt
# Test model architecture
python model.py
# Train on toy data (CPU, ~1 min)
python train.py
# Run MCP server (connect to Claude Desktop)
python mcp_server.py
# Run signaling server for P2P
python signal_server.py
P2P RAG Flow
Device A (your phone)
ββ TRIDENT_P2P.query("fibonacci")
ββ WebRTC DataChannel β Device B (iPad)
ββ Device B searches local chunks
ββ Returns top-K matches
ββ Device A merges local + peer results
ββ Top-K embeddings β RAGFusionGate β model generates
MCP Tools
| Tool | What it does |
|---|---|
trident_generate |
Generate text, pick head, optionally RAG |
trident_add_chunk |
Add knowledge to RAG store |
trident_search_rag |
Search chunks without generating |
trident_router |
Predict which head fits a query |
trident_list_chunks |
List all stored knowledge chunks |
Architecture
Input β [Embedding + PositionalEncoding]
β [Backbone: 4x TransformerBlock] β shared
β [HeadRouter] β softmax weights
β βββββββββββββββββββββββββββββββ
β RAGFusionGate (cross-attn) β β per head
β 1x TransformerBlock β
β LM Head β logits β
βββββββββββββββββββββββββββββββ Γ 3 heads
β weighted ensemble OR forced single head
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support