Inside three.ws: The Open-Source Stack That Gives AI Agents a Body, a Brain, a Wallet, and a Job three-ws • 6 days ago • 44
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States KangLiao • 2 days ago • 18
Evaluating LLMs Under Production Parity: A Replay Pipeline for Safe Model Swapping in Conversational Agents TechforHumans • 3 days ago • 12
Internet-Scale Knowledge Retrieval: A Novel Vector Search Dataset at 10B Scale Qdrant • 2 days ago • 12
Cosmos-Reason2 on a Quest 3: Performance, Cost, and What Compression Actually Buys You ErenAta00 • 1 day ago • 7
Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC ibm-granite • 10 days ago • 32
DeepSeek-R1 Dissection: Understanding PPO & GRPO Without Any Prior Reinforcement Learning Knowledge NormalUhr • Feb 7, 2025 • 299
A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyond karina-zadorozhny • Jan 19 • 45
Inside three.ws: The Open-Source Stack That Gives AI Agents a Body, a Brain, a Wallet, and a Job three-ws • 6 days ago • 44
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States KangLiao • 2 days ago • 18
Evaluating LLMs Under Production Parity: A Replay Pipeline for Safe Model Swapping in Conversational Agents TechforHumans • 3 days ago • 12
Internet-Scale Knowledge Retrieval: A Novel Vector Search Dataset at 10B Scale Qdrant • 2 days ago • 12
Cosmos-Reason2 on a Quest 3: Performance, Cost, and What Compression Actually Buys You ErenAta00 • 1 day ago • 7
Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC ibm-granite • 10 days ago • 32
DeepSeek-R1 Dissection: Understanding PPO & GRPO Without Any Prior Reinforcement Learning Knowledge NormalUhr • Feb 7, 2025 • 299
A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyond karina-zadorozhny • Jan 19 • 45