Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers nvidia • 4 days ago • 77
NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • 5 days ago • 53
Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 4 days ago • 53
Aether-7B-5Attn: A 100% Open-Source Sovereign Foundation Model — and a Controlled Experiment in Heterogeneous Attention FINAL-Bench • 2 days ago • 21
One Adapter, Both Modalities: Field Notes from Building and Serving a Multimodal Reranker lightonai • 5 days ago • 15
Be Ready Before the Attack: A Practical Guide to Self-Hosting an Open Model for Cyber Defense jeffboudier • 1 day ago • 5
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and Lessons sherryxychen • Sep 30, 2025 • 76
A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyond karina-zadorozhny • Jan 19 • 39
Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers nvidia • 4 days ago • 77
NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • 5 days ago • 53
Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 4 days ago • 53
Aether-7B-5Attn: A 100% Open-Source Sovereign Foundation Model — and a Controlled Experiment in Heterogeneous Attention FINAL-Bench • 2 days ago • 21
One Adapter, Both Modalities: Field Notes from Building and Serving a Multimodal Reranker lightonai • 5 days ago • 15
Be Ready Before the Attack: A Practical Guide to Self-Hosting an Open Model for Cyber Defense jeffboudier • 1 day ago • 5
How I Trained Action Chunking Transformer (ACT) on SO-101: My Journey, Gotchas, and Lessons sherryxychen • Sep 30, 2025 • 76
A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyond karina-zadorozhny • Jan 19 • 39