Youssef Amrani
youssefam
AI & ML interests
Efficient LLM inference, KV cache optimization, quantization, speculative decoding, model pruning
Recent Activity
upvoted a paper less than a minute ago
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression liked a model about 20 hours ago
KVCache-ai/Qwen3-30BA3B-GGUF upvoted a paper about 20 hours ago
Fathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV CachesOrganizations
None yet