Quantization as cache amplification: expert offloading, sub-2-bit codecs and cache policy for MoE inference on a single laptop.
π€ Open to Collab
Kavinkumar M
mkvn
Β·
AI & ML interests
Researching and building systems in Large Language Models, Multimodal AI, Vision Language Models, LLM Quantization, Efficient Inference, RAG, Document Intelligence, Information Extraction, AI Agents, and Applied Generative AI. Interested in improving the efficiency, accuracy, and practical deployment of large AI models.
Recent Activity
updated a collection 3 days ago
MoE inference on commodity hardware updated a collection 3 days ago
MoE inference on commodity hardware updated a collection 3 days ago
MoE inference on commodity hardware