view article Article Bekko Embedding: how small can a multilingual retrieval model be? hotchpotch • 7 days ago • 5
view article Article Fine-tuning LLMs to 1.58bit: extreme quantization made easy +4 medmekk, marcsun13, lvwerra, pcuenq, osanseviero, thomwolf • Sep 18, 2024 • 282
view article Article Experimenting with the proposed Cross-Origin Storage API in Transformers.js tomayac • Jun 23 • 8
view article Article Profiling in PyTorch (Part 1): A Beginner's Guide to torch.profiler +3 ariG23498, sayakpaul, sergiopaniego, ror, pcuenq • May 29 • 155
view article Article Two Years of Local AI on a Laptop: When Open Models Outpaced Moore's Law mishig • May 11 • 25
view post Post 5407 Qwen3.6-27B is out now! Run it locally on 18GB RAM. 💜Qwen3.6-27B surpasses Qwen3.5-397B-A17B on all major coding benchmarks.GGUFs to run: unsloth/Qwen3.6-27B-GGUFGuide + MLX: https://unsloth.ai/docs/models/qwen3.6 See translation 🔥 29 29 + Reply
CohereLabs/cohere-transcribe-03-2026 Automatic Speech Recognition • 2B • Updated Jun 10 • 1.01M • • 1.07k
LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models Paper • 2310.08659 • Published Oct 12, 2023 • 30