Revisiting Residual Connections: Orthogonal Updates for Stable and Efficient Deep Networks Paper • 2505.11881 • Published May 17, 2025 • 5
Qwen 3.6 - Reg/Uncensored 9b, 12b, 21b, 27b, 40B Collection Fine tuned Qwen 3.6/3.8 models, including source and GGUF from 9B and up. 9B,12B, 21B and 40B are custom built by me. Tuning : Unsloth / COLD FUSION. • 22 items • Updated 4 days ago • 55
Heretic - Abliterated, Uncensored, Unrestricted POWER. Collection Models that have be abliterated using the HERETIC method. Done properly, this completely removed almost all censorship with no damage to the model. • 134 items • Updated 4 days ago • 152
200+ Roleplay, Creative Writing, Uncensored, NSFW models. Collection Oldest models listed first, with Newest models at bottom of the page. Most repos have full examples, instructions, best settings and so on. • 290 items • Updated 4 days ago • 1.03k
100 Coder/Programming - MOE, Reasoning, Reg, Imatrix, Fused. Collection Models (0.8B to 87B) in regular, "reasoning", "Brainstorm", MOE (1x to 8x / 128 experts), and expanded to create better and stronger code, faster. • 71 items • Updated 4 days ago • 60
SWE-bench Collection SWE-bench (Lite, Verified, Multimodal, Multilingual) all in one place! • 5 items • Updated Dec 14, 2025 • 16
view article Article Welcome Inkling by Thinking Machines +3 burtenshaw, merve, pcuenq, ariG23498, andito • Jul 15 • 166
view article Article Native-speed vLLM transformers modeling backend hmellor, lysandre • Jul 8 • 75
view article Article Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel nvidia • Jun 24 • 39
view article Article Ulysses Sequence Parallelism: Training with Million-Token Contexts kashif, stas • Mar 9 • 33
view article Article Welcome Gemma 4: Frontier multimodal intelligence on device +5 merve, pcuenq, sergiopaniego, burtenshaw, Steveeeeeeen, alvarobartt, SaylorTwift • Apr 2 • 928
view article Article Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers tomaarsen • Apr 16 • 81
view article Article Hugging Face and Cerebras bring Gemma 4 to real-time voice AI +2 A-Mahla, andito, lvwerra, vyassaurabh • Jul 1 • 101
view article Article Welcome Gemma 2 - Google’s new open LLM +4 philschmid, osanseviero, pcuenq, lewtun, tomaarsen, reach-vb • Jun 27, 2024 • 133
Nemotron-Cascade 2 Collection Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation • 4 items • Updated 25 days ago • 53
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch Paper • 2311.03099 • Published Nov 6, 2023 • 37
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time Paper • 2203.05482 • Published Mar 10, 2022 • 9