Search Articles — Sudonull

Search Results

In this project

FLUX.2-dev GGUF Q4_K_M: 29 GB VRAM on Apple Silicon

https://sudonull.com/flux-2-dev-gguf-q4-k-m-29-gb-vram-on-apple-silicon

Breakdown of 29 GB VRAM consumption by FLUX.2-dev GGUF model on M3 Pro. MPS-overhead, dequantization, optimizations reduce peak by 1.5 GB. Instructions for developers.

Compression of GLM-5.1 on 16 GB VRAM: GQA hacks

https://sudonull.com/compression-of-glm-5-1-on-16-gb-vram-gqa-hacks

Breakdown of compressing 744B MoE model GLM-5.1 to 388 MB for T4. Fix GQA errors, config code, inference. For middle/senior dev — repeat on your hardware.

Console AI Agent "Botinok": Automation of Linux Servers

https://sudonull.com/console-ai-agent-botinok-automation-of-linux-servers

Learn about "Botinok" — a local console AI agent for Linux that automates administration tasks via SSH, using Ollama and minimum VRAM. Ideal for DevOps and system administrators.

ACE-Step 1.5: music neural network outperforms Suno locally

https://sudonull.com/ace-step-1-5-music-neural-network-outperforms-suno-locally

Open-source model ACE-Step 1.5 XL generates music locally, outperforms Suno in SongEval. LM+DiT architecture, from 4 GB VRAM. Installation, benchmarks, limitations. Test it yourself for IT projects.

Qwen 3.6 vs Gemma 4: testing on a local machine

https://sudonull.com/qwen-3-6-vs-gemma-4-testing-on-a-local-machine

Performance comparison of Qwen 3.6 and Gemma 4 on a laptop with RTX 4070. LM Studio setup, VRAM optimization, test results. Read the guide.

Trending Now