Search Articles — Sudonull

Search Results

In this project

Compression of GLM-5.1 on 16 GB VRAM: GQA hacks

https://sudonull.com/compression-of-glm-5-1-on-16-gb-vram-gqa-hacks

Breakdown of compressing 744B MoE model GLM-5.1 to 388 MB for T4. Fix GQA errors, config code, inference. For middle/senior dev — repeat on your hardware.

Text compression Brentwick-7 to 50 tokens

https://sudonull.com/text-compression-brentwick-7-to-50-tokens

Learn how Cambridge compresses texts to prompts with 98% accuracy. Brentwick-7 method for developers: latent reduction, embeddings, markets. Test it yourself.

TurboQuant: 6x compression of LLM KV-cache

https://sudonull.com/turboquant-6x-compression-of-llm-kv-cache

Break down Google's TurboQuant: PolarQuant + QJL for 6x KV-cache compression without accuracy loss. Inference acceleration up to 8x. Benchmarks, principle of operation for LLM developers.

LLM Hallucinations: Data Compression Artifacts

https://sudonull.com/llm-hallucinations-data-compression-artifacts

Breaking down LLM hallucinations through lossy compression: why they occur, how to minimize with RAG and fine-tuning. For developers: Shannon's theory and examples. Learn how to work with artifacts.

TurboQuant: lossless KV-cache compression for AI

https://sudonull.com/turboquant-lossless-kv-cache-compression-for-ai

Learn how Google's TurboQuant compresses transformer memory down to 3 bits with PolarQuant and QJL. Benchmarks on Gemma, Mistral. Optimization for AI developers.

From other projects

Rarity of heart cancer: mechanical force of contractions

https://ymaho.com/rarity-of-heart-cancer-mechanical-force-of-contractions

Science reveals centuries-old mystery: why the heart does not get cancer. Compression suppresses proliferation through Nesprin-2. Learn about a new class of drugs.

Inflation in Russia 5.47%: why this is not a reason for joy

https://ymaho.com/inflation-in-russia-5-47-why-this-is-not-a-reason-for-joy

Annual inflation slowed to 5.47%, but behind the statistics lies stagflation and demand compression. Find out what will happen to the Central Bank rate and the OFZ market. Risk analysis.

From the web

GitHub - 0xk1h0/ChatGPT_DAN: ChatGPT DAN, Jailbreaks prompt

https://github.com/0xk1h0/ChatGPT_DAN

ChatGPT DAN, Jailbreaks prompt. Contribute to 0xk1h0/ChatGPT_DAN development by creating an account on GitHub.

Download GitHub Desktop

https://desktop.github.com/download/

Simple collaboration from your desktop Download GitHub Desktop Focus on what matters instead of fighting with Git. Whether you're …

GitHub - basketikun/chatgpt2api: ChatGPT官网接口纯协议的逆向实 …

https://github.com/basketikun/chatgpt2api

Apr 19, 2026 · ChatGPT2API 主要是对 ChatGPT 官网相关能力进行逆向整理与封装,提供面向 ChatGPT 图片生成、图片编辑、多图 …

大家用 grill-me (拷问我) 这个 AI Agent Skill 的体验如何? - 知乎

https://www.zhihu.com/question/2054005413406946147

我经常用这个工具,也推荐给了很多人。 在有这个 skill 之前,我通常会在一段 instruction 的结尾要求 agent:“ 如果我的需求或设计有 …

Writing - Reddit

https://www.reddit.com/r/writing/

Your critique submission should be a top-level comment in the thread and should include: * Title * Genre * Word count * Type of …

如何看待 Codex 与ChatGPT 合并,实际使用体验有什么变化?

https://www.zhihu.com/question/2058849029027632241

Jul 10, 2026 · 到了ChatGPT Work、Codex和ChatGPT合并之后,几乎每天增长100万,曲线立即变得陡峭。 从用户增长策略来 …

United States Air Force Reddit

https://www.reddit.com/r/AirForce/

Community for current and past members of the US Air Force.

普通人要Codex有什么用? - 知乎

https://www.zhihu.com/question/2020702567198410280

这也算是一个能分享的小技巧。 毕竟如果一个人自己开了 20x ChatGPT Pro 订阅畅用,那至少这个人的钱包我觉得不是很普通,国内 …

GitHub - rasbt/LLMs-from-scratch: Implement a ChatGPT-like LLM in ...

https://github.com/rasbt/LLMs-from-scratch

This repository contains the code for developing, pretraining, and finetuning a GPT-like LLM and is the official code repository for the …

chatGPT在国内怎么使用? - 知乎

https://www.zhihu.com/question/1963658042986962975

ChatGPT是由OpenAI推出的一款AI聊天对话机器人,能够进行自然语言交互,帮助用户完成问答、写作、编程等多种任务。

Trending Now