News
Latest
Top
Search
Submit
Login
Search
▲
10
Integer Quantization: Deep Dive
(hello-fri-end.github.io)
by matt_d |
view
|
1 comments
▲
10
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
(arxiv.org)
by chrsw |
view
|
0 comments
▲
2
Faiss vs. Turbovec vs. Infino: Comparing 4-bit vector quantization
(infino.ai)
by ekechinwokah |
view
|
0 comments
▲
2
Quantization for Neural Networks
(leimao.github.io)
by eigenBasis |
view
|
0 comments
▲
2
Ask HN: Does treating Inflation as a "Quantization Snap" resolve slow-roll?
by aplowe |
view
|
0 comments
▲
1
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
by stared |
view
|
0 comments
▲
1
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
by stared |
view
|
0 comments
▲
1
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
by stared |
view
|
0 comments
▲
1
How to run Qwen3.8-27B on a single 16GB card: quantizations and Llama.cpp flags
(autodidacts.io)
by Curiositry |
view
|
0 comments
▲
1
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
by stared |
view
|
0 comments
▲
1
Quantization-Aware Healing:a compressed 4-bit model that outperf full-precision
(huggingface.co)
by grigio |
view
|
0 comments
▲
1
LLM Quantization Part 2: You're Gonna Need a Bigger VRAM
(lttlabs.com)
by rldjbpin |
view
|
0 comments
▲
1
LLM Quantization Part 2.5: Do Floats Dream of Real Numbers?
(lttlabs.com)
by rldjbpin |
view
|
0 comments
▲
1
LLM Quantization Part 3: Honey, I Shrunk the Numbers
(lttlabs.com)
by rldjbpin |
view
|
0 comments
▲
1
Compress and Forget: Bitsandbytes Quantization Amplifies Proactive Interference
(arxiv.org)
by sbulaev |
view
|
0 comments
▲
1
Hopf Spherical Compression and Quantization in C++23
(github.com)
by rfgplk |
view
|
1 comments
▲
1
Unsloth releases Dynamic v3 quantizations of Qwen3.8-27B
(huggingface.co)
by Curiositry |
view
|
0 comments
▲
1
GGUF Quantization Compared: Q4_K_M vs. IQ4_XS vs. IQ4_NL
(kaitchup.substack.com)
by peter_d_sherman |
view
|
0 comments
▲
1
DeepSeek V4 Pro at 207 tok/s with the full 1M context, no quantization
(runinfra.ai)
by OsamaJaber |
view
|
0 comments
▲
1
DeepSeek V4 Flash at 278 tok/s, full precision, no quantization
(runinfra.ai)
by OsamaJaber |
view
|
0 comments
▲
1
Muse Glimmer 30B unsloth GGUF format Q4, Q6, Q8 quantizations
(huggingface.co)
by walrus01 |
view
|
0 comments
▲
1
A Gentle Deep Dive into Quantization
(cuecloud.substack.com)
by krav-sales |
view
|
0 comments
▲
1
INT2 KV-cache quantization is getting surprisingly good
(arxiv.org)
by yunuyean |
view
|
0 comments
▲
1
A Visual Guide to Quantization – Demystifying the Compression of LLMs
(newsletter.maartengrootendorst.com)
by sebg |
view
|
0 comments
▲
1
Quantization hurts knowledge nonlinearly – Qwen3.6 27B case study
(quesma.com)
by stared |
view
|
0 comments
▲
1
Do Qwen 3.6 27B quantizations break the pelican?
(quesma.com)
by stared |
view
|
0 comments
▲
1
Show HN: Fast NF4 dequantization Triton kernel (1.41x faster than bitsandbytes
(github.com)
by Griffith-7 |
view
|
0 comments
▲
1
Scaling Law for Quantization-Aware Training
(arxiv.org)
by JumpCrisscross |
view
|
0 comments
▲
1
Negative squaring – pre-tilted 3-bit quantization beat naive 4-bit
(github.com)
by faDOOM |
view
|
0 comments
▲
1
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
(mixedbread.com)
by breadislove |
view
|
0 comments
▲
1
LLM Quantization Project Part 1: What Even Is an LLM?
(lttlabs.com)
by LabsLucas |
view
|
1 comments
▲
1
Mlx-optiq: per-layer mixed-precision LLM quantization for Apple Silicon
(mlx-optiq.com)
by codelion |
view
|
0 comments
▲
1
GGUF vs. GPTQ vs. AWQ: The Plain-English Guide to LLM Quantization
(vettedconsumer.com)
by ermantrout |
view
|
0 comments
▲
1
GGUF vs. GPTQ vs. AWQ: The Plain-English Guide to LLM Quantization
(vettedconsumer.com)
by ermantrout |
view
|
0 comments
▲
1
KVarN: Native vLLM KV-cache quantization back end by Huawei
(github.com)
by theanonymousone |
view
|
0 comments
▲
1
Spikes in LLMs Are Bias Vectors: Spike-Free Quantization
(arxiv.org)
by sbulaev |
view
|
0 comments
▲
1
Show HN: Glq LLM quantization using E8 lattice
(github.com)
by acd |
view
|
0 comments
▲
1
Emergent Quantization from a Dynamic Vacuum
(journals.aps.org)
by bookofjoe |
view
|
0 comments
▲
1
3.125-Bit LLM quantization bypassing tensor cores
(blog.djellalmohamedaniss.workers.dev)
by dmaniss |
view
|
0 comments
▲
1
Evaluation of Various MLX Quantizations
(github.com)
by d-_-b |
view
|
1 comments
▲
1
Scalar and Binary Quantization for Pgvector Vector Search and Storage (2024)
(jkatz05.com)
by eigenBasis |
view
|
0 comments
▲
1
Emergent Quantization from a Dynamic Vacuum
(journals.aps.org)
by davedx |
view
|
0 comments
▲
1
Quantization for Modern AI Systems (70-page free eBook)
(pawankjha.substack.com)
by pawanjha25 |
view
|
0 comments
▲
1
Advanced Quantization Algorithm for LLMs
(github.com)
by lastdong |
view
|
0 comments
▲
1
LLM Quantization
(huggingface.co)
by Anon84 |
view
|
0 comments
▲
1
A DuckDB extension for vector search indexes with pluggable quantization
(github.com)
by CheeseWang |
view
|
0 comments
▲
1
The Quantization Robustness of Diffusion Language Models in Coding Benchmarks
(arxiv.org)
by matt_d |
view
|
0 comments
▲
1
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
(arxiv.org)
by matt_d |
view
|
0 comments
▲
1
A guide to model quantization in fine-tuning (and how to pick the right GGUF)
(siquick.com)
by siquick |
view
|
0 comments
▲
1
Quantization, LoRA, and the 8% Problem Benchmarking Local LLMs for Production AI
(walsenburgtech.com)
by cowartc |
view
|
0 comments