Simulated layer-by-layer post-training-quantization error and run-level perplexity-degradation and quality-cliff labels for 70 model, method, and bit-width runs across RTN, GPTQ, AWQ, SmoothQuant, and QLoRA-NF4 on open-weight LLMs, grounded in the outlier-feature and bit-width scaling-law literature. Forward-looking task predicts the full-model quality cliff from only the first 25 percent of quantized layers.
This dataset simulates the layer-by-layer error introduced by five post-training weight-quantization methods, naive round-to-nearest, GPTQ, AWQ, SmoothQuant W8A8, and QLoRA NormalFloat4, applied to seven open-weight model architectures, Gemma 2B/7B, Gemma 2 9B/27B, Llama 3 8B/70B, and Mistral 7B, at bit-widths from 8 down to 3. Each layer's error combines standard uniform-quantization rounding noise with an outlier-feature range-crushing penalty modeled on the emergent-outlier phenomenon documented for models beyond roughly 6.7 billion parameters, a method-specific error-reduction factor reflecting each method's core mechanism, and a convex penalty below 4 bits matching the bit-width cliff finding in the k-bit scaling-laws literature. Per-run cumulative error is converted into an approximate relative perplexity-degradation percentage and a binary quality-cliff label at 5 percent. All formulas are documented and cited in quantization_theory.py and the README Research basis, ten sources: LLM.int8, GPTQ, AWQ, SmoothQuant, QLoRA, k-bit scaling laws, ZeroQuant, OmniQuant, AQLM, and the Gemma 2 technical report.
Lets Teams Self-hosting Open-weight Models Estimate, Before Running A Full Calibration And Evaluation Pass, Which Quantization Method And Bit-width Combination Is Likely To Cross A Quality Cliff, Using Only Early-layer Error Statistics As An Early-warning Signal Partway Through A Calibration Run.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.