<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Model Quantization on Shaowen Chen's Website</title><link>https://www.chenshaowen.com/en/tags/model-quantization/</link><description>Recent content in Model Quantization on Shaowen Chen's Website</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>&amp;copy;2016 - {year}, All Rights Reserved.</copyright><lastBuildDate>Sat, 06 Sep 2025 00:00:00 +0000</lastBuildDate><sy:updatePeriod>weekly</sy:updatePeriod><atom:link href="https://www.chenshaowen.com/en/tags/model-quantization/atom.xml" rel="self" type="application/rss+xml"/><item><title>What Is Model Quantization</title><link>https://www.chenshaowen.com/en/blog/what-is-model-quantization.html</link><pubDate>Sat, 06 Sep 2025 00:00:00 +0000</pubDate><atom:modified>Sat, 06 Sep 2025 00:00:00 +0000</atom:modified><guid>https://www.chenshaowen.com/en/blog/what-is-model-quantization.html</guid><description>1. What Is Model Quantization Model quantization is the process of converting the weights and activations of a high-precision model (usually 32-bit floating point FP32 or 16-bit floating point FP16) into a low-precision model (such as 8-bit integer INT8).
The value range of FP32 is -3.4*10^38 to 3.4*10^38, with 4 billion values.</description><dc:creator>微信公众号</dc:creator><category>AI</category><category>LLM</category><category>Model Quantization</category><category>Optimization</category><category>Operations</category></item></channel></rss>