<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Disaggregation on Shaowen Chen's Website</title><link>https://www.chenshaowen.com/en/tags/disaggregation/</link><description>Recent content in Disaggregation on Shaowen Chen's Website</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>&amp;copy;2016 - {year}, All Rights Reserved.</copyright><lastBuildDate>Sat, 17 Jan 2026 00:00:00 +0000</lastBuildDate><sy:updatePeriod>weekly</sy:updatePeriod><atom:link href="https://www.chenshaowen.com/en/tags/disaggregation/atom.xml" rel="self" type="application/rss+xml"/><item><title>Alibaba Cloud eRDMA Testing and PD Disaggregation Application Deployment</title><link>https://www.chenshaowen.com/en/blog/test-and-deploy-pd-disagg-app-with-erdma-on-aliyun.html</link><pubDate>Sat, 17 Jan 2026 00:00:00 +0000</pubDate><atom:modified>Sat, 17 Jan 2026 00:00:00 +0000</atom:modified><guid>https://www.chenshaowen.com/en/blog/test-and-deploy-pd-disagg-app-with-erdma-on-aliyun.html</guid><description>In a PD disaggregated deployment, heterogeneous GPU models are often used to deploy the model across machines, which multiplies the cross-machine communication pressure. RDMA devices are usually brought in to accelerate kvcache transfer between nodes so as to achieve a lower FTTL. This post describes how to test eRDMA devices</description><dc:creator>微信公众号</dc:creator><category>Alibaba Cloud</category><category>eRDMA</category><category>RDMA</category><category>PD</category><category>Disaggregation</category><category>AI</category><category>Application</category><category>Operations</category></item><item><title>Deploying PD-Disaggregated Applications with vLLM</title><link>https://www.chenshaowen.com/en/blog/using-vllm-to-deploy-pd-disagg-app.html</link><pubDate>Sat, 20 Sep 2025 00:00:00 +0000</pubDate><atom:modified>Sat, 20 Sep 2025 00:00:00 +0000</atom:modified><guid>https://www.chenshaowen.com/en/blog/using-vllm-to-deploy-pd-disagg-app.html</guid><description>1. Why Deploy LLM Applications with PD Disaggregation In the process of LLM inference, there are two serial stages: Process the entire input context and generate the KV Cache (Prefill stage) Incrementally generate new tokens (Decode stage) These two stages have different resource requirements. The Prefill stage has to compute</description><dc:creator>微信公众号</dc:creator><category>vLLM</category><category>Deployment</category><category>PD</category><category>Disaggregation</category><category>Application</category></item></channel></rss>