<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Fault-Tolerant Training on Shaowen Chen's Website</title><link>https://www.chenshaowen.com/en/tags/fault-tolerant-training/</link><description>Recent content in Fault-Tolerant Training on Shaowen Chen's Website</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>&amp;copy;2016 - {year}, All Rights Reserved.</copyright><lastBuildDate>Sat, 17 Aug 2024 00:00:00 +0000</lastBuildDate><sy:updatePeriod>weekly</sy:updatePeriod><atom:link href="https://www.chenshaowen.com/en/tags/fault-tolerant-training/atom.xml" rel="self" type="application/rss+xml"/><item><title>Elastic, Fault-Tolerant Training with DLRover-Managed Jobs</title><link>https://www.chenshaowen.com/en/blog/use-dlrover-to-manage-training-job.html</link><pubDate>Sat, 17 Aug 2024 00:00:00 +0000</pubDate><atom:modified>Sat, 17 Aug 2024 00:00:00 +0000</atom:modified><guid>https://www.chenshaowen.com/en/blog/use-dlrover-to-manage-training-job.html</guid><description>1. Problems Facing Distributed Training Estimating training resources is difficult and cannot be automated How much compute, how much time, how much bandwidth, how many CPUs, how much memory — without enough accumulated experience it is hard to estimate accurately. The result is over-requesting and over-allocation, causing enormous resource waste.</description><dc:creator>微信公众号</dc:creator><category>DLRover</category><category>AI</category><category>Training</category><category>Kubernetes</category><category>Elastic Training</category><category>Fault-Tolerant Training</category><category>Operations</category></item></channel></rss>