<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Data Lake on Shaowen Chen's Website</title><link>https://www.chenshaowen.com/en/tags/data-lake/</link><description>Recent content in Data Lake on Shaowen Chen's Website</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>&amp;copy;2016 - {year}, All Rights Reserved.</copyright><lastBuildDate>Thu, 12 Sep 2024 00:00:00 +0000</lastBuildDate><sy:updatePeriod>weekly</sy:updatePeriod><atom:link href="https://www.chenshaowen.com/en/tags/data-lake/atom.xml" rel="self" type="application/rss+xml"/><item><title>Processing Data on Kubernetes with Iceberg and Spark</title><link>https://www.chenshaowen.com/en/blog/use-iceberg-and-spark-on-kubernetes.html</link><pubDate>Thu, 12 Sep 2024 00:00:00 +0000</pubDate><atom:modified>Thu, 12 Sep 2024 00:00:00 +0000</atom:modified><guid>https://www.chenshaowen.com/en/blog/use-iceberg-and-spark-on-kubernetes.html</guid><description>1. Data Processing Architecture It is mainly divided into four layers: Processing capability layer: Spark on Kubernetes provides streaming data processing capability Data management layer: Iceberg provides dataset access operations such as ACID and tables Storage layer: Hive MetaStore manages Iceberg table metadata, PostgreSQL serves as the storage backend for</description><dc:creator>微信公众号</dc:creator><category>Spark</category><category>Iceberg</category><category>Kubernetes</category><category>Big Data</category><category>Data Lake</category><category>Operations</category><category>Data Processing</category><category>Learning</category></item></channel></rss>