# data lakehouse (blogs) — RSS Amplifier

Recent posts from the 3 feeds in the RSS Amplifier directory that cover data lakehouse.

Page: <https://rssamplifier.com/topics/data-lakehouse/blogs>  
Feed: <https://rssamplifier.com/topics/data-lakehouse/blogs.md>

---

## [AI Weekly: Four Frontier Models in Four Days](https://amdatalakehouse.substack.com/p/ai-weekly-four-frontier-models-in)

_2026-08-20 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

Week of August 11 to 18, 2026

## [Migrating to Apache Iceberg: Strategies for Every Source System](https://amdatalakehouse.substack.com/p/migrating-to-apache-iceberg-strategies)

_2026-08-19 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

This is Part 15, the final article of a 15-part Apache Iceberg Masterclass. Part 14 covered hands-on Dremio Cloud.

## [Hands-On with Apache Iceberg Using Dremio Cloud](https://amdatalakehouse.substack.com/p/hands-on-with-apache-iceberg-using)

_2026-08-18 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

This is Part 14 of a 15-part Apache Iceberg Masterclass. Part 13 covered streaming approaches.

## [The Plan and the Worker: Two Open Specifications for Agent Harnesses](https://amdatalakehouse.substack.com/p/the-plan-and-the-worker-two-open)

_2026-08-14 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

Every agent harness solves the same two problems, and almost every one of them solves both privately.

## [Apache Data Lakehouse Weekly: August 5 - August 12, 2026](https://amdatalakehouse.substack.com/p/apache-data-lakehouse-weekly-august)

_2026-08-14 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

This Week at a Glance

## [AI Weekly: GPT-5.6-Cyber, Muse Glimmer, and the Agent Browser](https://amdatalakehouse.substack.com/p/ai-weekly-gpt-56-cyber-muse-glimmer)

_2026-08-13 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

Week of August 5 to August 12, 2026

## [Approaches to Streaming Data into Apache Iceberg Tables](https://amdatalakehouse.substack.com/p/approaches-to-streaming-data-into)

_2026-08-11 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

This is Part 13 of a 15-part Apache Iceberg Masterclass. Part 12 covered Python and MPP engines.

## [Apache Arrow Flight and ADBC, and Why Database Connectivity Finally Went Columnar](https://amdatalakehouse.substack.com/p/apache-arrow-flight-and-adbc-and)

_2026-08-10 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

A data scientist runs a query against a warehouse.

## [Apache Data Lakehouse Weekly: July 29 to August 5, 2026](https://amdatalakehouse.substack.com/p/apache-data-lakehouse-weekly-july-7b2)

_2026-08-07 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

This was a week of decisions.

## [AI Weekly: The Week Frontier Pricing Broke](https://amdatalakehouse.substack.com/p/ai-weekly-the-week-frontier-pricing)

_2026-08-06 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

Two of the largest models ever released shipped inside five days of each other, and one of them costs a quarter of what the leader charges.

## [Auto-trade Cross-Exchange Gaps (Sponsored)](https://crawlproof.com/a/5kdokME1g0MF)

_2026-08-06 · **Sponsored**_

Find cross-exchange spreads, backtest strategies, and execute automatically with risk controls.

## [Apache Polaris 1.7.0 and the Quiet Work of Making a Catalog Trustworthy](https://amdatalakehouse.substack.com/p/apache-polaris-170-and-the-quiet)

_2026-08-04 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

A Spark job commits a table update.

## [Using Apache Iceberg with Python and MPP Query Engines](https://amdatalakehouse.substack.com/p/using-apache-iceberg-with-python)

_2026-08-04 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

This is Part 12 of a 15-part Apache Iceberg Masterclass. Part 11 covered metadata tables.

## [File Encryption for the Lakehouse: The Terminology, the Machinery, and the Hard Problem of Interoperable Encrypted Tables](https://amdatalakehouse.substack.com/p/file-encryption-for-the-lakehouse)

_2026-07-31 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

For years, the open lakehouse had an honest gap that practitioners whispered about and slide decks skipped: encryption.

## [AI Weekly: Opus 5 Lands, MCP Goes Stateless, and AMD Ships Helios](https://amdatalakehouse.substack.com/p/ai-weekly-opus-5-lands-mcp-goes-stateless)

_2026-07-31 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

Week of July 22 to July 29, 2026

## [Apache Data Lakehouse Weekly: July 21 to July 29, 2026](https://amdatalakehouse.substack.com/p/apache-data-lakehouse-weekly-july-9c1)

_2026-07-31 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

This was a week where the open lakehouse stack spent most of its energy on contracts.

## [A Deep Dive Into File Compression: How Data Gets Smaller, Why Codecs Differ, and What to Actually Use in the Lakehouse](https://amdatalakehouse.substack.com/p/a-deep-dive-into-file-compression)

_2026-07-30 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

Somewhere in your data platform right now, a single configuration property is quietly deciding a meaningful percentage of your storage bill, your query latency, and your compute spend.

## [A Reader's Guide to My Books: Which One to Pick Up, Depending on What You're Building](https://amdatalakehouse.substack.com/p/a-readers-guide-to-my-books-which)

_2026-07-29 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

The question I get most often after talks, after podcast episodes, and in newsletter replies is a simple one: where do I start with your books?

## [Apache Iceberg Metadata Tables: Querying the Internals](https://amdatalakehouse.substack.com/p/apache-iceberg-metadata-tables-querying)

_2026-07-28 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

This is Part 11 of a 15-part Apache Iceberg Masterclass. Part 10 covered maintenance operations.

## [Apache Iceberg v4: An Efficiency Rewrite of the Table Format](https://www.dremio.com/blog/apache-iceberg-v4-efficiency-rewrite/)

_2026-07-28 · will.martin@dremio.com · Dremio_

Iceberg 1.11.0, released in May 2026, contains the full implementation of the v3 spec(ification), including features like deletion vectors, the VARIANT type, and row lineage. But what comes next? V4, obviously, as four comes after 3. However, it's still early days on the v4 spec, so we are a ways off from the next version \[…\] The post Apache Iceberg v4: An Efficiency Rewrite of the Table Format…

## [The File Format Renaissance: Parquet, Lance, Vortex, Nimble, BtrBlocks, and the New Physics of Columnar Storage](https://amdatalakehouse.substack.com/p/the-file-format-renaissance-parquet)

_2026-07-27 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

For a decade, the file format layer was the most settled real estate in data.

## [Meet Claude, an AI assistant (Sponsored)](https://crawlproof.com/a/oK2Le3tD6jjN)

_2026-07-27 · **Sponsored**_

Chat with Claude at claude.ai — a conversational AI assistant.

## [Apache Data Lakehouse Weekly: July 16 to July 23, 2026](https://amdatalakehouse.substack.com/p/apache-data-lakehouse-weekly-july-2bb)

_2026-07-24 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

The lakehouse community spent this week deciding what belongs in the format and what belongs outside it.

## [Lakehouse Table Formats in 2026: Iceberg, Delta Lake, Hudi, Paimon, and DuckLake, How They Work, Where They Stand, and Where They're Going](https://amdatalakehouse.substack.com/p/lakehouse-table-formats-in-2026-iceberg)

_2026-07-24 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

The table format war is over, and the table formats are not.

## [AI Weekly: MCP Goes Stateless, AMD Ships 2nm Silicon](https://amdatalakehouse.substack.com/p/ai-weekly-mcp-goes-stateless-amd)

_2026-07-23 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

The plumbing of the AI industry got rebuilt this week.

## [The Filters We Build: How Every New Medium Rewires Our Defenses, From Radio Ads to AI Slop](https://amdatalakehouse.substack.com/p/the-filters-we-build-how-every-new)

_2026-07-23 · Alex Merced · Data, Lakehouse and AI with Alex Merced_

My grandparents’ generation learned to tune out the radio pitchman.

## [Maintaining Apache Iceberg Tables: Compaction, Expiry, and Cleanup](https://www.dremio.com/blog/maintaining-apache-iceberg-tables-compaction-expiry-and-cleanup/)

_2026-07-21 · Alex Merced · Dremio_

This is Part 10 of a 15-part Apache Iceberg Masterclass. Part 9 covered how tables degrade. This article covers the four maintenance operations that keep Iceberg tables healthy and the three approaches to running them. Table of Contents The Four Maintenance Operations 1. Compaction (File Rewriting) Compaction reads small files, merges them into optimally-sized files (128-512 MB), and \[…\] The post…

## [Apache Ossie (Incubating): The New Name for Open Semantic Interchange](https://www.dremio.com/blog/apache-ossie-incubating-the-new-name-for-open-semantic-interchange/)

_2026-07-13 · elise · Dremio_

Apache Ossie is currently undergoing incubation at The Apache Software Foundation (ASF). If you've been following the Open Semantic Interchange project — the open specification for semantic layer and ontology — there's an important update. The project has been accepted into the Apache Incubator under a new name: Apache Ossie (Incubating). The spec, the community, and \[…\] The post Apache Ossie…

## [How Data Lake Table Storage Degrades Over Time](https://www.dremio.com/blog/how-data-lake-table-storage-degrades-over-time/)

_2026-07-13 · Alex Merced · Dremio_

This is Part 9 of a 15-part Apache Iceberg Masterclass. Part 8 covered embedded catalogs. This article explains the five ways Iceberg table storage degrades and how to detect each problem before it impacts query performance. An Iceberg table that works well on day one will not work well on day 365 without maintenance. Every append, update, and \[…\] The post How Data Lake Table Storage Degrades Over…

## [What&#8217;s The Deal With Apache Parquet?](https://www.dremio.com/blog/apache-parquet-columnar-format-analytics/)

_2026-07-10 · will.martin@dremio.com · Dremio_

Apache Parquet is the recommended file format used in every modern data platform, and for good reason. But what are those reasons? And would it really matter if you stuck with CSV? The short answer is "YES". The slightly longer answer is "Yes, because columns". The full answer is below, so keep on reading to \[…\] The post What’s The Deal With Apache Parquet? appeared first on Dremio .

## [When Catalogs Are Embedded in Storage](https://www.dremio.com/blog/when-catalogs-are-embedded-in-storage/)

_2026-07-06 · Alex Merced · Dremio_

This is Part 8 of a 15-part Apache Iceberg Masterclass. Part 7 covered the traditional catalog landscape. This article examines a newer approach: embedding the catalog directly inside the storage layer. Traditional Iceberg architectures have three components: the query engine, a standalone catalog, and object storage. Embedded catalogs collapse the catalog into the storage layer itself, reducing…

## [The Semantic Layer: From Human Shortcut to Agent Guardrail](https://www.dremio.com/blog/modern-semantic-layer-ai-agents/)

_2026-07-02 · will.martin@dremio.com · Dremio_

For most of its history, the semantic layer was considered a solved problem. You built it once, business users queried it wherever it lived, and (hopefully) everyone would agreed on what "revenue" meant. However, much like the information in your data dictionary, the popularity of the semantic layer went stale and businesses turned to new, \[…\] The post The Semantic Layer: From Human Shortcut to…

## [Crush Every Threat in Real Time (Sponsored)](https://crawlproof.com/a/4EHJqM9MZsCC)

_2026-07-02 · **Sponsored**_

One agent for real-time detection, exposure tracking, and active defense.

## [Belgian Houses Fair Value](https://abdessettar.xyz/projects/belgian-houses-fair-value/)

_2026-05-14 · abdessettar · Abdessettar's Blog._

The webapp can be accessed through this link The code related to this article can be found in the following repository . Feel free to reach out for any questions or suggestions. I. Introduction I.A. The problem I.B. What we deliver I.C. Why we wrote this I.D. What this article is not II. Related Work III. The Data III.A. Sources III.B. Collection and integration III.C. Cleaning and the scraping…

## [Visualizing Insights from My Spotify Data Lakehouse](https://abdessettar.xyz/projects/spotify-gold-eda/)

_2026-04-20 · abdessettar · Abdessettar's Blog._

After building a fully automated data lakehouse on Azure to ingest, transform, and enrich my entire Spotify listening history , including new data continuously pulled via the API, I wanted to take a deeper dive into the actual insights buried within that pile of data. While my initial exploration started with a few ad-hoc charts in a Jupyter Notebook just to understand the raw schema, this post…

## [Designing and Building a Personal Spotify Data Lakehouse](https://abdessettar.xyz/projects/designing-building-personal-spotify-data-lakehouse/)

_2026-04-09 · abdessettar · Abdessettar's Blog._

The code related to this article can be found in the following repository . Feel free to reach out for any questions or suggestions. I. Introduction II. The Data Challenge III. Architecture and Implementation IV. Pipeline Design V. What's Next I. Introduction The Problem With Spotify Wrapped Every December, the internet loses its mind over Spotify Wrapped where people share their top artists,…

## [Building a Scalable, Privacy-Preserving Intelligent System for Legal Contracts](https://abdessettar.xyz/projects/private-ai-system-for-legal-contracts/)

_2026-02-21 · abdessettar · Abdessettar's Blog._

The code related to this article can be found in the following repository . Feel free to reach out for any questions or suggestions. The Context I. The Cost of Silence in Entreprise Data I.A. The Public AI Shortcut and Its Risks I.B. Why Off-the-Shelf AI Often Falls Short I.C. A Private-First Approach II. The Architecture II.A. The Interface II.B. The Worker II.C. The Inference Engine II.D. The…

## [Serverless at Scale: Building a Resilient Real Estate Scraper on AWS](https://abdessettar.xyz/projects/serverless-at-scale-real-estate-scraper/)

_2026-01-28 · abdessettar · Abdessettar's Blog._

The code related to this article can be found in the following repository . To understand how real estate market behaves over time, simply browsing listing websites is not enough. Relying solely on published statistics and market analysis is not sufficient either since these are released with a delay and therefore reflect past situation rather than current market climate. To accurately track…

