Latest Articles: page 15

692 posts from Alex's blogs and newsletters, newest first. Thoughts on tech, data, policy, and philosophy.

Alex Merced's Blog

Approaches to Streaming Data into Apache Iceberg Tables

Stream data into Iceberg with Spark Structured Streaming, Flink, or Kafka Connect. Here is how each works and the trade-offs between latency and maint...

Alex Merced's Blog

Using Apache Iceberg with Python and MPP Query Engines

Access Iceberg tables from Python with PyIceberg, DuckDB, and Polars, or through MPP engines like Dremio, Spark, and Trino. Here is how each approach ...

Alex Merced's Blog

Apache Iceberg Metadata Tables: Querying the Internals

Iceberg metadata tables let you query snapshots, files, manifests, and partitions using SQL. Here is every metadata table and how to use them....

Alex Merced's Blog

Maintaining Apache Iceberg Tables: Compaction, Expiry, and Cleanup

Keep Iceberg tables fast with compaction, snapshot expiry, orphan cleanup, and manifest rewriting. Here is when and how to run each operation....

Alex Merced's Blog

How Data Lake Table Storage Degrades Over Time

Iceberg tables degrade through small files, orphan files, metadata bloat, sort order decay, and partition skew. Here is how to diagnose each problem....

Alex Merced's Blog

When Catalogs Are Embedded in Storage

S3 Tables and MinIO AI Stor embed the Iceberg catalog directly in the storage layer. Here is when embedded catalogs make sense and when they do not....

Alex Merced's Blog

What Are Lakehouse Catalogs? The Role of Catalogs in Apache Iceberg

Lakehouse catalogs store metadata pointers, manage namespaces, and enforce access control. Here is the complete catalog landscape from Polaris to Glue...

Alex Merced's Blog

Writing to an Apache Iceberg Table: How Commits and ACID Actually Work

Here is exactly how an engine writes to an Iceberg table, step by step, from data files through the atomic commit that makes ACID guarantees possible....

Alex Merced's Blog

Hidden Partitioning: How Iceberg Eliminates Accidental Full Table Scans

Iceberg's hidden partitioning separates physical layout from user queries using transform functions....

Alex Merced's Blog

Partition Evolution: Change Your Partitioning Without Rewriting Data

Iceberg lets you change partition schemes without rewriting data. Here is how partition evolution works internally and why Hive-style partitioning cou...

Alex Merced's Blog

Performance and Apache Iceberg's Metadata

Iceberg's three-layer metadata tree eliminates directory listing and enables multi-level data skipping. Here is how scan planning actually works....

Alex Merced's Blog

The Metadata Structure of Modern Table Formats

Iceberg uses a metadata tree, Delta Lake uses a transaction log, Hudi uses a timeline. Here is exactly how each format organizes metadata and why it m...

Alex Merced's Blog

What Are Table Formats and Why Were They Needed?

Table formats like Apache Iceberg solved the ACID, schema, and performance problems that turned data lakes into data swamps. Here is how each one work...

Alex Merced's Blog

2025 Year in Review Apache Iceberg, Polaris, Parquet, and Arrow

A look back at key developments in Apache Iceberg, Polaris, Parquet, and Arrow in 2025....

Alex Merced's Blog

dremioframe & iceberg - Pythonic interfaces for Dremio and Apache Iceberg

Discover DremioFrame and IceFrame, two new Python libraries that simplify working with Dremio and Apache Iceberg. Learn how these tools streamline dat...

Alex Merced's Blog

Introducing dremioframe - A Pythonic DataFrame Interface for Dremio

Discover dremioframe, a new Python library that offers a DataFrame-like experience for interacting with Dremio's data lakehouse platform. Learn how to...

Alex Merced's Blog

Comprehensive Hands-on Walk Through of Dremio Cloud Next Gen (Hands-on with Free Trial)

Walkthrough with the new trial of the Dremio Cloud Platform...

Alex Merced's Blog

2025-2026 Guide to Learning about Apache Iceberg, Data Lakehouse & Agentic AI

A curated guide to mastering Apache Iceberg, data lakehouse architectures, and the emerging field of Agentic AI for data professionals....

Alex Merced's Blog

An Exploration of the Commercial Iceberg Catalog Ecosystem

Dive into the world of commercial Iceberg catalogs and discover how they enhance data lakehouse architectures for modern data engineering....

Alex Merced's Blog

Building a Universal Lakehouse Catalog - Beyond Iceberg Tables

Exploring paths to a universal lakehouse catalog that supports multiple data formats and engines, building on Apache Iceberg's success....

Alex Merced's Blog

Intro to Apache Iceberg with Apache Polaris and Apache Spark

Learn how to leverage Apache Iceberg with Apache Polaris and Apache Spark to build scalable and efficient data lakehouses....

Alex Merced's Blog

The State of Apache Iceberg v4 - October 2025 Edition

What's Coming in Apache Iceberg v4: A Deep Dive into the Future of Open Table Formats...

Alex Merced's Blog

The Ultimate Guide to Open Table Formats - Iceberg, Delta Lake, Hudi, Paimon, and DuckLake

Understanding Iceberg, Delta Lake, Hudi, Paimon, and DuckLake...

Alex Merced's Blog

The 2025 & 2026 Ultimate Guide to the Data Lakehouse and the Data Lakehouse Ecosystem

What is the Data Lakehouse and the Data Lakehouse Ecosystem? This comprehensive guide covers everything you need to know about the Data Lakehouse arch...