Apache Iceberg

Apache Iceberg is the open table format Alex writes about more than anything else.

His books on it include Apache Iceberg: The Definitive Guide from O’Reilly and Architecting an Apache Iceberg Lakehouse from Manning.

Below are his articles, videos and podcast episodes that cover Iceberg, newest first, pulled from his feeds.

Articles

Alex Merced's Blog

Keeping Audit Snapshots Alive While Iceberg Snapshot Expiration Runs Every Night

How Iceberg snapshot tags keep audit snapshots alive through nightly expiration: retention calendars, RETAIN semantics, compaction cost, and erasure c...

Alex Merced on Medium

Apache Data Lakehouse Weekly: September 15 to 23, 2026

Release managers ran the show this week, and license files kept tripping them. Iceberg 1.12.0 went to a second release candidate after… Continue readi...

Alex Merced on Medium

Fast Classification Models, LLMs, and the Apache Iceberg Lakehouse

Open Data Lakehouses for Everyone

Fast Classification Models, LLMs, and the Apache Iceberg Lakehouse

Open the bill for any team that has put a large language model into a data pipeline and look at what the calls are doing....

Alex Merced's Blog

Why the Iceberg DataFusion Integration Is Moving to Apache DataFusion

Why the Iceberg DataFusion integration moved to the DataFusion project, and what the split means for users, Comet, and iceberg-rust contributors....

Alex Merced's Blog

What Iceberg v4's Proposed FILE Type Means for Multimodal Tables

Iceberg v4's proposed FILE type brings first-class media references to tables, via Parquet's FILE logical type, ranges, checksums, and pre-signed URLs...

Alex Merced's Blog

Fast Classification Models, LLMs, and the Apache Iceberg Lakehouse

How fast classification models like Jev alongside open alternatives such as GLiClass compare with LLMs, and how to run both together inside an Apache ...

Alex Merced's Blog

CVE-2026-73334 and the Trust Boundary Inside an Encrypted Parquet File

CVE-2026-73334 lets a tampered Parquet footer route a reader's KMS token to an attacker. Here's the fix, Iceberg's safe path, and how to audit your la...

Alex Merced's Blog

Parquet Page Indexes and the Last Mile of Pruning in Apache Iceberg

Parquet page indexes can cut selective Iceberg scans by an order of magnitude on sorted data. How they work, what they cost, and how to lay out tables...

Alex Merced's Blog

How Apache Polaris Plans to Share Iceberg Tables Across Organizations

Apache Polaris's Open Sharing proposal adds first-class shares, external consumers, and listings so any Iceberg REST engine can read shared tables....

Alex Merced's Blog

Vector Search Directly Over Iceberg Tables, and When You Still Need a Vector Database

Embeddings are just columns. When exact vector search over Iceberg scans beats a vector database, when it does not, and how to lay out tables....

Alex Merced on Medium

Apache Data Lakehouse Weekly: September 9 to 17, 2026

Alex Merced on Medium

Apache Data Lakehouse Weekly: September 3–9, 2026

Alex Merced on Medium

The Open Lakehouse Explained, Then Built on Your Laptop with Dremio and MinIO

Alex Merced's Blog

Migrating Into Iceberg Without Moving Data

add_files, snapshot, and migrate compared: the three in-place paths into Iceberg, the reconciliation each requires, the layout traps, and the rollback...

Alex Merced's Blog

What Iceberg Table Maintenance Actually Costs

A cost model for compaction, snapshot expiry, orphan cleanup, and manifest rewriting: what each operation spends, on which meter, and how to set a sch...

Alex Merced's Blog

Running an Iceberg Lakehouse on Kubernetes

Catalog, maintenance, and compaction as Kubernetes workloads: scheduling classes, job structure, credential flow, and the failures that come from the ...

Alex Merced's Blog

Partition Statistics Files in Apache Iceberg

The underused Iceberg metadata for planning: what the partition statistics file holds, what the spec guarantees, how to write one, and when it earns i...

Alex Merced's Blog

What to Assert When You Test an Iceberg Pipeline

Fixtures, in-memory catalogs, and golden metadata: the assertions that catch wrong rows, unsafe reruns, schema drift, and concurrent-write corruption ...

Alex Merced's Blog

The 2026 Iceberg REST Catalog Compatibility Report

A repeatable test for what an Iceberg REST catalog actually serves, a scoring scheme that separates design from breakage, and the 2026 evidence across...

Alex Merced's Blog

Serving Iceberg Tables From Two Regions

Three multi-region topologies that work and one that mostly does not, what an Iceberg commit costs across regions, and where the catalog has to live....

Alex Merced's Blog

How Iceberg Catalogs Hand Engines Storage Access

Credential vending end to end: the wire protocol, scoped access on each cloud, remote signing, credential lifetime on long jobs, and failures that loo...

Alex Merced's Blog

Kafka Connect to Iceberg: How the Commit Actually Works

Exactly-once semantics in the Iceberg sink connector: the coordinator, the control topic, offsets stored inside Iceberg snapshots, and where duplicate...

Alex Merced's Blog

The Open Lakehouse Explained, Then Built on Your Laptop with Dremio and MinIO

The five layers of the open lakehouse explained, then a lab: Parquet, Iceberg, Polaris, Arrow, and Ossie running in two containers on your own machine...

Alex Merced on Medium

Inside the Puffin File Format

Alex Merced's Blog

dbt on Iceberg: Incremental Models on Open Tables

How dbt incremental materializations map to Iceberg operations, and the configuration, predicates, and maintenance that keep them healthy....

Alex Merced's Blog

Disaster Recovery for Iceberg Tables: Replication, Backup, and Restore

Disaster recovery for Iceberg across four tiers: snapshots, object versioning, catalog backup, and cross-region replication....

Alex Merced's Blog

Deleting User Data From an Immutable Lakehouse: GDPR Hard Deletes on Iceberg

How to turn a logical delete on immutable Iceberg into a physical erasure across snapshots, versions, replicas, and downstream copies....

Alex Merced's Blog

Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet

How Iceberg v3 geometry and geography types, bounding boxes, and native Parquet types give spatial data first-class standing....

Alex Merced's Blog

Default Column Values and Field IDs: How Iceberg Schema Evolution Works at the Spec Level

How field IDs and initial and write defaults let Iceberg change schemas on large tables without rewriting data, at the spec level....

Alex Merced's Blog

The Iceberg Table Properties That Actually Matter

The Iceberg table properties that decide file count, pruning, write amplification, retention, and metadata growth, by workload....

Alex Merced's Blog

The Lakehouse Ingestion Tool Landscape: Fivetran, Airbyte, dlt, and CDC vs Batch

How Fivetran, Airbyte, dlt, and CDC and streaming tools land well-behaved Apache Iceberg tables, and how to choose and maintain them....

Alex Merced's Blog

Local Iceberg Development Environments: Docker, MinIO, and In-Memory Catalogs for CI

Local Iceberg development environments: in-process catalogs, a Docker Compose stack with MinIO, and CI configurations that run either....

Alex Merced's Blog

Moving Iceberg Tables Between Catalogs Without Rewriting Data

Why moving Iceberg tables between catalogs is a pointer copy, and the protocol that makes a cutover safe for one table or thousands....

Alex Merced's Blog

Logs, Traces, and Metrics as Tables: Building an OpenTelemetry Data Lake on Iceberg

Building an OpenTelemetry data lake on Iceberg: schemas for spans, logs, and metrics, ingestion, query patterns, and retention....

Alex Merced's Blog

Postgres Meets the Lakehouse: pg_lake, pg_duckdb, and When Postgres Is Enough

What pg_lake, pg_duckdb, and pg_mooncake do at the Iceberg level, and honest thresholds for when Postgres is enough....

Alex Merced's Blog

Schema Registries and Event Schemas: Avro, Protobuf, and JSON Schema on the Way Into the Lakehouse

How Avro, Protobuf, and JSON Schema evolve through a registry, and how that maps to the schema evolution rules of Iceberg....

Alex Merced on Medium

Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet

Open Data Lakehouses for Everyone

Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet

A logistics team stores 40 million delivery stops in an Apache Iceberg table....

Alex Merced's Blog

Agent-Driven Storage Tiering for Apache Iceberg: Moving Cold Data Without Breaking Queries

A background agent can move cold Iceberg partitions to cheaper tiers without breaking live queries. Heatmaps, path-safe moves, and restore paths....

Alex Merced's Blog

DataFusion Comet 1.0 and What Native Rust Scans Change for Spark on Iceberg

DataFusion Comet 1.0 replaces Spark Iceberg scans with native Rust. What speeds up, what still falls back to the JVM, and how to deploy it....

Alex Merced's Blog

FSST and ALP: The Two Encodings Fixing Parquet's Weakest Compression Cases

ALP and FSST target Parquet's worst cases: floats and high-cardinality strings. How they work and what they change for Iceberg tables....

Alex Merced's Blog

High-Throughput Branch Merging: Automating Concurrency and Conflict Resolution in Multi-Branch Iceberg Pipelines

High-throughput Iceberg branch merges need conflict detection and automation. How to reconcile concurrent writes without stalling pipelines....

Alex Merced's Blog

Multi-Cloud REST Catalog Topologies: Running Apache Polaris Across AWS, Azure, and GCP

Polaris can catalog Iceberg tables across AWS, Azure, and GCP. Four topologies, credential vending, and the tradeoffs of each design....

Alex Merced's Blog

Parquet-Only Manifests in Iceberg v4: Why the Metadata Layer Is Going Columnar

Iceberg v4 is moving manifests from Avro to Parquet so planners can read only the stats they need. Why the metadata layer is going columnar....

Alex Merced's Blog

Semantic Layer Federation: One Logical Model Over Data on Three Clouds

One logical model over Iceberg and databases on three clouds. Pushdown, egress, Reflections, and where semantic federation still breaks....

Alex Merced's Blog

Serverless Iceberg Ingestion with PyIceberg and DuckDB: Micro-Batches Without a Spark Cluster

Land small Iceberg micro-batches with PyIceberg and DuckDB in a serverless function. Commits, concurrency, and why Spark is the wrong default....

Alex Merced's Blog

Zero-Copy Warehouse Modernization: Migrating Legacy Databases to Apache Iceberg Without Downtime

Move a legacy warehouse to Iceberg without downtime by virtualizing first. Consumer cutover, parity checks, and background copy without double-ETL....

Podcast episodes

Alex Merced Tech Podcast
Alex Merced Tech Podcast

AI Updates, Lakehouse Streaming (Iceberg/Fluss), and more!

Episode Details
Alex Merced Tech Podcast
Alex Merced Tech Podcast

2025 Reflections, Google Antigravity, NotebookLM, Dremio AI Agent, Pangolin Catalog, Dremioframe & Iceframe Python Libraries for Apache Iceberg

Episode Details
DevRel & Developer Advocacy
DevRel & Developer Advocacy

2025 Reflections, Google Antigravity, NotebookLM, Dremio AI Agent, Pangolin Catalog, Dremioframe & Iceframe Python Libraries for Apache Iceberg

Episode Details
Alex Merced Tech Podcast
Alex Merced Tech Podcast

Lakehouse Catalogs Beyond Apache Iceberg, What could they Look Like?

Episode Details
Alex Merced Tech Podcast
Alex Merced Tech Podcast

Understanding the role of MCP, Langchain and Agent2Agent

Episode Details

Books