Apache Iceberg
Apache Iceberg is the open table format Alex writes about more than anything else.
His books on it include Apache Iceberg: The Definitive Guide from O’Reilly and Architecting an Apache Iceberg Lakehouse from Manning.
Below are his articles, videos and podcast episodes that cover Iceberg, newest first, pulled from his feeds.
Articles
Keeping Audit Snapshots Alive While Iceberg Snapshot Expiration Runs Every Night
How Iceberg snapshot tags keep audit snapshots alive through nightly expiration: retention calendars, RETAIN semantics, compaction cost, and erasure c...
Apache Data Lakehouse Weekly: September 15 to 23, 2026
Release managers ran the show this week, and license files kept tripping them. Iceberg 1.12.0 went to a second release candidate after… Continue readi...
Fast Classification Models, LLMs, and the Apache Iceberg Lakehouse
Fast Classification Models, LLMs, and the Apache Iceberg Lakehouse
Open the bill for any team that has put a large language model into a data pipeline and look at what the calls are doing....
Why the Iceberg DataFusion Integration Is Moving to Apache DataFusion
Why the Iceberg DataFusion integration moved to the DataFusion project, and what the split means for users, Comet, and iceberg-rust contributors....
What Iceberg v4's Proposed FILE Type Means for Multimodal Tables
Iceberg v4's proposed FILE type brings first-class media references to tables, via Parquet's FILE logical type, ranges, checksums, and pre-signed URLs...
Fast Classification Models, LLMs, and the Apache Iceberg Lakehouse
How fast classification models like Jev alongside open alternatives such as GLiClass compare with LLMs, and how to run both together inside an Apache ...
CVE-2026-73334 and the Trust Boundary Inside an Encrypted Parquet File
CVE-2026-73334 lets a tampered Parquet footer route a reader's KMS token to an attacker. Here's the fix, Iceberg's safe path, and how to audit your la...
Parquet Page Indexes and the Last Mile of Pruning in Apache Iceberg
Parquet page indexes can cut selective Iceberg scans by an order of magnitude on sorted data. How they work, what they cost, and how to lay out tables...
How Apache Polaris Plans to Share Iceberg Tables Across Organizations
Apache Polaris's Open Sharing proposal adds first-class shares, external consumers, and listings so any Iceberg REST engine can read shared tables....
Vector Search Directly Over Iceberg Tables, and When You Still Need a Vector Database
Embeddings are just columns. When exact vector search over Iceberg scans beats a vector database, when it does not, and how to lay out tables....
Apache Data Lakehouse Weekly: September 9 to 17, 2026
Apache Data Lakehouse Weekly: September 3–9, 2026
The Open Lakehouse Explained, Then Built on Your Laptop with Dremio and MinIO
Migrating Into Iceberg Without Moving Data
add_files, snapshot, and migrate compared: the three in-place paths into Iceberg, the reconciliation each requires, the layout traps, and the rollback...
What Iceberg Table Maintenance Actually Costs
A cost model for compaction, snapshot expiry, orphan cleanup, and manifest rewriting: what each operation spends, on which meter, and how to set a sch...
Running an Iceberg Lakehouse on Kubernetes
Catalog, maintenance, and compaction as Kubernetes workloads: scheduling classes, job structure, credential flow, and the failures that come from the ...
Partition Statistics Files in Apache Iceberg
The underused Iceberg metadata for planning: what the partition statistics file holds, what the spec guarantees, how to write one, and when it earns i...
What to Assert When You Test an Iceberg Pipeline
Fixtures, in-memory catalogs, and golden metadata: the assertions that catch wrong rows, unsafe reruns, schema drift, and concurrent-write corruption ...
The 2026 Iceberg REST Catalog Compatibility Report
A repeatable test for what an Iceberg REST catalog actually serves, a scoring scheme that separates design from breakage, and the 2026 evidence across...
Serving Iceberg Tables From Two Regions
Three multi-region topologies that work and one that mostly does not, what an Iceberg commit costs across regions, and where the catalog has to live....
How Iceberg Catalogs Hand Engines Storage Access
Credential vending end to end: the wire protocol, scoped access on each cloud, remote signing, credential lifetime on long jobs, and failures that loo...
Kafka Connect to Iceberg: How the Commit Actually Works
Exactly-once semantics in the Iceberg sink connector: the coordinator, the control topic, offsets stored inside Iceberg snapshots, and where duplicate...
The Open Lakehouse Explained, Then Built on Your Laptop with Dremio and MinIO
The five layers of the open lakehouse explained, then a lab: Parquet, Iceberg, Polaris, Arrow, and Ossie running in two containers on your own machine...
Inside the Puffin File Format
dbt on Iceberg: Incremental Models on Open Tables
How dbt incremental materializations map to Iceberg operations, and the configuration, predicates, and maintenance that keep them healthy....
Disaster Recovery for Iceberg Tables: Replication, Backup, and Restore
Disaster recovery for Iceberg across four tiers: snapshots, object versioning, catalog backup, and cross-region replication....
Deleting User Data From an Immutable Lakehouse: GDPR Hard Deletes on Iceberg
How to turn a logical delete on immutable Iceberg into a physical erasure across snapshots, versions, replicas, and downstream copies....
Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet
How Iceberg v3 geometry and geography types, bounding boxes, and native Parquet types give spatial data first-class standing....
Default Column Values and Field IDs: How Iceberg Schema Evolution Works at the Spec Level
How field IDs and initial and write defaults let Iceberg change schemas on large tables without rewriting data, at the spec level....
The Iceberg Table Properties That Actually Matter
The Iceberg table properties that decide file count, pruning, write amplification, retention, and metadata growth, by workload....
The Lakehouse Ingestion Tool Landscape: Fivetran, Airbyte, dlt, and CDC vs Batch
How Fivetran, Airbyte, dlt, and CDC and streaming tools land well-behaved Apache Iceberg tables, and how to choose and maintain them....
Local Iceberg Development Environments: Docker, MinIO, and In-Memory Catalogs for CI
Local Iceberg development environments: in-process catalogs, a Docker Compose stack with MinIO, and CI configurations that run either....
Moving Iceberg Tables Between Catalogs Without Rewriting Data
Why moving Iceberg tables between catalogs is a pointer copy, and the protocol that makes a cutover safe for one table or thousands....
Logs, Traces, and Metrics as Tables: Building an OpenTelemetry Data Lake on Iceberg
Building an OpenTelemetry data lake on Iceberg: schemas for spans, logs, and metrics, ingestion, query patterns, and retention....
Postgres Meets the Lakehouse: pg_lake, pg_duckdb, and When Postgres Is Enough
What pg_lake, pg_duckdb, and pg_mooncake do at the Iceberg level, and honest thresholds for when Postgres is enough....
Schema Registries and Event Schemas: Avro, Protobuf, and JSON Schema on the Way Into the Lakehouse
How Avro, Protobuf, and JSON Schema evolve through a registry, and how that maps to the schema evolution rules of Iceberg....
Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet
Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet
A logistics team stores 40 million delivery stops in an Apache Iceberg table....
Agent-Driven Storage Tiering for Apache Iceberg: Moving Cold Data Without Breaking Queries
A background agent can move cold Iceberg partitions to cheaper tiers without breaking live queries. Heatmaps, path-safe moves, and restore paths....
DataFusion Comet 1.0 and What Native Rust Scans Change for Spark on Iceberg
DataFusion Comet 1.0 replaces Spark Iceberg scans with native Rust. What speeds up, what still falls back to the JVM, and how to deploy it....
FSST and ALP: The Two Encodings Fixing Parquet's Weakest Compression Cases
ALP and FSST target Parquet's worst cases: floats and high-cardinality strings. How they work and what they change for Iceberg tables....
High-Throughput Branch Merging: Automating Concurrency and Conflict Resolution in Multi-Branch Iceberg Pipelines
High-throughput Iceberg branch merges need conflict detection and automation. How to reconcile concurrent writes without stalling pipelines....
Multi-Cloud REST Catalog Topologies: Running Apache Polaris Across AWS, Azure, and GCP
Polaris can catalog Iceberg tables across AWS, Azure, and GCP. Four topologies, credential vending, and the tradeoffs of each design....
Parquet-Only Manifests in Iceberg v4: Why the Metadata Layer Is Going Columnar
Iceberg v4 is moving manifests from Avro to Parquet so planners can read only the stats they need. Why the metadata layer is going columnar....
Semantic Layer Federation: One Logical Model Over Data on Three Clouds
One logical model over Iceberg and databases on three clouds. Pushdown, egress, Reflections, and where semantic federation still breaks....
Serverless Iceberg Ingestion with PyIceberg and DuckDB: Micro-Batches Without a Spark Cluster
Land small Iceberg micro-batches with PyIceberg and DuckDB in a serverless function. Commits, concurrency, and why Spark is the wrong default....
Zero-Copy Warehouse Modernization: Migrating Legacy Databases to Apache Iceberg Without Downtime
Move a legacy warehouse to Iceberg without downtime by virtualizing first. Consumer cutover, parity checks, and background copy without double-ETL....
Podcast episodes

AI Updates, Lakehouse Streaming (Iceberg/Fluss), and more!

2025 Reflections, Google Antigravity, NotebookLM, Dremio AI Agent, Pangolin Catalog, Dremioframe & Iceframe Python Libraries for Apache Iceberg

2025 Reflections, Google Antigravity, NotebookLM, Dremio AI Agent, Pangolin Catalog, Dremioframe & Iceframe Python Libraries for Apache Iceberg

Lakehouse Catalogs Beyond Apache Iceberg, What could they Look Like?
