Latest Articles: page 3
692 posts from Alex's blogs and newsletters, newest first. Thoughts on tech, data, policy, and philosophy.
Kafka Connect to Iceberg: How the Commit Actually Works
Exactly-once semantics in the Iceberg sink connector: the coordinator, the control topic, offsets stored inside Iceberg snapshots, and where duplicate...
What a Query Costs
Four meters, their real proportions, and how to attribute compute to a query, a table, and a team so a platform can answer what a dashboard costs to r...
The Open Lakehouse Explained, Then Built on Your Laptop with Dremio and MinIO
The five layers of the open lakehouse explained, then a lab: Parquet, Iceberg, Polaris, Arrow, and Ossie running in two containers on your own machine...
Where Lock-In Went
The format war ended and exit cost did not: where lock-in relocated after open tables won, how to measure it, and which costs are worth keeping down....
The Gig Economy and Labor Law: A Mismatch That Hurts Workers
TL;DR The gig economy is not a new development in American labor....
Apache Data Lakehouse Weekly: August 26 to September 2, 2026
By Alex Merced, Data Lakehouse and AI Evangelist...
AI Weekly: Cheap Tokens, Tight Safeguards, and a Two Million GPU Order
Week of August 26 to September 2, 2026...
The Housing Crisis Is a Government Problem
TL;DR The United States is in the middle of a genuine housing affordability crisis, and the primary cause is government policy at every level: zoning ...
Inside the Puffin File Format
A query joins a 2-billion-row fact table to a 40,000-row dimension table....
Data Quality Tooling Compared: Great Expectations, Soda, dbt Tests, and Anomaly Detection
A comparison of Great Expectations, Soda, dbt tests, and anomaly detection, and a layered design that uses each where it fits....
The Data Team of the Agentic Era: Generalists Owning End-to-End Workflows
The case for generalists owning end-to-end data workflows with agents, the counterargument, and how to make the transition work....
dbt on Iceberg: Incremental Models on Open Tables
How dbt incremental materializations map to Iceberg operations, and the configuration, predicates, and maintenance that keep them healthy....
Disaster Recovery for Iceberg Tables: Replication, Backup, and Restore
Disaster recovery for Iceberg across four tiers: snapshots, object versioning, catalog backup, and cross-region replication....
Deleting User Data From an Immutable Lakehouse: GDPR Hard Deletes on Iceberg
How to turn a logical delete on immutable Iceberg into a physical erasure across snapshots, versions, replicas, and downstream copies....
Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet
How Iceberg v3 geometry and geography types, bounding boxes, and native Parquet types give spatial data first-class standing....
Default Column Values and Field IDs: How Iceberg Schema Evolution Works at the Spec Level
How field IDs and initial and write defaults let Iceberg change schemas on large tables without rewriting data, at the spec level....
The Iceberg Table Properties That Actually Matter
The Iceberg table properties that decide file count, pruning, write amplification, retention, and metadata growth, by workload....
Inside the Puffin File Format
The Puffin file format inside out, byte by byte, covering Theta sketches for distinct values and deletion vectors....
The Lakehouse Ingestion Tool Landscape: Fivetran, Airbyte, dlt, and CDC vs Batch
How Fivetran, Airbyte, dlt, and CDC and streaming tools land well-behaved Apache Iceberg tables, and how to choose and maintain them....
Local Iceberg Development Environments: Docker, MinIO, and In-Memory Catalogs for CI
Local Iceberg development environments: in-process catalogs, a Docker Compose stack with MinIO, and CI configurations that run either....
Metadata Platforms in 2026: DataHub, OpenMetadata, Atlan, and Catalog Convergence
How the technical catalog and the metadata platform are converging in 2026, and how to arrange the two layers for a lakehouse....
Moving Iceberg Tables Between Catalogs Without Rewriting Data
Why moving Iceberg tables between catalogs is a pointer copy, and the protocol that makes a cutover safe for one table or thousands....
Logs, Traces, and Metrics as Tables: Building an OpenTelemetry Data Lake on Iceberg
Building an OpenTelemetry data lake on Iceberg: schemas for spans, logs, and metrics, ingestion, query patterns, and retention....
Orchestration in 2026: Airflow 3 vs Dagster vs Prefect vs Event-Driven
Where Airflow 3, Dagster, Prefect, and event-driven triggering stand for lakehouse pipelines in 2026, after the Prefect acquisition of Dagster....