Latest Articles: page 2
692 posts from Alex's blogs and newsletters, newest first. Thoughts on tech, data, policy, and philosophy.
Apache Data Lakehouse Weekly: September 9 to 17, 2026
Release managers had a humbling week....
Apache Data Lakehouse Weekly: September 9 to 17, 2026
AI Weekly: DeepSeek Cuts Prices as Agents Go Hosted
AI Weekly: DeepSeek Cuts Prices as Agents Go Hosted
Week of September 10 to 17, 2026...
The Fed and Monetary Policy: Who Pays for Inflation
TL;DR The Federal Reserve does not print money in the literal sense....
Apache Data Lakehouse Weekly: September 3-9, 2026
Three things ran through the dev lists this week....
Apache Data Lakehouse Weekly: September 3–9, 2026
The Open Lakehouse Explained, Then Built on Your Laptop with Dremio and MinIO
The Open Lakehouse Explained, Then Built on Your Laptop with Dremio and MinIO
Most people learn the open lakehouse backwards....
AI Weekly: Four Frontier Models in Seven Days
Four labs shipped flagship models inside one week, and not one of them cut its headline price....
AI Weekly: Four Frontier Models in Seven Days
Agentic Data Architecture
A six-layer reference architecture for agents on company data: planners, tool boundaries, identity, the semantic layer, and what breaks when a layer i...
Running Apache Polaris in Production
Apache Polaris past the quickstart: persistence backends, realm bootstrap, replica token signing, upgrades, backups, and which failures take the lakeh...
What the 2026 Consolidation Means for Open Formats
Tabular, Dremio, and a wave of data platform acquisitions: what changes for teams building on open formats and which guarantees survive a change of ow...
Context Engineering for Data Agents
Why text-to-SQL accuracy collapses on enterprise schemas, the five kinds of context an agent needs, where each one hides, and how to make the semantic...
Guardrails for AI on Company Data
The control surfaces that actually contain damage once an agent is fooled: identity, permissions, audit trails, and prompt injection at the query laye...
Migrating Into Iceberg Without Moving Data
add_files, snapshot, and migrate compared: the three in-place paths into Iceberg, the reconciliation each requires, the layout traps, and the rollback...
What Iceberg Table Maintenance Actually Costs
A cost model for compaction, snapshot expiry, orphan cleanup, and manifest rewriting: what each operation spends, on which meter, and how to set a sch...
Running an Iceberg Lakehouse on Kubernetes
Catalog, maintenance, and compaction as Kubernetes workloads: scheduling classes, job structure, credential flow, and the failures that come from the ...
Partition Statistics Files in Apache Iceberg
The underused Iceberg metadata for planning: what the partition statistics file holds, what the spec guarantees, how to write one, and when it earns i...
What to Assert When You Test an Iceberg Pipeline
Fixtures, in-memory catalogs, and golden metadata: the assertions that catch wrong rows, unsafe reruns, schema drift, and concurrent-write corruption ...
The 2026 Iceberg REST Catalog Compatibility Report
A repeatable test for what an Iceberg REST catalog actually serves, a scoring scheme that separates design from breakage, and the 2026 evidence across...
Serving Iceberg Tables From Two Regions
Three multi-region topologies that work and one that mostly does not, what an Iceberg commit costs across regions, and where the catalog has to live....
How Iceberg Catalogs Hand Engines Storage Access
Credential vending end to end: the wire protocol, scoped access on each cloud, remote signing, credential lifetime on long jobs, and failures that loo...