Latest Articles: page 7
692 posts from Alex's blogs and newsletters, newest first. Thoughts on tech, data, policy, and philosophy.
How Iceberg V3 Variant Shredding Changed Semi-Structured Data on S3 Tables
How Iceberg V3's Variant type and Parquet shredding turn JSON columns into prunable typed columns, with real benchmark tradeoffs and a migration path....
Reading the Apache Iceberg V4 Proposals Before They Land
A field guide to the Apache Iceberg V4 proposals: adaptive metadata trees, single-file commits, typed statistics, column families, and what is safe....
Building an Honest TCO Model for Open Lakehouses and Proprietary Warehouses
An honest TCO framework for open lakehouses versus proprietary warehouses: five cost categories, measured numbers, sensitivity analysis, and where eac...
Why Agentic AI Needs a Governed Semantic Layer Behind the Model Context Protocol
Why agentic AI needs a governed semantic layer behind the Model Context Protocol: metric consistency, access control, Apache Ossie for portable....
Moving From Supply Chain Dashboards to Decision Loops With the Model Context Protocol
Moving from supply chain dashboards to decision loops with MCP: sense, decide, act, and verify, with typed action tools, idempotency keys, and graduat...
Metric Contracts as the Interface AI Agents Actually Need
Metric contracts as the interface AI agents need: calculation, inclusion rules, grain, temporal semantics, ownership, semantic versioning, and testing...
Cross-Cloud Credential Vending in Apache Polaris and the End of Permanent Storage Keys
How Apache Polaris vends short-lived, prefix-scoped storage credentials across AWS, Azure, and GCP, and how to retire permanent storage keys for good....
Designing Policy-Aware Telemetry Tables for AI Systems in Apache Iceberg
Designing policy-aware AI telemetry tables in Apache Iceberg: what to log, tamper evidence, retention against conflicting deletion requirements....
Defending the Lakehouse Gateway Against Prompt Injection and Data Exfiltration
Defending the lakehouse gateway against prompt injection and data exfiltration: per-user identity, no-SQL tool surfaces, volume bounds, and detection....
How the Iceberg REST Catalog Turned Into the Lakehouse Control Plane
How the Iceberg REST catalog became the lakehouse control plane: multi-table atomic commits, credential vending, capability negotiation, and what stil...
A Migration Playbook for Moving Legacy Warehouses onto Apache Iceberg
A dependency-first playbook for migrating legacy warehouses onto Apache Iceberg: snapshot vs migrate vs add_files, four-level parity validation....
What Zero-Copy Data Sharing Actually Does Between Salesforce, Snowflake, and Databricks
What zero-copy data sharing actually does across Salesforce, Snowflake, and Databricks: query federation, file federation, catalog federation, and whe...
Apache Polaris 1.7.0 and the Quiet Work of Making a Catalog Trustworthy
Apache Polaris 1.7.0 deep dive: idempotent writes, semantic models, stricter credential vending, orphan cleanup, and what the upgrade asks of you....
Designing Batch Pipelines That Write Well Into Apache Iceberg
How to design batch pipelines that write well into Apache Iceberg: commit strategy, partitioning, sort order, write-audit-publish, and maintenance don...
Apache Iceberg Support Across the Major Hyperscalers
How AWS, Google Cloud, and Microsoft Azure actually support Apache Iceberg: storage, catalogs, maintenance, governance, and interoperability, layer by...
The Inner Work of Freedom: Confronting the Fear and Hate in Your Own Heart
TL;DR The case for a freer world is usually made in the language of constitutions, economics, and policy....
Guardrails for Analytics Agents That Do More Than Answer Questions
The risk isn't agents going rogue, it's agents acting correctly on bad input at machine speed....
Building Agent Telemetry Tables in Iceberg That Survive an Audit
A practical guide to building agent decision traces in Apache Iceberg that support audit reconstruction, governance review, and cost attribution....
What Agentic Analytics Actually Costs, and How to Keep It Bounded
Agent analytics generates two cost streams that scale on different variables. Here's the arithmetic, the levers that actually move the number, and how...
Running an Apache Iceberg Lakehouse With No Internet Connection
A practical guide to deploying an Iceberg lakehouse in air-gapped environments: component choices, artifact pipelines, identity without a cloud....
When the Query Optimizer Starts Managing Its Own Materializations
Autonomous materialized view management replaces quarterly review meetings with workload-driven scoring, and it's essential when AI agents generate....
Why AI Agents Fail on Raw Data, and What to Give Them Instead
Agents fail on raw lake data because business rules live in people's heads. Data products with semantic contracts fix this at the source....
Why Iceberg V4 Wants to Retire Equality Deletes, and What Streaming Teams Should Do About It
Equality deletes made streaming upserts into Iceberg practical at the cost of read performance....
The Five Layers Between Your Lakehouse and a Trustworthy Agent
Agent reliability is a property of the stack the model sits on. Five layers with distinct owners and failure modes turn the agent is unreliable....