Latest Articles: page 10
692 posts from Alex's blogs and newsletters, newest first. Thoughts on tech, data, policy, and philosophy.
The State of Apache Parquet in 2026: The Quiet Format Enters Its Loudest Decade
Apache Parquet in 2026, variant types, geospatial, ALP encoding, footer redesign, the versioning debate, and how the decade-old format is renovating....
The State of Apache Polaris in July 2026: From Incubating Catalog to the Governance Layer of the Open Lakehouse
Apache Polaris as a TLP, federation, credential vending, semantic layers, lineage, and how the open catalog became the governance plane....
The State of Streaming to Apache Iceberg in July 2026: Every Path, Its Latency, and What to Do When Seconds Are Not Fast Enough
Every path for streaming data into Iceberg in 2026, Flink, Spark, Kafka Connect, broker-native, managed pipelines, with honest latency numbers....
High-Performance Columnar Transfers: Combining Apache Arrow Flight and Iceberg REST Catalogs
Modern lakehouse architecture is easier to reason about when you separate two questions. The first question is how a system discovers and governs a......
Block vs. Object Storage: A Deep Dive Into the Foundation of Modern Data, and How the Lakehouse Made the Slow Option Fast
Here is one of the strangest and most consequential plot twists in the history of data infrastructure: over the past decade, the analytics industry......
The Buyer's Scorecard for Agentic Analytics: Evaluating Tooling in the Enterprise AI Era
Agentic analytics demos are easy to enjoy and hard to evaluate. A user asks a question, an assistant answers, a chart appears, and the room leans f......
Building Closed-Loop Decision Agents: Moving from Passive BI Dashboards to Active Goal-Directed Workflows
Dashboards are excellent at showing people what happened. They are less good at deciding what should happen next. That gap is where closed-loop de......
Conversational AI on Managed Iceberg: Exposing Amazon S3 Tables through MCP
The most interesting part of conversational analytics is not the chat box. The chat box is just the surface area. The harder question is what happe......
Decoupled Catalogs vs. Managed Tables: Architectural Freedom in the Age of Table Format Convergence
Open table formats have changed buyer expectations. A few years ago, the question was whether an organization should put more analytical data into......
Designing Your Own AI Harness: A Deep Dive Into the Architecture of Agent Loops, Tools, Context, and Control
A deep dive into custom AI harness architecture: model layers, tool design, context management, permissions, control budgets, persistence, orchestrati...
Deterministic Data Engineering With AI Harnesses: Using Claude Code, Codex, Antigravity, and OpenCode for Data Work You Can Actually Trust
How to use AI agent harnesses for data engineering without losing determinism, reproducibility, and trust in your data pipelines and analytics....
Preparing Your Data Lakehouse for the EU AI Act: Auditable Lineage and Data Provenance
The EU AI Act changes the conversation around AI architecture because it makes trust operational. It is not enough to say that an AI system is usef......
Federation and the Lakehouse: Two Roads to Unified Data Access, and How to Know Which One to Take
Every data strategy document written this decade contains some version of the same sentence: we need a single place to access all our data. The sen......
A Deep Dive Into File Compression: How Data Gets Smaller, Why Codecs Differ, and What to Actually Use in the Lakehouse
Somewhere in your data platform right now, a single configuration property is quietly deciding a meaningful percentage of your storage bill, your q......
The File Format Renaissance: Parquet, Lance, Vortex, Nimble, BtrBlocks, and the New Physics of Columnar Storage
For a decade, the file format layer was the most settled real estate in data. Apache Parquet held the analytical world, ORC held the Hive legacy es......
Enforcing Fine-Grained Security at Machine Speed: Dynamic Access Control for High-Frequency AI Agents
AI agents change the security model for analytics. A human user may run a handful of queries, pause, interpret the answer, and ask a follow-up. An......
Implementing Positional Deletes in Iceberg v3: Streamlining Merge-on-Read for Fast-Inbound Event Lakes
Event data has a way of humbling neat architecture diagrams. It arrives late. It arrives twice. It arrives with incorrect attributes. It needs priv......
Mapping the Variant Type in Iceberg v3: Standardizing Semi-Structured AI JSON Payloads
AI applications are messy data producers. They create prompts, completions, tool calls, retrieval traces, ranking signals, evaluation scores, safet......
Designing Idempotent Pipelines in the Agentic Lakehouse: Eliminating Double-Write Anomalies
Agents retry. Networks fail. Jobs time out after doing some work. APIs return ambiguous responses. Schedulers run the same workflow twice. A human......
File Encryption for the Lakehouse: The Terminology, the Machinery, and the Hard Problem of Interoperable Encrypted Tables
For years, the open lakehouse had an honest gap that practitioners whispered about and slide decks skipped: encryption. Not the checkbox kind, ever......
The 2026-07-28 Model Context Protocol Release Candidate: What the Stateless Spec Means for Data Platforms
The date in this topic matters. Today is July 6, 2026. A release candidate dated July 28, 2026 is still in the future. That means this article shou......
The Metric Contract Mandate: Standardizing Semantic Layers Before AI Agent Access
AI agents are very good at moving quickly. That is the opportunity and the risk. If an agent can inspect metadata, generate queries, compare result......
Multi-Engine Catalog Federation with Apache Polaris: Syncing Google Cloud, AWS, and Azure Metadata
Open table formats changed the data lakehouse conversation, but they did not finish it. A table can be stored in an open format and still be hard t......
Open Source Foundations, Explained: What Apache, Linux, Eclipse, and Their Peers Actually Do, and Why Governance Differences Matter
Writing about open data and AI means repeating the same phrases over and over: donated to the Apache Software Foundation, incubating at the Linux F......