Latest Articles: page 5
699 posts from Alex's blogs and newsletters, newest first. Thoughts on tech, data, policy, and philosophy.
DataFusion Comet 1.0 and What Native Rust Scans Change for Spark on Iceberg
DataFusion Comet 1.0 replaces Spark Iceberg scans with native Rust. What speeds up, what still falls back to the JVM, and how to deploy it....
FSST and ALP: The Two Encodings Fixing Parquet's Weakest Compression Cases
ALP and FSST target Parquet's worst cases: floats and high-cardinality strings. How they work and what they change for Iceberg tables....
Governance-as-Code for the Lakehouse: Managing REST Catalog RBAC and Masking in Git
Put REST catalog RBAC and masking in Git. How to review grants, apply them safely, and keep lakehouse access from drifting....
Metric Contracts in Code: Testing, Versioning, and Serving Business Logic to Multi-Agent Systems
Metric contracts in code let teams test, version, and serve business logic to multi-agent systems without each agent inventing its own SQL....
High-Throughput Branch Merging: Automating Concurrency and Conflict Resolution in Multi-Branch Iceberg Pipelines
High-throughput Iceberg branch merges need conflict detection and automation. How to reconcile concurrent writes without stalling pipelines....
Multi-Cloud REST Catalog Topologies: Running Apache Polaris Across AWS, Azure, and GCP
Polaris can catalog Iceberg tables across AWS, Azure, and GCP. Four topologies, credential vending, and the tradeoffs of each design....
Parquet-Only Manifests in Iceberg v4: Why the Metadata Layer Is Going Columnar
Iceberg v4 is moving manifests from Avro to Parquet so planners can read only the stats they need. Why the metadata layer is going columnar....
Query Routing at Machine Scale: Dynamic Workload Distribution Across Lakehouse Engines
Route each lakehouse query by shape, not by sender. Signals, rules, and how to keep dashboards, batch jobs, and agents from sharing one engine....
Semantic Layer Federation: One Logical Model Over Data on Three Clouds
One logical model over Iceberg and databases on three clouds. Pushdown, egress, Reflections, and where semantic federation still breaks....
Serverless Iceberg Ingestion with PyIceberg and DuckDB: Micro-Batches Without a Spark Cluster
Land small Iceberg micro-batches with PyIceberg and DuckDB in a serverless function. Commits, concurrency, and why Spark is the wrong default....
Zero-Copy Warehouse Modernization: Migrating Legacy Databases to Apache Iceberg Without Downtime
Move a legacy warehouse to Iceberg without downtime by virtualizing first. Consumer cutover, parity checks, and background copy without double-ETL....
The Agent Is Now a Named Coworker, and It Needs a File Format
Named, persistent agents need a file format. Open Agent Profile, Buzz, Grok Bot, and Hermes Bot Mode show why a portable agent identity matters....
Your Agent Should Answer the Phone: A Field Guide to AI Gateways on Slack, Discord, Telegram, Signal, and Teams
A field guide to AI gateways on Slack, Discord, Telegram, Signal, and Teams: architecture, auth, cost, and the failure modes that matter....
Graphs in AI Engineering Have Solved Three Problems. The Fourth Is the Plan.
Knowledge graphs, GraphRAG, and LangGraph solved three problems. The fourth is the work itself: a reviewable graph of bounded agentic loops....
The Hidden Cost of Tiny Iceberg Commits
Trace what one tiny Iceberg commit writes, then model hourly, per-minute, and per-second cadences so streaming costs become arithmetic, not adjectives...
Deletion Vectors vs Position Deletes vs Equality Deletes: The Iceberg Delete Story in 2026
Position deletes, equality deletes, and deletion vectors compared from the Iceberg spec: what each writes, how readers apply it, and when to use which...
Iceberg Is Becoming a Library, Not Just a Table Format
Iceberg is turning from a JVM table format into a library other systems embed. What that shift changes for engines, catalogs, and the spec itself....
Iceberg Is Escaping the JVM: Why Rust, Go, Python and C++ Implementations Matter
Rust, Go, Python, and C++ Iceberg implementations change who can write the format. Why multi-language clients matter more than another JVM engine....
The Iceberg REST Catalog Compatibility Test: One Suite of Operations Every Platform Should Pass
One suite of REST catalog operations every Iceberg platform should pass. What sameness means, where implementations diverge, and how to test it....
Iceberg REST Remote Scan Planning Changes More Than Query Performance
Remote scan planning moves Iceberg file selection into the catalog. What that changes for engines, governance, and operational cost beyond query speed...
Iceberg Row Lineage: The Feature AI and CDC Workloads Will Eventually Depend On
Iceberg row lineage gives rows a durable identity across rewrites. Why CDC pipelines and AI workloads will eventually depend on it....
Iceberg v4's Adaptive Metadata Tree, Explained From First Principles
Iceberg v4's adaptive metadata tree, explained from first principles: why commits rewrite too much today and how the tree makes change cheaper....
Why Iceberg v4 Is Really About Making the Cost of Change Proportional to the Change
Iceberg v4 is really about making the cost of a change proportional to the change. The principle, the current tax, and what the redesign pays down....
The Catalog Can Now Plan Your Iceberg Query: Inside REST Scan Planning
A mechanics walkthrough of Iceberg REST scan planning: client-side planning, remote endpoints, pagination, and where engine support stands in 2026....