Latest Articles: page 14

692 posts from Alex's blogs and newsletters, newest first. Thoughts on tech, data, policy, and philosophy.

Alex Merced's Blog

Legacy Warehouses to Open Lakehouses: A Step-by-Step Migration Playbook

Migrating from a legacy data warehouse to an open lakehouse? This step-by-step playbook covers assessment, phased migration, validation, and avoiding....

Alex Merced's Blog

Building the Brain of the Agentic Lakehouse: Designing an Open Catalog Architecture

The open catalog is the brain of the agentic lakehouse. Learn how Apache Polaris, Dremio's Open Catalog, and catalog-native governance enable reliable...

Alex Merced's Blog

Evaluating the TCO of an Open Lakehouse vs. Proprietary Data Warehouses

Open lakehouse vs proprietary warehouse: a comprehensive TCO breakdown covering storage, compute, engineering, and hidden costs to help you make the r...

Alex Merced's Blog

Real-Time BI: Enabling Sub-Second Queries on Apache Iceberg Data Lakehouses

Sub-second queries on Apache Iceberg are achievable with the right architecture. Learn how Reflections, C3 cache, and query acceleration close the BI....

Alex Merced's Blog

The Rise of Agentic Analytics: Shifting BI from Passive Dashboards to Goal-Directed Action

Agentic analytics replaces static dashboards with AI agents that pursue business goals autonomously....

Alex Merced's Blog

The Semantic Layer as a Translation Engine: Bridging Natural Language and SQL

The semantic layer translates business language into accurate SQL for AI agents. Learn how virtual datasets, metric definitions, and wikis power agent...

Alex Merced's Blog

Comparing the Top 2026 Agentic Analytics Tools: ThoughtSpot, Databricks, and Tableau

How do ThoughtSpot, Databricks, and Tableau compare as agentic analytics platforms in 2026?...

Alex Merced's Blog

Trustworthy AI in the Agentic Lakehouse: Reconciling Concurrency and Isolation Contracts

Hundreds of AI agents querying simultaneously create concurrency and isolation problems. Learn how Iceberg OCC, Dremio FGAC, and guardrail policies en...

Alex Merced's Blog

The 2026 Unified Data Architecture: Reconciling Multi-Cloud Data Lakehouses

Multi-cloud data lakehouses in 2026 run on Apache Iceberg, open catalogs, and zero-ETL federation. Here's what a composable, unified architecture look...

Alex Merced's Blog

Why Traditional Lakehouses Fail AI Agents: The Mathematical Case for the Agentic Lakehouse

Traditional lakehouses expose raw directories and ambiguous schemas to AI agents, causing hallucination....

Alex Merced's Blog

The Era of Zero-ETL Federation: Fueling AI Agents with Real-Time Cross-Enterprise Data

Zero-ETL federation lets AI agents join real-time CRM data with historical lakehouse tables instantly....

Alex Merced's Blog

Use Hermes Agent for Free With DeepSeek V4 and Slack

Install Hermes Agent, choose a current DeepSeek provider, and connect Slack through Socket Mode with the tokens, scopes, events, and allowlist Hermes ...

Alex Merced's Blog

Concurrency, Isolation, and MVCC: How Engines Handle Contention

Databases handle concurrent access using locks, MVCC, or optimistic concurrency control. Here is how each approach works and what tradeoffs each creat...

Alex Merced's Blog

Hash, Sort-Merge, Broadcast: How Distributed Joins Work

Distributed joins move data across the network using shuffle, broadcast, or co-location strategies. Here is how each works and when engines choose whi...

Alex Merced's Blog

Partitioning, Sharding, and Data Distribution Strategies

Hash partitioning distributes data evenly. Range partitioning enables fast range scans. Both create tradeoffs....

Alex Merced's Blog

Buffer Pools, Caches, and the Memory Hierarchy

Databases use buffer pools, column caches, and result caches to keep hot data in RAM. Here is how each caching strategy works and what happens when da...

Alex Merced's Blog

Volcano, Vectorized, Compiled: How Engines Execute Your Query

The Volcano model processes one row at a time. Vectorized execution processes batches with SIMD. Code generation fuses operators into compiled code....

Alex Merced's Blog

Inside the Query Optimizer: How Engines Pick a Plan

Query optimizers transform SQL into execution plans using rule-based rewrites, cost-based search, and adaptive runtime adjustments....

Alex Merced's Blog

B-Trees, LSM Trees, and the Indexing Tradeoff Spectrum

B-trees balance reads and writes for OLTP. LSM trees maximize write throughput. Bitmap indexes accelerate OLAP filtering. Here is when to use each....

Alex Merced's Blog

How Databases Organize Data on Disk: Pages, Blocks, and File Formats

Databases structure data on disk as heap files, sorted files, or LSM trees, then wrap it in formats like Parquet with metadata that lets engines skip....

Alex Merced's Blog

Row vs. Column: How Storage Layout Shapes Everything

Row stores keep records together for fast transactions. Column stores keep field values together for fast analytics....

Alex Merced's Blog

How Query Engines Think: The Tradeoffs Behind Every Data System

Every database is a collection of engineering tradeoffs. Learn the 9 design decisions that shape how query engines store, index, and process your data...

Alex Merced's Blog

Migrating to Apache Iceberg: Strategies for Every Source System

Migrate to Iceberg from Hive, data warehouses, or raw files using in-place migration, full rewrite, or the zero-downtime view swap pattern....

Alex Merced's Blog

Hands-On with Apache Iceberg Using Dremio Cloud

A practical walkthrough of creating, querying, and optimizing Iceberg tables on Dremio Cloud, from account setup to AI-powered analytics....