Latest Articles: page 14
692 posts from Alex's blogs and newsletters, newest first. Thoughts on tech, data, policy, and philosophy.
Legacy Warehouses to Open Lakehouses: A Step-by-Step Migration Playbook
Migrating from a legacy data warehouse to an open lakehouse? This step-by-step playbook covers assessment, phased migration, validation, and avoiding....
Building the Brain of the Agentic Lakehouse: Designing an Open Catalog Architecture
The open catalog is the brain of the agentic lakehouse. Learn how Apache Polaris, Dremio's Open Catalog, and catalog-native governance enable reliable...
Evaluating the TCO of an Open Lakehouse vs. Proprietary Data Warehouses
Open lakehouse vs proprietary warehouse: a comprehensive TCO breakdown covering storage, compute, engineering, and hidden costs to help you make the r...
Real-Time BI: Enabling Sub-Second Queries on Apache Iceberg Data Lakehouses
Sub-second queries on Apache Iceberg are achievable with the right architecture. Learn how Reflections, C3 cache, and query acceleration close the BI....
The Rise of Agentic Analytics: Shifting BI from Passive Dashboards to Goal-Directed Action
Agentic analytics replaces static dashboards with AI agents that pursue business goals autonomously....
The Semantic Layer as a Translation Engine: Bridging Natural Language and SQL
The semantic layer translates business language into accurate SQL for AI agents. Learn how virtual datasets, metric definitions, and wikis power agent...
Comparing the Top 2026 Agentic Analytics Tools: ThoughtSpot, Databricks, and Tableau
How do ThoughtSpot, Databricks, and Tableau compare as agentic analytics platforms in 2026?...
Trustworthy AI in the Agentic Lakehouse: Reconciling Concurrency and Isolation Contracts
Hundreds of AI agents querying simultaneously create concurrency and isolation problems. Learn how Iceberg OCC, Dremio FGAC, and guardrail policies en...
The 2026 Unified Data Architecture: Reconciling Multi-Cloud Data Lakehouses
Multi-cloud data lakehouses in 2026 run on Apache Iceberg, open catalogs, and zero-ETL federation. Here's what a composable, unified architecture look...
Why Traditional Lakehouses Fail AI Agents: The Mathematical Case for the Agentic Lakehouse
Traditional lakehouses expose raw directories and ambiguous schemas to AI agents, causing hallucination....
The Era of Zero-ETL Federation: Fueling AI Agents with Real-Time Cross-Enterprise Data
Zero-ETL federation lets AI agents join real-time CRM data with historical lakehouse tables instantly....
Use Hermes Agent for Free With DeepSeek V4 and Slack
Install Hermes Agent, choose a current DeepSeek provider, and connect Slack through Socket Mode with the tokens, scopes, events, and allowlist Hermes ...
Concurrency, Isolation, and MVCC: How Engines Handle Contention
Databases handle concurrent access using locks, MVCC, or optimistic concurrency control. Here is how each approach works and what tradeoffs each creat...
Hash, Sort-Merge, Broadcast: How Distributed Joins Work
Distributed joins move data across the network using shuffle, broadcast, or co-location strategies. Here is how each works and when engines choose whi...
Partitioning, Sharding, and Data Distribution Strategies
Hash partitioning distributes data evenly. Range partitioning enables fast range scans. Both create tradeoffs....
Buffer Pools, Caches, and the Memory Hierarchy
Databases use buffer pools, column caches, and result caches to keep hot data in RAM. Here is how each caching strategy works and what happens when da...
Volcano, Vectorized, Compiled: How Engines Execute Your Query
The Volcano model processes one row at a time. Vectorized execution processes batches with SIMD. Code generation fuses operators into compiled code....
Inside the Query Optimizer: How Engines Pick a Plan
Query optimizers transform SQL into execution plans using rule-based rewrites, cost-based search, and adaptive runtime adjustments....
B-Trees, LSM Trees, and the Indexing Tradeoff Spectrum
B-trees balance reads and writes for OLTP. LSM trees maximize write throughput. Bitmap indexes accelerate OLAP filtering. Here is when to use each....
How Databases Organize Data on Disk: Pages, Blocks, and File Formats
Databases structure data on disk as heap files, sorted files, or LSM trees, then wrap it in formats like Parquet with metadata that lets engines skip....
Row vs. Column: How Storage Layout Shapes Everything
Row stores keep records together for fast transactions. Column stores keep field values together for fast analytics....
How Query Engines Think: The Tradeoffs Behind Every Data System
Every database is a collection of engineering tradeoffs. Learn the 9 design decisions that shape how query engines store, index, and process your data...
Migrating to Apache Iceberg: Strategies for Every Source System
Migrate to Iceberg from Hive, data warehouses, or raw files using in-place migration, full rewrite, or the zero-downtime view swap pattern....
Hands-On with Apache Iceberg Using Dremio Cloud
A practical walkthrough of creating, querying, and optimizing Iceberg tables on Dremio Cloud, from account setup to AI-powered analytics....