An independent resource for the open lakehouse
Reference material on Apache Iceberg, lakehouse catalogs, the agentic lakehouse, and modern data architecture. It covers what table formats are, how to deploy Apache Polaris, and how to connect query engines to Iceberg tables. Written by a practitioner, free to read.
Not an Apache project. This is a personal site by Alex Merced. It is not affiliated with, endorsed by, or sponsored by the Apache Software Foundation or the Apache Iceberg project, whose official site is iceberg.apache.org.
- 425+
- Articles
- 200+
- Reference entries
- 14
- Subject areas
- Free
- No paywall, ever
Browse by topic
Nine long-form pillar guides that cover the lakehouse stack from the file format up to the agents querying it.
- Pillar guide
Apache Iceberg
Covers the metadata tree, snapshots, hidden partitioning, and the catalog API from end to end.
Read the guide - Deep dive
Iceberg Architecture
How manifest lists, manifest files, and data files fit together at query time.
Read the guide - Catalogs
The REST Catalog
What the Iceberg REST Catalog spec standardizes, and how engines authenticate against it.
Read the guide - Deep dive
Snapshots & Time Travel
Atomic commits, snapshot expiration, rollback, and querying a table as of any point in time.
Read the guide - Deep dive
Schema Evolution
Add, drop, rename, and reorder columns safely, and why Iceberg's field IDs make it work.
Read the guide - Comparison
Iceberg vs Delta Lake vs Hudi
A neutral comparison of the three open table formats across design, features, and ecosystem.
Read the guide - Pillar guide
The Data Lakehouse
What a lakehouse actually is, the layers it is built from, and how it differs from a warehouse.
Read the guide - Agentic AI
The Agentic Lakehouse
Semantic layers, MCP, and the architecture AI agents need to query your data reliably.
Read the guide - Foundations
Open Table Formats
Why table formats exist at all, and the problems they solved for data lakes.
Read the guide
Recent Posts
- 31 MIN READ•Aug 6, 2026
Apache Arrow Flight and ADBC, and Why Database Connectivity Finally Went Columnar
Arrow Flight and ADBC move database results as columnar data, ending the row-oriented bottleneck between engines and applications. Here's how.
Apache ArrowADBCArrow Flight - 21 MIN READ•Aug 4, 2026
Budgeting for Agentic Analytics When Every Question Costs Something Different
Budgeting for agentic analytics when every question costs something different: token economics, query economics, instrumentation, and the cost controls that actually return.
AI AgentsTCOCost Management - 21 MIN READ•Aug 4, 2026
The Five Layers of an Agentic Lakehouse and Where the MCP Server Sits
The five layers of an agentic lakehouse and where the MCP server sits: storage, catalog, semantic layer, MCP gateway, and agent surface, plus identity, session isolation, and budgets.
AI AgentsMCPAgentic Lakehouse - 21 MIN READ•Aug 4, 2026
Autonomous Table Optimization When Your Query Workload Stops Being Predictable
Autonomous table optimization when query workloads stop being predictable: observing file layout and query patterns, scoring compaction work, adaptive sort order, and cost discipline.
Apache IcebergTable OptimizationCompaction - 21 MIN READ•Aug 4, 2026
Building Apache Iceberg Lakehouses That Run Without an Internet Connection
How to build an Apache Iceberg lakehouse that runs fully offline: storage, catalog, compute, cross-zone transfer, compliance, and the failure modes that bite.
Apache IcebergAir-GappedOn-Premises - 21 MIN READ•Aug 4, 2026
Wiring Analytical Queries to Transactional APIs in Closed-Loop Decision Agents
Wiring analytical queries to transactional APIs in closed-loop decision agents: conditional writes, sagas with compensations, decision records, and blast radius controls.
AI AgentsDecision LoopsSaga Pattern
Learn alongside the rest of the community
A Slack workspace for lakehouse practitioners, plus a shared calendar of meetups, webinars, and Lakehouse Linkups.
Must reads on Iceberg, agentic AI, and the lakehouse
-
The Definitive Guide to the Semantic Layer
Understand what a semantic layer is, why it matters for modern data architectures, and how it creates a consistent, governed layer between raw data and business consumers.
Read article -
Apache Polaris: The Catalog Standard for Lakehouses and AI
A deep dive into Apache Polaris, the open-source catalog that is emerging as the standard for managing Iceberg tables across multi-engine Lakehouses and AI workloads.
Read article -
What Are Table Formats and Why Were They Needed?
Explore the history and motivations behind open table formats like Apache Iceberg, Delta Lake, and Apache Hudi, and why they solved critical problems in big data engineering.
Read article -
What is Dremio?
An overview of Dremio's Lakehouse platform: how it unifies data access, accelerates queries, and powers self-service analytics across cloud and on-premise sources.
Read article -
What Apache Iceberg Native Actually Means
Not all Iceberg integrations are equal. This article breaks down what it truly means for a platform to be 'Apache Iceberg native' and why the distinction matters for your architecture.
Read article -
Open Source and the Data Lakehouse
A survey of the open source ecosystem powering modern Data Lakehouses, from Apache Iceberg and Nessie to Apache Arrow and Spark, and how they work together.
Read article -
What is Agentic Analytics?
How AI agents are changing analytics pipelines by querying data, generating insights, and taking actions on their own, and what that means for the Lakehouse.
Read article
This site is an independent publication by Alex Merced. It is not affiliated with, endorsed by, or sponsored by the Apache Software Foundation or the Apache Iceberg project, whose official home is iceberg.apache.org.