SERIES BREC 34TYPE BRIEFgarnetgrid.com
Snowflake in 2026: Still the Gold Standard, or Yesterday's Architecture?
The question

Insights
Snowflake did something genuinely new, and most of it is now table stakes. The decision in 2026 turns on three quieter things: how much of your bill is awake-time rather than work, who would own the catalogue if you ever left, and whether governance and data sharing are actually on your critical path. This is what goes wrong in practice, and what we honestly don't know.
The interesting question is not whether Snowflake still works. It does, and the parts of it that were genuinely novel a decade ago have aged well. The question is narrower and harder: is the thing Snowflake optimises for still the thing your organisation is actually constrained by, and if you decide in two years that it isn't, what does it cost you to change your mind? Those are two separate questions, and most evaluations answer neither, because they get spent comparing benchmark times on queries nobody runs.
What you are actually buying
Strip the marketing away and Snowflake sells two things. The first is the separation of storage from compute: data sits in object storage as small immutable files with per-column metadata, and you rent short-lived clusters, called warehouses, that read it. The second, less discussed and more valuable, is the deletion of a tuning surface. No indexes to design, no vacuum to schedule, no shard key to regret, no rebalancing when you add capacity. The planner leans on file-level metadata to skip files it can prove are irrelevant. Concurrency is another warehouse pointed at the same data. Cloning a database is a metadata operation. Recovering a table someone truncated is a SQL statement rather than a restore ticket.
That set of properties was remarkable when it shipped. In 2026 most of it is table stakes: the other major warehouses separate storage from compute and have removed most of the knobs too. Snowflake's remaining advantage is not architectural. It is operational polish, governance, and the fact that SQL-literate analysts can run it with nobody on call. Judge it on that, not on the architecture diagram.
The bill is awake-time, not work
You are not billed for compute you use. You are billed for the time a warehouse is awake, and the gap between those two quantities is where every cost overrun lives.
The mechanisms are boring and specific. A dashboard tool polling every few minutes against a warehouse with a generous auto-suspend never lets it sleep, so you pay all day for a handful of queries. A warehouse that resumes constantly pays a minimum charge each time it wakes, which makes chatty small jobs expensive per unit of useful work. Each size step roughly doubles the burn rate, and that only pays for itself if a single query can use the extra parallelism. Spilling is the tell: a query spilling to local disk is memory-bound and may well go faster on a bigger warehouse, while a query that spills nothing and scans little just costs twice as much for the same wall-clock time. Queueing is a different problem entirely, and upsizing does not fix it.
The subtler one is pruning. Metadata skipping only works if the data landed in an order that correlates with how you filter it. Append-only, time-ordered ingestion prunes beautifully on a date predicate. The same table rebuilt nightly by a full-refresh transformation, in whatever order the join emitted rows, prunes on nothing, and every dashboard query quietly becomes a full scan. A clustering key is not a free fix either: on a high-churn table, background reclustering can cost more than the scans it saves.
Storage has its own trap, because retention and fail-safe windows mean churn is not free. A table that rewrites most of its rows nightly keeps paying for versions you no longer want. Clones are cheap until they diverge, after which they are simply tables.
Then attribution. The billing unit is the warehouse, not the query, so one shared warehouse gives you no way to say which team spent the money, and a warehouse per team gives you many idle tails and many minimum charges. Both are defensible. Choosing neither, which is the usual outcome, gives you a bill nobody can explain.
None of that is really a Snowflake defect. It is modelling and orchestration, and it migrates with you.
The moat moved to the catalogue
The real shift since Snowflake's design is that open table formats made object storage legible to engines other than the one that wrote it. Iceberg and Delta added manifest and snapshot metadata on top of Parquet, plus a catalogue naming the current snapshot. Both major vendors have publicly embraced Iceberg and shipped catalogue implementations, so the storage moat is largely gone: you can keep your own Parquet in your own bucket and read it from more than one engine.
What replaced it is the catalogue, because whoever owns the catalogue owns the write path. Transactions, compaction, snapshot expiry and the permission model all live there. "Open format" and "open catalogue" are separate claims, and it is worth establishing which one a vendor is making.
Two writers over one table is where the theory stops being clean: two compaction strategies and one metadata pointer is a coordination problem, and read support has matured ahead of write support more or less everywhere. Cross-engine reading is real today. Cross-engine writing with governance intact is something to test on your own workload rather than accept from a slide.
And the lock-in was never SQL. It is the streams, tasks, scheduled transformations, stored procedures, ingestion pipelines, masking and row-access policies, and the semantic layer built above all of it. Table-format portability does not touch a line of that. If Iceberg is your portability story, price the rewrite of everything that is not a SELECT.
Where it earns the money, and where it doesn't
Three things genuinely justify paying for a managed warehouse. Governance as a product: column masking, row-level access, tagging and a single coherent permission model, which is a config exercise here and a project anywhere else. Sharing without copying: granting a counterparty a live read of a table instead of a nightly file drop is the most underrated feature in the product and the hardest to rebuild. And elasticity for genuinely spiky work, where the alternative is a cluster sized for a peak that happens twice a week.
Notice that those are organisational strengths, not technical ones. If governance and external sharing are not on your critical path, you may be insuring a risk you do not carry.
Against that, a great deal of what runs on distributed warehouses would fit on one machine with fast local storage and a vectorised engine. DuckDB, ClickHouse and Postgres with a columnar extension have all become serious, and hardware has outrun many organisations' data growth. There is no honest general threshold for where the line sits: it depends on the largest table you actually scan rather than the one you store, on how many people query interactively at once, and on how much of the work is batch rather than serving. Anyone quoting you a figure in terabytes is guessing. The test is cheap, though. Take your heaviest real query, with its real filters, against your largest real table, and run it on one box.
The symmetrical error is just as common: replacing a managed warehouse with infrastructure nobody in the company is funded to operate is not a saving, it is a deferred outage.
The AI pitch, and our bias in judging it
The current pitch is inference next to the data, with model functions callable from SQL so the table never moves. The underlying argument is sound. Data gravity is real, and a second copy of your data in a second system with a second permission model is a cost people routinely ignore.
Two cautions. A model function invoked inside a query is metered per call, and a join can multiply those calls in ways nobody reviews before merging; that spend sits on top of warehouse awake-time, not instead of it. And if your constraint is that the data cannot leave your premises or a named jurisdiction, this is not a trade-off at all. You are not choosing between warehouses, you are choosing whether to use one.
We should be straight about our position. We build private systems that run on hardware the customer owns, so the cases that reach us are disproportionately the ones where the answer is "that data cannot go there". That is a selection bias and you should read this section knowing it. For an organisation whose data can legitimately live in someone else's cloud, this is a real trade-off and Snowflake is a reasonable answer to it.
What to ask before you sign, or before you leave
Four questions settle this faster than a benchmark.
Do not migrate for architectural fashion, and do not stay out of inertia. Spend a fortnight instrumenting instead: warehouse metering and query history, grouped by warehouse and by the transformation that issued the query, plus an honest measure of how much of the bill is idle time. Fix auto-suspend, fix the models that destroyed your pruning, and split or consolidate warehouses deliberately so the bill maps to a team. Land new immutable data as Iceberg in storage you control, so changing engines later is a decision rather than a project. Then make the call with numbers you own rather than benchmarks someone else published. The honest conclusion is usually not that Snowflake is yesterday's architecture. It is that the expensive part of your stack was never the warehouse.
Read next
- Databricks vs Snowflake vs Fabric
- Data Lake vs Lakehouse vs Warehouse
- Cloud Cost Optimization That Survives Contact
Related: Architecture Audit