> ## Content Index
> Fetch the complete content index at: https://stack-rundown.ghost.io/llms.txt
> Use this file to discover other available public pages before exploring further.

# Cloud Data Warehouse Cost Models: Guide
- URL: https://stack-rundown.ghost.io/cloud-data-warehouse-cost-models-guide/
- Published: 2026-08-26T01:32:36.000Z
- Updated: 2026-09-08T18:43:57.000Z
- Description: Choose a data warehouse by billing model—compute, storage, concurrency, and egress are the real cost drivers to control.
- Author: SR Staff
- Tags: Comparisons

**Most cloud warehouse bills come down to six things: compute, storage, serverless features, reserved capacity, concurrency, and data transfer.** In many cases, compute makes up **80% to 90%** of total spend, while storage and add-on services fill in the rest.

If I were trying to size up [Snowflake](https://www.snowflake.com/en/?ref=stack-rundown.ghost.io), [BigQuery](https://cloud.google.com/bigquery?ref=stack-rundown.ghost.io), [Redshift](https://aws.amazon.com/redshift/?ref=stack-rundown.ghost.io), or [Databricks](https://www.databricks.com/?ref=stack-rundown.ghost.io) fast, I’d focus on three questions:

- **What is the billing unit?** Credits, slot-hours, node-hours, DBUs, or TB scanned.
- **What shape is the workload?** Steady usage often fits reserved capacity. Spiky usage often fits on-demand or serverless.
- **Where does cost drift start?** Idle compute, full-table scans, poor partitioning, long retention, extra concurrency, and egress.

A few numbers stand out right away:

- **BigQuery on-demand:** about **$5.00 per TB scanned**
- **Snowflake credits:** often **$2.00 to $4.00+** each
- **Redshift managed storage:** about **$0.024 per GB/month**
- **Cloud egress:** often **$0.08 to $0.12 per GB**
- **Reserved capacity savings:** up to **75%** in some cases

The short version: *you are not just picking a platform*. You are picking a billing model. And that model affects whether your costs stay steady or drift month after month.

## The real cost of cloud data warehouses: how they bill you and which is the most cost-efficient

###### sbb-itb-fd683fe

## Quick comparison

| Platform       | Main billing unit        | Main cost pattern                                      | Common cost leak                 | Best fit                                     |
| -------------- | ------------------------ | ------------------------------------------------------ | -------------------------------- | -------------------------------------------- |
| **Snowflake**  | Credits                  | Time-based compute plus storage and serverless charges | Warehouses left running too long | Teams that want separate compute control     |
| **BigQuery**   | TB scanned or slot-hours | Query-scan pricing or reserved slots                   | SELECT \* and weak partitioning  | Light, uneven analytics or shared slot usage |
| **Redshift**   | Node-hours or RPU-hours  | Provisioned nodes or serverless capacity               | Fixed clusters with low use      | Steady reporting and set workloads           |
| **Databricks** | DBU-hours                | DBUs plus cloud VM and storage charges                 | Idle all-purpose clusters        | Mixed data and SQL workloads                 |

**My main takeaway:** before I commit to any warehouse plan, I’d model **12-month** and **36-month** cost, check how much data moves across regions, and review the top cost drivers every month in **USD**. That alone can stop a lot of bill shock.

## Core Cloud Data Warehouse Pricing Models

Those cost drivers tend to show up in six common billing models. Think of these as the frame for the platform-by-platform pricing breakdown that comes next.

| Pricing Model         | Billing Unit          | Best-Fit Workload             | Main Upside                       | Main Cost Risk                                 |
| --------------------- | --------------------- | ----------------------------- | --------------------------------- | ---------------------------------------------- |
| **Pay-Per-Query**     | TB/GB scanned         | Irregular or light analytics  | Pay only for what the query scans | Poor partitioning or broad queries spike costs |
| **On-Demand Compute** | Node-hour / Credit    | Variable or spiky workloads   | High flexibility; no commitment   | High unit cost during peak usage               |
| **Reserved Capacity** | Fixed monthly fee     | Steady, predictable reporting | Up to 75% cost reduction          | Paying for idle or unused capacity             |
| **Storage Tiers**     | GB-month              | Long-term data retention      | Lower cost for aged data          | Growth from time travel or backups             |
| **Serverless**        | Compute unit / second | Bursty or automated tasks     | Zero infrastructure management    | Hidden spikes during heavy automation          |
| **Data Transfer**     | GB egressed           | Single-region operations      | Low cost for internal movement    | High fees for cross-region or cross-cloud exit |

### Pay-Per-Query Pricing

With pay-per-query pricing, you pay for the data a query scans, not for cluster time. That makes this model a good fit for light or irregular analytics.

The downside is pretty simple: broad queries can get expensive fast. The same goes for weak partitioning. If a query touches far more data than it needs, the bill climbs with it.

### On-Demand Compute vs. Reserved Capacity

On-demand compute bills active resources by the second, minute, or hour, with no upfront commitment. That gives teams plenty of flexibility, especially when usage jumps around.

But there’s a tradeoff. Unit costs are high, and busy periods can turn into bigger-than-expected bills.

Reserved capacity lowers the unit price in exchange for a fixed commitment, often across one- to three-year terms. If your reporting workload stays steady, this can cut costs by as much as 75%.

The catch? You’re paying for that capacity whether you use it or not. So if usage drops, idle or underused commitments can start to drain budget.

### Storage Tiers and Retention Costs

Storage is billed per GB-month, so costs tend to grow as retention grows. That sounds manageable at first, but it adds up over time.

Features like Time Travel and backups can stretch retention windows and push storage use higher. So even if query volume stays flat, monthly storage charges can keep creeping up.

### Serverless, Concurrency, and Egress Fees

Serverless options remove infrastructure management, which is a big part of their appeal. They also work well for bursty jobs and automated tasks.

Still, spend can jump during heavy automation. A busy set of scheduled jobs, concurrent workloads, or frequent data movement can push total costs up in a hurry.

Egress fees come into play when data moves out of a cloud platform or across regions. These charges usually fall between **$0.08 and $0.12 per GB**, with AWS averaging about **$0.09 per GB** and Azure about **$0.087 per GB**. That means cross-region and cross-cloud movement deserves close attention, even when compute looks under control.

Next, these billing models map differently across Snowflake, BigQuery, Redshift, and Databricks.

## How [Snowflake](https://www.snowflake.com/en/?ref=stack-rundown.ghost.io), [BigQuery](https://cloud.google.com/bigquery?ref=stack-rundown.ghost.io), [Redshift](https://aws.amazon.com/redshift/?ref=stack-rundown.ghost.io), and [Databricks](https://www.databricks.com/?ref=stack-rundown.ghost.io) Price Usage

![Snowflake](https://assets.seobotai.com/stackrundown.com/6a8e313cf0ae24ed42a33241/a6b6955040946b291302773a283d4c07.jpg)

![Cloud Data Warehouse Cost Models: Snowflake vs BigQuery vs Redshift vs Databricks](https://assets.seobotai.com/undefined/6a8e313cf0ae24ed42a33241-1787707504223.jpg) 

Cloud Data Warehouse Cost Models: Snowflake vs BigQuery vs Redshift vs Databricks

Here’s how those cost drivers show up on each platform. This table helps tie each billing unit to the thing that usually pushes the bill up.

| Platform       | Main billing unit           | Storage Model                  | Reservation Option             | Concurrency Behavior             | Common Drift Source           |
| -------------- | --------------------------- | ------------------------------ | ------------------------------ | -------------------------------- | ----------------------------- |
| **Snowflake**  | Credits (time-based)        | Monthly average per TB         | Pre-purchased credits          | Multi-cluster warehouses         | Long auto-suspend windows     |
| **BigQuery**   | Bytes scanned or slot-hours | Active & Long-term tiers       | Slot reservations (Editions)   | Shared slot pool                 | Unoptimized SELECT \* queries |
| **Redshift**   | Node-hours or RPU-hours     | Managed Storage (RMS)          | Reserved Instances (1 or 3 yr) | Concurrency Scaling (credits/hr) | Underutilized fixed nodes     |
| **Databricks** | DBUs (processing power)     | Cloud object storage (S3/ADLS) | DBU pre-purchase commit        | Serverless scaling               | Idle all-purpose clusters     |

### Snowflake and BigQuery: Credits, Slots, Storage, and Serverless Charges

Snowflake charges compute in credits tied to warehouse size. Bigger warehouses burn through credits faster. Credits often land between **$2.00 and $4.00+** based on your edition, such as Standard, Enterprise, or Business Critical.

Storage can grow in ways teams don’t always spot at first. **Time Travel** and **Fail-safe** keep older versions of data, sometimes for as long as 90 days, which can push storage costs up. On top of that, serverless features like [Snowpipe](https://docs.snowflake.com/en/user-guide/data-load-snowpipe-intro?ref=stack-rundown.ghost.io) and auto-clustering show up as separate charges.

BigQuery works a bit differently. It charges either by bytes scanned for on-demand use or by slot-hour capacity. On-demand queries cost about **$5 per TB scanned**, while slot capacity through [BigQuery Editions](https://docs.cloud.google.com/bigquery/docs/editions-intro?ref=stack-rundown.ghost.io) starts at around **$0.04 per slot-hour**. Storage also gets cheaper after 90 days without changes, moving into a lower long-term rate. Streaming inserts and certain API calls can add extra line items too.

Same idea, different meter: each platform has its own unit, but the bill still moves based on compute time, scan volume, storage, and idle usage.

### Redshift and Databricks: Node-Hours, Managed Storage, and DBU-Based Compute

Redshift provisioned clusters are billed by node-hour, while Redshift Serverless uses RPU-hour. With [RA3](https://aws.amazon.com/redshift/features/ra3/?ref=stack-rundown.ghost.io), compute and storage are split apart. That matters because Managed Storage (RMS) is billed on its own at about **$0.024 per GB/month**, so storage can grow without forcing you to scale compute at the same time.

RA3.xlplus nodes start at about **$0.85 per hour**. If you query data in S3 through [Redshift Spectrum](https://docs.aws.amazon.com/redshift/latest/dg/c-using-spectrum.html?ref=stack-rundown.ghost.io), that adds about **$5 per TB scanned**. Teams with stable workloads can cut hourly rates quite a bit by using Reserved Instances on one- or three-year terms.

Databricks pricing stacks a few things together. You pay for DBUs, plus cloud VM and storage charges. DBU rates change by workload type. For example, SQL Pro compute runs at roughly **$0.22 per DBU**. Serverless SQL is closer to **$0.55 per DBU**. That’s more per unit, but it helps avoid paying for clusters that sit around doing nothing. Idle all-purpose clusters, by contrast, keep charging even when nobody’s using them.

### Which Billing Levers Teams Can Control on Each Platform

The main question isn’t just *how* each platform bills. It’s which settings help stop waste before it piles up.

A few levers matter more than most:

- **Snowflake:** tighten auto-suspend settings so warehouses don’t sit idle
- **BigQuery:** use partitioning and clustering to cut scanned data
- **Redshift:** match RA3 with reservations for steadier workloads
- **Databricks:** use serverless SQL for ad hoc work to avoid idle cluster spend

## Where Costs Drift in Real-World Use

List pricing is neat and easy to follow. Actual bills usually aren’t.

Costs drift because a handful of repeat patterns stack up over time. These platforms stay predictable only when teams keep a close grip on compute, retention, and data movement. Once that slips, bills start to creep - often in ways that tie straight back to the six cost drivers covered earlier.

### The Most Common Billing Leaks Across All Four Platforms

The table below shows the drift scenarios that show up most often, how they hit the bill, and the simplest fix.

| Drift Scenario                      | Snowflake                                                                                                      | BigQuery                                                                   | Redshift                                                                       | Databricks                                                                                                    | Practical Fix                                                                                      |
| ----------------------------------- | -------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- | ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| **Idle compute**                    | Warehouses left running after ETL jobs keep consuming credits because auto-suspend is disabled or set too high | Unused slot reservations can still bill even when demand is low            | Idle reserved capacity racks up node-hours during idle periods                 | All-purpose clusters left running between jobs drive DBU charges                                              | Set Snowflake auto-suspend to 5–10 minutes or less; enforce auto-termination on all clusters       |
| **SELECT \* and full-table scans**  | More micro-partitions are scanned, increasing credit use                                                       | Bytes processed increase directly, raising on-demand query cost            | Higher disk I/O and CPU per query, longer runtimes                             | More data shuffled per job, increasing DBU consumption                                                        | Select only needed columns; filter early                                                           |
| **Poor partitioning or clustering** | Micro-partition pruning fails, excess data scanned per query                                                   | Unpartitioned tables incur full-scan bytes processed charges               | Suboptimal sort/distribution keys cause heavy I/O and network redistribution   | Non-partitioned [Delta](https://delta.io/?ref=stack-rundown.ghost.io) tables increase shuffle and job runtime | Partition on the most-used filter column; use clustering, sort keys, or Z-Ordering where supported |
| **Unnecessary retention**           | Stale clones, Time Travel history, and old snapshots inflate storage                                           | Rarely queried tables accumulate active and long-term storage charges      | Large historical tables increase managed storage and backup size               | Old Delta snapshots grow object storage and add metadata overhead                                             | Archive or delete unused data                                                                      |
| **Uncontrolled concurrency**        | BI dashboards refreshing every few minutes trigger multi-cluster scale-up                                      | Frequent dashboard refreshes re-scan large tables under on-demand pricing  | Misconfigured WLM queues trigger unnecessary concurrency scaling charges       | Multiple users sharing large clusters multiply DBU usage                                                      | Stagger refreshes; cap concurrency per role or workload queue                                      |
| **Repeated data movement**          | Cross-region replication and frequent COPY exports generate egress and extra compute                           | Exporting data across regions or to external tools triggers egress charges | Copying data between S3 buckets in different regions drives egress and compute | Continuous cross-region pipeline copies add network costs on top of DBUs                                      | Use sharing instead of physical copies                                                             |

The biggest leaks usually begin with query design and table layout. That’s where small habits turn into monthly cost creep.

### How SQL Design and Data Layout Affect Your Bill

A `SELECT *` against a large fact table forces every platform to read columns the report never touches. On BigQuery, that means more bytes processed and a higher on-demand charge. On Snowflake, it means more micro-partitions scanned and more credit use. The fix is simple: list only the columns you need and push filters as early as possible.

Partition filters matter just as much. BigQuery can prune whole partitions, and Snowflake can skip micro-partitions when filters line up with clustering. But one small SQL choice can ruin that. Wrap a partition column in a function - like `DATE_TRUNC('month', event_date)` \- and pruning may fail, forcing a full scan. It’s the kind of mistake that hides in plain sight until you check query history.

Data layout adds another layer. On Databricks, uncompacted tables can create too many small files, which adds planning overhead and slows job startup. Wide tables with dozens of rarely used columns also push up storage and scan costs across every platform. And if those wide reporting tables stick around too long, they can drag down query speed too - not just storage charges. Narrower reporting tables or incremental materializations for common dashboards help keep both storage and refresh compute easier to track.

### What Small Teams Should Check Every Month

If no one owns FinOps, cost drift can sit there for weeks before anyone spots it. A simple monthly review goes a long way.

- Review spend by compute, storage, serverless, and egress.
- Inspect the 10 costliest queries.
- Check whether reservation utilization is underused.

That routine catches most drift before it snowballs.

Once you know where bills drift, the next step is choosing the pricing model that fits your workload shape.

## How to Pick the Right Cost Model for Your Workload

Once you know what drives cost, the next step is simple: **match pricing to the shape of your workload**.

If demand stays steady, reserved capacity often makes more sense. If usage comes in waves, on-demand or serverless pricing is usually a better match. It’s less about which model sounds better on paper and more about how your team actually works day to day.

### Best Fit for Steady Reporting and Repeated Queries

If your team runs the same dashboards on a set schedule, handles scheduled ETL jobs, and supports a steady number of concurrent users, a **reserved capacity model** is often the better fit.

Why? Because stable usage is exactly what reserved pricing is built for. You’re trading flexibility for lower cost over time. But don’t rush into a commitment based on a good month or two. Check that your baseline usage has stayed stable for at least 12 months.

That 12-month view matters. If demand slips later, you could end up paying for capacity you’re not using.

### Best Fit for Spiky Analytics and Small Teams

On-demand or serverless pricing works well for teams with uneven usage. That includes infrequent exploratory queries, bursty ad hoc analysis, or smaller teams that don’t want to deal with capacity planning.

The upside is clear: you pay when work runs, not for idle capacity.

The catch? Peak periods can get expensive fast. And monthly spend can be harder to pin down once API charges, data transfer, and admin costs start piling on in the background.

### Workload Patterns and Best-Fit Models

As a quick rule of thumb, **predictable workloads** usually line up with reserved capacity, while **bursty workloads** tend to fit on-demand or serverless models.

| Workload Pattern                           | Best-Fit Model            | Main Reason                                                       | Main Caution                                                     |
| ------------------------------------------ | ------------------------- | ----------------------------------------------------------------- | ---------------------------------------------------------------- |
| **Steady dashboard reporting**             | Reserved capacity         | Predictable costs; up to 75% savings with 1 to 3 year commitments | Risk of paying for idle capacity if demand drops                 |
| **Bursty ad hoc analysis**                 | On-demand / serverless    | Pays only for active use; no upfront commitment                   | Higher unit costs during peak usage periods                      |
| **Mixed analytics and pipeline workloads** | Serverless / auto-scaling | Flexibility to handle varying data types and query volumes        | Multiple billing levers can make monthly costs harder to predict |

Before you commit, model both 12-month and 36-month TCO. Include storage growth, egress, and concurrency-driven scaling so you’re not comparing sticker prices and missing the parts that hit later.

## Conclusion: Choosing a Cost Model You Can Predict and Control

Once you line up pricing with your workload, the next job is keeping the bill under control. Cloud data warehouse costs usually come from a mix of items: compute, storage, concurrency, egress, and commitments. And the cheapest posted rate on the pricing page doesn’t always lead to the cheapest bill in practice. The best model is the one you can forecast with some confidence and keep in check.

That’s why small leaks matter so much. A few wasteful patterns can turn a manageable bill into an inflated one. In most cases, overruns come from idle compute, repeated scans, retention creep, and data movement.

Use that lens before signing any commitment. Ask a few blunt questions first: Which limit gets hit first - compute, concurrency, or storage? Is the workload steady enough for a 1- to 3-year commitment? And if data moves between services or regions, how much will egress add?

### Key Points to Take Away

In practice, this comes down to four rules:

- Know the billing unit before you buy.
- Match the pricing model to your workload shape - reserved for steady use, on-demand for spiky use.
- Use governance features as cost controls.
- Review spend monthly in USD.

## FAQs

### How do I choose between on-demand and reserved pricing?

Choose **on-demand** when your workload is hard to predict or when you need more room to adjust. You pay only for what you use, instead of locking yourself into capacity ahead of time. That can help you avoid paying for more reserved capacity than you need.

Choose **reserved pricing** when demand stays steady and you can forecast usage with some confidence. The tradeoff is commitment, but the payoff is lower per-unit costs. Committed use can cut costs by **up to 75%** compared with on-demand over roughly **1–3 years**.

### What usually drives the biggest warehouse cost overruns?

The biggest cost overruns usually come from **hidden infrastructure costs** and **poor resource management**.

In plain English, spending tends to drift when teams don't keep a close eye on compute usage, use more storage than they need, or underestimate maintenance and training. Those costs can sneak up fast.

High data volumes, frequent retrievals, and egress fees can push total spending even higher. And without clear governance and controls like **rightsizing** and **data lifecycle management**, budgets can get drained before anyone sees it coming.

### How should I estimate 12-month and 36-month warehouse costs?

Estimate **total cost of ownership**, not just the subscription price. Start with your expected data growth, storage tiers, and compute demand. Then layer in the costs that tend to sneak up on teams: data ingestion, ETL, maintenance, training, plus any integration or egress fees.

For longer forecasts, include committed-use discounts for 1- to 3-year agreements. Those deals can cut costs by **up to 75%** compared with on-demand pricing. It also helps to add a **30% to 40%** buffer for the hidden stuff, like API calls and data movement.

## Related Blog Posts

- [Cloud Storage Pricing: AWS S3, GCP, Azure, B2](https://stack-rundown.ghost.io/cloud-storage-pricing-aws-s3-gcp-azure-b2/)
- [How to Choose Billing Software for SaaS](https://stack-rundown.ghost.io/how-to-choose-billing-software-for-saas/)
- [Future of Workflow Automation: AI and iPaaS](https://stack-rundown.ghost.io/future-workflow-automation-ai-ipaas/)
- [AWS vs. Azure vs. Google: Archive Storage](https://stack-rundown.ghost.io/aws-vs-azure-vs-google-archive-storage/)

---

## More on StackRundown

Continue on the [Software Comparisons hub](https://stack-rundown.ghost.io/comparisons/), or read next:

- [Cloud Storage Pricing: AWS S3, GCP, Azure, B2](https://stack-rundown.ghost.io/cloud-storage-pricing-aws-s3-gcp-azure-b2/)