Indonesia Singapore ไทย Pilipinas Việt Nam Malaysia မြန်မာ ລາວ
← Back to Blog

Data Costs Are Rising: What Your Warehouse Isn't Telling You

Unattributed warehouse costs aren't a finance problem — they're a data governance problem, and fixing them starts with lineage visibility.

An abstract editorial illustration showing tangled data pipelines converging into a single glowing cost meter
Illustrated by Mikael Venne

Warehouse bills are climbing and the culprits are hidden in plain sight. Here's how to find them — and build leaner, smarter data pipelines.

Your data warehouse bill went up again last quarter. You know it. Finance knows it. And nobody in the room can point to exactly why.

This isn’t a budgeting problem — it’s an observability problem. And in Southeast Asia, where data teams are often lean and platform ecosystems are sprawling (think multi-channel stacks spanning Shopee, Lazada, LINE, and homegrown CRMs), the gap between what your warehouse costs and what it explains is wider than most teams want to admit.

The Real Reason Your Warehouse Costs Keep Climbing

Monte Carlo’s release of their Cost Agent puts a useful frame around a problem most data teams feel but can’t fully articulate: the data you’d need to diagnose rising warehouse costs — query history, lineage, schema metadata, usage patterns — is scattered across systems that don’t communicate with each other. According to Monte Carlo, finding the specific tables and pipelines responsible for cost spikes, at scale, typically requires a dedicated analyst doing forensics work that most teams simply don’t have bandwidth for.

The practical consequence is that organisations keep paying for pipelines nobody is actively using, tables that haven’t been queried in months, and redundant transformations that accumulated as the stack grew. In a region where many brands are still consolidating first-party data programmes post-cookie, this kind of invisible bloat compounds fast. Every unused pipeline carrying user consent data is not just a cost liability — it’s a compliance surface you don’t need.

Lineage Is the Unlock — If You Can Actually See It

The Databricks Data + AI Summit brought some useful signals from the Fivetran and dbt Labs partnership. Specifically, dbt’s introduction of dbt State — a mechanism for tracking what has and hasn’t changed across model runs — and the early previews of dbt Core v2.0 point toward a future where transformation pipelines are leaner by design, not by accident.

dbt Wizard, their AI-assisted development tool, is worth watching for teams managing complex Southeast Asian data stacks: multi-language product catalogues, platform-specific schemas, and fragmented identity graphs all create transformation overhead that compounds quickly. When a single SQL model is touching four different source schemas because your Thai and Indonesian storefronts were onboarded separately, the cost isn’t in the compute — it’s in the unmanaged complexity.

The strategic move here is treating lineage documentation not as a nice-to-have but as a cost control mechanism. If you can’t trace a table back to a business question, you probably shouldn’t be storing it.


First-Party Data Programmes Need a Cost Discipline Too

Here’s the angle that rarely makes it into warehouse cost conversations: first-party data programmes are not free just because you own the data. The collection infrastructure, consent management layers, identity resolution pipelines, and activation workflows all run on compute. And because these programmes are often built incrementally — one consent touchpoint here, one enrichment feed there — the architectural debt accumulates quietly.

The discipline I’d advocate for is what I think of as consent-proportionate infrastructure: the cost and complexity of your data pipeline should scale with the value and consent signal behind the data, not with the volume of data you happen to have collected. A high-consent, high-intent dataset from a loyalty programme deserves robust, well-maintained pipelines. A low-signal behavioural feed from a third-party integration that predates your current consent framework probably doesn’t.

This isn’t just good governance — it’s good economics. Brands that have audited their data infrastructure through this lens consistently find 20–30% of their storage and compute is supporting data assets that either lack proper consent documentation or haven’t driven a single activation in over six months.

What to Actually Do This Quarter

Cost visibility tools like Monte Carlo’s Cost Agent are genuinely useful, but they’re diagnostic, not prescriptive. The strategic work is deciding what to do with what you find. A practical starting point:

Audit by activation rate, not just query frequency. A table that gets queried weekly by an automated pipeline but has never fed a campaign segment, a personalisation model, or a business report is not an asset — it’s infrastructure theatre. Map every major data asset to at least one activation outcome in the last 90 days. What can’t be mapped gets flagged for deprecation review.

Separate your consent tiers in your data architecture, not just your CMP. If your warehouse treats opted-in loyalty members and cookied anonymous browsers as equivalent rows in the same table, you’re creating both compliance risk and cost inefficiency. Structuring your data model around consent levels makes it easier to run proportionate retention policies and simplifies compliance responses when regulators in Thailand, Indonesia, or the Philippines ask questions.

Use dbt State or equivalent change-tracking to stop rebuilding what hasn’t changed. This is table-stakes hygiene that many Southeast Asian data teams skip because their stacks grew fast and documentation lagged behind. dbt Core v2.0’s direction on this is the right one — incremental, state-aware transformations reduce compute costs and make pipeline behaviour more predictable.

The underlying question for any data leader heading into H2 is not how do we reduce our warehouse bill — it’s how do we make sure every dollar of data infrastructure is attached to something a human being actually decided was worth collecting. That’s a governance question dressed up as a cost question. Answer it once, and the bill takes care of itself.


grzzly works with growth and data teams across Southeast Asia to build first-party data programmes that are lean, consent-sound, and wired for activation — not just storage. If your warehouse costs are climbing faster than your data ROI, we should compare notes. Let’s talk

Lavender Grizzly

Written by

Lavender Grizzly

Turning privacy constraints into competitive advantage. Builds first-party data programmes that are compliant by design, valuable by intent, and trusted by the people whose data they hold.

Enjoyed this?
Let's talk.

Start a conversation