Indonesia Singapore ไทย Pilipinas Việt Nam Malaysia မြန်မာ ລາວ
← Back to Blog

Precision Data Collection: Why Signal Quality Beats Volume

Signal precision beats data volume — audit what you collect before expanding what you store, or your CDP will choke on its own noise.

Editorial illustration of a data architect sorting signal from noise in a customer data pipeline
Illustrated by Mikael Venne

Collecting more data rarely solves the real problem. Here's how precision signal architecture and smart data pruning build CDPs that actually perform.

Most CDP implementations fail before a single segment is activated. Not because the platform was wrong, or the integration was botched — but because the data going in was contaminated at the source. Garbage in, expensive garbage out.

The Recycling Bin Problem in Data Collection

Tealium’s analysis of BBVA Technology’s data strategy uses a quietly brilliant analogy: collecting data badly is like sorting recycling incorrectly. Get it wrong, and you don’t just waste the bad data — you contaminate the good data around it. BBVA’s approach to solving this was architectural, not cosmetic. Rather than layering more collection tags on top of broken signal structures, their team rebuilt around precision signals — defining exactly what behavioural and transactional events matter, and ensuring those events are captured cleanly, consistently, and in a format that downstream systems can actually use.

This distinction matters enormously in Southeast Asia, where brands often operate across Shopee, Lazada, their own web properties, and a LINE official account simultaneously. Each platform emits its own event taxonomy. Without a disciplined signal definition layer sitting upstream of your CDP, you end up with four incompatible versions of “product view” and no reliable way to stitch them into a unified profile. BBVA’s lesson: define what a signal means before you decide how to collect it.

More Context Isn’t Always More Intelligence

There’s a parallel problem emerging on the AI side of CDP infrastructure that most teams aren’t thinking about yet. As brands pipe customer data into LLM-powered personalisation or segmentation tools, the assumption is that richer context produces smarter outputs. Towards Data Science’s analysis of prompt engineering in production systems challenges that directly: LLMs don’t fail because they lack information — they fail because they’re given too much of the wrong kind.

Research into prompt-pruning architectures shows that as conversation history and context windows grow, redundant and low-value tokens accumulate. The result isn’t neutral — it actively degrades output quality while increasing cost and latency. The fix isn’t a bigger context window. It’s a deterministic pruning layer that strips noise while preserving dependency chains.

Applied to CDP strategy, this is more than an analogy. If you’re feeding behavioural data into AI-driven segmentation, recommendation engines, or next-best-action models, the quality of what you send matters more than the quantity. A unified customer profile bloated with duplicate events, misfired pixels, and ambiguous session data will produce worse model outputs than a leaner, well-governed dataset — even if the leaner set contains fewer raw signals.


Building the Signal Governance Layer Teams Skip

The practical fix sits upstream of both your CDP and any AI layer sitting on top of it: a signal governance framework that most brands skip because it’s unglamorous work. Here’s what it actually involves.

First, build a signal registry — a documented taxonomy of every event your business cares about, with agreed definitions across product, marketing, and analytics. “Add to cart” should mean the same thing whether it fires on your Shopee store, your Grab merchant page, or your owned app. Second, implement validation rules at the collection layer itself — reject or flag events that don’t conform to schema before they enter your CDP. Tealium’s approach with BBVA included exactly this kind of upstream quality control, treating data collection as a discipline with defined standards, not an instrumentation free-for-all.

Third — and this is where teams consistently underinvest — build signal pruning logic into your CDP ingestion pipeline. Not every event that passes validation should persist at full fidelity forever. High-frequency, low-signal events (scroll depth on a page someone bounced from in 0.8 seconds, for instance) can be aggregated or discarded without losing meaningful profile intelligence. What you preserve should be intentional, not accidental.

For Southeast Asian brands managing multilingual interfaces — Bahasa Indonesia, Thai, Vietnamese, and English often coexisting in a single app — this governance layer also needs to handle language-variant events consistently. A product detail page view should resolve to the same signal regardless of which locale triggered it.

The Licence Fee Question Nobody Asks Upfront

CDP contracts are large. Most are renewed on the assumption that the platform is working because data is flowing into it. The harder question — whether the right data is flowing in, and whether downstream activation is producing measurably better outcomes — rarely gets asked until renewal, when it’s too late to fix the architecture.

BBVA’s signal precision work is a useful benchmark precisely because it treats data quality as a first-order business problem, not an analytics team concern. When your CDP’s unified profile is built on clean, governed, pruned signals, segmentation becomes more reliable, personalisation becomes more defensible, and the models sitting on top of your data stop hallucinating patterns that don’t exist.

The brands earning their CDP licence fee in 2026 aren’t the ones with the most data. They’re the ones who decided, clearly and early, what a signal actually means — and built infrastructure to enforce that decision at every collection point.


Key Takeaways

  • Define your signal taxonomy before you instrument — ambiguous event definitions contaminate profiles faster than missing data ever will.
  • Apply upstream validation and pruning logic at ingestion, not as a retrospective data-cleaning exercise; the BBVA model demonstrates this works at enterprise scale.
  • If AI-driven segmentation or personalisation sits on your CDP, treat signal quality as a model performance input — more context actively degrades outputs when that context is noisy.

The uncomfortable truth is that most CDP roadmaps prioritise more — more integrations, more signals, more AI features — when the actual constraint is signal discipline. The question worth sitting with: if you audited every event type flowing into your CDP today and graded each one on whether it meaningfully improves profile fidelity, how many would you keep?


At grzzly, we work with growth and data teams across Southeast Asia on exactly this — auditing what’s flowing into their CDPs, rebuilding signal governance frameworks, and making sure the unified customer profile is actually unified. If your platform is live but your activation results don’t reflect it, that’s a conversation worth having. Let’s talk

Velvet Grizzly

Written by

Velvet Grizzly

Architecting the unified customer profile — stitching together behavioural, transactional, and declared data into platforms that actually earn their licence fee.

Enjoyed this?
Let's talk.

Start a conversation