Data Product Heuristics

How to decompose data into data products of the right size

How big should a data product be? There is no single right answer. However, there are good questions that can help you find good boundaries. This page presents 15 heuristics for sizing data products: Five fundamentals that apply to every data product, and specific heuristics for each of the three archetypes – source-aligned, aggregate, and consumer-aligned.

Heuristics are guiding questions, not rigid rules. Use them to start the right conversation.

The heuristics are based on the article The right size of a Data Product . We also provide a one-page checklist with all the guiding questions.

Why boundaries matter

When data products have the wrong boundaries, the same problems show up again and again.

Stitching
Consumers combine several data products to answer simple questions.

Ownership vacuum
Nobody is responsible for meaning, quality, or changes.

Duplicate logic
Teams rebuild the same logic, which is costly and inconsistent.

Bloated data products
Data products become overly complex and difficult to understand and maintain.

Incidents as routine
Failures become normal instead of rare.

Example: an online shop

We use the same example, an online shop, to explain each heuristic. Throughout the customer journey, there are four domains, each with its own operational system. On the other hand, the business functions of Finance, Operations and Purchasing want to use the data.

The online shop. Four domains along the customer journey, each with its own system: Order with the Order Management System, Catalog with the Catalog Service, Fulfillment with the Warehouse Management System, and Shipping with the Shipping Service. Below them, three business functions that use the data: Finance, Operations, and Purchasing.

Fundamentals

The five fundamentals apply to all data products, regardless of their archetype.

#1 Clearly defined consumer & use cases

Describe its main purpose in one sentence.

A data product requires real consumers with genuine requirements. Without a use case, there can be no data product. A data contract can make the expectations explicit; the Data Product Canvas helps plan these from the early stages.

Example: "The purchasing team needs replenishment signals from inventory and product data." This sentence leads to restock_alerts. The source-aligned data products it uses have their own purpose: stock_movements "shows every stock change per product and warehouse" and product_catalog "provides current product information such as names, categories, and prices". Purchasing is one of their consumers.

Guiding questions

  • Can you describe the main purpose in one sentence?
  • Are there any specific teams or roles that want to use this data product right now?

#2 Stable ownership

One business domain, one team.

One team is responsible for meaning, quality and operation. Without a stable owner, there can be no data product. The team that triggers a process is not necessarily its owner.

Example: The Order Management System triggers shipments, but the shipping team owns shipments, including the business logic and the software systems behind it.

Guiding questions

  • Is one specific domain or team accountable for semantics, quality, and operations?
  • Would the owner credibly handle future changes?

#3 Consistent data quality

Quality attributes consistent across all output ports.

All output ports of a data product must share the data quality characteristics, except for technical reasons caused by the technical characteristics of the underlying technologies. Inconsistent semantics between output ports are never acceptable.

Example: stock_movements is offered as both Kafka stream and a daily batch export. Only the freshness differs. Both ports have the same accuracy (the quantities are correct), the same completeness (every stock movement is included), and the same conformity (same schema, units, and product identifiers).

Guiding question

  • Are data quality attributes consistent across output ports?

#4 Low integration burden vs. #5 Bounded scope

Smallest useful standalone unit.

A scale from high integration burden on the left to bloated scope on the right. On the left, order line items alone are too granular. In the middle, orders, product catalog, and stock movements are each the smallest useful standalone unit. On the right, orders and shipments combined in one data product are too broad.

Aiming for low integration and staying within a bounded scope are two opposing heuristics. If it's too small, consumers have to stitch data products together. If it's too big, the scope becomes bloated. Aim for the sweet spot: the data product must be valuable on its own.

Example: order_line_items would be too specific to provide value without the order context. orders captures the state at time of purchase, so no join with live data is needed. product_catalog and stock_movements share product_id, but stay separate on purpose. orders_and_shipments is too broad because it is too generic to be of direct use to end users, while spanning two business domains.

Guiding questions

  • Is this the smallest useful standalone unit that does not force consumers to stitch data products together?
  • Can a typical consumer immediately start using this data product meaningfully on their own?
  • Does the data product include only what is needed for its purpose?
  • Is it limited to only including things that are useful in the present?

Source-aligned data products

Source-aligned data products form the basis of all our data use cases. They provide domain-specific data that is close to the operational truth, remaining within a single business domain.

#1 Semantic coherence

All data required for interpretation.

Three options compared. Too fragmented: dates, states, and totals as separate data products. Coherent: one orders data product with order lines, status, and currency and tax. Mixed meanings: one shop data product that contains orders, returns, and reviews.

The data product must be self-explanatory, with each concept having exactly one meaning. This does not imply that it supports every analysis, merely that the data itself can be interpreted. This emphasises the importance of fundamental #4.

Example: orders with order lines, status, and currency & tax is coherent. However, splitting off dates, states, and totals renders it unusable. Merging returns and reviews into shop_data makes fields such as date or status ambiguous.

Guiding questions

  • Does the data product make sense on its own, or does it require the other parts of the source data?
  • Does it feel like a cohesive, integrated whole rather than a random collection of related items?

#2 Single business domain

One module. Not all data.

Create boundaries between business domains rather than entire systems. Different topics require different areas of expertise, and people may not want access to all the information available in one piece of software. The decomposition should not be technology-driven.

Example: product_catalog and customers contain distinct data, with both coming from SAP. Creating an all_sap_data data product would result in a messy, hard-to-maintain data set with unclear responsibilities.

Guiding questions

  • Does the decomposition follow meaningful domain modules rather than whole systems?
  • Does the data contain only internal or also cross-domain context?

#3 Source data blast radius

Different source, different data product.

The counterpart to #2: limit the blast radius. Changes to a source dataset or a software application's business logic should not affect multiple source-aligned data products. Otherwise, maintenance can become significantly more time-consuming.

Example: A change to stock_movements from the warehouse system has no effect on product_catalog from SAP.

Guiding question

  • Do changes on the data source impact only this data product directly?

Aggregate data products

Aggregates combine data from several aligned data products and require strict governance. The first question should be "Do we need it at all?", not "How big should it be?" Keep the number of aggregate data products to a minimum.

#1 Value & reuse vs. #2 Cost & complexity

Build it when 3+ teams need it.

Using an aggregate avoids duplicate work and ensures consistent terminology. However, it also incurs costs, requires governance and an owner must be appointed. A good rule of thumb is to have three or more teams. What if there is only one consumer? Build a data product aligned to the needs of the consumer. Do consumers have completely different needs? Shift the effort towards consumer- or source-aligned data products.

Example: order_lifecycle combines orders, returns, and shipments into a single, comprehensive view of an order, from purchase to delivery and return. The finance, operations and customer service departments all need this view to have the same meaning.

Guiding questions

  • Are there more than two teams that need the same derived view with identical meaning?
  • Would teams repeatedly build the same integration or calculation without the aggregate?
  • Does value emerge only after combining sources?
  • Is the derivation expensive (feature engineering, entity matching, deduplication, cross-source joins)?
  • Is there someone in the company willing to bear the costs of this data product?

#3 Scope & governance discipline

General-purpose. Strictly scoped.

Keep the scope tight to avoid it turning into a mini data warehouse. Logic specific to consumers should stay in data products aligned to consumers, and every addition increases the effort required for governance. Avoid feature creep, otherwise the data product will become expensive and unmaintainable.

Example: order_lifecycle contains only general-purpose, coherent content that is required by all its consumers consumers: combined order events, linked returns, and shipment dates. Logic specific to consumers is excluded. Only the finance department needs the net revenue calculation, and they can build it themselves. Late-delivery rules belong to the operations team, who may not even be familiar with them. Customer segments are specific to marketing and customer service.

Guiding questions

  • Is the scope tight enough so the data product is not drifting toward a mini data warehouse?
  • Is the outcome valuable enough to justify the required strong governance?
  • Can the owning team maintain the integrated semantics, despite spanning multiple sources?

Consumer-aligned data products

Consumer-aligned data products are designed for specific users and serve a specific purpose. They are most effective when consumers can use them without having to build their own integration layer.

#1 Clear purpose

One sentence. Verb plus object.

Three consumer-aligned data products, each with its purpose as verb plus object. Net revenue report: report monthly net revenue. Delivery performance: monitor delivery SLA. Restock alerts: trigger replenishment.

The data product should have a clear purpose and be intended for a specific audience. Additional details for further examination or interpretation are welcome.

Guiding question

  • Can this purpose be expressed as a verb + object sentence?

#2 Natural, focused size

Sized by the job – not the artifact.

One data product can generate several related outputs. If a data product is used in many different decision contexts, it should be split.

Example: delivery_performance feeds an SLA dashboard, a carrier report, and a reverse-ETL feed.

Guiding question

  • Does the data product support a single decision context that results in one or more related dashboards, reports, or outputs?

#3 Meaningful boundaries

Process lines, not system lines.

Don't define boundaries based on where the data comes from. The data product reflects consumer behaviour and decision-making.

Example: order_to_cash can be used right away because it follows the process across the shop, ERP, and payment systems, whereas payment_gateway_data would only reflect information from one system. While this is not a problem in itself, on its own it is probably not useful for consumers. First, they would need to combine it with shop and ERP data.

Guiding questions

  • Does the boundary follow a process and not a system boundary?
  • Does the decomposition reflect how a consumer acts or decides, not how data happens to be stored?

#4 Business consumers

Name the team. Then the data product.

You have to specify who will use it, and then include only what they need for their job, not everything that might be useful someday.

Example: Finance uses net_revenue_report for the monthly close. analytics_data for "Finance, Marketing, ad hoc, and future ML" is an untargeted catch-all that is almost certainly bloated while also containing unused data, which has a negative effect on documentation and maintenance.

Guiding question

  • Will your data product be used by business users, data analysts, data scientists, or applications? Who exactly, and for what?

The whole picture

When applied to the online shop, the heuristics result in nine data products. However, not every data product aligned with consumers needs an intermediate aggregate: restock_alerts builds directly on source-aligned data products.

We acknowledge that we disregarded our own three-consumer rule of thumb for the sake of the example's simplicity.

Nine data products for the online shop in three layers. Source-aligned: orders, returns, shipments, product catalog, and stock movements. Aggregate: order lifecycle, built from orders, returns, and shipments. Consumer-aligned: net revenue report and delivery performance, both built on order lifecycle, and restock alerts, built directly on product catalog and stock movements. On top, the consumers: Finance uses the revenue report, Operations the SLA dashboard, and Purchasing the reorder feed.

Author

Stefan Negele works as a senior consultant at INNOQ. He has worked as a software developer and architect since 2012. His focus is on data architectures, data governance, and distributed systems.