Progressive Data Modeling is a modern approach that enables enterprises to deliver analytics iteratively through horizontal slices of end-to-end pipelines, balancing rapid time-to-value with sustainable data architecture foundations.
One of the biggest hurdles for an enterprise to start a new initiative around data and insights is expediting time-to-market by positioning the data to answer business department questions and provide insights.
In the current changing dynamics of business, it's not easy for a program sponsor to get approved for substantial budgets without providing practical benefits to business departments. CXOs typically anticipate comprehensive feedback and measurable outcomes within a quarter, reviewed during Quarterly Business Reviews (QBRs).
Hence, program sponsors must showcase progress measured in department onboardings and analytics value realization. The actual change in department productivity, efficiency, and feedback to enterprise leadership serves as the real testimony to budget allocation.
If a program with these fast-paced needs takes a traditional route for modeling data all at once before starting analytics, it would neither sustain beyond a couple of QBRs nor help the organization at scale.
A new data modeling approach is needed that preserves Data Model benefits as the foundation, provides a Build-as-you-go Data Model, and delivers RAD (Rapid Application Development) chunks to the Analytics wing for realization and release to end users. This approach is coined as "Progressive Data Modeling".
Progressive Data Modeling expects the data solutions team to work in tandem with different program segments. Starting with business analysts, data engineers, and data analysts, continuous iterations of analytics are delivered to end users through horizontal slices of end-to-end pipelines.
Progressive Modeling starts with Business Analysts discussing with department users to understand their daily activities, pain points, and opportunities for improvement and analytics. The data engineering team, in collaboration with BAs, understands the data required for such requirements and works on bringing that data into the data ecosystem until ready for modeling.
Data is quickly modeled by the tag team of BA and DEs to join base entities (OLTP) and form a resultset for validation. This layer can be developed in SQL or a Viz tool modeler like Tableau Prep.
Tableau Prep models serving workbooks are smart choices for quick changes and expedited workbook development while adhering to business requirements. The final model developed in Tableau Prep can be reverse-engineered to form queries for loading OLAP entities.
If developers prefer SQL, the same can be achieved in a temporary layer through combinations of joins and native SQL clauses. The BI developer can then pick up the data, realize it in the dashboard as a prototype, and present to end users.
This process's iterative feedback and prototyping loop provides quick results to users and shapes their vision of expected analytics. A data platform like Snowflake Data Cloud fits best in such scenarios, with numerous native features including Unistore, a unified data platform framework for OLTP + OLAP.
The second step processes knowledge gained from the temporary analytics layer and, along with enterprise veterans, designs EDW (Enterprise Data Warehouse) level entities. These entities form a foundation for the data model like strong RCC of a building.
Frequent data model review calls with Data Engineers, Architects, BAs, and veterans help design the data model perfectly.
A typical model layer architecture includes:
Analytics – Another layer of downstream exposure focusing on requirement specifics and data sourcing from Core Analytics. Can hold second-degree aggregations.
Core Analytics – A downstream exposure layer at detail level of entities, joined as required, or aggregated to first-degree.
EDW – The bottom layer with strong entities.
This step provides opportunities to govern the model and all related changes using governing tools or custom processes, preserving change history and introducing better control regarding data security, stewardship, and performance.
The above two steps form a defined, neat model for Analytics to migrate to, providing a canvas to paint the analytics picture and tell stories to gain insights. The analytics team leverages reusability benefits of Core Analytics and Analytics layer views formed from larger entities, catering to multi-attribute needs of similar granularities that can be rolled up as required.
The above process can be followed in an iterative loop:
1. Assess
For every new requirement, the data team should first assess whether it can be satisfied from existing objects.
2. Expand
If requirements can be greatly satisfied with base object entities but are missing attributes, such entities can be expanded. A caution to follow is validating Horizontal expansion vs Vertical expansion.
Horizontal expansion is adding attributes by selecting additional attributes of existing base tables or joining new base tables. When joining additional base tables, existing granularity and record count should remain intact without resulting in vertical expansion.
In case of inevitable vertical expansion, the data model review committee should make calculated decisions guaranteeing that changes will be absorbed without affecting end reports.
3. Design
If requirements cannot be satisfied with existing entities or vertical expansion is beyond absorption, a new entity can be created and parsed through the Progressive Model process.
This process can be efficiently implemented using transformation tools like DBT in the form of a metric layer. Leveraging GIT integration guarantees a versioned model and increases deployment speed following a Data Ops approach.
A powerful combination of tools like DBT and platforms like Snowflake Data Cloud offers fast GTM data models, answering varied business questions with velocity.
Despite the merits of progressive data modeling, occasions and scenarios exist where traditional Data Models will be more beneficial and appropriate. This choice can be influenced by factors such as:
One such scenario is a data model around a product serving as an application with a defined set of entities and serving as a hybrid OLTP and OLAP model. However, it's worth evaluating the choice based on mentioned factors.
Progressive Data Modeling emerges as a modern solution adept at addressing complexities of accelerating data-driven initiatives in today's dynamic business environment. This approach ensures data models evolve in tandem with business needs by facilitating continuous and iterative development, thereby providing timely and actionable insights. The emphasis on RAD and horizontal slices of end-to-end pipelines allows organizations to quickly showcase tangible progress and value realization, crucial for maintaining executive support and securing budgets.
By integrating Business Analysts, Data Engineers, and Data Analysts, Progressive Data Modeling fosters a collaborative ecosystem that enhances productivity and efficiency across departments.
While Progressive Data Modeling offers significant advantages, it is important to recognize scenarios where traditional data models might be more suitable, particularly for well-defined, stable data environments. Ultimately, the choice of data modeling approach should be informed by specific needs, urgency, and adaptability of underlying data infrastructure within the enterprise.
By embracing Progressive Data Modeling, organizations can more effectively navigate challenges of modern data initiatives, ensuring they remain agile, responsive, and competitive in an ever-evolving marketplace.