
Data unification started with monolithic architectures like data warehousing and later on with data lake. Implementing these tightly coupled architectures takes time, and they are not easy to maintain as well. With the current pace of data growth and competition, time to market becomes a key factor. Companies have adopted agile methods to develop software; however, the critical issue is how we can implement agile methodologies for data projects and create a loosely coupled architecture that is scalable and easy to maintain.
With all these modernization projects, we also need to modernize the underline outdated architecture and processes. Again, scaled architecture can be a good candidate for it. I will go over briefly how you can implement this architecture; for detailed understanding, you can read the Domain-Driven Design book by Eric Evans and Data management at Scale book by Piethein Strengholt.
As with any data integration/unification project, you commence with your source data and manufacture it to create a curated data layer. We use the schema-on-read method in the data lake, and for the data warehouse, we use the schema-on-write method. With a data warehouse, you use dimensional modeling and create fact and dimension tables.
With scaled architecture, you work with your business architect to define the business operation and business domains. This operation can act as a bounded context, a logical group for your business domains. Now to design domains, use business workflow, and create domains supporting that. Domains can also have sub-domains. Now we are not talking like an object model or highly normalize model here. Instead, domains can be denormalized and becomes a data product.
For large systems, you can have each domain owned by a team; this way team will have a deeper understanding of the business domain, data, and fewer dependencies on other tasks or teams. Each domain team can work independently using an agile process and reduce time to market with clear ownership of data. Since data is in domains, it is accessible to implement APIs, streaming, or any std ways of accessing it to create data as a service.
As you can see, Scaled Architecture helps to deliver data products faster, defining clear data ownership, aligning with business capabilities, and implementing data as a service.