Skip to content

Master data after twenty years: the same problem, different tools

From SAP MDM on the NetWeaver stack to a master data layer you can stand up in weeks. Why most MDM programmes fall short, what CVI actually does before an S/4HANA conversion, and how to measure your own duplicates before you invite anyone in. A column by Jacek Bugajski.

A buyer orders a part that is already sitting on the shelf. Not out of carelessness: the material number exists twice in the system, under two names, and there is no reason for anyone to notice before the delivery arrives. The same business partner appears in the ERP under its full legal name, in the CRM as an abbreviation somebody typed by hand, in the purchasing system without a registration number, and in the warehouse with an address from before the move. Each of those systems is locally right. None of them is completely right.

I have been working on this problem for twenty years. In that time almost everything in the technology has changed, and nothing in the problem itself has. What has changed is how much you have to invest to solve it. That is what this text is about.

When the architecture diagram did not fit on two sheets of paper

I sold my first master data implementation back when everything was a stack. NetWeaver, with SAP MDM inside it, Portal next to it, PI next to that, every component with its own installation, its own expert and its own set of problems. The architecture diagram did not fit on one sheet of A4, and it did not fit on two either.

Several institutions in Poland bought that solution at the time and got exactly what they needed: one place where you settle who is who and what is what. I still consider that sale a good one, because the value was real. The price, on the other hand, was high, and it was not about licences.

To run a single function, the client’s team had to master ten technologies around it. Six months of the implementation went into learning the stack rather than into cleaning up data. During that time the project changed owners, sometimes sponsors, and occasionally the priorities of the whole organisation changed too. Anyone who has worked on implementations like that knows the data model is not the hardest part. The hardest part is holding the organisation’s attention for longer than two quarters.

Back then there was no alternative. Today there is, and that is the heart of the whole change.

Why this topic always lost the budget fight

Master data has a congenital flaw: you cannot demo it. There is no screen that impresses a board, no chart pointing upwards, no effect inside a quarter. On project plans it lands where documentation and performance tests land, marked “phase two”. Phase two rarely starts.

The cost of the delay does not disappear. It scatters across the organisation so that nobody sees it whole.

Diagram: the cost of inconsistent master data spread across five departments - finance reconciles balances by hand, procurement verifies the supplier before granting a rebate, the warehouse accepts a second delivery of the same part, controlling builds the report from three sources, IT maintains a script that merges two files

Each of those items looks trivial on its own. Together they are one of the most expensive pieces of work in the company, and at the same time the one nobody accounts for, because it is spread across a dozen people in several departments. No budget has a line called “reconciling data”, so formally the cost does not exist.

What changed in a single year

The market stopped treating this layer as back office, and not on the strength of declarations but of transactions.

Salesforce acquired Informatica; the deal closed on 18 November 2025. SAP acquired Reltio, closing on 7 May 2026, and put it plainly in the announcement: the point is to make SAP and non-SAP enterprise data ready for artificial intelligence. In April 2026 Gartner brought back the Magic Quadrant for Master Data Management Solutions after a five-year gap, and analysts do not return to an abandoned category without a reason.

Three independent signals in one year look like a repricing rather than a fashion. The master data layer has moved from the category of back-office cost to the category of operating precondition. For companies in this region that has one practical consequence: a topic you previously had to explain internally from scratch now has external justification at board level.

Artificial intelligence raised the stakes

The reason for that repricing is practical, and you see it in every project where an agent touches operational data.

An agent works with whatever it is given. Faced with four versions of the same business partner it will answer immediately and with complete confidence, including when the answer is untrue. It will not ask which version is the correct one, because it has no way to settle that. A language model does not distinguish incomplete data from complete data; it only distinguishes data it has from data it does not have.

Gartner forecasts that through 2026 organisations will abandon 60 per cent of artificial intelligence projects that are not supported by AI-ready data (Gartner press release, 26 February 2025). That forecast is not about models. It is about the fact that automation does not tidy data up; it increases the speed and the reach of whatever is already in it.

The same applies to the automation we build ourselves. A robot that copies data faster than a person will also copy an error faster and into more places. That is why, on every agentic project, we ask about the state of the master data before we agree the scope.

Why most MDM programmes fall short

Gartner estimated that through 2025 more than 75 per cent of Master Data Management programmes would fail to meet business expectations. That figure says less about tools and more about how projects are run. I have seen three recurring variants of the failure.

Variant one: scope without a boundary. It starts with a decision that sounds reasonable: since we are cleaning up data, let us do it across every domain at once. The scope requires agreement between departments, so a committee is formed. The committee appoints a working group. The working group runs workshops on field naming that last longer than building the solution would have. After a year there is no result, there are slides, and there is a candid internal opinion that “MDM does not work here”.

Variant two: cleansing without rules. Apparently cheaper and faster. The company orders a one-off clean-up of the master file, receives a clean file and closes the project. Nobody has changed the way data enters the system, however, and nobody is accountable for its correctness. A quarter later the file is back where it started, and the organisation has proof that “this cannot be sustained”.

Variant three: an owner without a mandate. The project has a sponsor in IT and nobody on the business side who can settle a disputed record. Every doubt goes back to the committee, the committee does not want to carry responsibility for the consequences of the decision, so the records stay in the queue. The queue grows and after a while becomes an argument against the project in its own right.

The common denominator of all three: no result that can be measured before the work and after it. If you do not know what is supposed to improve and by how much, there is no way to defend the project halfway through, and every data project gets attacked halfway through.

How we got into this: from GDPR and DORA to a product

At SNOK we have been working on master data for three years, and we did not start from an MDM strategy or from a product idea. We started from a question clients kept bringing us: where in the source systems does specific data live, who has access to it, and how do we demonstrate that to an auditor when GDPR or DORA comes calling.

It sounds like a week’s work. In practice it means going through every master file, every copy of the same records, and every export to a spreadsheet that somebody once made “just for now” and that a process then settled on permanently. We did this several times and at some point we noticed we were building the same mechanism over and over: record comparison, change history, an audit trail, a decision queue for the business.

That was the moment it stopped being a set of scripts and became a product. That is how SNOK MDM came about: out of an audit requirement, not out of a vision deck. I consider that a better origin than most product roadmaps I have seen, because the first user was real and had a deadline.

Proof from an implementation. In the employee master data domain we work with Medicover: 18 countries, roughly 45,000 employees and around 100,000 employee data events a year handled by the platform. That implementation covers this one domain and that is how we describe it, without stretching it to business partners, materials or finance, which we clean up for other clients.

What a golden record is and what it will not do for you

A golden record is a single master record describing a real object: a specific business partner, material, product or employee. It comes from rules, which means normalising the way values are written, validating critical fields and resolving conflicts between sources, rather than from agreements reached in a meeting. It has an owner on the business side, a change history and an audit trail. The source systems keep running; nobody switches them off or rebuilds them.

Alongside it sits the reference data layer, RDM for short. Put simply: MDM answers the question of who and what, while RDM provides the shared language, meaning code lists, classifications and organisational units. Without that second layer, master records get described differently in every subsidiary and the reports still diverge, only at a more advanced level.

Independence from the ERP and the CRM. The master layer does not have to live inside a transactional system. When it stands separately, changing the ERP, rolling out a new CRM or acquiring a company does not invalidate the order in your data. In groups where the number of systems runs into three figures, that is the difference between a project and a permanent renovation.

A human decision on merges. Language models recognise variants of the same company name accurately and can justify the suggestion. The result is probabilistic, though, so a data steward approves the merge. Automatically joining two separate legal entities is the kind of error that surfaces months later, usually during debt collection or an audit, and it is harder to repair than a duplicate. A machine suggestion plus human sign-off is currently the best ratio of quality to speed.

What a golden record will not do. It will not replace the business decision about which variant of a name is binding. It will not fix a process in which five people can create a new business partner without validation. It will not change a situation in which nobody is accountable for data quality. The technical layer enforces order only where somebody has defined that order.

Six domains and the order worth cleaning them in

In practice the whole MDM conversation covers six domains: business partners and customers, suppliers, materials and item numbers, products, employees, and financial accounts. There is no single correct order for everyone, but there is a question that sets it: which domain generates the most manual work and the most arguments about numbers today.

Two identical cardboard boxes on a warehouse rack - the image of a duplicated material number: the same part ordered twice, because it exists in the system under two names

A few typical calls from projects:

  • A group after acquisitions usually starts with business partners and suppliers, because that is where consolidated reporting hurts and where the effect shows fastest at group level.
  • Manufacturing and maintenance start with material numbers, because a duplicated item number means ordering a part that is already in the warehouse, and downtime when the right part is not.
  • An organisation ahead of an S/4HANA conversion starts with business partners, because their condition translates directly into the number of exceptions during migration.
  • A services company with distributed HR starts with employees, because that is where regulation is most demanding and the volume of events highest.
  • A financial institution under DORA starts with whatever lets it answer the auditor’s question: where is the critical data and who has access to it.

There is one rule: the first domain has to have an owner, a measurable success criterion and real pain in the process. A domain chosen because it is “technically the simplest” will not convince anyone to fund the second one.

Before the S/4HANA conversion, not after it

If you are looking for the moment when this work pays back fastest, today it is set by the migration timeline.

In S/4HANA the Business Partner model is mandatory. Customer and vendor stop being separate objects maintained independently, and Customer Vendor Integration performs the move to the new model. It is a precondition of the S/4HANA conversion, not an option to consider in a later phase.

CVI’s job is to move data into the new structure. Improving its quality is not its job and does not happen by itself. In practice that means three things:

  1. Duplicates carry over. If the same business partner exists in three variants, the conversion will not join them. Three records stand every chance of becoming three business partners, unless somebody merges them first.
  2. Gaps stall records. A record missing a field the new model requires comes back to the team as an exception for manual handling. The extent of those exceptions depends on configuration and mapping, which is why you establish the scale by measuring your own master file rather than by assumption.
  3. Rulings need the business, not IT. The question of whether two similar records are the same entity requires knowledge from procurement, sales or finance. During the cutover window those people have other work, and the exception list does not wait.

Diagram: the order of work before an S/4HANA conversion - measuring duplicates in weeks 1-2, deduplication with data steward sign-off in weeks 3-8, and only then Customer Vendor Integration in the cutover window

The whole thing can run in parallel with preparing the S/4HANA project, without extending the timeline by a single week. On the other side, after go-live, the same work costs more and gets done under production pressure.

What the regulations require

Three obligations touch master data directly, and none of them waits for an internal timetable.

GDPR, Article 5 requires personal data to be accurate. The data of people representing business partners is personal data, so the business partner master file is also a set of personal data, with an obligation to keep it up to date and to be able to satisfy data subject rights.

KSeF, the Polish national e-invoicing system, introduces a structured invoice with hundreds of fields. Not every error in partner data causes a technical rejection of the document, but an inconsistent master file raises the risk of rejections and practically guarantees manual exception handling, in a process where nobody planned for that work. It is worth checking the deadlines and the scope of the obligation at source, because they have changed several times.

DORA requires financial entities to demonstrate where critical data resides and who has access to it. Without a master layer and without a record of data lineage, that is a manual inventory before every audit, repeated from scratch. Those are exactly the questions that started our road to a product.

One domain instead of a year-long programme

Our method follows directly from watching why large programmes fall short. Five steps, in this order.

Step one: pick the domain and the success criterion. One domain and one sentence that can be checked. Not “we will improve data quality”, but for example “the number of business partner records describing the same entity drops below an agreed threshold, and new records pass registration number validation”.

Step two: measure the current state. On the organisation’s real data, not on a market benchmark. The number of duplicates in the domain, the share of records with gaps in critical fields, the number of discrepancies between systems. That measurement turns the scope conversation from an exchange of opinions into a decision based on a number, and it is the only argument that genuinely works on a board.

Step three: the master record model and the rules. Normalisation, rules for resolving conflicts between sources, versioning, a data owner on the business side, the set of mandatory fields. This is the stage that decides whether the order survives the project.

Step four: a pilot with human oversight. Standing up the master layer for the chosen domain, a merge queue with reasoning, sign-off by a data steward, feeding the target systems. We usually close a single-domain pilot in four to eight weeks, depending on the state of the data and the number of systems to feed.

Step five: metrics and extension. Comparison against the baseline measurement, a data quality report as part of the routine, and only then the next domain. Without a confirmed result we do not extend the scope, even if the client wants to.

What stays on the client’s side is what cannot be outsourced: knowledge of their own business and decisions on disputed cases. The technology stack stays on ours, and that is the only change compared with the projects I ran twenty years ago.

How to measure your own duplicates before you invite anyone in

You can run this measurement yourself, and it is worth doing, because it gives you a reference point in every conversation that follows. Four steps on one domain, for example on the business partner master file.

  1. An extract from one system. Name, legal form, registration number, address, creation date, active status. No integration, an ordinary export.
  2. Normalisation for comparison. Strip legal forms, punctuation and letter case from the names, standardise how addresses are written, clean spaces and dashes out of registration numbers.
  3. The tests that are enough to see the scale. How many records share an identical registration number under different names. How many share an identical normalised name under different numbers. How many have no number at all.
  4. A sample for manual verification. Fifty random pairs from the results and a human decision: same entity or not. The hit rate shows how many of the automatically flagged cases are real.

Our projects suggest a simple rule of thumb: once the result exceeds a few per cent of the master file, you have a case for a separate project. If you would rather not do it yourself, we run the measurement on a sample from one domain and hand over the result within five working days, without access to the production environment, on data that may be pseudonymised.

When MDM is not the answer

Not every data problem calls for a master layer.

If you have one system and one master file, and the problem is a lack of validation on data entry, you need rules in that system, not a new platform. If the data is consistent and the reports diverge, the problem sits in metric definitions and in the analytical layer. If the organisation has nobody who can settle a disputed record, implementing MDM will produce a decision queue nobody works through; in that case the first step is to establish ownership, not to run a technical project.

A master layer makes sense where the same object lives in several systems at once and where somebody is prepared to take responsibility for which version is binding.

What has not changed

After twenty years the problem is the same. The same business partner appears in four systems in four versions, a buyer orders a part that is sitting on the shelf, and the management report is produced in a file called “final version revised 3”.

What has changed is how much you have to invest to fix it. Back then it took a programme, a technology stack and six months of learning. Today it takes one domain, one success criterion and a few weeks of work, and the result is visible before the project changes owners.

This week and in the weeks to come I am writing more about this: what exactly happens to the business partner master file before an S/4HANA conversion, what measuring duplicates looks like step by step, and why one domain beats a four-quarter programme.

Frequently asked questions

How does MDM differ from a data warehouse or a lakehouse? A warehouse and a lakehouse are responsible for collecting, modelling and analysing data. They do not settle which business partner record is the correct one or who is accountable for its correctness. Modern data stack platforms such as Snowflake or Microsoft Fabric have no built-in master data layer; you bring one in separately, through your own configuration or a partner solution. Without it, analytics consolidates inconsistencies instead of removing them.

How long does an MDM implementation take? A single-domain pilot usually closes in four to eight weeks. Classic programmes are planned over a horizon of several quarters to several years. The difference is not the pace of work but the order: first a confirmed result on one domain, then extending the scope.

We already have SAP MDG under licence. Why would we need anything else? If MDG covers your need, we recommend MDG and help you use it, because that is part of our SAP practice. The question is whether the scope also covers data from non-SAP systems, and whether the layer should stay independent of the ERP when the system landscape changes. If so, you need a layer outside SAP.

Does the data have to leave our infrastructure? It does not. SNOK MDM can run in the client’s environment where the security policy or the IT architecture requires it, and the source code can be placed in escrow for business continuity purposes. Processing that uses language models can be run locally or in the client’s cloud.

What does security look like when language models are used for deduplication? The model suggests that two records describe the same entity and justifies the suggestion. A data steward approves the decision, because the result is probabilistic. The scope of data passed to the model is limited and described in the data processing agreement, and the processing can take place without data leaving the environment.

What drives the cost of a project? The number of domains in scope, the number of systems to read from and feed, the state of the input data, and the requirements around data governance and the audit trail. We bill per phase with a defined scope and a defined deliverable. We set the price after the measurement, because before it any figure would be guesswork.

How do we measure the effect after the pilot? Against metrics agreed before the start: the number of duplicated records in the domain, the share of records with gaps in critical fields, the number of discrepancies between systems, and the time it takes to process a change in master data. We compare the result against the baseline measurement.


Jacek Bugajski is the CEO of SNOK, a Polish IT consulting company specialising in SAP, cybersecurity, automation and data.

Sources

Found this useful? Please pass it on:

Get in touch