In the space of twelve months Salesforce paid around USD 8 billion for Informatica, SAP acquired Reltio, and Gartner brought back its Magic Quadrant for Master Data Management after a break of more than four years. The market put a multi-billion price tag on a category that had spent years being treated as plumbing. Over the same period EY found that only 9 percent of medium and large companies in Poland have a data infrastructure ready to feed AI models.
We put that gap into a report: “Master data in Polish enterprises” - 28 pages, ten chapters, every figure with a named source and a date. It is written from the Polish market, which makes it useful well beyond Poland: the country is running the largest mandatory e-invoicing rollout in the European Union right now, and the rest of Europe is next in line under ViDA.
One company, four versions of the truth
Start with a scene we know from projects rather too well.
An accounts specialist issues an invoice to a long-standing business partner. In the ERP that partner carries the full legal name and the tax identification number from the day the contract was signed. In the CRM a salesperson saved an abbreviation, with the address of an office the company left two years ago. In the purchasing system the partner exists a third time, with no tax number at all, because that was quicker. Each version is locally correct. None of them is complete.
Until January 2026 that situation mostly cost patience: reconciling balances, clarifying payment terms, assembling a board report from three sources. The cost was real, but spread across enough people and departments that it never had a line of its own in any budget. Today the same discrepancy has both an addressee and a deadline.
Reason one: mandatory e-invoicing checks your records every working day
KSeF (Krajowy System e-Faktur, Poland’s national e-invoicing system) became mandatory for large taxpayers in February 2026. Between 1 February and 26 May more than 289 million e-invoices passed through it, from over 2 million issuers - figures from the Polish Ministry of Finance. It is the first mechanism in the history of the Polish economy that tests the customer records of more than two million taxpayers simultaneously, every working day.
It tests them differently than many companies assume. KSeF validates that the file matches the schema. It does not check whether the buyer’s tax identification number actually belongs to the party the transaction concerns. An invoice with a number off by a single digit still receives a KSeF number and is delivered to whichever company holds that number; if no such number exists, it reaches nobody. That position was set out by the Director of the KIS (Krajowa Informacja Skarbowa, the national tax information authority) in an individual ruling issued in March 2026.
The repair path changed as well. Buyer-issued correction notes no longer exist, so a buyer can no longer fix somebody else’s mistake with a single document. A wrong tax number means the issuer must correct the invoice to zero and issue a new one. For one mistake that is an inconvenience. For a customer file holding hundreds of stale records and thousands of documents a month, it becomes a permanent exception-handling process that somebody has to run by hand.
The first months confirmed this is not a marginal problem. Tax advisers described EU VAT numbers entered in place of domestic tax numbers and documents circulating twice; according to data from one software vendor, 18 percent of correction invoices retrieved from KSeF contain errors. Treat that last figure with care - it comes from one vendor’s customer base - but the direction is unmistakable. The penalty-free period ends with 2026. The mechanism stays for good.
Why this matters outside Poland: under the EU’s VAT in the Digital Age package, structured e-invoicing and digital reporting move across the single market in the years ahead. Poland is the working prototype, and the failure mode it exposes - clean schema, wrong counterparty - is not specific to Polish law.
Reason two: a migration calendar that will not wait
The second process comes with exact dates. SAP mainstream maintenance for ECC ends on 31 December 2027, and according to the February 2026 survey by DSAG, the German-speaking SAP user group, 54 percent of the organisations polled still run ECC or older Business Suite releases. Microsoft has announced the end of mainstream support for Dynamics GP at the close of 2029. In Poland, where Comarch and domestic vendors hold strong positions alongside SAP, the same moment arrives with a version change or a post-acquisition consolidation.
Every one of those operations is a data migration in practice. And here sits the trap we keep finding in project plans: a conversion moves data into a new structure, it does not improve its quality. Three variants of the same business partner will land in the new system as three separate records unless somebody merges them first. A record missing a field the new model requires comes back to the team as an exception to handle manually - usually during the cutover window, when the team has the least time and every hour of downtime costs the most.
Market numbers confirm this is not a paper risk. ISG reports that 58 percent of SAP migration programmes exceed budget and schedule, while 15 percent land inside both. The right order is the reverse of the one we see most often: first measure duplicates and gaps, then deduplicate with the data owner in the room, and only then migrate. The same work done after go-live costs more, because it happens under production pressure.
Reason three: AI exposed the state of the records faster than any audit
The third process is the youngest and moves quickest. Language models and agents do not check whether data is true - they answer on the basis of what they are given, with equal confidence on consistent and contradictory records. Ask an agent for a partner’s balance when that partner exists in four variants across your systems and the answer arrives instantly. It will not mention that it picked one of four candidates.
Automation built on inconsistent data does not remove the mess. It increases its speed and its reach.
Analysts have been saying so for two years: Gartner projected that through the end of 2026 organisations will abandon 60 percent of AI projects unsupported by AI-ready data. The Polish figures fit that projection without any stretching. AI technologies are used by 8.7 percent of enterprises here (Statistics Poland, 2025) against an EU average of 20 percent (Eurostat), and the complete AI-ready data infrastructure mentioned above is claimed by 9 percent of medium and large firms. Models are available off the shelf. Prepared data is not.
Poland in context: the systems are there, the order in the data is not
Put public statistics next to market research and a picture emerges that explains the daily experience of Polish teams rather well. The infrastructure exists and is growing: 40.5 percent of enterprises use an ERP system and more than half buy paid cloud services (Statistics Poland, 2025). The work done on data looks thinner: 24.5 percent of Polish firms perform data analysis, against 60 percent in Denmark (Eurostat, 2025).
The averages hide a sharp split. AI technologies are used by 42 percent of large firms, 15.6 percent of medium ones and just 6.1 percent of small ones (Statistics Poland, 2025). Large organisations experiment, mid-sized ones are only now arriving - and they are the ones who most often discover that the first obstacle is not the budget for a model, but the state of their own records.
The most telling numbers are the organisations’ own declarations. In a study by Algolytics and SW Research covering more than 700 companies, 73 percent admitted they lack the competence to make use of their data, and 66 percent do not use data in decision-making at all. Those are not numbers about technology. They are numbers about a missing layer between the systems and the decisions - the layer that settles which record is authoritative and who answers for its accuracy.
There is one more Polish thread rarely thought of in data terms: mergers and acquisitions. The Polish market recorded 330 such transactions in 2025 (M&A Index Poland), with the sale of a 49 percent stake in Santander Bank Polska to Erste Group as the largest. Every acquisition means merging partner catalogues, material indices and charts of accounts maintained under different rules. Public research measuring the scale of this in Poland does not exist - we say so plainly in the report - but from projects in capital groups after acquisitions, including a chemical group consolidating data from five companies across three countries, we know the pattern repeats: the first weeks of integration go on establishing which company uses which catalogue, and how the definitions of apparently identical terms differ.
Why clean-up programmes fail
“Let us run a big data clean-up project” has poor statistics behind it. Gartner projected that more than 75 percent of master data management programmes would fail to meet their business objectives. From our own projects and pre-sales conversations, the causes rarely sit in the technology. Four patterns recur: a scope covering every domain at once, a committee instead of a data owner with a mandate to decide, fully automatic record merging, and a one-off clean-up with no rules to police the quality of new entries. That last one is the most treacherous, because it produces an effect for a quarter - after which the file returns to its original state and the organisation is left convinced that “we tried this and it did not work”.
In master data pilots the technology is what surprises us least. What surprises us most is how fast decisions get made once a measurement, rather than an opinion, is on the table.
That is why the report describes the opposite of the standard market approach: one data domain, one success criterion agreed before the start, and a result measured on the organisation’s real data in weeks rather than quarters.
In practice the method has four steps. First, choose the domain generating the most manual work and the most arguments about numbers - usually business partners or material indices. Then measure the state: number of duplicates, gaps in critical fields, discrepancies between systems. The measurement is what lets the decision rest on figures instead of opinions. Third come the rules for the master record, with name normalisation and a data owner holding the mandate to settle disputes. Only the fourth step is merging - and here one detail matters: language models are good at suggesting that two entries may describe the same entity, but a human approves the merge, because wrongly merging two separate companies is more expensive to repair than a duplicate. Every change stays in the audit trail.
We usually close a single-domain pilot in four to eight weeks. This is how we work with Medicover in the employee data domain: 18 countries, roughly 45,000 employees and around 100,000 data events a year, with full change history and an approval flow - a project SAP recognised as an official reference, one of three in Poland.
When master data is not your priority
We would be misleading you if we claimed this work is urgent for everyone. If you run a single system, the customer file holds a few hundred partners, and invoicing is handled by one person who knows all of them, you will get through mandatory e-invoicing on operational discipline alone, without a separate data layer.
We also advise against starting when the organisation expects a one-off clean-up with no data owner named on the business side, plans to cover every domain at once before the first confirmed result, or has nobody able to settle a disputed record on the merits. Under those conditions the project ends exactly as Gartner’s statistic predicts - and that wastes both money and trust. Honest qualification saves both sides a quarter of conversations with nothing at the end.
Priority belongs to organisations preparing a system migration, consolidating companies after acquisitions, launching automation or AI on operational data, or needing to demonstrate to an auditor where the data sits and who has access to it.
Measure it yourself before you spend anything
The most practical part of the report requires no purchasing decision at all. You can size the problem in one domain with your own people in a few days. All it takes is an extract from the customer file and four questions:
- How many records share the same tax identification number under different names or addresses?
- How many records have gaps in the fields your processes treat as critical?
- How many entries differ between the ERP and the CRM for the same entities?
- Who decides which version is authoritative, and how long does that take?
“We do not know” is also a measurement result on any of these, and the most instructive one. If you would rather someone external ran that measurement, we do it on a sample from a single domain and hand over the result with a recommended scope within five working days. Details on the duplicate measurement page.
Download the report
SNOK report · August 2026
Master data in Polish enterprises
PDF, 28 pages, about 0.9 MB - no form, no details to hand over. Pass it around your team freely - that is what it was made for.
Download the report (PDF)Inside you will find what this article could not hold: the full chapter on the market and acquisitions, the detailed mechanics of KSeF with the KIS ruling, the DSAG and ISG data on migrations, four anti-patterns of clean-up programmes, the incremental method step by step, and a methodological note with the complete list of sources.
From the report to a first result
If the subject turns out to concern your organisation, we suggest the route we take with clients most often. Start with a duplicate measurement on a sample from one domain - you receive the result and a recommended scope within five working days. If the numbers confirm the problem, we run a proof of concept on the SNOK MDM platform: a single-domain pilot on your real data, with the success criterion agreed before the start and a result you can compare before and after. Would you rather see the platform first? Book a SNOK MDM demo or simply write to us.
A week before this report we published a column on twenty years with master data - a good complement, this time from a personal angle. Questions to the authors go straight to them: Jacek Bugajski (jacek.bugajski@snok.ai) and Michał Korzeń (michal.korzen@snok.ai).
