FM Talk

If I had £1 for every time someone asked which datasets we use…

By EMC Associates 7 July 2026 6 min read
Share
If I had £1 for every time someone asked which datasets we use…
Benchmarking • FM Talk

.. I’d probably have enough to buy another benchmarking database.

But I still wouldn’t know whether the benchmark was right.

That’s because the biggest mistake in FM benchmarking isn’t using the wrong dataset.

It’s believing the dataset is the benchmark.

It isn’t.

And that misunderstanding is costing organisations far more than the price of the benchmarking exercise itself.

Because in FM benchmarking, the source of the data tells you very little unless you also know whether the data is genuinely comparable.

And that is where many benchmarks start to fall apart.

The FM market is under real pressure. Operating costs remain difficult to control, labour remains a dominant cost driver in cleaning, catering and security, and clients are asking for sharper evidence before they accept price increases, retenders, extensions or savings claims. At the same time, technology, dashboards and AI are making benchmarking look more sophisticated than ever.

But a better dashboard does not make weak data reliable.

It simply makes weak data look more convincing.

The Dataset Is Not the Benchmark

In FM, we have become slightly obsessed with scale.

How many contracts are in the database?

How many buildings?

How many square metres?

How many cleaning hours?

How many meals served?

How many sites?

The assumption is understandable:

More data must mean better benchmarking.

Unfortunately, it does not.

A large dataset can still produce a poor benchmark if the contracts inside it are not comparable. In fact, the bigger the dataset, the easier it can be to hide weak assumptions behind an impressive sample size.

Take two office cleaning contracts.

Both cover 20,000 square metres.

Both cost £450,000 per year.

On paper, they look similar.

But one includes washroom consumables. One does not.

One operates five days a week. One operates seven.

One includes laboratory areas. One is standard office space.

One requires enhanced vetting. One does not.

One includes periodic deep cleans, pest control and window cleaning. The other excludes them.

Same area. Same headline cost. Completely different service.

If those two contracts are treated as equivalent, the benchmark is not evidence. It is arithmetic with a blindfold on.

Bad Inputs Create Confident Conclusions

This matters more now because benchmarking is becoming more visible in commercial decision-making.

The Procurement Act has increased the emphasis on transparency, performance information and contract management across public procurement. PFI expiry and handback are putting pressure on authorities to understand asset condition, lifecycle obligations and service continuity. Across FM, rising labour and maintenance costs mean clients and suppliers are having harder conversations about value, risk and affordability.

In that environment, a benchmark is not just a spreadsheet exercise.

It can influence budgets, procurement strategy, supplier negotiations, savings targets, contract extensions and board-level decisions.

So the question should not be:

“Which dataset did you use?”

It should be:

“How do you know this comparison is fair?”

That is a much better question.

Because unless the underlying scope, service levels, operating conditions and cost drivers are understood, the benchmark may only prove that two numbers can be placed beside each other.

It does not prove they mean the same thing.

Benchmarking Is a Data Quality Problem

After years of forensic benchmarking, one lesson stands out clearly:

Benchmarking is not really a data problem.

It is a data quality problem.

The organisations that get the most value from benchmarking are not necessarily those with access to the biggest databases. They are the ones prepared to interrogate the evidence before drawing conclusions.

They ask:

Is the specification actually the same?

Have service levels drifted?

Are frequencies comparable?

Has inflation been normalised?

Are exclusions consistent?

Are occupancy levels aligned?

Are operating hours comparable?

Are we comparing outputs, inputs or simply costs?

Has geography, risk, asset condition or building use been properly considered?

These questions are not always glamorous. They rarely make the front page of a benchmark report.

But they decide whether the report is useful.

The EMC DICE Method

This is why we developed the EMC DICE Method.

Not because the industry needed another acronym.

Because it needed a clearer way to test whether benchmarking evidence is defensible.

Every benchmark should pass four tests.

D – Data Integrity

Is the source data complete, accurate, current and traceable?

If the baseline data is out of date, incomplete or poorly coded, every conclusion built from it is vulnerable.

I – Input Consistency

Are we comparing like with like?

That means checking scope, frequencies, service levels, operating hours, inclusions, exclusions and contract structure before drawing conclusions from unit rates or annual costs.

C – Context Validation

What operational factors explain the numbers?

A hospital, laboratory, school, civic building, corporate HQ and mixed-use estate may all have “cleaning costs”, but the operational reality behind those costs can be completely different.

Context matters: occupancy, risk, asset condition, geography, compliance, security requirements, workplace expectations and service intensity all shape cost.

E – Evidence

Could the benchmark be defended in front of a Finance Director, procurement board, client team or supplier?

If the answer is no, it is not evidence.

It is opinion dressed up as data.

What FM Leaders Should Ask Instead

The biggest benchmarking mistake is believing the dataset is the benchmark.

It is not.

The benchmark is the conclusion drawn from the data.

And that conclusion is only as strong as the questions asked before the analysis begins.

So next time someone says they benchmark against thousands of contracts, do not stop at the headline number.

Ask:

How many were truly comparable?

What was excluded, and why?

How were scope differences normalised?

Were service levels, occupancy and operating hours aligned?

How was inflation treated?

Were contract-specific risks separated from market norms?

Could the conclusion be defended commercially?

Those answers matter far more than the size of the database.

Because in FM, the danger is not having too little data.

The danger is having enough data to sound certain, but not enough discipline to be right.

Final Thought

If I had £1 for every time someone asked which datasets we use, I would probably have a decent holiday fund.

But I would much rather be asked this:

“How do you know your benchmark is right?”

That is the question that protects budgets.

It improves procurement decisions.

It strengthens negotiations.

It avoids false savings targets.

And ultimately, it leads to better FM outcomes.

Hi, I’m Ernie, Founder and Consultant Partner at EMC, a specialist FM advisory consultancy helping organisations procure, benchmark, tender and manage FM services with more control, better value and fewer headaches.

Our forensic benchmarking approach, including the EMC DICE Method, focuses on data integrity, context and defensible evidence, not just bigger datasets.

FAQ

What is FM benchmarking?
FM benchmarking compares the cost, performance or specification of facilities services against relevant comparators. It can cover services such as cleaning, catering, security, maintenance, helpdesk, waste, grounds maintenance and total FM contracts.

Why is a large dataset not always better?
A large dataset is only useful if the records inside it are comparable. If contracts differ by scope, service levels, geography, operating hours, risk profile or exclusions, a large sample can produce misleading conclusions.

What makes two FM contracts comparable?
Comparable contracts have enough alignment in specification, service frequency, operating environment, occupancy, asset type, contract scope and performance expectations to allow a fair comparison. Where differences exist, they should be normalised or clearly explained.

How often should FM contracts be benchmarked?
Most organisations should benchmark at key commercial decision points: before tendering, before contract extension, during major scope changes, ahead of price negotiations, or when service quality and cost appear misaligned. For high-value or complex contracts, periodic review during the contract term is also sensible.

What is forensic benchmarking?
Forensic benchmarking goes beyond headline rates. It tests the underlying data, contract scope, assumptions, exclusions and operational context to establish whether a benchmark can be commercially defended.

Can AI improve FM benchmarking?
Yes, but only if the underlying data is structured, accurate and comparable. AI can help identify patterns quickly, but it cannot automatically fix poor scope definition, missing exclusions or inconsistent service data.

Is your FM contract delivering what it should?

Book a free discovery call with an EMC consultant. Evidence led, no obligation.