Back to Blog
Guides

The Data You Already Own: How UK SMEs Turn Decades of Records into Revenue

A
Arun Godwin Patel
September 7, 202611 min read

Your archive is not just history, it is the raw material for products you could sell. A complete guide to finding, valuing, cleaning and commercialising the data your business already holds.

Shelves of archived records on the left, with three routes out of them: running the business better, selling a product, and licensing.

There is a particular kind of British business that has been trading for twenty or forty years, is quietly profitable, and is sitting on something it has never thought of as an asset.

A publisher with a back catalogue nobody can search. A law firm with thirty years of drafted clauses. A tour operator with two decades of itineraries that worked. A specialist manufacturer with drawings and tolerances that took a generation to get right. A recruitment firm that knows which placements lasted.

These are information businesses. What they sell is knowledge, judgement, provenance or access, and the record of all that knowledge is usually in a shared drive, an email archive and several people's heads.

This guide is about turning that into something the business gets paid for. It covers how to find it, how to value it, what it costs to make usable, what you may legally do with it, and the three routes to revenue that actually work for a business of this size.

Why this is different from "doing analytics"

Analytics asks what happened. This asks what you know that other people do not, and whether it can be sold.

The distinction matters because it changes what you look for. An analytics project looks at the systems of record: sales, finance, operations. A value project looks at the places where judgement was exercised and recorded. Quotes and their outcomes. Notes explaining why something was decided. The correspondence around a job that went wrong. Almost none of that is in a database, which is exactly why it has been overlooked.

Roughly 55 per cent of business data is never used for analysis or decisions. The proportion is higher for this kind of material, because it was never structured to begin with.

Step one: find it

Five places, in order of how often they turn out to hold the valuable material.

The email archive. For most professional and service businesses this is the single richest source. It contains the reasoning, the negotiation, the exceptions and the apologies. It is also the hardest to use, which is why nobody has.

Quotes, proposals and tenders. Every one records what you thought a piece of work was worth, and the outcome records whether the market agreed. See your pricing history is a dataset.

Job, case or project records. What was planned, what happened, what it cost, what went wrong.

Document archives. Scanned files, drawings, contracts, reports. Frequently technically present and practically invisible, because a scan is a photograph of information rather than information. See scanned, filed and forgotten.

People's heads. The largest and least protected store. Covered in technology and succession.

Spend a week on this and write down what exists, roughly how much, in what format, and covering which years. That inventory is the whole foundation, and a fuller method is in dark data: how to audit what your business is already storing.

Step two: work out what is actually valuable

Four tests. Data has to pass all four, and most fails on the last one.

Specific. About your customers, your jobs, your failures. Not the sector generally.

Verified. What happened, not what was planned. Outcomes, not intentions.

Exclusive. Not purchasable by a competitor. Industry data is an input, not an asset.

Connected. Joined to something. A customer list is mildly useful. A customer list joined to purchase history, price paid and what went wrong is a different asset entirely.

The joins are where the value sits and where the work sits. In most businesses the facts exist in three systems and have never been put in the same place.

Step three: the three routes to revenue

For a business of ten to five hundred people, there are three that work. They are in descending order of how often they pay off.

Route one: price better

The most reliable and the least glamorous. Your archive contains the answer to what you should have charged, and almost nobody has looked.

What it reveals is consistent across sectors: certain job types are systematically underpriced, certain customers absorb a disproportionate amount of unbilled time, and discounting is concentrated in a small number of relationships that would have proceeded anyway.

Worked example. A 45-person engineering consultancy, £4.2m turnover, twelve years of proposals in email and a job costing spreadsheet. Extracting quoted versus actual across 1,800 projects cost £14,000 and took ten weeks. It showed that one service line ran 19 per cent over on average and had been quoted from the same template since 2019. Correcting it added roughly £71,000 of annual margin on existing volume.

That is the shape of it. No new customers, no new product, no new headcount.

Route two: productise the judgement

Turn something you currently sell as hours into something you sell as a thing. A diagnostic, an assessment, a benchmark, a report, a tool.

This works when you have enough history to say something a client cannot say about themselves. A recruiter who can tell a client how their offer compares to 400 comparable placements has a product. One who can only describe the market does not.

Covered in turning expert knowledge into a product you can sell twice.

Route three: make the archive itself the product

Narrower, and where it applies it is the largest of the three. This is the publisher whose back catalogue becomes searchable and licensable, the manufacturer whose drawings become a spares business, the travel operator whose itinerary history becomes a configurator that sells at three in the morning.

The test is whether the archive answers a question somebody outside the business wants answered and would pay for. Most do not. Where they do, it is usually the most valuable thing the business owns. See publishing and media and travel and tourism.

Step four: the legal position

Owning the data is not the same as being allowed to use it however you like, and this is where enthusiasm most often meets a wall.

Personal data needs a lawful basis for the new purpose. Collecting customer records to fulfil orders does not automatically permit using them to train a model. Purpose limitation under UK GDPR is the specific obstacle, and legitimate interests requires a documented balancing test rather than an assertion.

Client confidentiality may bite before data protection does. For professional services, the engagement letter and professional rules often restrict use more tightly than GDPR. Aggregate insight derived from client work is usually fine. Anything traceable to a client is usually not.

Check who owns what. Data created under a client contract may belong to the client. Data from a platform may be governed by that platform's terms. Data from an acquisition may carry restrictions from the original collection.

The workable answer in most cases is aggregation and anonymisation, which removes most of the obstacles and preserves most of the value. Detail in can you legally use your own customer data to train AI.

Step five: what it costs

Realistic UK ranges.

Inventory and feasibility: £3,000 to £8,000, or free if you do the week yourself. Always do this first.

Extraction and structuring: £8,000 to £40,000 depending on format. Structured systems are cheap. Documents are moderate. Paper and handwriting are expensive, and worth it only when the answer is genuinely valuable.

Cleaning and joining: £5,000 to £25,000. Consistently underestimated, and it is where projects overrun. See data quality before AI.

Building something on top: £15,000 to £80,000 for a working internal tool or a first commercial product.

Ongoing: 10 to 20 per cent of build cost annually, plus the cost of keeping the data current. An archive that stops being updated stops being valuable within about two years.

A route-one pricing project typically lands between £12,000 and £30,000 all in. A route-three archive product is a real investment starting around £50,000.

The order that works

Name the decision first. Not "we want to use our data". "We want to know which job types we underprice." If nobody can name a decision that would change, stop here and save the money.

Inventory for a week. Yourself.

Test on a slice. Take one year, or one service line, and answer the question on that alone. Three to six weeks, a few thousand pounds. This tells you whether the data supports the answer before you commit to the full extraction, and it is the step most often skipped by suppliers who would rather sell the whole thing.

Then extract properly, if and only if the slice worked.

Build last. Tools built before the data is understood automate a misunderstanding.

What goes wrong

Starting with the technology. A platform bought before the question is defined will answer a question nobody asked.

Assuming the paper is worth digitising. Sometimes it is. Often the useful facts are already in a system and the paper is ballast. Check before scanning forty boxes.

Ignoring the legal position until the end. The most expensive failure, because it can invalidate the whole project after the money is spent. Establish it during feasibility.

Treating it as a one-off. An archive is a snapshot. The value comes from the flow continuing, which means changing how records are captured going forward. That change is usually the cheapest part and the most often forgotten.

Key Takeaways

  • The valuable material is where judgement was recorded: quotes and outcomes, email, job records, notes. Not the systems of record.
  • Data must be specific, verified, exclusive and connected. Most archives fail on connected, and the joins are where both the value and the work sit.
  • Three routes to revenue: price better, productise the judgement, or make the archive itself the product. The first is the most reliable by a distance.
  • Owning data is not permission to use it. Purpose limitation under UK GDPR and client confidentiality both bite. Establish the position during feasibility, not at the end.
  • Name the decision, inventory for a week, test on one year, then extract. Building first automates a misunderstanding.

Frequently Asked Questions

How much data do we need for this to be worth doing?

Enough to cover a full cycle of whatever you are studying, with enough instances that patterns are not noise. As a rough guide, a few hundred comparable records. Below that you can read them yourself and reach the same conclusions for nothing.

Our records are patchy and inconsistent. Is it hopeless?

No, and it is the normal starting position. The question is never whether the data is clean, it is whether it is good enough for the specific decision. Patchy quotes can still show that one service line runs over, even if they cannot support anything more precise.

Can AI just read our archive and tell us what is in it?

It can do a great deal of the extraction and summarising, and that has genuinely changed the economics of this work over the last two years. What it cannot do is decide which question is worth asking, or tell you when the underlying records are unreliable. Those are the two things that determine whether the project is worth doing.

Should we hire a data person?

Usually not as the first move. The first project is a defined piece of work with an end. Hiring before you know whether the data supports anything gives you a salaried person with no brief. If three or four projects prove out, that is the point to consider it.

How long before we see anything?

The slice test gives you an answer in three to six weeks. A full route-one pricing project delivers something actionable in three to four months. Archive products are a year or more. Anyone promising commercial returns in weeks is describing the slice test and charging for the full project.


Sitting on decades of records and wondering whether there is anything in them? Talk to Halo Technology Lab. Our strategy and scoping service starts with a week of inventory and a slice test, and we will tell you plainly if the answer is not there.

Share this article

Enjoyed this? Get the next one by email

Practical AI playbooks, build logs and tool teardowns. One email a week, free, unsubscribe in one click.

See what’s in it first

Have a project in mind?

Let's discuss how we can help bring your ideas to life.

Get in Touch