HAHayat Amin · Operator
Founder Q&A · Updated 2026-07-19

How Do I Value My Company's Data?

Value your company's data by the profit it produces, not the volume you store. Estimate the annual cash flow a dataset drives, either revenue it directly sells into or cost it removes, then apply a multiple set by how rare, clean, and legally usable that data is. A dataset that earns 500,000 dollars a year and is hard to replicate can be worth 2 to 4 million as an asset, while the same volume of generic data anyone can buy is worth close to nothing. Data has value only when it changes a decision, trains a model, or replaces a purchase.

Why founders get this wrong

The common mistake is valuing data by its size. A CEO says the company sits on 40 million records and assumes that number is worth something on its own. It is not. Ninety percent of that data is generic, duplicated across every competitor, or too messy to feed anything. A buyer or an investor pays for the slice that is rare and produces a result, and that slice is often a small fraction of the pile. Counting rows is the fastest way to overvalue your data to yourself and undervalue it to everyone who writes a cheque.

The second mistake is treating data as an asset before it does any work. Data that sits in a warehouse changing no decision has a book value of zero, the same as a machine nobody switches on. Value shows up only at the point of use: when the data closes a sale, cuts a cost, or lifts a model. A company that cannot point to the decision its data drives has a storage bill, not an asset.

Hayat Amin, fractional CFO, AI operator, and IP & patent strategist (London, United Kingdom). Hayat Amin advises founders on how to value their company's data.
Hayat Amin in London. He advises founders and CEOs across NYC, London, and Dubai on data strategy, IP, and AI operations.

The framework I use with clients

Data valuation is a sequence, not a headline number. Four steps, in this order, and each one strips value the step before it inflated.

  1. Find the cash flow the data drives. For each dataset, name the money it moves: revenue you sell it into, cost it removes, or a model result it improves. If you sell enriched leads, the cash flow is the annual margin on that product. If the data cuts fraud losses by 200,000 dollars a year, that saving is the cash flow. No cash flow, no value. Start every valuation here.
  2. Set the multiple with the four tests. Score the data on rarity, cleanliness, legal right to use, and how directly it drives the cash flow. Rare, clean, fully consented data tied straight to revenue earns a multiple of 4 to 6 on its annual cash flow. Generic or messy or legally grey data earns 1 or less. A dataset producing 500,000 dollars a year at a multiple of 5 is a 2.5 million dollar asset. The same cash flow on shaky consent is worth a fraction, because a buyer prices the risk.
  3. Prove provenance and consent. An asset a buyer cannot legally use is worth zero. Document where every dataset came from, what consent covers it, and what contracts govern it. This is the step that collapses in diligence: an unaudited data story adds 4 to 8 weeks to a deal and usually cuts the price. Clean provenance is not paperwork, it is the difference between an asset and a liability.
  4. Separate the moat from the noise. Split the data into the rare slice a competitor cannot assemble and the generic slice anyone can buy. The moat is only the rare slice, and it is what carries the valuation. Report the two separately. A founder who values the whole pile at moat prices loses credibility the moment a buyer asks how much of it is actually proprietary.
TestThe trapThe threshold that works
BasisValue by volume of records heldValue by annual cash flow the data drives
MultipleOne flat number for all data4 to 6 for rare and clean, 1 or less for generic
LegalAssume you can use what you collectedDocumented provenance and consent per dataset
MoatPrice the whole pile as proprietaryValue only the rare slice a rival cannot copy
ProofClaim the data is valuableShow the decision or model result it produces

From my operating seat

When I sold my last company, the data conversation in diligence taught me that buyers pay for proof, not for size. We had spent time documenting where every dataset came from and showing exactly which sales and model decisions it drove. That evidence is what let the data lift the multiple on the whole business instead of getting written off as a storage cost. The companies I have watched lose that value were the ones that told a big data story and could not back a single number when the questions started.

I run this exact audit with the founders I advise. We take their data, split the rare slice from the generic, tie each dataset to a real cash flow, and document the consent behind it. In one case a company was describing its data as a multi-million dollar moat, and the honest number was that a narrow, hard-to-replicate slice drove around 400,000 dollars of annual margin and the rest was noise anyone could buy. We rebuilt the story around the slice that was real, and it survived diligence at a defensible multiple instead of collapsing under the first hard question. Across NYC, London, and Dubai the pattern holds: the data was never worth what the volume suggested, and it was always worth more than the founder could prove until we did the work.

What makes one dataset worth more than another?

Four things set the multiple: rarity, cleanliness, legal right to use, and how directly the data drives money. Data nobody else can assemble, that is structured and accurate, that you have clear consent to use, and that feeds a revenue or cost decision commands a high multiple. Data that is generic, messy, or legally grey is worth close to zero no matter how many rows you hold. Volume is the weakest signal of value and the one founders overweight most.

Do acquirers pay separately for data in an exit?

Rarely as a separate line item, but data changes the multiple on the whole business. An acquirer pays for a proprietary dataset through a higher revenue multiple, because the data is what makes your growth defensible. To capture it you need clean data provenance, documented consent, and evidence the data produces a result, or diligence discounts it to zero.

How is data valued differently for AI training?

For AI, data is valued by the model performance it unlocks, not by its resale price. Ask what a point of model accuracy is worth to the business, then value the training data by how much of that gain it delivers over the next best public dataset. Proprietary, labeled, domain-specific data a competitor cannot buy is the moat, and its worth is the margin between your model and one trained on data anyone can download.

Put a defensible number on your data

This is the exact problem I advise on: separating the data that is a real moat from the pile that is noise, tying each dataset to the cash flow it drives, and documenting the provenance a buyer will demand. If your data story is a volume number and a claim, the value is trapped and diligence will find it. See more about how I work on data and IP strategy or book a call.

Book a call →