20sources
The Empty Fields Got Filled. So Did the Ones That Were Already Right.
Most of the blanks came back full, which is what you paid for. In one record the number a client gave you herself had been replaced, and nothing in the row said what was there before, where the new one came from, or when either was true. What an appended field actually asserts, why no honest decay rate exists, and the two columns that make all of it manageable.
You ran the pass because two thirds of the database had no phone number in it. That is a real problem and enrichment is a real answer to it, and by the afternoon most of those blanks were full.
Somewhere in the middle of the file is a woman you sold a house to three years ago. She gave you that number herself, standing in her own kitchen, and you have texted her on it since.
Her record now has a different number in it.
Nobody decided that. The pass filled in what was empty, and where a field was not empty it wrote anyway, because that is what the default was and nobody was asked. There is nothing in the row that says what used to be there, nothing that says where the new one came from, and nothing that says when either of them was true.
The blanks getting filled is the part everybody talks about. The overwrite is the part nobody mentions, and it is the more expensive half.
In short
- An appended field is somebody else's assertion about a person, and it arrives without the two things that would let you weigh it: when it was true, and who said so. The Federal Trade Commission put compulsory orders to nine of these companies and reported that most of their data comes from other companies like them rather than from an original source.
- There is no honest published rate for how fast contact data goes stale. We followed the circulating figures rather than assuming where they came from, and they end at a press release with a distribution disclaimer on it, at vendor blog posts, and at an aggregator citing a benchmark whose original is not linked. On the pages we opened, the same claim is quoted at thirty percent, at twenty two and a half, at twenty to thirty, and at up to seventy for email addresses.
- And the expensive failure is not the record that comes back empty. It is the field you already knew, quietly replaced by a value you cannot check, in a row that does not record what was there before.
Three things this is not
The article next door owns half of what you came here for.
Not the trace
Starting from a property and finding contact details for an owner who has never spoken to you is a different act with two federal statutes attached to it, and it has its own long article on this site. That question is about acquiring something. This one starts from a person who is already in your database, usually because they came to you, and asks what you are entitled to believe about the row you now hold.
Not the duplicate
Deciding whether two records in your own systems describe one person is a question you can answer, because you can open both. There is a published model for it, and it has a third outcome worth knowing about. Here the other record is inside a company you have no access to, and you cannot see how it decided anything.
Not the old list
Whether you may still contact somebody who went quiet three years ago is a question about permission, and permission has dates in it. This article assumes you are allowed to make the call. It is about whether the number you would dial is the one that reaches her, and about what happened to the number that used to be in that field.
What data enrichment actually is, and what it is not
Three articles on this site are neighbours to this one and one of them is very close indeed, so it is worth drawing the lines before anything else.
Data enrichment is the pass that completes and corrects records you already hold. A contact came in with a first name and an email and no phone. Another has a mailing address and no idea which property behind it they own. A third is the same person as a record you took two years ago under a different email. Enrichment sends what you have to an outside provider and writes back what comes home: a number, an email, the property detail behind the address, a merge.
Two words in that description are doing all the work, and almost nothing written about this subject examines either of them. "Corrects" assumes the outside answer is better than yours. "Writes back" assumes there was nothing there.
The narrow thing worth understanding is that enrichment is not a lookup and it is not a fix. It is the arrival of an assertion from a company you have never spoken to, about a person you have, into a system where it will be indistinguishable from something you knew.
What comes back is a claim rather than a fact
Think about what actually has to happen for a phone number to appear in that column.
You send an identifier: a name, an address, an email. Somebody else's system then decides which of its own records describes the same human being, on partial information, with no way to ask anybody. Then it returns the value it holds against whichever record it picked.
So the number in your CRM is the end of a chain of at least two guesses: this record is about your person, and this number belongs to that record. Neither guess is shown to you. What arrives is a bare string in a field, formatted exactly like the numbers your clients typed in themselves.
That matching problem has its own long piece on this site, the one about keeping two systems in step, and it is worth reading because the published model for it has a third outcome that most builds throw away. One difference changes the whole picture here. In a sync, both systems belong to you and you can open both. In an enrichment response the other system is a black box, and the threshold it used, the fields it weighed and the confidence it settled on are all facts about a company you are merely a customer of.
What arrives in the column
Four things an appended value does not tell you.
It does not say when
A phone number in an enrichment response is the best answer in somebody's file on the day you asked, and the file does not say when it was assembled. Two numbers that look identical in your CRM can be a fact confirmed last month and a fact that was true in a different decade, and nothing in the row distinguishes them.
It does not say who
The name on your invoice is the company you bought from, and that company mostly bought it as well. So the truthful answer to where a value came from is the last name in a chain rather than the name of whoever first wrote it down, and nobody publishes the chain. That is not a criticism of any particular supplier. It is the shape of the market, and a federal regulator has said so on the record.
It does not say observed or inferred
Some of what sits in these files was written down by somebody who saw it happen. Some of it was worked out from other things, by a company that has never met the person it describes. Both come back through the same interface in the same shape, so a conclusion that has been rounded into a value is indistinguishable from a measurement by the time it reaches your CRM.
It does not say what it replaced
This is the one that costs money and it is entirely under your control. If the pass wrote into a field that already held something, the row should say so. Whether yours does is a thing you can check this afternoon, and the answer decides whether a bad pass is reversible. Where it does not, the old value is simply gone, nobody chose to discard it, and the one person who could have confirmed it was right is the client whose number you no longer have.

Where an appended field actually comes from
There is one primary document on this and it has held up for twelve years because of how it was made.
In December 2012 the Federal Trade Commission issued compulsory orders under section 6(b) of the FTC Act to nine named data brokers, requiring them to file special reports on where their data comes from, what they do with it and what rights consumers have over it. The resulting report, Data Brokers: A Call for Transparency and Accountability, was published in May 2014 and covers their practices from January 2010. Nine companies, named, under compulsion, describing themselves to a regulator. There is nothing else like it in this subject and everything written since leans on it.
The evidence
Where nine data brokers got their data, under compulsory FTC orders
Counts out of the nine named companies the Federal Trade Commission ordered to file special reports in December 2012 under section 6(b) of the FTC Act, covering their practices from January 2010. The three rows overlap: most of the nine do all three. The axis is nine because that is the whole population studied, not a sample of the industry. Source: Federal Trade Commission, Data Brokers: A Call for Transparency and Accountability, May 2014.
This report is from 2014 and the industry has changed a great deal since, so do not read the bars as a description of whoever supplies your own enrichment today. Read the middle one as a structural fact, because that is what it is: when companies in a market mostly buy from each other, no single one of them can tell you where a value originally came from. The chart cannot show you the thing that follows from it, which is that the answer to "where did this come from" stops being a name and becomes a direction of travel. Nothing about that arrangement has become simpler in the years since.
The finding that reorganises how you should think about an enrichment response is the middle bar. These are not nine companies each independently observing the world. They are, to a substantial degree, one market trading the same records among themselves, and the Commission states the consequence plainly: the nine "obtain most of their data from other data brokers rather than directly from an original source", and one of them draws consumers' contact information "from twenty different sources".
That is the honest picture behind the field in your CRM. Not a company that knows something about your client, but the last company in a queue, passing on what it was passed.
It would be virtually impossible for a consumer to determine how a data broker obtained his or her data; the consumer would have to retrace the path of data through a series of data brokers.
Read that as a statement about you rather than about consumers. If the person the record is about cannot retrace it, neither can you, and you are the one who is going to be asked.
Observation and inference arrive in the same column
The same report describes two different kinds of content in these files, and it is the distinction that should change what you do with the output.
There is raw data, which the report describes as things "such as a person's name, address, home ownership status, or age". And there is derived data, "which they infer about consumers". The report gives its own examples of how inference works: a data broker "might infer that an individual with a boating license has an interest in boating, that a consumer has a technology interest based on the purchase of a 'Wired' magazine subscription, or that a consumer who has bought two Ford cars has loyalty to that brand".
Those are marketing categories rather than contact details, and it would be a mistake to say your appended phone number is an inference. What is not a mistake, and is the point, is that both kinds of value come back through the same interface, in the same shape, with no marking to say which is which. A file that contains observations and guesses in the same schema, sold through one API, will land both of them in your database as facts, because a field has no way of holding the difference.
The practical version of this is a question with a short answer, and it is worth putting in writing to whoever supplies you: for each field you buy, is this something the source observed or something the source concluded. A provider who can answer that field by field is telling you a great deal about how carefully the product was built.
The same file, sold under three different names
There is a second finding in that report, about what the same nine companies sell, and it is the one that connects this subject to the law.
The evidence
What the same nine companies sell the same underlying data as
Counts out of the same nine companies, from the same report. They add up to more than nine because several of them sell in more than one category, which is the finding rather than an error: the same underlying records are packaged as a marketing list, as an identity check and as a public lookup page. Source: Federal Trade Commission, Data Brokers, May 2014, on the three product categories the nine sell.
The counts are from 2014 and the industry has changed since, so read them as a shape rather than as a market share, and note that the three overlap because several of the nine sell in more than one category. What the chart cannot draw is the part that matters, which is that nothing in the underlying records changes as they move between these three shelves. The same rows, the same fields, the same unknown ages, sold three times under three descriptions to three kinds of buyer, and the description is picked at the moment of sale rather than at the moment anybody wrote the data down.
One set of underlying records, three shopfronts. The identity check and the marketing list and the people search page are built out of the same material, and what separates them is the purpose the buyer had.
American law works the same way round, and this is where this article stops and hands you to a different one. Two federal statutes decide what may be done with contact information about a person, they turn on where it came from and on what you intend it for rather than on which fields are in the file, and the liability lands on the buyer rather than on the seller. All of that is worked through at length in the article on skip tracing, including the questions to put to any provider in writing, and none of it is repeated here.
What belongs in this article is the narrower half: the same data, relabelled at the point of sale, and a label that is a fact about the transaction rather than about the record.
Nobody has published an honest decay rate, and we went and looked
Every page selling this quotes a figure for how fast contact data goes bad. Thirty percent a year is the usual one. We did not want to assert that those figures are unsourced, because asserting a reason without checking it is a mistake this project has made before, so this round the trail got followed.
Here is where it goes.
The most prominent recent version of the thirty percent claim is a press release from a company that sells contact data, carried on a newspaper's website under a notice stating that it is "press release content distributed by XPR Media" and that the paper's editorial staff "were not involved in the creation of this content". No sample, no method, no population.
Below that are vendor blog posts. Data quality companies and enrichment providers, each stating a rate, none stating what was measured or on how many records. The most useful thing on any of them is not a statistic: one recommends taking a random sample of a hundred to two hundred of your own oldest contacts and verifying them by hand, which is the correct answer and is the one thing on the page that nobody is charging for.
Below those are the aggregator pages, which cite each other and eventually cite a benchmark attributed to a marketing research publisher whose original study is not linked from any of them.
And the spread is the finding. On the pages we opened, the same claim about the same thing is quoted at thirty percent, at twenty two and a half percent, at twenty to thirty percent, and, for email addresses specifically, at up to seventy. Those are not measurements that disagree with each other. They are a number that has come loose from whatever produced it and is now being cited by people who are citing each other.
So there is no decay rate in this article, none in the calculator, and none on our service page. There is something better, which is the reason a rate cannot be a single number in the first place.
A rate is a property of the people, not of the data
Contact details do not rot on their own. They stop being true when something happens to a person: a move, a job change, a marriage, a new phone, a business closing. So the speed at which a field goes wrong is the speed at which that thing happens to the people in your database, and different people are not the same.
Which means the useful question was never how fast data decays. It is what sits underneath a particular field, and how often that thing changes for the particular people you hold. One government survey measures exactly that shape, for exactly one field, and the field is employment.
The evidence
How long people had been with their current employer, by age, January 2024
Median years of tenure with the current employer, from a supplement to the Current Population Survey, a monthly sample of about sixty thousand households. Median means the point at which half of all workers had more and half had less, so half of the youngest group had been in the job under two years and nine months. Source: U.S. Bureau of Labor Statistics, Employee Tenure in 2024, Table 1, released September 2024.
This is about employment, not about the mobile number of somebody who bought a house from you, and it is here for one reason: it is the only measurement in the whole of this subject that names its survey, its sample and its definition. Use it as an argument rather than as a number. Whatever underlies a field decides how fast that field goes wrong, and the rate is different for different people, so a single percentage covering everybody in your database is describing a population that does not exist. We went looking for an equivalent measurement for a personal mobile number, a personal email address or a homeowner's mailing address and did not find one with a stated method, which is worth holding next to the fact that every figure circulating about data decay is quoted as though somebody had done exactly that study.
The Bureau of Labor Statistics release for January 2024 puts the overall figure at "3.9 years", the lowest since 2002. The chart above is Table 1 of the same release, and the spread across it is the whole argument.
Now put both of those people in one database with a work email address each, and run a single percentage across the pair of them. Whatever that percentage is, it is wrong about both of them in opposite directions, and it will be quoted as though it described the database.
Which is why the number worth having is not a benchmark at all. Take two hundred records at random from the part of your database you would really work, and check them by hand. It takes an afternoon, it costs nothing, and what comes out is measured on your own people in your own market at their own ages, which is the only version of this figure anybody can defend.
What to do when two sources disagree
This is the question the whole subject turns on, so it is worth being concrete about it.
A pass runs. For a given contact, your record says one thing and the response says another. There are four possible behaviours and every build has one of them. Find out which before a pass runs, not afterwards.
The first is that the newest value wins, where newest means the value that arrived most recently rather than the value that most recently became true. A number your client confirmed to you last month gets replaced by a number a provider has held since some date nobody recorded, purely because the query happened today.
It is worth seeing how that behaviour is presented by the software, because it tells you which way the tooling leans. HubSpot's documentation for importing records describes protecting what you already hold as something you switch on. There is an advanced option called Prevent property overwrite, which the page explains as: "if you're updating existing records, prevent the import from overwriting records' existing property values for the row". When it is selected for a property, "the import will update the property for new records and existing records that have never had a value for the property. It won't update the property for existing records that have a value or had a value in the past, even if currently empty."
What matters there is the shape and not the detail. On one major CRM, per property, at import time, keeping your own value is a checkbox. That is one platform and not a survey of the field, and it is exactly why the four choices below are worth settling in writing before a pass runs instead of discovering afterwards which one you got.
The second is that yours wins and the response is discarded. Safe, and it quietly turns enrichment into a fill-the-blanks exercise, which is often exactly what you wanted and should be a decision rather than an accident.
The third is that both are kept, in separate fields, with the outside one clearly marked as a suggestion. This is the one that costs least to be wrong about, and it costs one extra column.
The fourth is a review queue: disagreements go to a short list and a person settles them. That is right when the field matters and the volume is small, and it is worth knowing that it is the same shape as the clerical review step in the published record-linkage model that the CRM sync article covers, which exists for exactly this reason.
Whichever you pick, one habit does more than the choice itself. Write down, on every enriched row, where the value came from and when it was written. Not in a log somewhere. In the record, beside the value, where the person about to dial it can see it. That single column is the difference between "I do not know where that number came from" and a one sentence answer, and it costs nothing on the day the build is done.

The one law that describes this is about whoever holds the data
There is one place where a legislature has written down what happens when a business is holding something wrong about a person, and it is not a federal statute and it probably does not apply to you. It is worth knowing anyway, because it says whose problem this is.
California's consumer privacy law, at Civil Code 1798.106, gives a consumer "the right to request a business that maintains inaccurate personal information about the consumer to correct that inaccurate personal information, taking into account the nature of the personal information and the purposes of the processing". A business receiving such a request "shall use commercially reasonable efforts to correct the inaccurate personal information as directed by the consumer".
Read who that is addressed to. Not the data broker. The business that maintains the information, which in the scenario at the top of this article is you.
Whether it applies to you specifically is a separate question, and the thresholds are published, so it takes a minute rather than a lawyer. The same law defines a covered business at Civil Code 1798.140 as one doing business in California that also meets one of three: annual gross revenues over twenty five million dollars, or annually buying, selling or sharing the personal information of a hundred thousand or more consumers or households, or deriving half or more of its revenue from selling or sharing personal information. Read those three against your own year and you will know where you stand.
So this is not a compliance obligation for most readers of this page. It is a description of the right shape, written by people who thought hard about it, and it says two things worth adopting whether or not anybody is making you. The duty attaches to whoever holds the record rather than to whoever supplied it. And answering it requires knowing where a value came from, which is the column this article keeps coming back to.
The same FTC report is blunt about how rarely that is possible today. Of the nine companies studied, it found that "only two of the data brokers allow consumers to correct their personal information for marketing purposes". The report's sentence does not make its denominator explicit, so it is quoted rather than drawn as a chart, and the direction of it is not in doubt.
What New York's law actually covers, and it is not this
People reach for New York's SHIELD Act here, and it is worth saying plainly why it does not cover most of what an enrichment pass adds.
The New York Attorney General's own description of the law it enforces says the Act requires "any person or business that maintains private information to adopt administrative, technical, and physical safeguards", with no revenue threshold attached, which is a genuinely wide obligation. But "private information" is a defined and narrow term. On the Attorney General's page, it means personal information combined with a Social Security number, a driver's licence number, or an account number with its security code, and the Act extended it to biometric information and to a username or email address together with a password.
A telephone number is not on that list. A mailing address is not on that list. An email address on its own is not on that list.
So the New York statute people mention in this context is about security and breach notification for a specific set of high risk identifiers, and it is mostly not about the fields enrichment appends. That is a checked absence rather than an omission, and it is here because the alternative is a page implying a duty that does not exist.
Nothing above is quoted from the statute itself. It all comes from the Attorney General's published summary of the law that office enforces, which is a step further from the text than this article would like to be, and worth saying out loud, because the whole finding here is what one defined term does and does not include.
The system
From your own record back to your own record.
Six hops, and it is a loop rather than a line, which is the thing that makes this different from every other data project in your business. The output lands on top of the input. Three of the six happen inside a company you have no access to, and the only two you control are the first and the last, which is where all of the available honesty lives.
Scroll to follow the chain
- Your record: What you hold now
- The query: What you send out
- The match: Their guess, unseen
- The claim: A value with an age
- The write: What it lands on
- The row: What it says after
In your numbers
How many fields would a pass replace that you already knew?
The number you would actually send, which for most people starts as the whole database and should not be.
Not the share you would call. The share where the field you are about to enrich is not empty. Sort by the column and look.
This one is yours to supply and it is worth measuring rather than guessing: run two hundred records where you already know the answer and count the disagreements. Nobody can quote it to you honestly before seeing your list.
Opening the record, seeing where your own value came from, deciding. Seconds when you remember the person and much longer when you do not.
Fields where somebody has to decide which answer is yours
200disagreements
- Records in the passyour 2,000
- 2,000records
- Where the field is not emptyyour 40%
- 800filled fields
- Where an outside file says something elseyour 25%
- 200disagreements
- At your settling timeyour 3
- 600minutes
- In hours60 minutes in an hour
- 10hours
The headline is the third row rather than the hours, because the hours are the affordable half and the count is the one nobody asks about before buying. Every one of those is a moment where something will choose on your behalf if nobody has chosen deliberately, and the choosing happens silently. Shares of records produce fractions, and half a disagreement is not a thing, so read anything with a decimal in it as a rough count. Three things this deliberately refuses. There is no rate of decay anywhere in it, because the figures that circulate do not survive being followed, which is worked through in the section above. There is no figure for how often the appended value turns out to be the correct one, because the only person who can measure that is you, on a couple of hundred of your own records where you already know the answer, and that measurement is worth more than anybody's benchmark. And there is no money in this calculator at all: a reachable client is not a commission, the loss from a number you can no longer reach is not payable on any date, and the arithmetic that turns either into a dollar figure does not exist.
What it costs, and how long it takes
The largest line on this bill is not ours and never passes through us. A supplier charges for each record it is asked about, so that part of the spend rises directly with how many rows you hand over.
Which makes the first cost decision the one nobody treats as a decision: how many records to run. The default is the whole database, because it is one click. The number worth having instead is the count of rows somebody in your business would actually call this quarter, and producing it is an hour of thinking rather than a purchase. Whatever it comes to is the honest size of the job, and the gap between it and the whole database is money about to be spent making people reachable that nobody was going to reach.
The second thing to settle before signing is whether the charge lands on every record you submit or only on the ones that come back with something. That distinction changes the arithmetic completely, and in an area where not much resolves, the first arrangement can spend a budget and leave you with almost nothing to show for it.
What we would build around that is the smaller half and it is where the value is. Deciding which records qualify. Writing the source and the date onto every enriched row. Making the disagreement behaviour explicit rather than inherited. Putting the flagged rows somewhere a person will look. Every one of those is cheap and dull, and they get skipped because nobody puts them on the list at the beginning.
On time, running a pass is fast. What takes the time is the conversation about the four disagreement behaviours above and about which fields you actually want touched, and that conversation is worth having before the first record moves, because retrofitting provenance onto a database that has already been enriched twice is a genuinely miserable job. If your contact records already carry a source and a date, this is straightforward. If they do not, adding that column is the project, and the enrichment is the easy part bolted on afterwards.
What it does not do, and should not pretend to
It does not verify a person. What verification establishes, on our own page and on every provider page we have read, is that a number is well formed, is in service and is not a duplicate. Those are useful checks and none of them is the check people assume, which is that this number reaches this person. Our own service page uses the word "verified" and this is the sentence that qualifies it.
It does not tell you how old a value is unless the provider is asked for it and passes it through. Freshness is the single most useful attribute an appended field could carry, and it is worth asking for by name, because a response that does not carry it looks exactly like one that does.
It does not promise a match, and a rate quoted before anybody has seen your list is a rate about somebody else. How much resolves varies with the town, with how thin the county file is behind a given address, and with how recently anybody moved. Two hundred of your own rows, run as a sample, settles the question for your own book in half a day.
It does not make a record legal to contact. Do-not-call registrations and consent rules attach to the call rather than to the data, they are not affected by how the number was obtained, and they are covered in detail in the article on reactivating an old database and in the skip tracing article above.
It does not tell you anybody wants to hear from you. A complete record is a reachable record. Interest is a different question, it is not in any file anybody can sell you, and a pass that makes four thousand people reachable has not produced four thousand conversations or any evidence about whether they would be welcome.
And it does not clean up after itself. If a pass writes a value that turns out to be wrong, the wrongness is now yours, sitting in your system of record with your business's authority behind it, and the provider's involvement ended at the response.
Three ways a pass that worked leaves you worse off
All three happen after the data arrives.
The database looks finished
Before the pass, the gaps were visible and everybody handled the list accordingly. Afterwards every row is full, and nothing in the interface distinguishes a number a client typed in herself from a number a file suggested. The database now carries the confidence of its best row and the accuracy of its worst, and the only person who could tell the two apart has stopped being able to.
Nobody can undo it
Six weeks later somebody notices a run of wrong numbers and wants to know which pass introduced them. If the write did not record the previous value, the date and the source, that question has no answer, and the only remaining option is to run another pass over the top of it and hope.
It is spent on people you will never call
Enrichment is charged per record, and a database that has been collecting for years holds a great many rows nobody has any intention of calling. Running the whole thing is the default because it is one click, and it converts a budget into fullness rather than into conversations. The version worth buying is the one where somebody chose the rows first, and choosing them costs an hour.
The honest read
Export twenty contacts you would swear you know, phone numbers and all, and send them to us with the numbers removed. We will run them and send back what an enrichment pass would have written into those fields, so you can see for yourself how many agree with what you already had, how many disagree, and how many come back empty.
It is twenty rows, it costs nothing, we do not need access to your CRM, and you keep the only copy that has the real numbers in it.
How to find out what your last pass actually did
Twenty minutes, on the database you already have. No tool and no subscription.
- Open five contacts you personally remember. People you have spoken to. Not the first five, and not the ones you have been working this week.
- For each one, ask where the phone number came from. Not whether it is right. Where it came from. If the record cannot tell you, you have already found the finding, and you have found it on the five people you know best.
- Look for a date beside any enriched value. Anything at all: an appended-on date, a source field, a note. If there is nothing, then every value in the database is the same age as every other value, which is to say unknown.
- Find a contact you know moved. Somebody who sold and left the area. Look at what the record says now, and decide whether the record knows they moved or is quietly still describing the person who lived there.
- Sort by the phone column and look at the gaps. Count how many rows are empty. That number is the actual size of the problem enrichment solves, and it is worth holding next to the number you were about to send.
- Ring one number that you did not put there yourself. Not to sell anything. Just to find out who answers. One call tells you more about your own data than any report, and if the answer surprises you, that is the finding.
- Ask whoever ran the last pass what the overwrite setting was. Not what should have happened. What the setting was, per field. If nobody can answer, treat that as an open question rather than as a reassurance.
Common questions, answered honestly
What is data enrichment, in plain terms?
It is a pass over contact records you already hold, which sends what you know to an outside provider and writes back what it holds: a missing phone number or email, the property detail behind an address, a merge of two records that turn out to be one person. It is sold as completing your database. What it actually does is add somebody else's assertions to your database, which is a genuinely useful thing and is not the same thing.
How is enrichment different from skip tracing?
One begins at an address and works towards a stranger. The other begins at somebody already sitting in your database and completes what you hold about them. The suppliers overlap and the technology is the same. What separates them is whether there was ever a relationship, that is the half carrying the legal weight, and it is set out in the skip tracing article rather than here.
Will enrichment overwrite the data I already have?
That depends on how the pass is configured, and it is worth establishing rather than assuming. On HubSpot, for instance, not overwriting a value you already hold is a per-property checkbox selected at import time, which tells you which way the tooling leans. Decide it deliberately: fill blanks only, keep both values in separate fields, or send disagreements to a person. Whichever you choose, insist that the previous value, the source and the date are written to the record, because without those there is no way back from a bad pass.
How accurate is appended contact data?
Nobody can tell you before running your list, and the figures that circulate do not survive being followed, which is worked through above. What you can do is measure it: take two hundred records where you already know the answer, run them, and count how many agree, disagree and come back empty. That takes an afternoon and it is a number about your own market rather than somebody's benchmark.
How fast does contact data go stale?
There is no published rate worth quoting. The circulating figures for contact data run from twenty to thirty percent, with up to seventy quoted for email addresses on their own, and every one of them traces back to a press release or a vendor page rather than to a study. What is knowable is the mechanism: a field stops being true when something happens to a person, so the rate is a property of the people in your database rather than of the data, and it will be different for a first time buyer in her twenties and a couple downsizing in their sixties.
Does an enriched list mean I can call it?
No, and the two questions are entirely separate. Whether a record is reachable is a data question. Whether you may contact it is a rules question with dates and registries in it, and the answer does not change because the number was appended rather than volunteered. Both of the neighbouring articles on this site cover that ground properly.
What should I ask a provider before signing?
Five short questions. Does the charge land on every record submitted or only on the ones that resolve. Does the response carry a date or an age for each value. For each field, is it observed or inferred. If a person tells you never to contact them again, what changes in your system. And will you run two hundred records from my own area before I commit to anything. Written answers to all five tell you what kind of company you are dealing with. A rate card tells you nothing.
What is the one thing that makes all of this manageable?
A source column and a date column on every contact record, filled in from the day the record is created, whether the value came from a form on your website, a conversation, or a pass. It costs nothing to add before you have data and it is close to impossible to reconstruct afterwards. Every hard question in this article becomes easy once those two columns exist.
What to do about it
Nothing in this article argues against enrichment. Two thirds of a database with no phone number in it is a real problem, and there is no version of solving it that does not involve buying somebody else's assertions.
What it argues is that the assertion should arrive wearing a label. Where it came from. When. What it landed on. Those three facts turn an appended value from something you have to trust into something you can weigh, and they cost one afternoon of build time to capture and are unrecoverable once a pass has run without them.
This sits beside everything else we build, on the RealtyLT AI page; what actually gets appended and checked is on the data enrichment page. If you would rather somebody looked at what your existing records already say about themselves before anything is bought, that is the AI audit.
A fuller database is easy. A database that can tell you where it got something is the one worth having.
Pick one contact you are certain about, somebody whose number you have used, and open their record. Then find the answer to a single question: where did that number come from, and when. If your CRM cannot tell you that about the one person you are sure of, it cannot tell you about the four thousand you are not, and every argument in this article is an argument about a column you have not added yet.
There is no figure here because the largest part of the bill is paid to somebody else, per record, and it moves with how many records you send and whether you are charged for attempts or for successes. What we would build around it is the smaller and duller half: deciding which records are worth running at all, writing the source and the date onto every row, and making sure a value that lands on a field which already had one goes somewhere a person can see. The AI audit is an hour, done with you, and for this topic it starts by looking at what your existing rows already say about themselves.




