Skip to content

14videos

Fourteen Videos Went Out in Your Face. You Have Watched None of Them.

Fourteen statements were published in your name to fourteen people who now believe you said them. Whose face and voice a business may reproduce, why New York's oldest privacy statute makes the wrong version a misdemeanour, what happened when 315 people tried to tell synthetic faces from real ones, and the cost of a digital twin that nobody quotes.

Levan Tsiklauri24 min readUpdated

On Tuesday afternoon fourteen short videos went out, one to each person who had asked about the new listing on the ridge. Every one of them opened with that person's name. Every one of them was in your face and your voice, standing in a room you have been in, saying things you would say.

You were at a closing in Poughkeepsie for most of the afternoon and you have not watched any of them.

Nothing went wrong. That is worth sitting with for a second, because the thing that makes this topic difficult is not the disaster, it is the ordinary Tuesday. Fourteen statements were published in your name, to fourteen people who now believe you said them, and the only person who could have caught a wrong one was in a title company's conference room signing things.

This article is about what a digital twin is allowed to do, which turns out to be a much more useful question than what it is able to do. The ability is not in doubt any more. Whether a particular use of somebody's face is lawful, whether the viewer has to be told, and who is left holding it when a sentence is wrong, are all questions with actual answers, and most of them are shorter than you would expect.

In short

  • There are two completely different products under the same word. One reproduces your own face and voice, with your own written permission, on your own content. The other reproduces somebody else, and in New York that second one is not a grey area: using a living person's name, portrait, picture, likeness or voice for advertising or trade without their written consent first is a misdemeanour under a section of this state's Civil Rights Law.
  • Federal law is narrower here than it looks. The Federal Trade Commission's impersonation rule, in force since April 2024, prohibits falsely posing as a government body or as a business or an officer of one. It does not cover individuals. The Commission has proposed adding them, and has proposed making it a violation to supply goods or services knowing they will be used to impersonate, which would reach the people who build these things as well as the people who use them. Neither addition has been adopted, so today the federal rule leaves individuals to state law.
  • Whether the viewer can tell is not a matter of opinion. In a study of 315 people classifying real and synthetic faces one at a time, average accuracy was 48.2 percent against a chance level of 50. A second group of 219 people, given training and told after every single answer whether they were right, reached 59.0 percent and got no better with practice.

What a digital twin actually is, and the line everything rests on

A digital twin is two separate things sold as one word. There is a model of a face, built from a recording of somebody sitting still and talking, that can afterwards be made to say new words. And there is a model of a voice, built the same way from speech, that can afterwards read a script that person never read. Both are ordinary technology now, both are available to anybody, and neither is interesting on its own.

What is interesting is the second word. A twin of WHOM.

Almost every difficult question in this subject collapses into that one, and it separates cleanly into four cases that behave completely differently in law and in practice. Confusing them is how a business acquires a problem it cannot buy its way out of, because the cheap version and the illegal version look identical from the outside and come out of exactly the same software.

Four likenesses

Only one of these is a product.

Your own face, on your own content

You sat in front of a camera, you agreed in writing what may be made from the recording, and what goes out is signed with your name on a page you control. This is the ordinary case and it is the only one this business builds. The interesting questions about it are not legal ones, they are about quality and about who checks what goes out.

A colleague's face, on the brokerage's content

Also ordinary, and it needs one more thing than people expect: their written permission, separately from their employment, and an answer to what happens to the model when they leave. An agent who moves to another firm has not stopped owning their own likeness, and a library of videos of a person who no longer works for you is a problem you built yourself.

Somebody who never agreed

A client, a seller, a person on the other side of a transaction, a public figure whose endorsement would be useful. There is no version of this that is a product. It is the thing several statutes exist to stop, and one of them attaches a criminal penalty to it in this state.

Somebody who has died

The one people assume is free, and in New York it is the most precisely regulated of the four. A separate statute in force since 2020 defines a digital replica, gives the right to whoever inherited it, and lets an action be brought for up to forty years after the death. There is a public register of who holds those rights.

New York's answer is a statute, and it is a criminal one

This is not a doctrine that grew up quietly in the courts and has to be inferred from a line of cases. It is Civil Rights Law section 50, the whole of it fits in a single sentence, and that sentence is short enough to read in a breath.

A person, firm or corporation that uses for advertising purposes, or for the purposes of trade, the name, portrait, picture, likeness, or voice of any living person without having first obtained the written consent of such person, or if a minor of such minor's parent or guardian, is guilty of a misdemeanor.
Large letterpress type blocks lying face up and packed tightly together, photographed at a low angle so the rows recede, a few of them pale bare wood and most of them darkened with ink, with lowercase t, a, o and p legible across the middle of the frame, a question mark among the blocks along the top edge, a row of figures at the right hand edge, and a fold of pink and white checked cloth at the bottom right corner
Every one of these blocks is a copy. Somebody cut a master letter once, and everything printed afterwards came from a duplicate of it, which is why the same q could be set in a thousand shops at once and still be that q. What the case does not settle is who is allowed to pick a block up. That has never been a property of the type. Photograph by Kyle Van Horn, CC BY 2.0.

Three words in that sentence do the work, and every summary of it softens at least one of them.

The consent has to be WRITTEN. Not implied by a working relationship, not implied by somebody having been happy about it last year, not implied by an email that says sure. A written permission is the only kind this section recognises, and the reason that matters for a twin is that a twin is a thing you keep. Permission given for one video is not permission for a model that can make a thousand.

The consent has to be obtained FIRST. You do not get to let the video go out on Tuesday and have the paperwork catch up on Friday, because the offence is complete at the moment of the use.

And the section is CRIMINAL. It says guilty of a misdemeanour, which is not how most people picture a marketing dispute. The civil half sits next door in section 51, which lets the person go to the supreme court of this state for an injunction and for damages, and adds that where the use was knowing, the jury in its discretion may award exemplary damages. Knowing is not a hard standard to meet when the whole product is a deliberate reproduction of a specific person.

Two practical readings for somebody running a brokerage. If the face is yours, this section is a paperwork exercise you do once and file, and the paperwork is worth having anyway because it forces the questions about scope that nobody asks otherwise. If the face belongs to anybody else in your office, it is the same exercise with a second signature on it, and it should say what happens to the model when they leave, because section 50 does not stop applying on the day somebody changes firms.

Nothing in this article is legal advice, and this is exactly the paragraph to take to somebody whose advice it is.

What happens to a likeness after the person has died

The assumption is that death ends it. In New York the opposite is closer to true, and the statute that says so is recent enough to be easy to have missed.

Civil Rights Law section 50-f has been in force since 2020, on the version history its own page carries, and is titled, plainly, right of publicity. It creates a property right in a deceased personality's name, voice, signature, photograph and likeness, defining a deceased personality as a person domiciled in this state at death whose likeness had commercial value at the time of, or because of, their death. And it does something the older sections never had to: it defines the thing this article is about.

A digital replica, in the statute's own words, is a newly created, computer generated, highly realistic electronic representation that is readily identifiable as the voice or visual likeness of an individual, embodied in a sound recording, image, audiovisual work or transmission, in which either the individual did not actually perform, or did perform but the fundamental character of the performance has been materially altered. That is a careful definition and the second half is the part people miss. Altering what somebody actually said, past the point where it is still their performance, is inside the definition as surely as inventing it from nothing.

The mechanics matter less than the fact that there are any, and three of them are worth carrying. The rights are property, so they are inherited, sold and licensed like anything else. Anybody claiming to hold them can record that claim on a public register kept by the secretary of state, which means there is somewhere to look. And the right runs for forty years after the death, with damages starting at two thousand dollars or the actual loss, whichever is larger.

There are broad exemptions, and they are the reason a documentary or a satire is not caught: works of political or newsworthy value, parody, satire, commentary, criticism, biographical work with some degree of fictionalisation. What is not exempted is the case that would tempt a business, which is using a likeness to sell something.

So the honest summary for a brokerage is short. A dead person's face is not free, in this state it is somebody's property for forty years, there is a register you could check, and none of that is a project you want to be near.

The federal rule that covers businesses and does not yet cover people

There is a federal rule about impersonation and it is newer than most of the software. It is also narrower than almost everybody assumes, and the gap is worth knowing precisely rather than roughly.

16 CFR part 461 was published in March 2024 and took effect on the first of April, under the title Rule on Impersonation of Government and Businesses. It has three short sections. It makes it an unfair or deceptive act to materially and falsely pose as a government entity or officer, or as a business or officer of one, in or affecting commerce, and equally to materially misrepresent affiliation with, endorsement by or sponsorship by either. Its definitions section is where the reach is: officer includes executives, officials, employees, and agents.

Read that against your own trade for a moment. A video that appears to be an agent of a brokerage, made by somebody who is not, is already within the plain words of this rule, because an agent of a business is an officer of it for these purposes.

What the rule does not cover is an individual, and that is not an oversight. On the same day the final rule was published, the Commission published a supplementary notice of proposed rulemaking proposing to add exactly that. The proposal would rename the rule to cover individuals, define an individual as a person, entity or party, whether real or fictitious, other than a business or government, and add a new section making it a violation to materially and falsely pose as an individual in or affecting commerce.

The second half of that proposal is the one this business has to read carefully, because it is about us rather than about you. The Commission also proposed a means and instrumentalities section: that it would be a violation to provide goods or services with knowledge or reason to know that those goods or services will be used to impersonate. A rule in those terms would put the people who build a likeness inside the same rule as the people who publish one, and the standard is not intent, it is reason to know.

One thing has to be said plainly because it is the sort of claim that ages badly on a website. As of this writing, that proposal is still a proposal. The text in force at 16 CFR part 461 today is the three sections about government and businesses, and it has no section 461.4 in it. That was checked against the Code of Federal Regulations itself rather than against an article about it, and it is the kind of thing that could change between this being written and you reading it.

Nobody can reliably tell, and that is measured rather than assumed

There is a comfortable belief in this industry that a client would know. It is the belief a great deal of the informal ethics of AI video quietly rests on, and unlike most of what gets said in this area it has actually been tested.

Two researchers at Lancaster University and the University of California, Berkeley ran a set of perceptual studies with real participants and published them in the Proceedings of the National Academy of Sciences in 2022. They took four hundred synthesised faces and matched each one to a real photograph of a similar person, then asked people to sort them.

The evidence

How often people correctly sorted synthetic faces from real ones

How often people correctly sorted synthetic faces from real ones315 people, no training: 48.2%. 219 people, trained and told the answer every time: 59.0%. Percentage of faces correctly classified as real or synthesized, from two experiments on the same set of 800 faces, half of them generated and half of them real photographs matched to them for age, gender and appearance. Chance performance is 50 percent and is not drawn, because a coin is not a measurement. The first group saw the faces cold. The second group was trained first and told after every single answer whether they had been right, and the paper records that they got no better across the session.315 people, no training48.2%219 people, trained and told the answer every time59.0%

Percentage of faces correctly classified as real or synthesized, from two experiments on the same set of 800 faces, half of them generated and half of them real photographs matched to them for age, gender and appearance. Chance performance is 50 percent and is not drawn, because a coin is not a measurement. The first group saw the faces cold. The second group was trained first and told after every single answer whether they had been right, and the paper records that they got no better across the session. Source: Nightingale and Farid, AI-synthesized faces are indistinguishable from real faces and more trustworthy, Proceedings of the National Academy of Sciences, 2022.

Two honest limits on what this can be used for. These were still photographs of people who do not exist, not video of somebody the viewer knows personally, and a face you have met every week for three years is a different task from a stranger's. And the same paper found the synthetic faces were rated slightly MORE trustworthy than the real ones, 4.82 against 4.48 on a seven point scale, which is a small effect and an uncomfortable one. What none of this measures is whether your own clients would spot a video of you, and nobody has published that.

The result is not that people are bad at this. It is that there is nothing there to be good at. The trained group had been shown what to look for and were told after every single answer whether they had got it right, which is about as favourable a condition as anybody could set up, and they ended the session no better than they started it. The paper attributes that to some of the synthetic faces simply containing no perceptible artefact to find.

The natural response is that a machine should do the checking instead. That has been tested too, at a scale no individual company could manage.

In 2020, Facebook AI built and released a dataset of over one hundred thousand video clips made from three thousand four hundred and twenty six paid actors, and ran a public competition on it. The dataset is worth a sentence of its own for a reason that belongs on this page: the authors record that all recorded subjects agreed to participate in and have their likenesses modified during the construction of it, and note in the same paper that many previously released datasets in this field did not guarantee that. Two thousand one hundred and fourteen teams entered.

The evidence

The winning detector's precision on real-world fakes, at three settings

The winning detector's precision on real-world fakes, at three settingsSet to catch 1 in 10 of them: 98.0%. Set to catch 3 in 10: 76.1%. Set to catch 9 in 10: 53.9%. Precision, meaning the share of the videos it flagged that really were fakes, for the first placed entry in the DeepFake Detection Challenge, measured on real videos gathered outside the competition's own dataset rather than on the ones it was trained against. The three bars are the same model at three sensitivities, reported by the organisers at recall levels of one tenth, three tenths and nine tenths. Turning it up to catch more fakes is what makes it flag more things that were not fakes.Set to catch 1 in 10 of them98.0%Set to catch 3 in 1076.1%Set to catch 9 in 1053.9%

Precision, meaning the share of the videos it flagged that really were fakes, for the first placed entry in the DeepFake Detection Challenge, measured on real videos gathered outside the competition's own dataset rather than on the ones it was trained against. The three bars are the same model at three sensitivities, reported by the organisers at recall levels of one tenth, three tenths and nine tenths. Turning it up to catch more fakes is what makes it flag more things that were not fakes. Source: Dolhansky and others, The DeepFake Detection Challenge (DFDC) Dataset, Facebook AI, 2020.

The shape is the finding, not the decimals. A detector you can tune has a dial on it, and turning the dial toward catching more fakes is the same movement as turning it toward accusing more honest videos. At the setting where it caught nine in ten, about half of what it pointed at was innocent. Two things this cannot be stretched to say. It is a 2020 competition against 2020 fakes, and both sides of that race have moved since, in directions this article has no measurement of. And on the hidden test set the organisers report that 60 percent of all submissions scored at or better than predicting a coin flip on every video would have scored, with many of them simply random. That says more about the difficulty than any single number here does. The reason it is on this page at all is that it removes an excuse: you cannot leave the disclosing to a detector.

The organisers explain why they report precision rather than accuracy, and their reason is the useful part for a business. In realistic distributions, they write, the ratio of faked videos to real ones may be less than one in a million. When almost everything is genuine, a detector that is very accurate still produces far more false alarms than real catches, because there is so much more innocent material for it to be wrong about. That is a permanent property of the arithmetic rather than a temporary weakness in the models.

Put the two studies together and one conclusion falls out that nothing since has softened. Neither the audience nor the software can be relied upon to work out what a video is. Which leaves exactly one party who reliably knows.

So the telling has to come from you

Once you accept that the viewer cannot tell, the disclosure question stops being a matter of taste and starts being the only mechanism there is. The good news is that it is cheap. The interesting news is that there is a technical standard for the durable version of it, and reading what that standard says about itself is more instructive than reading anything written about it.

The Coalition for Content Provenance and Authenticity publishes an open technical specification for attaching signed, tamper evident provenance to a media file. It is a serious piece of engineering with serious companies behind it, and the useful thing about reading the document itself is how carefully it describes its own limits.

Provenance marking

A signature on a file is not a statement that the file is true.

It records what was done, cryptographically

A Content Credential is a set of signed statements travelling with the file: what created it, what was done to it afterwards, and whether any of that has been altered since it was signed. There is a specific value for a file that came out of a generative model, and it is a machine-readable string rather than a phrase somebody chose, which is the part that makes it checkable at all.

It does not tell anybody the video is honest

The specification refuses that job in its own guiding principles, and the wording is worth reading twice. It says the specifications should not provide value judgments about whether a given set of provenance data is good or bad, merely whether the assertions included within can be validated as associated with the underlying asset, correctly formed, and free from tampering. Signed and true are different words.

And it can simply be absent

The trust decision rests on the identity of whoever signed the claim, so a file with no credential at all is not evidence of anything, and a great deal of ordinary honest video has none. Anybody stripping provenance on purpose is not going to be stopped by a standard. Which is why the useful version of this is not detection, it is you saying so in the first frame.

The sentence quoted in the middle of that scene is from the specification's own scope section, and it is worth noticing what a standards body chooses to refuse. It will not tell you whether provenance data is good or bad. It will tell you whether the assertions in it are correctly formed and have not been tampered with, and it says elsewhere that the basis for any trust decision is the identity of whoever signed.

That is a standard describing its own limits accurately, and it is the reason the technical answer and the practical answer are different. The technical answer is a credential that only works if somebody chooses to inspect it. The practical answer is a sentence at the start of the video in which a person says, in their own words, that this was recorded once and assembled by software. It costs four seconds, it cannot be stripped, and it converts the entire problem from something a viewer might discover into something you told them.

What we will and will not build

Everything above is about the world. This section is about us, and it is the one thing on this page that needs no citation because it is a commitment rather than a claim.

We build a likeness of the person who sat in front of the camera and said in writing that we could. We do not build a likeness of anybody else, living or dead, for any reason, at any price, and we would rather lose the work than argue about it.

The /ai page already says that we do not build agents that pretend to be a specific human being, and that was written about voice on a telephone. It applies with more force to a face. A likeness of the person who sat in front of the camera and agreed in writing is a tool. A likeness of anybody else is the thing four separate bodies of law are pointed at, and no amount of the client being a good one moves that answer.

What that means concretely is a short list of things that are not negotiable, and the useful way to read it is as an order rather than as a policy. Each step below only makes sense if the one before it happened, which is why the ones at the end are the ones that get skipped.

What a twin honestly does for a brokerage

Strip out the excitement and there are three real jobs, and all three are smaller and more specific than the pitch.

It removes the recording session from things that were always scripted anyway. A market note, a new listing walkthrough, a short explanation of what happens after an offer is accepted. These are things where the words matter and the performance does not, which is exactly the case where a twin loses nothing.

It makes an individually addressed version affordable. Not better than a personal video, just possible at a count where a personal video is not. This is the honest version of the pitch: fourteen people get something with their own name in it instead of one blast, and the alternative was never fourteen real recordings, it was one email.

And it gets a face onto material that would otherwise have been text. A page of written follow up and a person saying the same words are not equally likely to be read. That is why so many agents intend to record more than they do.

What none of those three needs is for anybody to be deceived. Every one of them survives being labelled, which is a good test of whether a use is a legitimate one: if telling the viewer would ruin it, the use was never about saving you a recording session.

The cost nobody quotes, which is watching them

Here is the part that does not appear in any pitch, and it is the reason the fourteen videos at the top of this page are the story rather than an anecdote.

A twin does not reduce the amount of judgment a business has to apply. It moves that judgment from before the recording to after it. When you record something yourself, the checking happens automatically, because you cannot say a sentence without hearing it. When a script is generated and a model reads it, nothing about the process forces anybody to listen to the result, and the volume that makes the thing worth having is precisely the volume that makes reviewing it feel disproportionate.

In your numbers

How long would it take to watch everything that goes out in your face?

New listings, price changes, open houses, market notes, anything you would currently record yourself for if you had the time.

40

One general version, or one addressed to every buyer who asked. This is the number the whole category is sold on.

6

Watching it, not skimming the script. The point of checking is to catch the sentence that reads fine and is wrong.

2

Hours a year to watch everything that goes out in your face

8hours a year

Things you would make a video aboutyour 40
40a year
Videos going out in your faceyour 6
240videos
At your review timeyour 2
480minutes
In hours60 minutes in an hour
8hours a year

This is the only number in the whole topic that is definitely yours, and it is deliberately small at the settings it opens with. That is the argument rather than a weakness in it. A few hours a year is not a reason to say no to anything, which means there is no honest excuse for the second row going out unwatched, and the second row is the one that matters: those are statements published in your name to people who will remember them as yours. Watch the second row rather than the headline as you move the sliders. Four things this refuses to put a number on. There is no response rate, no reply rate and no conversion figure for personalised video anywhere on this page, because every figure of that kind that could be found is published by a company selling the software and none of them states a sample. There is no comparison against how long it takes to record one yourself, because nobody has timed that either. There is no dollar value, because the value depends entirely on what the video is for. And there is no estimate of how many viewers would notice, because the research on this page measures a different task.

What actually happens to that number is quieter than a refusal. The reviewing gets done for the first fortnight, then it gets done sometimes, then somebody says the last forty were all fine. Nobody ever decided to stop.

The fix is not technology and it is not a bigger budget. It is naming a person and a moment. Somebody watches, before it sends. Where no one has been given that duty by name, the honest response is to produce fewer videos rather than to pretend the watching is happening.

What nobody has measured, and what this page will not print because of it

Every article about personalised video carries a number about how much better it performs. This one does not, and the absence is deliberate enough to be worth explaining.

The figures in circulation for what a personalised or AI generated video does to a reply rate, a click rate or a conversion rate come, without exception among the ones that could be traced, from companies that sell video software, and none of them states a sample, a method or a control. That is the same pattern as the data broker figures another article on this site had to refuse, and the same answer applies: a number with no method behind it is not a small number, it is not a number.

There is a second gap and it is more specific. The perceptual research quoted above measures whether people can sort synthetic faces from real ones. It does not measure whether your own clients, who have met you, would recognise a video of you as synthetic. Familiarity is a different task and it might cut either way. Nobody has published it, so this page says nothing about it.

And there is a third thing this article deliberately does not describe, which is how any particular avatar is produced, what software makes it, or how convincing the result is. Those are questions about a specific build rather than about the subject, and an article that answered them would be a product sheet wearing an argument.

How to test one before you commission it

Four questions, none of them technical, and the first two are about paper.

Ask to see the consent document. Not a description of it, the document. It should name the person, say what may be made from the recording, say how long that permission lasts, and say what happens when they revoke it. If the answer is that consent is handled in the terms of service of a tool, that is not a consent document, that is a company protecting itself.

Ask what happens when somebody leaves. A brokerage that builds twins of four agents has built four assets that belong to four people, and two of those people will work somewhere else within a few years. The right answer involves deletion and it involves somebody being responsible for it.

Ask who decides what it is allowed to say. Watch for whether the answer describes a person or a prompt. Both are real answers, but only one of them is a person, and the sentences that get a brokerage in trouble are about schools, boundaries, taxes and permits rather than about anything the model would consider risky.

Ask how the viewer finds out. The correct answer is a sentence in the video. Anything about metadata or credentials as the primary mechanism is a description of something no viewer will ever see.

The honest read

Tell us the three things you would put in front of a camera every week if the recording took no time at all. We will tell you which of them actually needs your face, which would be better as two lines of writing, and which one is the sort of thing you should never let a script write on your behalf.

Send us the three

It is a short reply from a person, it costs nothing, we do not need a recording session to answer it, and the answer is often that one of the three is not worth building.

An antique wind-up gramophone standing against a plain yellow wall, its wide metal horn opening toward the left of the frame and mottled with age, the horn's neck curving down to a metal fitting on a reddish wooden box with a winding crank projecting from the right of it, all of it standing on a darker cabinet with canted corners and two small metal fittings on its front
This is a machine for sending a voice somewhere its owner has never been. Nobody thought the singer was in the room, and nobody was fooled, because the horn is enormous and the whole object announces itself. That is the part worth keeping rather than the technology. A reproduction that is obviously a reproduction has never needed anybody's permission to be honest about what it is. Photograph by Vince Alongi, CC BY 2.0.

What it costs, and how long it takes

This divides into a part that is easy to quote and a part that decides whether the money was well spent, and only the first part ever appears in a proposal.

The predictable half is the recording session and the model. It is a short session, it happens once, and the ongoing cost is a per minute or per video charge from whichever platform renders the result. Vendor prices in this category move quickly and are quoted per product tier, so no figure is printed here, and what actually drives yours is volume and length rather than anything clever.

The unpredictable half is everything either side of it. Deciding what the twin may say, writing the script boundaries so they hold up when a listing has something unusual about it, and building the review step into the day so it survives month three. That work is a conversation rather than a build, and it decides whether you end up with a library of useful videos or a library nobody has watched.

The recurring cost that is genuinely easy to underestimate is attention, and the calculator above is the honest way to size it. Everything else in this topic scales with volume and gets cheaper. That one scales with volume and does not.

What it does not do, and should not pretend to

It does not make you present. A twin can deliver information and it cannot notice that somebody is worried, which is most of what an agent is actually for in the moments that matter.

It does not speak for anybody who has not agreed in writing. That is a legal boundary in this state before it is a policy of ours, and the statute attached to it is a criminal one.

It does not decide what is safe to say. That list has to exist before anything is recorded, and writing it is a job for a person who has been in the transactions.

It does not remove the need for somebody to watch what goes out. It increases that need, because it increases the volume, and a build that is working perfectly goes wrong here more often than it goes wrong anywhere else.

And it does not stay honest by itself. A twin used with a sentence of disclosure is a tool. The same twin without that sentence is the same object doing something else entirely, and nothing in the software knows the difference.

Three ways a working twin costs you something

None of them are the technology failing.

Nobody watches them any more

The first fortnight, every video gets checked. By the second month it is a pipeline, and a pipeline is exactly the thing whose output stops being read. The failure is not that the twin says something outrageous. It is that it says something slightly wrong about a property, in your face, to a person who now believes you said it.

It answers a question it should have refused

A script generated from a listing will happily state a school district, a tax figure, a boundary or a permit status, and every one of those is a thing an agent gets asked and answers carefully. A twin has no sense of which sentences are the expensive ones, so the sentences it should decline to say have to be decided in advance by a person.

The library outlives the arrangement

A colleague's model, a client's testimonial, a video made for a listing that has since sold twice. Somebody has to own the question of what gets deleted and when, and if nobody does, the answer is nothing, forever, including the material of people who have long since left.

Common questions, answered honestly

What is an AI clone for a real estate agent?

It is a model of your face and a model of your voice, built from a recording of you, that can afterwards deliver new scripts as video without you being in front of a camera. It exists to remove the recording session from material that was always going to be scripted anyway. What it does not do is give you a second person: everything it says still has to be decided, bounded and checked by you.

Yes, and the paperwork is worth doing properly anyway. New York Civil Rights Law section 50 requires written consent obtained in advance before a living person's likeness or voice is used for advertising or trade, and when the person is you that is a document you sign once. Do it properly rather than informally, because a consent for one video is not a consent for a model that can produce a thousand, and the useful part of the exercise is being forced to write down what may be made.

Do I have to tell people the video was made with AI?

Treat it as required regardless of what any particular rule says today, because the research is clear that the viewer cannot work it out. A single spoken sentence at the start costs four seconds and converts something a person could discover into something you told them. There are technical provenance standards for marking a file as machine generated, and they are useful, but they are not a substitute: no client is going to inspect a credential.

Can I use a video of an agent after they have left my brokerage?

Ask, and get the answer in writing before you need it. Their likeness is theirs, not the firm's, and section 50 keeps applying after they change firms. The practical version is that the consent document should have said what happens on departure, and if it did not, the safe assumption is that the permission ended with the relationship.

What about a testimonial in a client's voice or face?

A real client, recorded with their written permission, is ordinary marketing. A synthesised version of a client, or words they did not say assembled into their voice, is not a grey area in this state. The digital replica definition in section 50-f explicitly covers materially altering a performance somebody actually gave, so editing what they said into something better is inside the same territory as inventing it.

Will my clients be able to tell it is not really me?

Assume not. In published research, 315 people sorting synthetic faces from real ones averaged 48.2 percent against a 50 percent coin flip, and 219 people who were trained and told the answer after every attempt reached 59.0 percent and got no better with practice. Those were still images of strangers rather than video of somebody familiar, so the transfer is not exact, and there is no published measurement of the familiar case. Plan for the version where nobody notices.

Who owns the model of my face and my voice?

Split that in two, because the halves have different answers and running them together is how this gets sold wrong. The avatar and the voice engine belong to the vendors who built them, and a brokerage licenses the use of a platform the same way it licenses everything else it did not build. Nobody is selling you a model you own. What is yours is the material at both ends: your likeness, the recording the model was built from, the scripts, and the finished videos. So the three lines to get in writing before the recording happens are not about owning the software. They are that your likeness and your recording are used only on your own content, that they are not licensed on to anybody else or used to build anything for anybody else, and that both are deleted on request. A vendor unwilling to write those three lines down has told you something useful.

Is this worth it for a one or two person brokerage?

Sometimes, and the test is not size. It is whether you already have material that is scripted, repeated and currently not being recorded because the recording is the bottleneck. If the honest answer is that you would not have made these videos at all, a twin is worth considering. If the answer is that you would have recorded them yourself and just have not, the twin is being asked to solve a discipline problem, which is not what it is.

What to do about it

Take twenty minutes before you take a camera.

Write down the three things you would put on video every week if recording cost nothing. Then write, beside each one, the single sentence in it that would be expensive to get wrong. A tax figure, a school, a boundary, a timeline, a condition of the property.

That second column is the actual specification for the build, and it is the part no vendor will write for you. It says which sentences a script may generate freely, which ones have to come from a source rather than from a model, and which ones a person has to approve before anything sends.

Then decide who watches. Not whether. Who.

Before you record anything, write two lists. On the first, every sentence you are happy for a machine to say in your voice without you hearing it first. On the second, every sentence you would want to be in the room for. The second list is usually the shorter one, and it tends to contain the only sentences that were ever worth saying on camera.

There is no price on this page and the reason is unusually specific to this topic: the recording session and the model are the small, predictable part, and the work that decides whether the thing is any good is the script boundary and the review step, which are yours to set and take a conversation rather than a quote. The AI audit is an hour, done with you, and for this topic it starts with the two lists above rather than with a camera.

Know somebody who would argue with this? Send it to them.

  • August 25, 2026

    Nine Town Pages. The Only Thing That Changed Was the Town.

  • August 25, 2026

    Three Businesses Show Up. Yours Is Not One of Them.

  • August 25, 2026

    Twelve Five-Star Reviews. The Newest One Is From 2023.