Skip to content

10mornings

Four Assistants Ran Overnight. Nobody Read What They Did.

Nine good mornings, and on the tenth an assistant confirmed a showing you had already moved. What an AI agent workforce actually is, why an assistant that is right most of the time is a different product from one that is right every time, where multi-agent systems really fail, and who is accountable when one of them is wrong.

Levan Tsiklauri22 min read

The inbox assistant had drafted your morning replies for two weeks. Nine working days, nine sets of drafts, and every one of them was fine. Around day four you stopped reading them properly, which is not laziness. It is what anybody does with something that has been right nine times.

On the tenth morning it confirmed a Thursday walkthrough to a buyer's agent. You had moved that walkthrough to Friday the previous evening, in a text message, from the car. The assistant read the calendar, and the calendar said Thursday, because the calendar was where the walkthrough had been before you moved it.

The draft went out at 6:40am, in your voice, from your address, and it was polite and well written and completely wrong.

Nothing malfunctioned. The assistant was not confused and did not hallucinate anything. It read what it had access to, did the job it was given, and produced a piece of work indistinguishable in tone from the nine that were correct. That last part is the whole subject of this article, because it is the reason nobody caught it.

In short

  • An assistant that is right most mornings is not an assistant that is right every morning, and the gap between those two is much larger than it feels. The benchmark that measures it found a leading agent solving over 60 percent of tasks on one attempt and under 25 percent when all eight attempts had to be right.
  • Most multi-agent failures are not the model being stupid. In 1,642 annotated runs across seven frameworks, 44.2 percent of what went wrong came from how the system and the instructions were specified, and the single commonest mode was an agent repeating a step it had already completed.
  • Nothing you delegate to an assistant changes who is accountable for it. New York's licensing law lists the parties that may hold a real estate licence and they are all people and companies, and the provision on a broker's responsibility turns on what the broker knew and what the broker kept.

The run that reported success

Everything it read was true an hour earlier.

Staged for illustration. No real client, no real address and no real transcript: this is the shape of the failure, written out so the two tracks can be read side by side.

The morning, in order

  1. Confirming we are on for the walkthrough. What time works Thursday?

    The buyer's agent / Wed 5:12pm

  2. You, by text, from the car: Thursday is out, let us do Friday morning instead.

    Your side / Wed 9:40pm

  3. Just checking, are we all set? My clients are driving up.

    The buyer's agent / Thu 6:31am

  4. Sent in your name: Confirmed for today at 11. Looking forward to meeting them.

    Your side / Thu 6:40am

What the assistant did

  1. Read the thread

    The two emails in the thread, and neither of them mentions Friday. The change was made by text, in another app, and it never arrived here.

    6:38am
  2. Read the calendar

    One event, Thursday at 11, unchanged. The message from the car never reached it either.

    6:39am
  3. Checked the brief

    Confirm known appointments. Escalate anything ambiguous. Nothing here looked ambiguous.

    6:39am
  4. Reported success

    One reply sent, one thread closed, no warnings and no flags. The run log is green and it is accurate.

    6:40am

What an agent workforce actually is, and what it is not

The pitch is easy to say and it is broadly true. Instead of one general chatbot you have to brief every time, you set up several assistants, each pointed at one recurring job, each with access to the systems that job needs. One reads the overnight email and drafts replies. One pulls comps and builds the deck for tomorrow's listing appointment. One watches every open file for the signature nobody chased. They run at the same time, they do not stop at five o'clock, and adding another one is a configuration change rather than a hire.

What that description leaves out is the second half of the sentence, and the second half is where all the money and all the risk are. You have not removed work. You have changed what kind of work it is. Producing has become reviewing, and reviewing four streams of output is a real job with real hours in it, done by you, in the morning, before anything else.

That is not an argument against doing this. It is an argument for knowing what you are buying, and it is the thing every page in this category, including our own, has historically been vague about.

What you are actually buying

Two halves, and the pitch only contains the first one.

One assistant, one job, already briefed

The difference from a general chat window is not intelligence, it is standing context. An assistant configured for one recurring task holds the description of that task, the tools it needs and the standard you want, so you are not re-explaining your business every morning. That part of the pitch is true and it is worth having on its own.

Several at once, on a trigger

You can only do one thing at a time. Assistants have no such limit, they do not stop at five o'clock, and they start because something happened rather than because somebody remembered. Adding another is a configuration change. That part is true as well, and it is the half every page in this category leads with.

What you actually removed was the producing

The work did not disappear, it changed category. Drafting became reading. Building became checking. That is very often a good trade, because reading is faster than writing and it can be done in one sitting instead of scattered through a day. It is a trade rather than a saving, and a page that describes it as a saving is describing half of it.

And reviewing four streams is a job

One assistant is a habit. Four is a morning routine with real hours in it, done by the one person who can tell whether the output is right. Nobody costs that in, which is why the calculator further down this page works out the reviewing rather than the saving. It is the only number in this subject that is entirely yours.

Right once and right every time are different products

Here is the distinction that reorganises the whole subject, and it comes from the best public measurement of this that exists.

In June 2024 a team at Sierra published a benchmark called tau-bench. It is not a quiz. It puts a language agent into a simulated business with a real database, a set of tools that can change that database, and a written policy it has to follow, and then has a second language model play a customer who wants something. Two domains: a retail one with five hundred customers, a thousand orders and a hundred and fifteen tasks, and an airline one with three hundred flights, two thousand reservations and fifty tasks. At the end of each conversation the benchmark compares the actual state of the database against the one correct outcome. Not the transcript, not the tone. What ended up in the system.

The authors also proposed a measurement that nobody had been using, and it is the important part. Everybody had been reporting whether an agent succeeds at a task. They asked instead how often an agent succeeds at the same task every single time it is attempted, and they named it pass hat k: the chance that all k independent attempts are successful, averaged across tasks.

Read that against your own morning. You do not need an assistant that can draft a good reply. You need one that drafts a good reply on Monday and Tuesday and Wednesday and the Thursday you were in the car.

What happens when you run the same job twenty times

The best model in that paper solved more than sixty percent of the retail tasks on a single attempt. Run the same tasks eight times each and require all eight to be right, and the paper reports that the figure drops below twenty five percent.

Sit with the shape of that rather than the numbers, because the numbers are from June 2024 and the models have moved since. Something that succeeds most of the time on any given morning succeeds every morning far less often than most of the time, and the gap widens the more mornings you ask about. That is not a flaw anybody introduced. It is what happens when you multiply a probability by itself, and it is the reason a demo is such a poor guide to a purchase. A demo is one attempt. A business is a hundred attempts in a row.

The same paper is worth reading for one more reason: it looked at what the failures actually were. Of thirty six failed runs it examined by hand, the largest group was the agent calling the right kind of tool with the wrong values in it. Not a refusal, not an error message, not an apology. The correct action, confidently, on the wrong record. Which is what happened at 6:40 on the tenth morning.

Where these systems actually go wrong, and it is mostly not the model

The other paper worth your time is more recent and it is about exactly the thing the service page is selling, which is several agents working at once.

A group at UC Berkeley collected 1,642 annotated execution traces from seven different multi-agent frameworks, built a taxonomy of what went wrong by having six human experts read a hundred and fifty traces closely, and then checked that the taxonomy was reliable by having independent annotators apply it and measuring how often they agreed. Their agreement measure came out at 0.88, which is high, and it matters because a taxonomy nobody applies the same way twice is an opinion rather than a finding.

Fourteen distinct failure modes, in three groups. Their headline number is worth knowing before anybody quotes you: across the seven systems they measured a failure rate between 41 percent and 86.7 percent.

The evidence

Where the failures came from, in 1,642 annotated multi-agent runs

Where the failures came from, in 1,642 annotated multi-agent runsSystem design and specification: 44.2%. Agents misaligned with each other: 32.3%. Verification of the result: 23.5%. The share of classified failures falling into each of the three categories, across 1,642 annotated execution traces drawn from seven multi-agent frameworks running coding, mathematics and general agent tasks. The taxonomy underneath it was built by six human experts reading 150 traces closely and then tested for consistency between independent annotators, which is the step that separates a taxonomy from an opinion.System design and specification44.2%Agents misaligned with each other32.3%Verification of the result23.5%

The share of classified failures falling into each of the three categories, across 1,642 annotated execution traces drawn from seven multi-agent frameworks running coding, mathematics and general agent tasks. The taxonomy underneath it was built by six human experts reading 150 traces closely and then tested for consistency between independent annotators, which is the step that separates a taxonomy from an opinion. Source: Mert Cemri and twelve co-authors, UC Berkeley, Why Do Multi-Agent LLM Systems Fail?, arXiv:2503.13657, 2025.

These are research frameworks running research tasks, mostly programming and mathematics, and not four assistants in a brokerage. Two limits matter. The bars are shares of the failures that happened, not a probability that anything will fail, so a system with very few failures and a system with very many can produce the same chart. And the taxonomy was applied at scale by a language model calibrated against the human annotations rather than by the humans themselves, which the authors state plainly. What transfers is the ranking rather than the percentages: the largest category is how the system and its instructions were specified, which is the half a buyer controls and the half nobody quotes on.

The chart is the finding. The largest group is not the model being stupid. It is system design: the job being specified badly, the agent repeating a step it had already done, the agent not knowing when it was finished. Their single most common individual mode, at 15.7 percent of everything, is step repetition. The second is the agent's stated reasoning not matching the action it then took. The third is not recognising the conditions under which it should stop.

None of those are fixed by a better model, and all of them are fixed by somebody thinking harder about the instructions and the checks. The paper's own observation about this is the practical one: the systems in their sample that had explicit verification steps built into them showed fewer failures overall.

The brief is the product

The tau-bench authors ran one more experiment which almost nobody talks about, and for a reader about to pay for any of this it is the result in either paper that matters most.

They took the written policy out of the agent's instructions and ran everything again. In the simple domain, where the rules are close to common sense, performance fell from 61.2 to 56.8 percent, which is barely anything. In the complicated domain, where the rules are specific and arbitrary in the way real business rules are, it fell from 33.2 to 10.8 percent.

The evidence

What the same agent scored with its written rules taken away

What the same agent scored with its written rules taken awaySimple task set, rules provided: 61.2%. Simple task set, rules removed: 56.8%. Complex task set, rules provided: 33.2%. Complex task set, rules removed: 10.8%. The share of tasks completed correctly on a single attempt by one leading model, measured by comparing the state of the database at the end of each conversation against the one correct outcome. All four figures come from the paper's ablation table, in which the written policy is removed from the agent's instructions. The simple task set is a retail domain of 115 tasks; the complex one is an airline domain of 50 tasks with rules that vary by membership tier and cabin class.Simple task set, rules provided61.2%Simple task set, rules removed56.8%Complex task set, rules provided33.2%Complex task set, rules removed10.8%

The share of tasks completed correctly on a single attempt by one leading model, measured by comparing the state of the database at the end of each conversation against the one correct outcome. All four figures come from the paper's ablation table, in which the written policy is removed from the agent's instructions. The simple task set is a retail domain of 115 tasks; the complex one is an airline domain of 50 tasks with rules that vary by membership tier and cabin class. Source: Shunyu Yao, Noah Shinn, Pedram Razavi and Karthik Narasimhan, tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, arXiv:2406.12045, 2024.

This was measured in June 2024 on models that have all been replaced, and every absolute number here would be different today. Read the pairs rather than the heights. Two more limits are worth knowing: the customer on the other side of every conversation is another language model rather than a person, which makes the conversations tidier than real ones, and the domains are simplified versions of real businesses rather than real ones. What survives all of that is the comparison the ablation was built to make, which is how much of an agent's usable ability is coming from a document somebody wrote rather than from the model underneath it.

Take the second pair seriously. Two thirds of what that agent could do came from a written document, not from the model. Which means the thing you are actually buying, when you buy an assistant, is the document: the description of the job, the rules, the exceptions, the things that must never happen. The model is a commodity and it improves every few months without you doing anything. The brief is yours, it is specific to your business, and nobody else can write it.

This is also why the honest version of the sales process is slower than the exciting one. "Tell us the job and we will build the assistant" sounds like a five minute conversation, and it is a two hour one, because most recurring jobs have never been written down and the first hour is spent discovering the exceptions that live only in somebody's head.

An open plan room with exposed timber ceiling beams, a long polished wooden dining table on the left and an orange fronted kitchen island with metal stools on the right, and in the centre a white cabinet whose front is a dark chalkboard panel carrying a handwritten list in pale green chalk reading milk, dog food, coffee, bread, cheese and soap
Six lines on a board and anybody in this house can do the shopping without asking a question. That is what a brief is, and it is the entire difference between an assistant that is useful and one that is fast and plausible and slightly wrong. Nobody can write yours except you. Photograph by Jeremy Levine Design, CC BY 2.0.

Why the second assistant costs more than the first

Everything above is about one assistant. Running several is not the same thing repeated, and the difference is worth being clear about because it is the difference the word "workforce" hides.

Two assistants that never touch each other are genuinely just two assistants, and the cost is roughly double the cost of one. That describes a lot of useful setups and there is nothing wrong with it.

But the moment one assistant's output becomes another assistant's input, you have built a system, and the Berkeley taxonomy has a whole category for what goes wrong there: information one agent held and did not pass on, an agent carrying on with an assumption instead of asking, an agent quietly drifting off the task it was given. Almost a third of everything they classified sat in that group. A handover between two pieces of software is not free, and it is exactly the place where an error stops being visible, because the second agent receives a confident summary rather than the thing itself.

The practical rule that falls out of this is dull and it is worth more than any feature list. Keep the assistants independent unless there is a specific reason not to, and where one has to feed another, make the handover something a person can read.

The system

Six hops, and two of them are not software at all.

The first and the last hop are the ones that decide whether this works, and neither of them is anything you buy. Writing the job down is where most of the assistant's usable ability comes from, and the read at the end is the only control in the whole diagram, because nothing else in it can tell you that the output was wrong.

The path from a job written down to a draft that a person reads before anybody sees it: The job, The brief, The access, The run, The draft, The readThe jobWritten down firstThe briefRules and exceptionsThe accessOnly what it needsThe runOn a trigger, aloneThe draftMade, never sentThe readA person, on a fixed day

Scroll to follow the chain

  1. The job: Written down first
  2. The brief: Rules and exceptions
  3. The access: Only what it needs
  4. The run: On a trigger, alone
  5. The draft: Made, never sent
  6. The read: A person, on a fixed day

Where the money actually goes when you run several

The cost of running an agent is not what you would guess, and the tau-bench paper measured it, which almost nobody does.

For each task their best setup handled, the agent cost 38 cents and the simulated customer on the other side cost 23 cents. That is not your price list and it should not be read as one. The number underneath it is the one that transfers: of what the agent cost, the input took 95.9 percent and everything the agent actually wrote took 4.1 percent.

In plain terms, almost the entire running cost of an assistant is it re-reading its own instructions, its tool definitions and the conversation so far, over and over, before every single thing it says. It is not being paid to write. It is being paid to remember.

That has three consequences you can act on. A longer brief costs money every time the assistant runs, so the discipline is a brief that is complete rather than a brief that is long. An assistant that is given access to ten tools it never uses is paying to read the descriptions of ten tools it never uses. And an assistant handling a long conversation gets more expensive with every turn, which is why a job that ends is cheaper than a job that lingers.

In your numbers

How many hours a year would you spend reading their work?

One per recurring job. Count the jobs you could describe to a new starter on their first day without stopping to think.

3

A draft, a summary, a deck, a chase. Count things somebody would have to look at, not steps the assistant takes.

10

Be honest about the number you would still be hitting in week six rather than the one you intend in week one.

60%

Properly means against the brief rather than for tone. Skimming for tone is what stops catching anything by week two.

3

Hours a year reading what they produced

47hours a year

Assistants runningyour 3
3assistants
Pieces of work eachyour 10
30a week
Over a year52 weeks
1,560pieces a year
That you would actually readyour 60%
936pieces to read
At your reading timeyour 3
2,808minutes a year
In hours60 minutes in an hour
47hours a year

This is the only calculator on this website that works out what the service costs you rather than what it saves you, and that is deliberate: the saving is the part everybody already estimates and the reviewing is the part nobody does. There is no hours-saved row because it would be a guess. Nobody has published a measurement of how long an assistant's draft takes a person to check in this industry, and every input above is therefore yours. There is also no row comparing this with a salary, and there was never going to be one. The published median wage for an administrative assistant is a real number and it is quoted further down this page, but it buys accountability and judgement and somebody who notices that the job has changed, and dividing it by anything here would be arithmetic on two things that are not the same purchase.

You did not remove the work. You changed it from producing into reading, which is usually a good trade and is never a free one, and the reading is done by the only person who can tell whether it is right.

What a person costs, and why you cannot divide by it

Every page in this category eventually reaches for a salary, and ours did too. So here is a real one, from the only source for it that publishes its method.

The United States Bureau of Labor Statistics reports that the median annual wage for secretaries and administrative assistants was $47,460 in May 2024, which is $22.82 an hour. Median means what the Bureau says it means and it is worth quoting, because half of this category's arithmetic depends on people not knowing: the median wage is the wage at which half the workers in an occupation earned more than that amount and half earned less. For context, the median for real estate sales agents was $56,320 over the same period, and the median for all occupations was $49,500.

Now the part where the arithmetic stops. That $47,460 buys something with properties an assistant does not have. It answers the phone when the caller is upset. It notices that the job it was given last March is no longer the job that needs doing. It can be told once. It can be held responsible. And it is a whole person rather than a set of tasks, so removing four tasks from that job does not remove four quarters of the salary.

There is also a number in the same table that nobody selling this will mention. The Bureau's projection for that occupation between 2024 and 2034 is zero percent growth, a change of minus 12,400 jobs out of roughly three and a half million. Not a collapse. Essentially flat, in the published forecast of the agency whose job is forecasting it.

So this page will not divide one of those numbers by the other. The honest comparison is not assistant against employee, because they are not substitutes. It is your morning with the assistants against your morning without them, which is a question about your own time and not about anybody's salary, and it is the question the calculator above is asking.

Who is responsible when an assistant is wrong

The email that went out at 6:40 was signed with your name. Everything else follows from that.

This industry is unusual in having already written down what happens when work is delegated, because delegating licensed work is already a regulated activity in New York. It is worth reading two provisions in the Department of State's own Real Estate License Law booklet, because neither is about artificial intelligence and both are about you.

Section 442-c deals with what a salesperson's misconduct means for the broker. A broker is not automatically on the hook for what an associate did. But there are two ways they become so, and the second is the one to read twice: a broker is exposed where they had actual knowledge of the violation, or where they retain the benefits, profits or proceeds of a transaction wrongfully negotiated by their salesperson or employee after notice of the misconduct. Keeping what the conduct earned is the thing that attaches you to the conduct.

Then read Section 440-a, which is the requirement to be licensed at all. It lists who may hold a licence: a person, a co-partnership, a limited liability company, a corporation. That list is a list of parties that can be disciplined, sued and struck off. It is not a list a piece of software is on, and nothing here is a prediction about future law. It is a description of the present one, and the description is that when an assistant working for you says something to a client, there is exactly one licensed party in the conversation and it is you.

What supervision looks like when the thing you are supervising is software

There is a second document worth borrowing, and this one is borrowed openly as an analogy rather than applied as a rule. It is about supervising people and it says nothing whatever about software.

Section 175.21 of the Secretary of State's regulations defines what supervising a salesperson actually consists of, and rather than leaving it to judgement it writes it down: regular, frequent and consistent personal guidance, instruction, oversight and superintendence, with respect to the brokerage business and all matters relating to it. The next paragraph requires written records of what the salesperson actually did.

Nobody is claiming that provision governs an inbox assistant. What it does is describe, in a document your regulator wrote, the standard this industry already applies to work done in your name by somebody who is not you. Regular. Frequent. Consistent. Written down.

Set that beside four assistants running overnight with nobody reading the output after day four, and you have the honest specification for what running this well requires. Not a dashboard. A habit, with a time in the diary, and a record of what was produced.

A living room with a wooden fireplace mantel carrying two large carved elephants and a row of turned candlesticks under a tall mirror, a fern in a white pot on the hearth in front of a black firebox whose surround is painted white, a five panel door standing open onto a further room with a window and a sofa in it, and glazed French doors at the right throwing sunlight across a polished wood floor
The door is open and the light is on the floor and nobody has walked through yet. Everything an assistant produces sits exactly like this until a person goes and looks at it, and the going and looking is not a temporary precaution for the first month. It is the shape of the job now. Photograph by smoMashup1, CC BY 2.0.

The honest read

Send us one recurring job in your own words, however roughly. We will send back the brief we would write for it, the exceptions we think it needs, and an honest opinion on whether it is delegable at all yet.

Send us one job

It is a short reply from a person, it costs nothing, and not yet is a perfectly good answer to give about a job.

What it costs, and how long it takes

Nobody can quote this from an article, because three separate things drive the cost and only one of them is the software.

The first is the brief, and it is the slow part. Writing down a job properly, exceptions included, is a sitting rather than a message, and it goes faster when a second person keeps pushing back on the first version you offer. The tau-bench ablation is the argument for spending that time rather than skipping it.

The second is access. An assistant that can read your calendar and your CRM is worth several times one that cannot, and the work is connecting it safely: the right permissions, nothing wider than the job needs, and a way to switch it off.

The third is the running cost, which is usage rather than a seat, tracks how much the assistant has to read rather than how much it writes, and is genuinely small per task and genuinely unbounded if nobody watches it.

What this page will not print is a per-model price, and the reason is not that the numbers are hidden. They are published, they are readable, and you can look them up in a minute. It is that they are quoted per million tokens, which means nothing until somebody knows how many tokens your job takes; that they change several times a year; and that which model sits behind an assistant is a build decision that can be changed without anything visible happening at your end. A figure typed into an article would be stale before the article was, and it would not have answered the question you asked. What can be said is the shape: the cost per piece of work is in cents rather than dollars, it is driven by the length of the instructions rather than the length of the answer, and the honest budget line is the review time above it rather than the compute.

What it does not do, and should not pretend to

It does not take responsibility. An assistant cannot be told off, cannot learn from being told off in any way that persists unless somebody edits the brief, and cannot be the person a client complains to. Every consequence lands on a licensed human being, and that human being is you.

It does not notice that the job has changed. This is the quietest failure of the four. A person who has been drafting your listing emails for a year will eventually say that the market has moved and the second paragraph now reads badly. An assistant will produce that second paragraph forever, with perfect consistency, until somebody rewrites the brief.

It does not do a job nobody has written down. A vague brief does not produce vague output, which would at least be a visible signal. It produces confident, fluent, plausible output that is subtly not what you wanted, and you find out three weeks later from a client.

It does not remove the reading. Anything that reaches a client should be read by a person first, an assistant that drafts is worth more than one that sends, and the review is not a temporary safety measure for the first month. It is the job now.

And it does not scale the way the word workforce suggests. Four independent assistants are four times the review. Four assistants feeding each other are four times the review plus a category of failure that only exists because they are connected, and the published taxonomy has six named modes inside it.

Three ways a good build produces nothing

None of them are the model.

The review that stopped happening in week two

It does not stop because anybody decided to stop. It stops because nine good mornings in a row teach you that the tenth will be fine, which is exactly the lesson a system with a high single-attempt success rate is built to teach. The control that matters is not a decision to be careful, it is a slot in a diary that survives a busy fortnight.

A brief nobody has reopened since the build

The document is the thing that makes the assistant useful and it is the only part of the setup that goes stale, because your business moves and the brief does not. A year-old brief produces a year-old standard with total consistency, and nothing in the output looks old, which is why nobody notices until a client mentions it.

Assistants wired to each other for no reason

Chaining one assistant's output into the next feels like progress and it buys a whole category of failure that independent assistants do not have: information held and not passed on, an assumption carried forward instead of questioned, a task quietly drifting. Unless there is a reason for the handover, two separate assistants and a person in the middle is the cheaper machine.

How to test one assistant before you run four

Do this with one assistant, on one job, before anybody builds you a set of them. It runs over a couple of weeks, it costs you nothing but attention, and it will tell you more than any demonstration.

Pick the dullest job you have that repeats, and write the brief before you look at any software. Include the exceptions. If you cannot write it, you have learned the most useful thing available today, which is that the job is not yet delegable to anybody, software or human.

Then run it against work you have already done. Take ten pieces of last month's output that you produced by hand, give the assistant the same inputs, and compare. This is the only honest accuracy test, because you already know the right answer and you are not grading it on how confident it sounds.

Then run the same task ten times and count how many times all ten are right. Not the average. All ten. That is the pass hat k measurement from the research above, done by hand, and it is the single number that predicts whether you will still be reading the output in week six.

Then break it on purpose. Give it an input with a contradiction in it, or a case the brief does not cover, and find out whether it stops and asks or whether it decides. An assistant that decides in the ambiguous cases rather than stopping will go on deciding, and one of those decisions reaches somebody outside your office before you have seen it.

Last, put a time in your diary. Fifteen minutes, the same slot every week, to read a sample of what it produced and check it against what you asked for. If you cannot find that slot for one assistant, you have your answer about four.

Common questions, answered honestly

What is an AI agent workforce, in plain terms?

It is a set of AI assistants, each configured for one recurring job and each connected to the systems that job needs. Rather than one general chatbot you brief from scratch every session, you have several that already know your business and their own task, and they run in parallel, on a schedule or a trigger, without you opening anything.

How is this different from workflow automation?

Workflow automation joins the software you already pay for, so the end of one step is what begins the next, and most of the steps in a good automation carry no judgement in them at all. An agent is what you reach for when a step genuinely needs a decision made from context. Both are often assembled from the same parts, and putting an agent where a simple rule would have done buys unpredictability nobody asked for.

How is this different from just using ChatGPT?

Three things, and the second is the one that matters most. A general chat session starts empty and cannot reach your systems. An assistant carries a written brief for one job, which the research above suggests is where most of its usable ability comes from. And it runs on a trigger rather than waiting for you to remember it, which is the difference between a tool and a member of staff.

How many assistants can I actually run at once?

Technically as many as you have jobs for, because they do not queue behind each other. Practically the limit is not the software, it is how many streams of output one person can review before the reviewing stops happening. It is a number worth working out rather than assuming, and the calculator above is there to let you find yours before you commit to it.

Do I need technical skills?

No, and the skill you do need is not technical. You need to be able to describe a job precisely, including what should happen in the cases that are not the normal case. That is a writing and thinking exercise rather than a software one, and it is the part that cannot be outsourced, because the exceptions live in your head.

What happens when one of them is wrong?

It produces something wrong that reads exactly like everything it produced when it was right, which is why the answer has to be structural rather than attentive. A build that is serious about this stops and asks rather than guessing on anything ambiguous, keeps a readable record of what it did and why, drafts rather than sends anything client-facing, and has a review step somebody actually performs.

Is it cheaper than hiring somebody?

That is the wrong comparison and we are not going to make it. The published median wage for an administrative assistant buys accountability, judgement and somebody who notices when the job changes, and an assistant provides none of those. The comparison that is real is your own week with and without, including the review time, which is what the calculator above works out.

How do I know it is working after the first month?

Not from a dashboard. Count how many pieces of its output you actually read last week and be honest about it, because the review is the control and a review nobody performs is not one. Then take a sample and check it against the brief rather than against your impression, and every quarter re-read the brief itself and ask whether it still describes the job you have now.

What to do about it

Go and read yesterday's output. Not the summary of it, the actual pieces, all of them, one after the other, with the brief open beside you. It takes twenty minutes and it is the only way to find out whether you have four assistants or four unread inboxes.

Then write one brief for one job, properly, before anybody sells you anything. The research says most of what an assistant can do for you comes out of that document, and the document is the part nobody else can write, and you can write it today for nothing.

The assistants are listed on the RealtyLT AI page; what each one is pointed at is set out on the AI agent workforce page. Working out which of your recurring jobs are genuinely delegable is what the AI audit does: we take one real job, write the brief with you, and build that one first.

The individual jobs are written up on their own: answering the website at midnight, picking up the phone at 9:42 on a Sunday, and the wiring between the tools that makes any of it possible.

Nine good mornings are not a track record. They are nine mornings.

Take one job you keep redoing and write the brief for it this week, before anybody sells you anything. Include what should happen in the cases that are not the normal case. If you cannot finish it, you have learned the most valuable thing available today, which is that the job is not yet delegable to anybody at all.

There is no price here because three separate things drive it and only one of them is software: the hour or two of writing each brief properly, the work of giving an assistant safe access to the systems its job needs, and a running cost that tracks how much it has to read rather than how much it writes. The AI audit is an hour, done with you, and it ends with one job written down properly rather than with a document.

Know somebody who would argue with this? Send it to them.

  • August 27, 2026

    The Answer Was Wrong in March. It Was Still Wrong in October.

  • August 26, 2026

    It Ran Every Morning for Two Years. Then a Field Came Back With a New Word in It.

  • August 26, 2026

    You Had Eleven Ideas. The Hour Crossed Four of Them Off.