Shadow mode: proving an agent works before it touches anyone

Most agent projects that die do not die in the build. They die in the week after go-live, when the owner asks whether it worked and nobody can answer. The agent sent some emails. Two placements happened. Would they have happened anyway? Nobody wrote anything down before it was switched on, so the conversation becomes a matter of taste, and taste loses to the next priority.

Shadow mode fixes that. The agent runs on real data on the real schedule, decides what it would do, and writes that decision to a table instead of acting on it. Nobody is emailed, no ATS record is changed, no candidate is scored for a live decision. After a few weeks you have a log you can count, and the go/no-go argument is about numbers instead of impressions.

This walkthrough builds the shadow harness for a back-office agent — a redeployment radar or a timesheet chaser — and the four numbers you read off it.

Step one: write down the decision you are testing

One sentence, before any code. Not "automate redeployment". Something a row in a table can represent:

Each weekday, for every contractor whose assignment ends within 45 days and who has had no recruiter contact logged in 14 days, draft outreach and name the two best-matched open reqs.

That sentence fixes three things: the trigger, the population, and the action. It also fixes what a wrong answer looks like, which matters more. A draft for a contractor who resigned last Tuesday is a false positive. A contractor who finished and left unworked is a false negative. You cannot measure either until the sentence exists.

Step two: one row per would-have-done action

The shadow table is the whole experiment. Keep it append-only and keep the inputs with the output, because in week three you will want to know why the agent liked a match you do not.

create table shadow_decisions (
  id             bigserial primary key,
  run_id         uuid        not null,
  agent_version  text        not null,   -- git sha of the rules and prompts
  decided_at     timestamptz not null default now(),
  subject_type   text        not null,   -- 'placement' | 'timesheet' | 'credential'
  subject_id     text        not null,
  action         text        not null,   -- 'draft_outreach' | 'chase' | 'no_action'
  reason         text        not null,   -- why it fired, in words
  suppressed_by  text[],                 -- rules that stopped it, if any
  inputs         jsonb       not null,   -- the fields as of decided_at
  proposed       jsonb,                  -- the draft text, the matched req ids
  label          text,                   -- filled in later by a human
  labelled_by    text,
  labelled_at    timestamptz
);
create index on shadow_decisions (subject_type, subject_id, decided_at);

Two details that are easy to skip and expensive to add later. Record no_action rows, not only the firings: half of the useful evidence is what the agent declined to do. And stamp agent_version on every row, because you will change the rules mid-experiment and you need to be able to count the periods separately.

Step three: backtest first, but distrust it

Before running forward, replay the last quarter. It is cheap and it kills bad ideas in an afternoon. It also lies to you in one specific way: leakage. Your ATS holds today's values, not January's. If a placement's end date was quietly extended in March, a naive replay sees the extended date and congratulates itself for a prediction it never made.

So replay from snapshots, not from current rows. If you have been landing daily snapshots of placements, you already have what you need (that pipeline is the one in pulling end dates from the JobAdder API). If you have not, your backtest is limited to fields that do not get rewritten — timesheet submission times, credential expiry dates as originally recorded, message timestamps — and that is fine. Say so in the write-up rather than quietly overclaiming.

A backtest answers one question honestly: how many events would this have fired on? If the answer is eleven in a quarter, stop. No agent pays back eleven drafts.

Step four: run forward, and label a sample

Now run it live-but-silent for four to six weeks on the real schedule. Then the part everyone tries to automate and should not: a human labels a sample.

Twenty rows a week, picked at random from the firings, reviewed by the recruiter who owns those contractors. Three labels, no more:

LabelMeaning
usefulI would have sent this, or acted on it, roughly as drafted
editRight call, wrong content — the match or the wording needed work
noiseI would not have done this at all

Twenty a week is about fifteen minutes of a recruiter's time and gives you 100-ish labelled rows by the end. Resist labelling everything; a complete review is a review nobody finishes. Also pull a false-negative sample: list contractors who finished during the window with no firing against them, and ask the same recruiter which ones should have been caught.

Step five: the four numbers

Read these off the table, and write them down in the same place every week.

Coverage. Firings divided by eligible events. If 60 assignments ended and the agent fired on 22, coverage is 37%. Low coverage usually means data, not logic — missing end dates, contact notes kept in someone's inbox.

Useful rate. useful plus edit over all labelled firings. Under about 60% and recruiters will stop reading the queue, which is a slower, more expensive failure than the agent being switched off.

Lead time. Days between the agent's firing and the date a human did the same thing unaided, when they did. This is where the redeployment case actually lives — not "the agent found it" but "the agent found it nineteen days earlier".

Hours. Minutes saved per useful row, times useful rows per month. Ask the recruiter for the minutes; do not model them. Then put it against the run cost from your metering table (what an agent run costs) and the build hours.

select agent_version,
       count(*) filter (where action <> 'no_action')                   as fired,
       count(*) filter (where label in ('useful','edit'))              as useful_ish,
       count(*) filter (where label = 'noise')                         as noise,
       round(100.0 * count(*) filter (where label in ('useful','edit'))
             / nullif(count(*) filter (where label is not null), 0), 1) as useful_pct
from shadow_decisions
where decided_at >= now() - interval '42 days'
group by agent_version
order by agent_version;

Step six: write the go/no-go rule before you start

Pick the thresholds in week zero, while nobody is attached to the result. Ours usually look like: coverage above 70% of eligible events, useful rate above 70% of labelled firings, and a defensible hours number that clears the monthly run cost at least three times over. Numbers that clear the bar mean you turn writes on, through an approval queue, not straight to send — that queue is the next build, not an afterthought.

Numbers that miss get one of three answers, and "not yet" is a legitimate one:

  • Fix the data first. Low coverage from unrecorded extensions or contact notes living in email is a data project, and it pays back on its own.
  • Narrow the scope. One desk, one skill vertical, one client program. Agents that fail across the whole firm often work on the healthcare desk where the end dates are actually maintained.
  • Drop it. Eleven firings a quarter does not become a business case because the build is already half done.

What shadow mode does not tell you

It does not tell you whether recruiters will act. A queue everyone ignores scores beautifully in shadow and does nothing for revenue; that only shows up after go-live, which is why the post-launch dashboard matters as much as this one.

It does not excuse you from records. If what you are shadowing is candidate scoring rather than back-office chasing, the shadow run is still processing candidate data, and the scores must not reach anyone making a decision — not in a spreadsheet, not "just to see". The moment they do, it is a tool used in an employment decision and the notice, audit and retention obligations in AI screening under NYC LL144 and Illinois HB 3773 attach. Shadow mode is a measurement technique, not a shelter.

And it does not replace scoping. It checks a ranked candidate workflow that a workflow assessment already said was worth building. Shadowing six ideas at once is just a slower way of not deciding.

Four to six weeks and about an hour a week of recruiter time buys you a sentence with numbers in it: on 100 labelled drafts, 74 were useful, average nineteen days earlier than the desk, about nine hours a month. That sentence survives a budget conversation. "It seems to be working" does not.

If you want the harness built against your ATS and your snapshots, tell us which system you run and how many contractors are on assignment.