This post is about rolling out AI-assisted development on a large legacy codebase: what moved, what didn’t, and what I got wrong. Most of it applies to any team trying to make AI coding tools work on code that’s older than some of the people working on it.

The product was a twenty-year-old C++ and C# application partway through a migration from a classic desktop architecture to a modern web one. That was deliberate. Greenfield AI adoption already works. The open question was whether it works on the code that pays the bills.

The approach was to build a harness for that codebase and hand it over to the team. A harness is the set of rules, skills, hooks and evaluations that sit in the repository and shape how Claude Code behaves inside it: which patterns it follows, which exemplars it copies, which checks it runs before it calls something done. A generic coding agent knows how to write C#. A harness tells it how this team writes C#.

I kept a log throughout, one row per result with a before value, an after value, and who did the work. Several rows carry a correction note where I got something wrong and went back. Those corrections taught me more than the wins did, so they get a lot of space below.

Mine the reviewer, not the style guide

Every change to this product passed through one required reviewer. That person’s review queue was the team’s biggest bottleneck, and the comments repeated from PR to PR.

The obvious move is to write a coding standards document and hope the agent follows it. I took a different route. I pulled 459 of the reviewer’s comments across 300 pull requests, clustered them into recurring themes, ranked them by frequency, and turned the top ones into a pre-review checklist with a small set of hard gates. The harness rules were derived from what the reviewer actually corrects, not from what a style guide says they should care about.

Then I measured. Before adoption, the reviewer left 5.3 comments per PR across 45 PRs by the four developers using the harness. After adoption, across 31 PRs by the same four people on the same repository, they left 0.35. Twenty-five of the 31 were approved with zero comments.

The first post-adoption sample was only 13 PRs, which is small enough to wave away. So I re-pulled once the sample had more than doubled. The rate improved instead of drifting back toward the old number. The least favourable reading I could construct, counting two boundary PRs reviewed the day before the harness existed, was still about a tenfold drop.

The comments also changed in kind. On the reviewer’s first pass over the new code, the developer reported that the questions weren’t about conventions anymore. They challenged whether the functionality was complete. That’s the review you want a senior engineer spending time on.

I want to be precise about attribution here. The developer hadn’t yet run the pre-review skill on those PRs. The improvement came from the harness rules shaping the code as it was being written. That’s a better result than the one I designed for, and it would have been easy to claim the wrong cause.

Later the reviewer said in a group meeting, with their manager present, that the code arriving for review was what they were looking for, that the comments were minimal, and that they could see a clear difference in their workload. They also started asking for their own preferences to be added to the harness as rules. When the reviewer begins maintaining the harness, it has stopped being an outside tool.

The bottleneck moved, and the team moved with it

One developer delivered an entire backend module in two days: 22 subtasks across 8 stacked PRs. When I asked how long it would have taken working the way the team normally works, the developer’s own estimate was 7 to 10 days. That figure is theirs, not mine. It’s net of a half hour where the agent went the wrong way, and the two days include integration tests they would normally have left to QA. Every one of the module’s 11 events was exercised end to end before QA saw it, and two real defects surfaced along the way, one of them long-standing.

The week that happened, two developers independently told me the same thing: code was being produced much faster, and review took as long as it always had. The constraint didn’t disappear. It moved.

Leadership had asked for this kind of shift to be surfaced rather than absorbed quietly, so I wrote it down and watched. I didn’t have to do anything else. Within a week the team had escalated for a second reviewer on their own. A developer proposed splitting the role so the existing reviewer covered frontend and someone else covered backend. Their managers routed it.

That’s the outcome I’d want from any enablement work. The team saw the new constraint and reorganised around it themselves, which means they’ll keep doing it without anyone prompting them.

A smaller version of the same lesson came earlier. A defect PR had been stuck for two weeks over a reviewer disagreement. The technique wasn’t the problem. The argument had moved to a call and never came back to the PR thread, so the thread had no record of the decision. Once that was fixed, the PR closed. AI speeds up writing code. It does nothing for a decision that never got written down.

The budget wasn’t the constraint. Fear was.

The early adopters had a monthly AI spend allowance. They were using a fraction of it, and I assumed they didn’t yet have enough work that justified more.

In a one-to-one I found out the real reason. They weren’t pacing themselves because of the cap. They were worried about being audited on spend that produced nothing.

I raised it with the engineering leader backing the rollout, who posted in the team channel encouraging people to use it and doubled the early adopters’ allowance. One of them used the entire raised budget in two days, the same two days the backend module above was delivered.

Spend limits get treated as a statement about what’s acceptable, whatever the policy says. If you want people to experiment, someone senior has to say so in public, specifically, and more than once. A policy document won’t do it.

A harness is a list of claims, and some of them are wrong

This is the section I’d most want anyone building a harness to read.

A harness rule is a confident statement about a codebase. Some of mine were confident about things nobody had measured. Here are the ones that got caught:

  • An over-strict data access ban, a testing rule that assumed code was mockable when it wasn’t, and a workflow command the team expected that didn’t exist. All three were found by the team within weeks and fixed within a day.
  • Two UI helper methods that don’t exist. My frontend rules told developers to call them. Following my instruction produced a call that throws.
  • Deprecated components recommended as the default. I’d built the component catalogue by ranking components on how often they appeared in templates. Frequency is historical. It can’t tell you what’s currently correct. The reviewer caught it when reviewing the harness itself, and the answer removed three of the mistakes we’d been logging, because the component they pointed us to picks the right control from metadata on its own.
  • The wrong test framework. My testing rules said the solution used one framework. When I counted, 48 test projects used three different ones, and the one I’d named was the rarest. An agent reading that rule would have added the wrong kind of test project.
  • “The legacy C++ can’t be unit tested.” This one did the most damage. It started in my own testing conventions file early on and spread into the governance framework and then into a formal exemption mechanism I proposed. An engineer on the team challenged it. I checked the repository instead of defending the claim. The core abstraction was a genuine interface with dozens of pure virtual methods, a test project already existed, and the real obstacle was one line that took the interface by injection and immediately cast it back to the concrete type. That’s fixable code, not an architectural limit. I withdrew the exemption mechanism before it was formalised.

What these have in common: a rule stated as fact gets followed as fact, by the agent and by every human who later copies it into another document. A wrong rule doesn’t stay in the file where you wrote it.

What I’d do differently from day one:

  • Back each rule with evidence from the repository: a count, a file, a review comment. If you can’t point to it, label it as an assumption.
  • When the mined data and the written standard disagree on how severe something is, go with the data. One of the developers withdrew a rule after finding that nothing in the 459 review comments backed it up. That made something else visible: our invariants file only had two settings, hard gate or absent, with no way to say “probably right, weakly evidenced”. Every team writing these files will run into that.
  • Expect the team to correct you and make that easy. The single best early signal was the team finding and fixing defects in my harness. One of them was serious: the verification script reported a pass without checking anything when it was run from outside the repository root, and that had been true since the first commit.

Watch for checks that pass without checking

That verification script turned out to be the first of several things that looked like checks and weren’t.

  • A silent install. One developer had copied the harness files in by hand. The slash commands never registered, the pre-review step did nothing, and the developer shipped without seeing any review output, with no error anywhere. The project’s own setup docs had told new developers to copy files by hand. I fixed the docs and built an onboarding skill whose first job is to confirm the harness is actually installed and working before it teaches anything.
  • A test suite nobody ran. I later added CI to the organisation’s standard harness, which already had a decent test suite and evaluation corpus. Nothing ran them automatically. The first CI run went red: one test had been failing for 20 days with 60 commits landed on top. The guard was correct. Nothing ever invoked it. The harness team fixed all four defects that CI surfaced within 26 hours.
  • My own CI matrix. I’d tested the oldest and newest Python versions the package claimed to support. When I installed the versions in between, two of them failed. A test that only covers the ends of a range can’t see a defect in the middle. I reported it and the harness team fixed it, with a better root cause than mine.
  • A tag nothing enforced. The QA team’s biggest pain was a group of payment tests that failed at every regression run. Sixty-seven tests were tagged as not safe to run in parallel. Nothing enforced the tag: the configuration allowed parallel files, the pipeline ran four workers, and the code that would have serialised them was bypassed by the way CI invoked the runner. All four files modified the same financial records on a shared staging database. In the same config file, two timeouts were multiplied by 60,000 where a third used 1,000, a units mistake, so a single hung click could hold a worker for half an hour. I also corrected my own first diagnosis, which had blamed transient slowness and recommended retries. If the cause is data collisions, retries hide the bug.

The pattern is the same each time: the protection exists and looks fine, and nothing exercises it. Whenever someone tells me a check exists, I now ask when it last failed. If nobody knows, it may not be running at all.

Measure before you adopt

As the harness pattern spread to other parts of the organisation, sixteen or more products onboarded the organisation’s standard harness, and none had recorded a before-state. Once a team has adopted, the clean before-window is gone, and per-PR review comments can’t be rebuilt from git history alone.

So when a second product line was about to adopt, I spent the days before it started pulling 456 pull requests and 2,731 review events from the API, and saved the extraction script so the after-measurement would use the same method.

The first thing the baseline showed was that the result I was proudest of wouldn’t carry over.

Review load at the second product had already fallen 2.9x, or 2.4x once normalised for PR size. Five of seven engineers were already using AI informally. Their real starting point was 2.30 comments per PR, not 6.78. Reviews were also spread across eight people, with no single required reviewer, so the strongest argument from the first product didn’t apply at all.

I also had to correct my own first reading. One person performed 80% of merges, which looked like a review bottleneck. They ranked sixth by review volume. They were integrating, not reviewing.

There’s one more caveat in that data: 38% of recent PRs had zero comments, against 26% before. That’s either cleaner code or lighter review, and the numbers can’t tell you which. I wrote that caveat next to the figure so nobody quotes the figure without it.

Something similar happened with CI. I sampled a month of runs on the second product and none had passed. I wrote it up as a finding. At the next standup the team explained that they already knew. A known set of long-failing tests was tech debt they compared against by hand on every PR. The measurement was correct. The recommendation changed: automate the comparison against the known failing set instead of reporting a problem the team already understood.

Measuring first stopped me setting an expectation I couldn’t meet. If I’d carried the first product’s tenfold result into the second, I’d have spent the following weeks explaining a shortfall that existed before any of the new work started.

The most valuable opportunities weren’t in the plan

Several of the findings that mattered most came from looking somewhere nobody had asked me to look.

  • A live production risk. While mapping what the legacy module did so the migration could match it, I found a trigger that never read the configuration setting it depended on. For some customers, notification emails may already have been failing silently, unrelated to the migration.
  • Gaps in work that was already done. Running a parity check against a module the team had already migrated turned up about nine real gaps, including a missed financial control, with no false positives once the check was hardened.
  • Duplicate testing driven by trust. After QA signed off a release, the support team re-tested it for about two weeks, because things had broken after testing before. The fix they wanted, folding support’s scenarios into QA’s automated suite, was blocked by one thing. Support’s scenarios weren’t written down anywhere. They were held by people with decades on the product. Getting that knowledge into an automatable suite is worth roughly two weeks of expert time per release. No tooling decision addresses it.
  • Teams already ahead. The QA automation team had built its own Claude harness alongside delivery, independently of this work. Their manager reported that the regression cycle had gone from 2.5 to 3 weeks with double the current headcount down to 4 to 5 days. The suite also grew over that period, so it isn’t all AI. Test authoring went from about 20 cases a month to more than 200 in two sprints. That’s their result, not mine, and it’s useful precisely because I had no hand in it.

The scope also shrank in places. Three of five planned governance deliverables turned out smaller than expected, because the teams had already solved them or because the premise was wrong. I reported that as a reduction instead of padding the remaining work to look even. One colleague’s correction stayed with me: governance is an organisational problem, not a technical one. You can’t settle with a template a question that’s really about who owns what.

The handover is the deliverable

A harness that depends on the person who built it doesn’t count as a result. These were the things that made this one stay.

Ship it on the main line, not as an opt-in. The team decided to merge the harness into the release branch every team branches from. Nobody has to install it, and there’s no separate copy to keep in sync. That beat every option originally on the table, all of which left it opt-in.

Remove objections before they’re raised. The reviewer had asked whether the harness would interfere with other teams using AI in the same repository. We’d already changed the skills so they only run when invoked explicitly. When the question came up in the decision meeting, it was already answered, and the decision went through without debate.

Commit the AI working files. The reviewer’s first position was that work-specific markdown files from AI sessions shouldn’t be committed. They changed their mind when the case was put as caching: those files are the spec the code was built from, so the next person fixing a defect starts there instead of re-reading the codebase.

Name an owner from the team. Succession came up directly, not as an afterthought. The developer who had contributed most to the harness was proposed as owner, their manager endorsed it, and maintenance was folded into the normal release cycle so it doesn’t need its own process. I wrote a maintenance guide covering who owns it, what belongs in the harness versus what’s a product decision, and every case where the harness had asserted an ideal instead of the team’s reality.

Within two days, the new owner was talking about onboarding as many other teams as possible to get better feedback, and planned to review the merge PR personally to find where it could fail. Nobody asked them to.

What I’d tell anyone starting one of these

  1. Pick the hard codebase. Greenfield proves little. Legacy proves the method.
  2. Build rules from real review history, then measure the reviewer’s load before and after, on the same people and the same repository.
  3. Expect the bottleneck to move to review, and let the team reorganise around it themselves.
  4. Ask about spend fear, not just spend limits. Get someone senior to give permission in public.
  5. Treat every harness rule as a claim that needs evidence. Label assumptions as assumptions. Make it easy for the team to prove you wrong.
  6. Ask when each check last failed. If nobody knows, run it.
  7. Capture the baseline before adoption. It can’t be reconstructed later, and it may tell you something you’d rather not hear.
  8. Write the correction next to the original claim. A log that only records wins isn’t evidence.
  9. Land it on the main line and name an owner from the team.