Let’s talkLet’s talk

July 29th, 2026

Define the failure condition before the AI project starts

Kamil Pyszkowski

Kamil Pyszkowski

10 mins

A success criterion that cannot return a verdict of "no" is not a criterion. Medicine fixed this problem with trial pre-registration. AI pilots need the same thing - a baseline, a target, a decision rule, and a date, all written down before the tool arrives.

Define the failure condition before the AI project starts

Most AI pilots that die don't fail. They go quiet. Six weeks in, the tool stops coming up in standup. A quarter later somebody asks whether we're still paying for it, and nobody can say whether it worked, because nobody wrote down what working would have looked like.

An earlier post here argued that the AI ROI crisis is a discipline problem rather than a capability one: teams adopted the technology before deciding what it should achieve. Which invites the question clients actually ask us in the room. Fine - how? What do you write down, and when?

Here is the whole answer in one line, and the rest of this post is the mechanics. A success criterion that cannot return a verdict of "no" is not a criterion. If there is no way for your pilot to come back a failure, there is also no way for it to come back a success. You have a slogan with a budget attached.

The hostile reviewer test

Look at the goals AI projects actually launch with:

  • Improve developer productivity
  • Explore agentic workflows
  • Reduce support load
  • Modernise our internal tooling with AI

Every one of those is unfalsifiable as written. There is no number, no starting point, no date, and nobody who owns the measurement. Six months later the project can be described as a success by anyone who wants it to be one, and as a waste by anyone who doesn't, and both of them will be arguing from the same evidence.

The test I use is adversarial. Imagine a reviewer who wants this project killed and has read your criterion. Can they use it, plus the data you will actually have on the review date, to kill it? If yes, you have a criterion. If they'd have to argue about interpretation, you have a slogan.

One warning about the number you pick. Optimise for speed alone and speed is what you get, including the parts you didn't want. Google's 2024 DORA report found delivery throughput and stability both moving the wrong way as AI adoption rose, stability by around 7% for a 25% rise in adoption. Teams shipped more and broke more. So a target needs a guardrail beside it: the number that must move, and the number that must not get worse. The guardrail is the one you'd rather not write down.

Medicine already fixed this exact problem

Clinical research had our problem first, and worse. A trial would measure a dozen outcomes, and the paper would report whichever one looked best once the data came in. Not fraud. Ordinary optimism applied to a large set of numbers, and it produced a literature full of drugs that succeeded at something other than what they were tested for.

The fix was pre-registration. Since 2005, the ICMJE has required trials to be registered in a public registry such as ClinicalTrials.gov before the first patient is enrolled, naming the primary outcome up front, as a condition of being published later. Declare the target before you see the results, or don't get to claim it.

Then someone checked whether it worked. The COMPare project at Oxford read 67 trials in five leading medical journals and compared each paper against its own registry entry. Fifty-eight of the 67 had discrepancies. Across them, 354 pre-specified outcomes went unreported and 357 outcomes that had never been pre-specified turned up in the papers as though they had been. Registration didn't eliminate outcome switching. It made outcome switching visible, which is a different and more useful thing.

That's the part worth stealing. Writing the criterion down in advance does not force honesty. What it does is leave a record that makes the swap detectable, including to yourself, eight months later, when you are the one hoping the pilot succeeded.

Your version needs no registry. It needs a document with a date on it, stored where you can't quietly edit it, that someone other than the project's owner has read.

The baseline is most of the work

You cannot claim a delta you never measured. Obvious, and still plenty of teams deploy first and reconstruct the "before" number at review time, from memory, under pressure to justify the spend. Whatever gets reconstructed that way will be the number that makes the result look good. Not deliberately. That's what memory does when it's asked a leading question.

So the first phase of an AI project is measuring the current state with the tool switched off, for long enough to cover normal variation. A single sprint's cycle time tells you about that sprint. This phase is also where projects get skipped, because measuring the status quo feels like not doing the project. It is the project.

Two practical notes from doing this on real teams. Prefer metrics you already collect - cycle time from first commit to merge, review turnaround, tickets closed without escalation, cost per resolved ticket with model spend included. A metric that needs new instrumentation is a metric you will not have when the review date arrives. And measure the work, not the tool. An agent that made 10,000 API calls has proven that it made 10,000 API calls.

If you're tempted to skip the baseline and rely on what the team reports instead, METR's randomized trial is the argument against it. Sixteen experienced open-source developers, 246 real tasks on their own repositories, AI allowed or forbidden per task. They expected a 24% speedup going in and estimated about 20% afterwards. Measured, they were 19% slower. The self-reports were not slightly off. They pointed the wrong way, from experts working in code they knew well.

Before-and-after is not evidence

Comparing this quarter to last absorbs everything else that changed in between: hiring, attrition, seasonality, a reorg, one difficult customer leaving. It also absorbs the fact that a team being measured on a new tool works differently from one that isn't.

What you want is a slice of comparable work that the tool doesn't touch. Split by ticket type, by repository, by team, or stagger the rollout and let the not-yet-migrated group be the control. It doesn't need to be a clean experiment. It needs to be a comparison you didn't construct after the fact.

Here's the honest limit on that advice. When METR re-ran their study in early 2026, the data came back unusable, partly because 30 to 50% of developers refused to do some tasks without AI. Once people depend on a tool, taking it away to preserve a control group stops being a measurement decision and becomes a fight about working conditions. Which means the window for a cheap baseline is early, before the tool is load-bearing, and it closes faster than you'd expect.

Write the decision rule and the date at the same time

A target without a rule attached is a wish. The rule says what happens on a specific date if the number hasn't moved: revert and stop paying, extend once by six weeks under the same rule, or ship it to the rest of the org.

Barry Staw's work on escalation of commitment in the 1970s described the pattern everyone recognises and nobody thinks applies to them: as sunk cost rises, people fund the losing course harder, while believing they are being rational. A date chosen before any money was spent is the cheapest defence anyone has found.

That third branch is the one that makes the other two mean something. Recall from the earlier post that S&P Global measured companies abandoning most of their AI initiatives before production at 42% in 2025, up from 17% the year before. That number gets read as evidence of failure, but abandoning a project on a date you set in advance is the process working. Abandoning it after eighteen months and three re-scopings is the same decision, taken later, for much more money.

The block, in one table

Six rows, one page, dated, circulated to at least one person who is not invested in the answer.

FieldWhat goes in itExample
MetricOne number, already collected, owned by a named personMedian first-response time on tier-1 support tickets
BaselineIts current value, measured, with the window stated4h 10m, median over the 60 days before rollout
TargetWhere it must get toUnder 2h
GuardrailWhat must not get worseReopen rate stays at or below its baseline of 11%
Decision dateA calendar date, not a phase15 October
RuleWhat happens on that date in each caseTarget and guardrail both met, roll out to tier 2. Otherwise revert and cancel the subscription.

Note what it leaves out. No room for a narrative, since the narrative is what gets written to explain away the number, and no list of secondary metrics, since eight things to watch is a licence to report whichever one moved.

What you should not pre-register

Not every use of AI is a pilot with a return to prove, and forcing an ROI target onto real exploration produces theatre. Trying a new model to find out what it can do is research. Give it a time box and a spend cap instead of a target, and let the criterion be whether you can name what you learned in a page and say what it changes about what you build next.

Keep the two budgets apart. Mixed together, exploration gets defended with numbers it was never designed to produce, and deployments get judged on how interesting the demo felt.

What we can and cannot prove about ourselves

Note

By the standard in this post, our own adoption doesn't qualify. We use AI heavily to build our software, we think it made us considerably faster, and we cannot prove it, because we never recorded a baseline or kept a control. There's no version of our last two years without these tools to compare against. The belief is sincere and the evidence is a feeling, which is exactly what this post argues is worthless. GitHub's controlled trial found a real 55% speedup on a scoped task, so the effect exists. That's their measurement, not ours.

The reason to be blunt about it: writing down a criterion is easy when someone else's budget is in question and gets skipped when the project is your own enthusiasm. What this post describes is the default behaviour of people who know better, and the only defence against it is having written the thing down before there was anything to defend.

An afternoon spent writing six rows costs an afternoon. The alternative is a year of work that ends without a verdict - no evidence, no lesson, and a line item nobody can defend or kill.


AKENA is a blockchain engineering studio. We build AI and blockchain infrastructure - agent-facing RPC and MCP endpoints, on-chain data pipelines, and the products that run on top of them. If you are scoping an AI build and want the criterion written down before the code, we should talk.

Let’s talk

Bring us your problem, we’ll help design the system. No hype, just engineering.