Design for the 850th time, not the first
Field notes from an engineering manager on the gap between a feature that ships and a feature that’s operable.

Halfway through the demo, I stopped taking notes and started counting.
I had come into that room to say thank you.
The cleanest delivery we’d ever run
We had just come off the biggest event of our year. Eight weeks of continuous peak, a hundred-odd individual events running back to back, several brand-new products — and not one of them late.
By every internal measure it was the best delivery my team has ever put together. So the debrief with the operations team was, in my head, a formality. We’d show what we built. They’d show what it did. Everybody would feel good.
Then somebody suggested we skip the slides and just hand over the laptop.
So we did. We asked the person who had actually spent eight weeks inside the thing to build one in front of us, narrating as they went.
Ninety minutes later I had a page of notes and a slightly different understanding of my job.
Here’s the part that took me longest to accept: nothing they showed us was a bug. Nothing had failed a test. Every item on my page was something my team had designed correctly, shipped on schedule, and would have defended in review — and I’d have backed them.
The gap was somewhere else entirely, and I couldn’t name it for about a week.
A note on specifics: I’m keeping our domain out of this. The details are ours; the pattern isn’t — I’ve described it since to people in logistics, clinical scheduling and internal platform work, and watched them nod before I finished the sentence. So: a consumer product, a live operations team running it in real time, a peak season that compresses a year into two months. The mechanics below are exact. Only the nouns are generic.
Watch me build one
The feature was a targeted promotional offer. Pick an item, improve its terms, aim it at a customer segment, schedule it to go live at a chosen moment.
We’d built it configurable per item and per event, with tier grouping so higher-value customers got a better version. Good build. Did exactly what the spec described.
The spec described it like this:
“An operator can configure a targeted offer on any item.”
One sentence. One action. One user story.
Now here’s the demo.
- One. They build the offer in the back office. Fine — that’s the feature. That’s what we shipped.
- Two. They manually re-rank the item, because publishing an offer doesn’t surface it. I hadn’t known that. Skip this step and the offer exists, is live, is correctly configured, and nobody ever sees it.
- Three. They set the tier rules.
- Four. They go hunting for the right threshold value, which requires an internal identifier that appears nowhere in the tool. In practice you guess, then check it against a different screen, then adjust.
- Five. If the offer is for one of our regulated jurisdictions, they calculate the payout ceiling by hand and convert it into dollars — on a configuration screen that’s in the local language. Which they may not read.
- Six. They watch it live, because if conditions change mid-event they’ll want to move fast.
Somewhere around step four I started doing arithmetic in the margin of my notebook instead of writing sentences.
Because I knew roughly how many of these we’d run. And it was around then that the operator said the thing that reorganised the meeting for me. Not as a complaint — just as narration, the way you’d describe your commute:
“And then you do that again for the next one.”
Then do it 850 more times
Over those eight weeks the team configured more than 850 of these.
One on every single event. Then a second one partway through every event — which was the whole competitive point of the feature, and also exactly why the number got so large. The thing we were proudest of was the thing generating the volume.
Now go back and read those six steps as a unit of work that repeats 850 times.
I want to be careful here, because the instinct is to read that list as a list of defects. It isn’t. Every single item on it is a reasonable thing to leave out of a first version.
Re-ranking manually after publish? Fine. Looking up an identifier yourself? Fine. Doing a currency conversion in your head? Fine.
They’re all fine at N=1.
They’re only expensive at N=850. And N=850 appears in none of the documents I wrote.
The one product nobody had ever researched
Here’s the part that bothered me most, and it took me the longest to see.
We research everything customer-facing. Discovery interviews, prototype tests, session recordings, funnel analysis, time-on-task, tests on the placement of a single button. If a customer hesitates for four seconds on a screen, we know, and somebody has a plan by Thursday.
The back office had never been researched. Not once. No usability test, no walkthrough, no stopwatch. In eight weeks it absorbed 850 configurations, and we had less evidence about how those went than we had about a signup form nobody has ever complained about.
And it isn’t as though the back office isn’t a product. It has users — a small, expert, captive population who sit in it eight hours a day, which is a research population most consumer teams would envy. It has flows. It has failure modes.
One of those flows ends two steps before the user’s job does, and the product gives no signal that anything is unfinished.
That isn’t an engineering oversight. Publishing works exactly as specified. It’s a design gap with an ordinary name: there’s no feedback loop. The user completes a task and has no way — from inside the tool or from the customer-facing side — to confirm that what they configured is actually real.
We would never ship that customer-side. We’d catch it in the first test we ran.
We just never ran one, because internal tools don’t get tested. They get tickets.
And here’s the uncomfortable version. In the absence of a feedback loop, the operator became the feedback loop. They built a habit — check it manually, every time, all 850 times — and their diligence is exactly what kept the gap invisible to us. Good users hide bad design. Expert users hide it completely.
That’s the lesson, and I could stop there. But those same eight weeks produced four more instances of the identical problem wearing four different costumes — and the costumes turn out to be the useful part, because you won’t recognise the second one from having seen the first
The same gap, four more costumes
A vendor said no, and it quietly became a roster
We moved one long-running product surface onto the homepage. Small piece of work. Enormous result — placement alone turned a sleepy corner of the product into one of our highest-volume lines, with well over a hundred thousand people using it.
One problem. The upstream vendor who supplies that service wouldn’t keep it available during events. And availability during events is the entire reason it’s valuable, because that’s the window when customers want to change their minds.
So the operations team ran it themselves. By hand. Twenty-four hours a day, across two offices in different time zones, for eight weeks. Somewhere around thirteen hundred hours of continuous coverage — because a third party had declined to do something.
Here’s what gets me. Everyone knew the vendor had declined. It wasn’t a surprise, a discovery or an incident. It just never made the trip from known limitation to product constraint requiring either a build or a costed staffing decision.
It travelled instead to the one place in the chain that can’t say no, and arrived there as work.
A constraint that was fine on paper and awful on a Tuesday
We built a curated shelf: a handful of named, hand-assembled bundles — human-picked combinations with a bit of personality — that a customer can take in one tap. Three live slots at a time. Available only before an event starts, so once any component’s event begins, that bundle drops off the shelf.
Both of those decisions are defensible. Three slots keeps it curated instead of a catalogue. Restricting it to pre-event avoids a whole category of edge cases. I’d make both calls again.
Together, though, they mean this: on a quiet day, one start time can empty most of the shelf. At a moment set by an external calendar rather than by anybody’s roster. And the only remedy is a human hand-assembling replacements, component by component, right now.
There was no way to stage anything ahead. The request I got was almost apologetic. Could we maybe queue up seven days in advance?
The abuse surface that lived in someone’s head
A few weeks before the peak we delivered a second promotional mechanic: the more components a customer added to a bundle, the better their reward. It performed extremely well, and covered three times as many categories as our nearest competitor.
What I learned in the debrief was that operations had spent, in their words, “many hours” testing and configuring it to be as abuse-resistant as possible.
After we shipped it. Under time pressure.
We shipped a mechanic; they closed the hole. And the reasoning that closed it now lives in two or three people’s heads rather than in limit templates or validation inside the tool — which matters, because we’re extending that mechanic to three more categories. Same reasoning. From scratch. Three more times.
The team that wasn’t in the room
Late in the peak we ran an aggressive promotion that resolved a customer’s transaction in their favour early, before the event had finished. We built the instant-resolution tooling fast, it worked, and a great deal of money went out the door correctly and immediately.
The support team got, in their phrase, an avalanche of tickets.
Of course they did. If you settle someone’s transaction twenty minutes into an event, under a promotion they’ve never encountered before, they will contact you to ask why.
That wasn’t an unknown risk. It was a completely predictable second-order consequence of the design, and it was budgeted precisely nowhere — because the team that absorbs it wasn’t in the conversation where the design happened.
Five seams
Once you have five examples the taxonomy writes itself. Each one comes with the question that would have caught it, and none of those questions takes more than ten minutes.
- Singular to plural. The spec describes the action once; the operator performs it N times, so small per-instance friction multiplies invisibly.
Ask: how many times will this be performed in its first season? Write the number down.
2. Vendor boundary unpriced. A third-party limitation is known to engineering but never converted into either a build or a staffing decision, so it lands as labour.
Ask: which part of this depends on someone outside the company, and who is staffed if they withdraw it?
3. Design constraint versus operating rhythm. A constraint that’s entirely sensible in isolation produces an unmanageable cadence once it meets a real calendar.
Ask: what does the operator’s day look like when this is live and the week is quiet?
4. Orphan work absorbed by gravity. Translation, abuse testing, QA, data cleanup — work with no named owner lands on whoever sits closest to the tool.
Ask: what non-engineering work does this create, and who is named for it? “Ops will pick it up” is not an owner.
5. Downstream cost externalised. The team that receives the consequence isn’t in the planning conversation, so the consequence isn’t in the plan.
Ask: whose queue grows when this launches, and are they reviewing this document?
Every one of those five gaps was knowable in advance by anyone who thought to ask one of those five questions.
Nobody asked. Including me.
The place neither of us was looking
Here’s the part I find genuinely interesting.
I couldn’t see any of this from inside engineering. My view stops at the edge of the artefact — it works, it’s tested, it matches the spec. From where I sit, “configure an offer” is one action, because in the code it is one action.
But operations couldn’t see it either. Not in a form anyone could act on.
They hit every one of those frictions hundreds of times and absorbed them, competently, without complaint — because absorbing friction is what a good operations team does. What they didn’t have was any forum in which “this costs me an extra sixty seconds” became “this costs the company two working weeks.”
Nobody ever performed the multiplication out loud.
So the gap wasn’t in engineering, and it wasn’t in operations. It sat at the interface between them, which is the one place neither team is looking, because it isn’t anybody’s surface.
That’s true of most of the expensive problems I’ve dealt with, now that I have language for it. A assumes B knows. Product assumes engineering understood. The customer asks for X and needs Y.
The technically difficult problems are the ones I’m actually equipped for. The costly ones live in the seams — and they’re costly precisely because no dashboard covers a seam and no team owns one.
What to change
I’m not going to pretend we’ve fixed this yet. What follows is what I’m taking into our next planning cycle, and I’d rather publish it as a proposal than as a victory lap.
Three habits. Deliberately no new process — new process would be the wrong response to a problem this cheap to fix.
Add an Operability section to the spec template. Short. And — this is the load-bearing part — filled in by whoever will operate the feature, not by whoever writes the document. Recurrence count. Steps per operation, multiplied by that recurrence. Anything required after the primary action, and what breaks silently if it’s skipped. Any value the operator needs that lives outside the tool. Which parts depend on a third party. Named owners for translation, abuse testing and QA. Whose queue grows at launch.
Most of those are one-line answers. The point isn’t rigour, it’s that each line maps to one of the five seams above.
Extend the definition of done. To something like: one operator, working unaided, completes the task in the target time on the first attempt — and can confirm from the customer-facing side that what they configured is actually live.
That last clause comes straight out of the demo. The operator had no way to check their own work from the outside. Which is a strange thing to learn about software you built.
Rank small tooling fixes by recurrence rather than by build size. I’ve since put our candidate list in that order, and the exercise was uncomfortable in a useful way: the two fixes that come out on top are the two smallest builds anyone has proposed all year. They’re on top because each of them sits behind 850 repetitions — and both had been sitting quietly at the bottom of the backlog, correctly ranked by effort and completely mis-ranked by value.
Whether any of this survives contact with a roadmap conversation, I don’t know yet. I’ll find out shortly.
Keep the demo — it was a usability test
If you take one thing from this, don’t take the checklist. Take the format.
We could have collected all of this in a survey. It would have come back as a bulleted list of five minor irritations, which I would have read, believed, and correctly deprioritised. A list understates friction. It has no way to convey volume.
Watching someone do it does.
There’s a specific moment — somewhere around the third repetition — when an extra click stops looking like an inconvenience and starts looking like an arithmetic problem. And you do the multiplication yourself, unprompted, while you’re sitting there.
Ninety minutes. One laptop. One user. The people who built the thing, watching quietly and not helping.
We already had a name for that. We’d just never once applied it to a tool of our own.
The features all worked. That was never the question.
The question was whether anyone had counted.
Design for the 850th time, not the first was originally published in Bootcamp on Medium, where people are continuing the conversation by highlighting and responding to this story.