“How long will this take?”
Estimating under uncertainty: the pitfalls, a five-step method, and what changes with AI
Every software project starts with this question, and almost every answer to it turns out wrong. Last week I gave a talk at work on estimation. It was mostly for our sales and tech folks, who answer this question together before anyone signs a contract. I called it An estimate they can’t refuse, a nod to The Godfather’s “I’m gonna make him an offer he can’t refuse.”
When estimating, the goal isn’t the lowest number, or the number the client wants to hear. It’s a number with enough reasoning behind it that neither side feels the need to argue with it.
In this post
Why estimates miss, even good ones
This is the cone of uncertainty. It shows how far off even a good estimate can be at each stage of a project. At the start, very little is decided, so the range is huge. As you write requirements, agree designs and build code, each decision removes some uncertainty. The range narrows, and it only closes when the software is done.
On day zero, when all you have is an idea, the actual cost can land anywhere from a quarter to four times your estimate.
Steve McConnell, The Cone of Uncertainty
That is a 16x range, and it is the best case, for skilled estimators.
We write most proposals and SoWs in that grey band on the left. Sometimes a little further right, if the client has done a lot of homework. But mostly there.
Don’t be scared of the range, but know it’s there. A single number on day zero is a guess with false precision. Estimating well means narrowing that range as you learn things, not squeezing it to look sure.
Next, look at how overruns spread across projects. A study of 1,471 IT projects found:
The average cost overrun was 27%.
Flyvbjerg and Budzier, Why Your IT Project May Be Riskier Than You Think, HBR (2011)
Sounds manageable, add a quarter and move on. Except the average hides the tail.
One project in six went about 200% over budget and 70% over time.
Same study
Projects can’t finish much earlier than planned, but they can finish far later. You can’t pad your way out of that. You have to spot the work likely to end up in the tail and treat it differently.
Three words we use interchangeably
We use estimate, target and commitment as if they mean the same thing. They don’t, and mixing them up is where many estimation arguments start:
What the work will probably take
A range, worked out from the work itself, by the people who'll do it. Say, 11 to 21 weeks. It changes only when you learn something new about the work.
What the business wants
A date or budget the client needs for their own reasons, like "we go live in 10 weeks because there's a launch event." It's a fair goal, but it says nothing about how long the work takes.
What you promise
A specific date and scope you agree to deliver, decided after looking at the estimate and the target together. You make this call after weighing both.
If the target sits outside your estimate, talk about scope or capacity. You can fix time, scope and team size, but only two of the three (the old project management triangle). Don’t change the estimate to match the target. That pushes the problem later in the project, where it’s harder and costlier to fix.
Four ways we make it worse
Anchoring
The first number you hear pulls yours towards it.
Fix: estimate before you hear the budget.
Narrowing the range to look confident
Nobody asks for a tight range, but we give one anyway.
Fix: widen it until you'd bet on it.
Estimating the code, not the project
Everyone forgets the work around the code.
Fix: a checklist, with a line for every item.
Cutting the number instead of the scope
Someone pushes, and the number shrinks. The work doesn't.
Fix: negotiate the scope instead.
1. Anchoring
This is my favourite study. Magne Jørgensen and Dag Sjøberg gave two groups of professionals the same spec. One group was told, in passing, that the client thought it would take 50 hours. The other was told 1,000 hours. Both were told the client knew nothing about software and to ignore the number.
Both groups estimated the same work. The second group’s answer came out eight times bigger. And when asked, the estimators said the client’s number hadn’t influenced them. Jørgensen has kept finding the same effect since. It even has a name: the anchoring effect.
Any number you hear first becomes the anchor: the client’s budget, a competitor’s quote, what sales is hoping for. So estimate before you hear the budget. Write it down. Then compare. Asking for the budget is still essential. It’s how you shape scope. Just don’t let it into the room before you have an estimate done.
2. Narrowing the range to look confident
McConnell’s book, Software Estimation, opens with a quiz: ten general-knowledge questions. For each, you give a range you’re 90% sure contains the answer. You should get nine right. The average is 2.8.
Notice what the quiz never asks for: a narrow range. People are free to make their ranges as wide as they like, and they still squeeze them. A tight range feels like expertise, and a wide one feels like admitting you don’t know. So we narrow ranges on our own, as a stand-in for confidence, when nobody asked us to.
The same thing happens in a proposal. “11 to 21 weeks” feels too vague to put in front of a client, so it gets trimmed to “14 to 16.” But at proposal stage, a wide range is the accurate answer, and a wide range is itself information. It tells you how many unknowns you still have.
3. Estimating the code, not the project
Engineers estimate what they picture: building features. They forget everything around it:
- Environments and access
- Integrating with the client’s systems
- Testing
- Security review
- Data migration
- Deployment
- Demos
- UAT support
- Waiting on client inputs
- Handover
Forgotten work like this is one of the most common sources of estimation error. On many of the projects I’ve seen, that “everything else” is close to half the effort. Keep a checklist like this one and put a line against every item, even if the line says zero. A zero you chose is fine. A zero you forgot is an overrun.
4. Cutting the number instead of the scope
“It’ll take 20 weeks.” “Can we do it faster?” There are only two real ways to go faster: less scope, or fewer unknowns. Shaving the number because someone pushed leaves the work the same size. It only feeds the planning fallacy. People haggle over a single number. A range with reasons behind it is much harder to haggle over.
How to estimate
Break it down and sort it
“Build a web service” is hard to estimate. “Build an endpoint that does these four things” is not. Then sort each piece into one of three kinds:
- Known: you’ve done this before, so the range is tight.
- Assumed: you believe it, but nobody has checked. The range widens, and the assumption goes in writing.
- Unknown: no estimate at all. Run a timeboxed spike to find out first.
Three numbers per piece
Best case, most likely, worst case. Then weight them, using the three-point (PERT) estimate:
expected = (best + 4 × likely + worst) / 6
The formula is a weighted average. The likely case gets four times the weight, because it’s what usually happens. The best and worst cases get one share each, so they still pull on the result.
Take a task where the best case is 10 days, the likely case is 15 and the worst case is 35. The best case is only 5 days better than likely, but the worst case is 20 days worse. Things can go a lot more wrong than they can go right. So the weighted answer comes out at 17.5 days, not 15.
That 2.5-day gap looks small on one task. Across fifty tasks it adds up to weeks. If you add up the “likely” numbers, your plan assumes nothing ever goes badly, and you’re late by design. Add up the expected numbers instead.
Alone first, then together
Do the first pass alone, in writing, with the codebase open. That way you build your own picture of the work before anyone else’s number can sway it. Then do a round with others, ideally with someone who’ll build it. This is Wideband Delphi; planning poker is a lighter version. When estimates spread from 6 to 28 weeks, don’t average them. Ask the highest and lowest to explain. They’re almost always picturing different work, and that conversation usually tells you more than the numbers do. Two or three rounds and you get to something like 12 to 17.
Check it against history
Keep a reference sheet of how long each kind of work has taken you before. Use the same pieces you broke the project into in step one: an API endpoint, an integration, a data migration, an agent. If a new estimate looks optimistic next to what that kind of work took before, trust the history. Flyvbjerg calls this reference class forecasting.
The sheet matters most for the kinds of work where gut feel is worst. Rescues and rewrites of existing systems are the usual example. They almost always take far longer than estimated, because the old code hides behaviour nobody thought to list. (More on that in From estimate to contract.) Your history shows that pattern long before your instinct accepts it. With agents in the loop, this history needs adjusting too, which the next section covers.
Quote a range and a confidence
Never a range without a confidence. “10 to 20 weeks” on its own doesn’t say how likely you are to land inside it. Said with 50% confidence, it’s as likely to miss as to hit. Said with 90%, it’s something the client can plan a launch around. The same range can be a safe promise or a reckless one, and only the confidence tells everyone which. It also tells you how much risk you’re taking on.
P50 is a coin toss. P85 is what you commit to. If you want to get rigorous, a Monte Carlo simulation over your three-point numbers gives you the whole curve. Troy Magennis has free spreadsheets for this. “14 to 18 weeks, and we commit to 18 for this scope.” My test for P85: would I bet a month’s salary on landing inside the range? If not, widen it.
Estimating for an AI-native SDLC
Most of my history, and most of the research above, comes from a world where humans typed every line. That’s not how we work anymore. An API that took me a day by hand now often takes a few hours with an agent doing the typing.
After adding everything up, I divide the build work by an AI multiplier. These days I usually use 1.5: if I'd have said 10 days by hand, it'll probably be 6 or 7.
It's a gut call, but it's a gut call I write down, so I can check it later.
How big the multiplier should be depends on a few things:
How well you know the stack
Agents amplify what you already know. In a stack you know well, you can review their output quickly and steer them away from bad ideas. In an unfamiliar one, you can't tell good output from output that only looks right, so the speed-up shrinks.
How new the problem is
CRUD endpoints, integrations with well-documented APIs and test scaffolding speed up a lot. New logic, tricky concurrency and anything nobody has written about before speed up much less.
How good your tooling and harness are
An agent with a fast test suite, clear conventions and context about the codebase is a different tool from one working blind. The better the harness, the bigger the multiplier you can justify.
How easy the output is to check
If a test or a quick demo proves it works, the agent's speed shows up as faster delivery. If a person has to read every line carefully, the review becomes the bottleneck.
Two cautions.
Only divide the build work. Agents speed up writing code. They don’t speed up the rest: waiting for client inputs, getting access to environments, a security review, UAT. Or the meeting where three stakeholders disagree on what “done” means. Remember the forgotten-work checklist from earlier. Apply the multiplier only to the pieces an agent will build. Apply it to the whole project and you’ll shrink the parts that were never about typing speed.
Track it separately. Your estimated-vs-actuals sheet from before AI is now skewed. Add a column for whether the work was agent-assisted, and what multiplier you assumed. In a few months you’ll know whether 1.5 was about right, too optimistic or too cautious, and you can stop guessing.
From estimate to contract
Everything so far applies to any team estimating its own work. If you build software for clients, there’s one more step. The estimate turns into a price, and the price turns into a contract. That’s where a wide range stops being an inconvenience. It becomes risk someone has to carry: you, the client, or both. So the rest of this post is about consulting. What do you do when the range is too wide to price? And how do you match the contract to the range you have?
If the range is wide, it’s because there are unknowns, and no technique will narrow it. The information doesn’t exist yet. So buy it. Sell a short, fixed-fee discovery and credit it against the build. Blair Enns calls this diagnosing before you prescribe.
A good discovery gives you five things:
- Exactly what you’ll hand over at the end of the build.
- A definition of done that both sides sign.
- A small spike on each unknown. Never connected Salesforce to Snowflake? Connect a toy instance. Never built an agent? Build a hello-world one.
- A starter eval set, if it’s AI work.
- A P85 estimate for the build.
Even if the client takes the plan elsewhere, they got their money’s worth.
Then let the width of the range pick the contract.
Narrow range, fixed price is fine. Wider, fixed price with checkpoints where both sides re-look at scope (the agile fixed price model). Wider still, target cost with overruns and savings split. Too vague, time and materials with a flexible scope.
A fixed price on a wide range isn’t a price. It’s a bet.
And watch out for scope that points at something instead of defining done. “Finish the build.” “Rewrite it in a modern stack.” “Just make it do what the current system does.” Nobody knows what the current system does, not you and not the client. The old system is the only complete spec of itself, and thanks to Hyrum’s Law, someone depends on every one of its quirks. Audit first, write down what it does, agree that list as the definition of done, then estimate.
Questions you’re probably asking
“Doesn’t estimating before the budget waste time on deals that won’t fit?”
It can. You don’t want to spend days estimating a deal that was never going to fit. Start with a SWAG: a quick range from someone experienced. Anywhere from 50k to 150k. Wide, but it tells you it isn’t 10k and it isn’t a million. That’s enough to know if the client’s appetite is in the neighbourhood. Then do the full estimate.
“How do I account for different speeds and skill levels on the team?”
Normalise to the middle of the team. Any team has a spread: some people finish a task in half a day, others take two or three. Estimate the work as a mid-level engineer on that team would do it, not your fastest person and not your newest.
Estimating to your best people is the common trap. They’re the ones most likely to get pulled onto something else, and then the estimate no longer matches who’s doing the work. A normalised estimate holds up when the staffing changes. If the team really is skewed, say mostly senior or mostly new, adjust for it openly and write down the staffing assumption next to the number.
Clients vary too: consider a 1.25x buffer for difficult clients. You can usually tell in the sales cycle.
“What if their budget is just too low?”
Every client has some constraint. Nobody has infinite time or money. If they’re at 50k and you’re at 100k, price the intangibles: a case study, a logo, a quote. Treat it as acquisition cost if the account could be worth millions later. But ask one more question: can we deliver this properly? If you’re taking the deal to get into a deep account, that first project has to blow them away. A tight budget that makes that impossible is worse than no deal.
“What if the client thinks we’re padding the estimate?”
An estimate higher than the client hoped for can look like sandbagging: a number inflated so the team looks good later. Arguing doesn’t fix that. Showing your work does.
Take a client who wants a full rewrite in two and a half months on a small budget, where the real estimate needs twice the team or twice the time. Don’t send a number. Send the breakdown behind it: every API and flow you’ve understood, the assumptions, and what each piece will take. Three pages is enough. Now the client can check the reasoning line by line, and a big number reads as thorough rather than padded.
The same applies to buffers. If you add contingency, label it and say why. Hidden padding is what makes estimates look dishonest, and once a client finds one, they stop trusting the rest.
“Should I leave any buffer?”
Leave room to be excellent. If the estimate only covers what you promised, you’ll deliver that and nothing more. Next time they’ll pick someone cheaper. Leave a little room to do something really well.
If you remember three words from this post: range, confidence, reasons.
The slides
Here’s the deck from the talk. Use the arrow keys or click to move through it, and press N for the speaker notes. Or open it full screen.