> ## Content Index
> Fetch the complete content index at: https://christingeorge.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# What Story Points Were Hired To Do
- URL: https://christingeorge.com/what-story-points-were-hired-to-do/
- Published: 2026-09-27T15:58:34.000Z
- Updated: 2026-09-27T16:08:28.000Z
- Description: In 2022 I coached my team into story points with LEGO. AI has changed what makes software expensive, and I'd now lead teams out of them, carefully.
- Author: Christin Emmanuel George
- Tags: Product, Agile, AI, Platform

In January 2022, while I was leading product at Kaleyra, I wrote a piece on Medium arguing for story points. Our sprints were planned the way most teams planned them then. Team leads estimated work in days and assigned it to developers, and progress was measured by counting stories, so a spelling fix and a new Voice API for Verify each counted as one. I wanted the whole team to size stories by complexity in Fibonacci points, plan from a prioritised list, and let velocity settle into a number stakeholders could rely on.  
I didn't want to hand the team a new process and a slide deck, so I brought in LEGO sets and \[confirm: Farmlands\] and ran coaching sessions where we \[confirm: what the exercise involved\], before any of it touched real work.  
It was the right call for that team at that time. I wouldn't make it today. If you lead product in an organisation that still runs on points, this isn't a piece about getting it wrong. The ground has moved under all of us, and I've changed my mind about what estimation is for.  
What points were hired to do  
When a practice stops working, I go back to the job it was hired to do. Story points had three, and they have aged very differently.  
The first was to start a conversation. The number was never the valuable part. The discussion that produced it did the work, whether that was the engineer who spots a dependency nobody else saw, the question that splits one story into three, or the moment someone admits they don't understand the requirement. Planning poker worked because it made disagreement visible. When one person holds up a 2 and another holds up a 13, the gap is the information. That job hasn't aged at all, and it's where I part ways with the #NoEstimates school. The number is mostly waste, but a team that drops estimation entirely usually loses the conversation too, without noticing.  
The second job was forecasting, and points always depended on two things most organisations can't supply. Velocity only means something if the same people do the work sprint after sprint, and teams re-form around new priorities all the time. It also has to stay inside the team. Once velocity sits on a leadership dashboard next to other teams' numbers, points inflate, a 3 becomes a 5, velocity rises, and nothing extra ships. The engineers notice long before leadership does.  
ZoomInfo showed me an alternative. Estimates there were done in days, with a multiplier on top for everything else developers carry, like support, reviews, production issues and meetings. It kept delivery steady across many teams without ever needing points, because a simple estimate with an honest buffer was legible to everyone from engineers to leadership. At that scale, legibility beat a better abstraction.  
For teams that want more rigour, flow metrics now do this job better. If work is sliced into small, similar pieces, throughput and cycle time forecast about as well as points, and a Monte Carlo simulation turns them into a date range with a probability attached, which is what stakeholders needed all along. Here I have to take back part of my 2022 argument. I dismissed counting stories because ours varied so much in size. The answer was smaller, more consistent stories. Better numbers would never have got us there.  
The third job was prioritisation, deciding whether something was worth building at its cost. That still matters, but it needs far less precision than points pretend to offer. Rough sizes, or a Shape Up-style appetite where the time is fixed and the scope flexes, are enough. The rule I give product managers is to estimate only when the estimate would change a decision.  
What changed my mind  
Everything so far was a reasonable argument in 2022\. AI is what moved me from thinking points had flaws to thinking they measure the wrong thing, and it's the part many engineering leaders are still sceptical about. I understand why. The gains are real but uneven, and the people closest to production systems see the uneven half first.  
Prototypes, new services and discovery work have become dramatically faster, and a working version of an idea often costs less than the spec for it. Then the work meets real data systems and slows right down.  
I ran into this myself on a much smaller scale. I used Claude to build a connector that logs my meals and health data into a self-hosted nutrition tracker. The connector came together quickly. The time went on a bug that only appeared once it was writing into a live database alongside the app's own client. Entries were created without a field the app relied on for syncing, so everything logged through the connector was silently duplicated the moment I saved anything in the browser. Generation speed didn't help. Understanding how the existing system behaved did.  
The same pattern shows up at every scale. In an early-2025 study by METR, experienced developers working in their own repositories were 19% slower with AI tools, yet estimated afterwards that AI had made them 20% faster. METR has since marked those results as out of date, and its newer data on late-2025 tools suggests the tools, and the people using them, have caught up a fair way. The perception gap is the lesson I'd keep. When a team says AI is making them faster, believe that it feels that way, and then look at what's shipping.  
What I most want other product leaders to see is that AI compounds. It multiplies an engineer who understands the system and knows exactly what they want from the tools. It multiplies everything else just as readily. Someone who doesn't understand the system now produces confident, plausible code faster than anyone can review it, and gaps in knowledge, unclear requirements and technical debt scale right alongside the good parts. The bottleneck moves to review, integration and verification, which haven't sped up.  
Google's DORA team reached the same conclusion in its 2025 report. AI amplifies existing strengths and weaknesses, and a high-quality internal platform is one of the key enablers for magnifying its effects across an organisation. That didn't surprise me. At Quintype, long before any of this, I was building entities and structured data into a CMS for newsrooms. We built it open-ended, so anything with inherent properties could be an entity, and entities could be linked into a graph. Rafael Nadal is a person, an athlete and a tennis player. Tennis is a sport. Roland Garros is the French Open, the championship he won fourteen times. All of it was connected.  
That is the kind of structure an AI can use without guessing. Clean, accessible data has never been useless, with or without AI. Every team building on a platform depends on it whether they think about it or not, and AI makes that dependency faster in both directions. Work that touches clean systems gets cheaper. Work that touches a mess gets messier, sooner.  
Put together, this breaks the unit points were built on. A point expressed how hard something was to build, and the build is now often the cheap part. The cost lives in how much existing system the work touches and how well the person steering understands it. The same story can take an afternoon or a month depending on those two things, which is a far wider spread than the gap between a 3 and a 5\. Agents stretch it further. Velocity assumed a team's capacity was bounded by headcount. On a lot of teams, agents now write much of the code, and the real limit is how much the team can review and verify.  
What I'd tell a product leader now  
When I brought points into Kaleyra, I taught them with LEGO before anyone estimated a real story. Taking them out deserves the same care.  
Don't rip them out on principle. They still teach new teams how to break work down, they still help in fixed-scope contract work where both sides need a shared unit, and large organisations running SAFe lean on them to coordinate. If they're working for you there, leave them alone.  
Where they aren't working, the order of the change matters more than where it ends up. Take velocity off any dashboard that compares teams, and do it first, because it costs nothing and removes the pressure that inflates the number. Change what stakeholders receive before you change what teams do. A forecast range lands better after a quarter beside the old number than as an overnight replacement. Keep the estimation conversation after the number goes, because teams often hear "no more points" as "no more planning". And invest as much in the platform and in the people directing the AI as you do in the tools, because their understanding and the quality of what they build on are what get multiplied.  
What I haven't worked out  
I don't have this solved. The theme this site runs on took more than thirty versions in about a week, most of them fixing things I'd broken myself. I couldn't have estimated that in points, and I'm not sure a range would have helped either. \[confirm: the open question you're still working through\]  
The case for predictability I made in 2022 still stands. I've stopped believing story points are how to get there