Data · · 8 min read
The AI budget is due this month. Price the data under it.
Most AI lines in next year's budget pay for the model, the licence and the integration, and nothing for the data the feature reads. How to size that missing line, one use case at a time.
By Precision Code Studios, Engineering team
Somewhere in your company this month, a spreadsheet called something like FY27 draft v4 is doing the rounds, and it has a new section near the top. An AI assistant for customer support. A renewal-risk model for the account team. A finance assistant that answers revenue questions in plain English. Each line has a model cost, a licence and a few engineer-months of integration. Almost none of them has a line for the data the feature will read.
That gap is where next year's AI projects stall. In February 2025, Gartner predicted that through 2026 organisations would abandon 60 percent of AI projects not supported by AI-ready data, and reported that 63 percent of the data management leaders it surveyed either lacked the right data practices for AI or were unsure whether they had them. The window that prediction covers closes in under twelve weeks. The budgets being signed now decide which side of that line your 2027 projects land on.
If you have watched a pilot stall, you know the shape of it. The demo ran on an export that someone cleaned by hand over two afternoons, and it was impressive. Production means the same data arriving fresh every day, with the right people seeing the right rows, long after the person who cleaned it has moved on. That is a different project, and nobody priced it.
Why the data line goes missing
Negligence rarely has much to do with it. The AI vendor's quote covers what the vendor sells: the model, the platform, the integration work. The data team, if there is one, is busy keeping last year's reports alive. And data work is invisible in a demo, so on a budget sheet it reads as housekeeping, which is the first thing cut when the total looks too big.
The usual correction is just as expensive. 'Fix the data first' becomes a two-year platform programme with no use case attached: hard to demonstrate, hard to measure and easy to cancel in the next budget round. Data platforms tend to die of abstraction rather than of bad engineering.
Forrester's 2027 budget planning guides, published in July, arrive at a similar place from the finance side. Their advice is to stop treating data cleanup as one big pot to fund or cut, and to aim the money at targeted fixes that make data usable for the AI work actually planned. Our position is the practical version of that: every AI line in the budget gets a data line beside it, sized for that use case and nothing else.
Five checks for every AI line
Sizing the data under an AI feature does not need a maturity model. It needs an hour with the person who will own the feature and someone who knows where the data actually lives. Ask five things.
- Sources. Which systems does the feature read, and how many are there? One system is an integration. Four is a data project, because each extra source brings a connector to build, a format to reconcile and another team to negotiate with.
- Definitions. Do those systems agree on what the important words mean? The CRM's customer is an account, billing's is a subscription and the product's is a workspace. Every disagreement is a decision a senior person has to make, and that takes meetings rather than code. This is the check that most often turns a small line into a large one.
- Freshness. How old can the data be when the AI uses it? Yesterday's data is a nightly batch job. Data that is minutes old is a different architecture with a different running cost. Teams routinely ask for real time when daily would do.
- Permissions. Who may see which rows, and will the AI respect that? An assistant that can read every contract will cheerfully summarise one for someone who should never have seen it. If access rules live inside the source application, they have to travel with the data, and carrying them across is work.
- Owner. Who gets the call when the feed breaks on a Sunday, and who decides when a definition changes? If the honest answer is nobody, the line is not finished, however small the rest of it looks.
Mark each check as ready, work or decision. Ready costs almost nothing. Work is engineering time you can estimate in weeks. Decision means a business argument has to be settled before engineering can start, and those take longer than anyone budgets for, because the people who must settle them have day jobs.
A worked example
Take a typical B2B software company of a couple of hundred people, with the three AI lines from the opening in its draft budget. Run the checks and the picture changes.
The support assistant is the easy one. It reads the help centre and the ticket history, both held in the help desk tool. Nobody disputes the definitions and daily freshness is fine. The one real piece of work is permissions: a customer must only ever get answers drawn from their own tickets. Its data line is small, and it does not need a new platform. A well-built pipeline into a search index will do.
The renewal-risk model is another matter. It needs the CRM, billing, product usage events and support tickets, and those four systems hold three different ideas of what a customer is. It needs a couple of years of history to train and test against, and nobody currently owns the usage data. This is where most of the money and most of the calendar will go, and the bulk of it is people agreeing on meaning before an engineer writes anything.
The finance assistant reads billing, the ERP and the CRM. It brings its own argument, about whether revenue means booked or recognised, and stricter access rules than anything else on the list.
Now read the second and third lines together. They share two source systems and the same unresolved question about customers. Funded as separate projects, perhaps with separate vendors, the company pays to settle that question twice and ends up with two answers to 'how many customers do we have?', which is precisely what a board member will ask. Funded together, the shared part becomes a foundation, and whatever comes third starts with much of its data already done.
When a lakehouse earns its line, and when it does not
A lakehouse is a single home for an organisation's data: files kept in open formats on inexpensive cloud storage, with warehouse-style tables, access control and SQL on top, so that dashboards, analysts and AI models all read the same tables. Open table formats such as Apache Iceberg and Delta Lake keep that data independent of any one query engine, which matters when the tools reading it change every year.
It is not the answer to every line in the budget. If you have one AI use case reading one or two systems, or an existing warehouse that already answers your reporting questions and holds what the new feature needs, you do not need one yet. Build the pipeline, write down the definitions and spend the rest on the feature.
The case for it appears when the work starts to overlap:
- Two use cases share sources. The second project that touches the same CRM and billing data is the point where separate pipelines start costing double.
- Dashboards and AI disagree. If the revenue figure in the board pack differs from the one the assistant quotes, you have two sources of truth, which in practice means none.
- Documents meet records. Contracts, call transcripts and tickets need to sit beside billing and usage data, with the same access rules, before an assistant can reason across them.
- The same data is copied everywhere. Each tool with its own copy brings its own licence, its own sync job and its own slightly stale version of the truth.
What to change in the draft this month
If the budget goes back to finance in the next few weeks, four changes make the AI section honest.
- A data line beside every AI line. Sized with the five checks, in the same quarter, so nobody can approve the feature without seeing what it stands on.
- A named owner for each definition. Someone in the business who can settle what customer or revenue means and make it stick. Put their hours in the plan; they are the scarcest input on the sheet.
- Running cost as well as build cost. Pipelines break when source systems change. Budget the people who keep them running in the second year, or the feature quietly degrades on stale data and nobody notices until a customer does.
- A gate in front of the AI spend. Release the model and integration money once the first use case's data passes its checks. This reorders the spending rather than adding to it.
When vendors pitch, whether they sell AI features or data platforms, put the same four questions to all of them. Which of our systems will this read? Who has confirmed that data is usable? What happens when a source system changes its schema? Can we query our own data without your product? Good vendors answer quickly. The ones who change the subject have shown you where the cost is hiding.
Of everything in the AI section, the model is the line most likely to cost less next October. The data under it only gets cheaper if you build it once, and whether you do is being decided in a spreadsheet this month.
The short film: transcript
Your AI budget needs a data line.: the short version (1:20)
It's budget season, and the AI section of most plans pays for the model, the licence and the integration. What the model actually reads is usually left out.
Once the person who tidied it moves on, nobody owns the feed. A month after launch the answers have quietly drifted out of date, and the first person to notice is usually a customer.
An hour with whoever will own the feature is usually enough to size the data under it. The one that most often turns a small job into a big one is about meaning, because three systems rarely agree on what a customer is.
The reason is rarely technical. Each vendor quietly settles what counts as a customer to suit its own feature, neither is wrong, and the argument lands back on your desk.
Every new feed built off the same C R M and billing records repeats work you've already paid for. Worse, when the board pack's revenue figure differs from the one the assistant quotes, people stop trusting both.
Of everything in the AI section, the model is the line most likely to cost less next year. The data under it only gets cheaper when the work is shared.
Read the full article at Precision Code Studios.