Deep|LLM: Enterprise AI Application Research (Vol.2) - Growth Keeps Flowing into Production Workflows; ROI Realization Is Gated by the Hours-Saved Threshold

Comparing Conclusions Across the Two Rounds, and What Is New This Round

This round covers 7 enterprise samples plus 1 cross-client enterprise AI transformation consultant, taking the two rounds together to 19 enterprise samples and 2 consultants. Both rounds asked the same set of questions, but of entirely different companies, roughly a week apart, so differences between the rounds reflect sample composition and do not constitute a time series. We therefore group the read-across into three parts: which conclusions the second set of samples independently corroborates, which mechanisms and constraints did not appear in the first round, and which differences should be attributed to sample type.

A. Conclusions independently corroborated by the second set of samples

1. AI spend is still growing, and the samples have already diverged internally. Both rounds contain expanding and tightening companies at the same time. The second set pushes the two ends further apart: large European insurance company A is guiding to 4–5x current consumption by March 2027, while very large Southeast Asian manufacturing company A may cut its allowance budget in half over the next six months. The dividing line is the same in both rounds, namely whether AI spend is tied to a specific production Use Case: those that keep adding, and those still sitting on per-head allowances are starting to be cut.

2. Growth is shifting from seats and subscriptions to APIs, tokens, and production workflows. This is the most robust conclusion across the two rounds, evidenced in the same direction by two non-overlapping sets of companies: in the first round, the subscription / API split at large US telecom operator A moved from 50 / 50 to 40 / 60; this round, the cross-client consultant sees the split move from roughly 90 / 10 a year ago to roughly 50 / 50 today, and business-application spend at large European insurance company A is roughly 70% Consumption.

3. The budget pool is migrating from central IT into the business units, which are starting to own both the cost and the outcome. Large US telecom operator A, large medical device company A, and large European automaker A in the first round, and large European insurance company A, large Canadian consumer goods company A, and large Southeast Asian retail company A this round, are two different sets of companies describing the same mechanism.

4. There is an order-of-magnitude gap between having a license and actually creating business value, and usage depth stays concentrated in engineering roles and a small group of heavy users. This holds in both rounds. This round pushes the gap to a more extreme point: Copilot coverage at large UK public sector organization A is close to 100% with 80%–85% monthly actives, while large European insurance company A has opened Copilot to all staff and estimates that under 2%–3% of employees derive clear business value from it.

5. Enterprises broadly run Budgeting rather than Tokenmaxxing; allowances are not hard caps, and high-value projects and heavy users can apply for more. Only 1 of the 12 enterprise samples in the first round leaned towards Tokenmaxxing; all 7 samples this round run Budgeting, in forms ranging from project approval to individual allowances to consumption-based chargeback.

6. The latest frontier models have not produced a broad step-up in budgets; their main effect is to raise the accuracy and complexity ceiling of existing use cases. Both rounds agree. The additional evidence this round is claims quality assurance at a large European insurance company A: switching models lifted accuracy by roughly 15 percentage points on the same data and context, but the use case itself is still the original pilot.

7. Falling cost per task and rising total AI spend can happen at the same time, with the savings reabsorbed by incremental usage and new use cases. Both rounds agree. The cross-client consultant this round puts the split at roughly 50 / 50, versus 60%–70% reinvested in AI from large US telecom operator A in the first round; the direction is the same, and the magnitude depends on where the company sits in its adoption cycle.

B. Mechanisms and constraints that did not appear in the first round

1. Hours saved carry a threshold effect. Large European pharma professional services company A notes that 5%–10% of scattered hours saved does not lower project cost, because that time is not enough for an employee to take on another piece of work; only when the effort on a class of work falls by roughly 30% or more can the freed capacity be reallocated. The first round discussed how high ROI is; this point explains why ROI is hard to realize.

2. Part of the efficiency gain is captured by clients. Clients of the same company have started asking for 10%–20% price reductions on projects, while the company estimates it can actually support only around 5%. The first round framed the destination of savings as two internal outlets, reinvestment or direct cost reduction; this round adds a third, external one.

3. The switching threshold for open-source models has been quantified for the first time. Cost has to be roughly 4x cheaper before the retesting and workflow rework are worth it; if GPU utilization on a self-built inference cluster falls below roughly 80%, the Self-host cost advantage largely disappears; and Hyperscaler hosting removes only around 20% of procurement concerns. The first round offered a directional TCO view but no decision thresholds that can be used directly.

4. The basis of the procurement decision is spelled out. The cross-client consultant’s read is that management now buys on vendor reputation and risk rather than model capability, and that none of the 50–60 large enterprises he works with has formally deployed an open-source model. The first round attributed the open-source blockage to infrastructure and compliance processes; this point points at the procurement logic itself.

5. Clear evidence of substitution has appeared in the funding source. After being told to hold the total budget flat, large US pharma company A has started assessing cuts to ServiceNow, Salesforce, and part of its on-premise data platform spend; the cross-client consultant also sees leading enterprises with up to roughly 40% of the IT budget already redirected to AI, leaving limited room for central IT to expand further. The first round went only as far as business units funding AI, without identifying what gets cut.

6. ROI has started to constrain resource allocation in reverse. Incremental allowances have to be tied to a specific project; business units that cannot articulate the expected benefit do not get incremental budget, and the CEO and CFO at large US pharma company A have used this to hold the total budget flat. ROI in the first round was still largely retrospective.

C. Observations that differ across the rounds but should be attributed to sample type

1. Open-source penetration in production. In the first round, 5 of the 9 quantifiable samples already ran open-source models in production, with a call share of up to roughly 40%; this round, only large European pharma professional services company A among the 7 enterprise samples runs them at scale. The two sample sets are very different in composition: the first round skewed towards telecom, e-commerce, automotive and IT consulting, all technically self-sufficient companies, while this round skews towards insurance, pharma, public sector, retail and manufacturing, all heavily regulated and procurement-driven. Read together, the two rounds suggest that the penetration boundary for open-source models currently sits inside technically capable companies, and that regulated enterprises have not yet cleared procurement and risk approval; reading it as a fall in penetration would be a misreading.

2. Adoption of a unified Router. The first round included samples that had already deployed one; none of the 4 samples that answered the question clearly this round has one in place. This is again a difference in company type. Taken together, a Router is mainly what technically capable companies do, while most enterprises still tier models through training, manual selection, and workflow evaluation.

3. How far cost-optimization headroom is quantified. The first round produced hard numbers such as 20%–30% and 8%–12%; this round, most experts could not give an overall estimate. This looks more like a difference in whether the interviewed company has built a cost model. Optimization this round is happening inside the workflow, and the headroom has not disappeared.

New Modules in This Round

Organizational and workflow change (now a standalone section). The first round touched on this only in scattered form across the sections. This round is the first cross-cut across all 19 enterprise samples from both rounds, covering the speed gap between tool rollout and employee adaptation, progress on non-coding workflow redesign, the sequence in which organizational and accountability boundaries change, and the constraints that actually bind today.

Six-month AI spend outlook across both rounds (Appendix 1). Places the six-month spend expectations of both rounds of samples on a single chart for comparison.

ROI measurement approaches (Appendix 2). Sorts all 21 samples into three groups by the measurement approach the expert disclosed, and summarizes the primary value sources. The classification reflects measurement approach only and says nothing about return levels.

New question dimensions added to the main text include the switching-cost threshold for open-source models and the economics of self-hosting, whether enterprises have actually deployed a unified Router, whether hours saved convert into reallocatable capacity, and whether clients have started asking for AI-driven price reductions.

1. Budget Growth and Outlook

We interviewed 7 enterprise samples plus 1 cross-client enterprise AI transformation consultant this round. The table below summarises each sample’s AI spend trajectory and budget expectations.

Enterprise AI spend is still growing overall, but the samples in this round are more dispersed than in the first round

Growth expectations vary widely across samples in this round, and the dispersion is more pronounced than in the first round.

  • Some enterprises are still in rapid expansion, with the increment coming mainly from more Use Cases moving into Production and from usage broadening from development teams to business users.
  • Agentic AI and business use case spend at large European insurance company A has nearly doubled over the past 3–6 months, with group-level spend now above $250k/mo, roughly 70% of which is API/Consumption. As projects developed during 2026 go live, the expert expects consumption to reach 4–5x current levels by March 2027, possibly more; roughly 70% of new projects target cost optimization and 30% revenue growth.
“It’s almost double, almost double now, and I am expecting it to go in much higher multiples in 2027 because we have a lot of use cases that are currently under development in 2026... By March 2027 we would be at least four times or five times, if not more than the current consumption... Once it goes into production and business users start using it, it’s going to explode.”
  • The Singapore-based enterprise AI transformation consultant (cross-client) says AI budgets at the enterprise and government clients he works with have roughly doubled over the past 12 months, and could double again by July 2027 before slowing to roughly 20%–30%. Growth over the past year came mainly from more employees being given paid AI tool access; usage by existing users has not risen materially. As individual AI tools finish rolling out across large enterprises, the incremental spend will depend more on Agentic AI and business workflow deployment.
“So I would say over the last kind of 12 months, budgets have doubled... It will double again to July next year. Yes, after that, so probably one more year it’ll slow down to about 20 or 30%.”“Not so much higher usage by existing users. The new users... only recently came onboard in the last 12 months.”
  • Other enterprises are still adding AI spend, but release budget step by step against defined Use Cases and pilot progress.
  • Large Southeast Asian retail company A currently spends ~$40–50k a month on its AI platform, excluding productivity tools such as Gemini Enterprise and OpenAI used by employees. Indexing January spend to 100, July is roughly 120 and December is expected to be roughly 140; the company keeps month-on-month growth within 3%–4% by optimizing Context and tokens. Total platform cost is expected to grow ~20% in 2027, and up to ~30% if more deployable use cases are identified during the year.
“It’s growing at a rate of not more than 3 to 4% month over month because we are also trying to optimize our context and tokens.”“We are expecting at least a 20% jump in the total AI, API, overall data processing and everything... The 20% can jump to 30% depending on the use case identification during the year.”
  • The sales and marketing AI budget at large Canadian consumer goods company A has risen from $250k/yr to $350k/yr, up ~40%. The original $250k funds 5 AI solutions already running, and the incremental $100k will go to two pilots over the next six months. The company plans to broaden usage once the pilots succeed, taking coverage from roughly 100 users today to 150 by June 2027; the total budget it intends to request for FY27/28 is ~$600k, with ~$400k for solutions already live and ~$200k for 4 new pilots.
“I just unlocked this 100000. My budget was 250000, so now it’s becoming 350... For the hundred thousand that I have, it will be only a few employees at first because this is a pilot project. It’s six months where maybe 10 people maximum will have access to these two AI solutions.”“By June 2027 my goal is to have 150 users... For this fiscal year, 350. For next year, I will be asking... something around 600K.”
  • AI subscription spend at some enterprises is starting to flatten for the rest of this year, while API spend and new solution development budgets keep growing. At some enterprises, API spend is also down versus the start of the year.
  • Within the roughly 6,000–7,000 person service delivery team at large European pharma professional services company A, monthly Coding Agent spend has risen from ~$1k or less in December 2025 to ~$18k today (on the expert’s stated basis). Over the next six months, existing Copilot and Coding Agent budgets are expected to stay broadly flat, with higher token usage offset by cheaper models and a lower Cost per Task; the team’s 2027 AI and automation development budget is expected to reach 2.5–3x 2026, mainly for new project development.
“That spend, let’s say that in December, was maybe 1000, or even less, and it grew exponentially all the way out to more or less 18000.”“We don’t expect to increase the budget. What we expect is that people will use more tokens, but they will use cheaper models... For next full year, 2027 compared to 2026, it’s probably somewhere around 2.5 to three times larger. But it’s mostly for development of new solutions.”
  • Large UK public sector organization A is one of the few samples where token spend is down YTD. The company re-architected its first AI product and reduced calls to Gemini, so Credit consumption fell slightly and token spend is now ~£10k/mo. As more products enter development and pilots widen, however, the expert expects token spend to roughly double over the next six months. The Microsoft Copilot license fee is ~£1.5m/yr and is expected to be broadly unchanged over the next six months. Token growth will come mainly from AI product development and pilot expansion: a pilot may start with 10–20 employees before rolling out to a business unit of roughly 600.
“The first product we launched used Gemini. We then rearchitected the product, which reduced the model consumption, so the credits actually dropped a little bit for us.”“Our license... that’s not going to go. That can be status quo. In the next six months, I can see our token usage doubling... A lot of it would be used in R&D or development because we’ll be piloting stuff before we scale out.”“When a pilot, we might be trying with, say, 10 people or 20 staff, there might be 600 in that unit, 600 staff that we’ll roll that out to.”
  • Some enterprises have started tightening budgets, requiring incremental AI spend to be funded by other cost reductions or supported by a clearer ROI; one expects AI spend to stall in the second half.
  • Total AI spend at large US pharma company A is up ~225% over the past six months, and token usage could still rise a further ~30%. The CEO and CFO have asked for the overall budget to stay flat and for AI projects to show clearer ROI and direct cost savings, so the expert expects AI spend to be flat to slightly down from here.
“It’s gone up 225%.”“It may actually shrink a little bit because my CEO and CFO are asking for better ROI metrics and direct budget reductions because of AI... Tokens will probably increase 30%.”“There is pressure from the CFO and the CEO to keep our overall budget flat, which means the rest of the organization has to cut, let’s say, people expenses by a little bit.”
  • From the second quarter of 2026, the local R&D subsidiary of a very large Southeast Asian manufacturing company set a $400 per person per month enterprise AI allowance for roughly 500 employees, equivalent to a nominal Credit cap of ~$200k/mo. Because most employees do not use the full allowance and non-engineering functions have yet to show a clear productivity gain, the expert expects the overall allowance budget to fall to about half of current levels over the next six months; allowances will be concentrated on engineers working on AI projects, Agents and Workflows, and reduced for non-engineering roles.
“After six months? Probably half, because when we realize that a lot of people are not using it maximum and a lot of productivity are not being born by using more AI, I think senior management will see this.”“The non-engineers will have less, and the engineers who are active on AI usage are having more access to it and creating more productivity.”

The AI budget continues to broaden from central IT into the business units, and incremental budget is starting to displace part of the existing IT and headcount spend

  • Roughly 80% of the AI budget at large European insurance company A is now requested and funded by the business units against specific Use Cases, with central budget covering only ~20%. The funding may come from incremental budget or be released by deferring existing IT projects. Large Canadian consumer goods company A also co-funds AI between central IT and the business units, currently split roughly 80% IT and 20% business; the company plans to take AI from ~3%–4% to 15% of the IT budget over the next three years, with part of the increment coming from cutting or consolidating other IT projects and software licenses.
‘For each of the use cases that comes through, we need to get the funding from the business... 80% is like that. 20% is something which we have centrally funded.’
  • The Singapore-based enterprise AI transformation consultant (cross-client) notes that some leading enterprises have already redirected up to ~40% of the IT budget to AI, leaving limited room for central IT to expand further. Incremental budget is expected to come increasingly from project budgets in sales, marketing, customer service, and other business functions, and from funds released by reducing headcount, MarTech, and Salesforce spend. Large US pharma company A is going through a similar shift: AI spend previously came mainly from incremental budget, but after the CEO and CFO asked for the overall budget to stay flat, the company has started assessing reductions in ServiceNow, Salesforce and part of its on-premise data platform spend.
‘The IT budgets for technologists are frozen... It’s all going to come from project-based CapEx, basically business unit heads going to the CEO and asking for a specific budget, or it’s going to be funded by cost savings.’

As in the first round, employee subscriptions are maturing, and incremental spend keeps shifting to APIs, tokens and production workflows

  • The Singapore-based enterprise AI transformation consultant (cross-client) sees the subscription-versus-API-and-token split in enterprise AI software budgets move from roughly 90% / 10% a year ago to roughly 50% / 50% today. Business-application spend at large European insurance company A is already API-led, currently ~70% consumption-based and 30% fixed fees; as projects in development move into production, the expert expects API calls to become the main source of increment.
‘A year ago, it would have been 90% subscription, 10% API and token. This year, the whole pie has increased in size, and it’s about 50% subscription, 50% token.’
  • Roughly 70% of the relevant spend at large Canadian consumer goods company A is subscription and 30% API. The company initially bought mostly off-the-shelf SaaS products, but once it confirmed some solutions would be used continuously, it began connecting them to internal sales and marketing data; the expert believes API integration is better for both business outcomes and cost, though data security and data ownership approvals typically take around six months.
‘We mainly pushed towards subscriptions and towards cloud... But now we have understood that business-wise and cost-wise, it’s better for us to work towards APIs and integration.’
  • Large UK public sector organisation A has already provided Copilot licences to roughly 3,000 employees at an annual cost of ~£1.5m, expected to be broadly unchanged over the next six months; token spend is ~£10k/mo today, but as projects widen from 10–20 person pilots to several hundred business users, the expert expects token spend to roughly double over the next six months. Large Southeast Asian retail company A likewise expects the API share of spend to keep rising; across the projects it has already deployed, cost after a model moves from development testing into production is typically ~40%–50% higher than in the development phase.

Paid AI tool coverage keeps rising, but active usage and actual business value still vary widely

  • Large European insurance company A has opened Copilot to all employees, but fewer than 2%–3% genuinely create clear business value with AI, concentrated in coding, and a small number of business applications already live; most of the rest still use it for email replies and document summarisation.
‘Copilot is live for everybody... Currently it would be less than 2% or 3% deriving actual business value out of it. The majority of these guys are on the tech side because they’ll be using it for coding.’
  • Roughly 30%–40% of the white-collar employees at large Southeast Asian retail company A have access to enterprise AI tools such as Gemini Enterprise and Claude, of whom ~60%–70% are weekly actives. Adoption was only ~15%–20% at the start; after sustained training, usage of some marketing, operations, and supply chain tools reached 90%–95%, showing that training and clear role-level Use Cases remain the key drivers of adoption.
‘Initially, our adoption rate was only 15% to 20%... Some of the functions are already at 90% to 95% in some tools.’
  • Copilot coverage at large UK public sector organization A has reached 100%, with monthly active usage of ~80%–85% and roughly 4 Prompts per active user per day, though the main use cases are still Chat, meeting summaries, and first drafts of reports. The local R&D subsidiary of a very large Southeast Asian manufacturing company A has also allocated enterprise AI Credits to roughly 500 employees, but fewer than 10% have actually applied for a higher allowance; future budget will shift towards engineers working on Agent and Workflow development.

The latest frontier models still have not produced a broad step-up in budgets, and enterprises are paying more attention to model tiering and cost efficiency

  • When the local R&D subsidiary of a very large Southeast Asian manufacturing company first opened up Fable, roughly 70%–90% of token traffic went to Fable, mainly because employees defaulted to the most capable model. Allowances were exhausted quickly, so the company ran model-selection training; roughly 60%–70% of traffic has now moved to Sonnet, and Fable's share is down to ~10%–20%. Fable still has an edge in a small number of use cases such as long context, multi-agent workflows and Judge Model work, but it is not suited to ordinary summarisation and document tasks.
‘In the first month... around 70% to 90% were using Fable... After that, we did a lot of training and demonstrations on how each model is different. Now around 60% to 70% are using Sonnet, and Fable is really only 10%.’
  • Large European pharma professional services company A believes mid-tier model capability has improved markedly over the past two months, and most of its consulting, software development and repetitive workflows no longer need the frontier. In high-frequency, repetitive Agentic Workflows, Open-weight models already account for ~30%–40% of usage; the expert expects token usage to keep rising, but a lower Cost per Task should keep the existing employee tool budget broadly unchanged.
‘Everything changed in the past two months, where we have a lot of good-enough models and we don’t need to go to the frontier.’
  • Large US pharma company A is moving towards a three-tier model structure: Opus for orchestration and complex reasoning, and Sonnet and Haiku for summarisation, document Q&A and Proposal generation. Fable added roughly 10% of incremental Workload after launch, but tasks only Fable can complete account for ~5%–7%, mainly highly specialized research; on price and data-retention grounds, the company allows only a small number of employees to use it.
  • Other enterprises are equally restrained on frontier models. Large Southeast Asian retail company A did not raise allowances when stronger models launched, and engineers are expected to choose between Claude, Gemini and other models by task. Large European insurance company A sees clear improvement on long-context tasks, but applications already in production are not switched lightly because a version change requires re-testing output stability and clearing governance approval; Fable also has not been approved on data retention and client PII grounds.

2. Open-Source Model Usage

Open-source penetration in production is lower in this round’s samples than in the first round, with only one enterprise using it at any scale

  • Among the 7 enterprise samples and 1 cross-client consultant in this round, only the expert at a large European pharma professional services company A says Open-weight models are meaningfully used in some production workflows. Roughly 10%–20% of its Coding Agent work runs on Open-weight models; in Agentic Workflows where the task is more stable and can be tested repeatedly, the Open-weight share has reached ~30%–40%.
‘For the coding agent... probably somewhere around 80% to 90% closed model still... for the agentic workflows... maybe 30%–40% open-weight models and 60%–70% closed models.’
  • Open-source models today are used mainly for simple, high-frequency, repetitive tasks. The company runs the same test set across Kimi, GLM, DeepSeek, and other models and picks the cheapest that meets the output requirement; complex reasoning, large-context work, and multi-Agent collaboration still run on closed frontier models.
The expert estimates that on some high-frequency repetitive tasks, Open-weight models can deliver the same output at roughly 3–4x lower cost, and up to 10x in some cases. Because switching models also requires retesting and reworking the workflow, however, the company will generally not migrate for a cost advantage of only ~2x; roughly 4x is needed to cover the switching cost and risk.‘For this high-volume repetitive task... at least three to four, up to ten times cheaper... the price difference would have to be at least fourfold or higher to consider... moving away from closed models.’
  • These Open-weight models are still consumed mainly through cloud platforms such as AWS and Azure, with very little deployed by the company itself. Self-hosting has been discussed at the company level but has not been scaled.

The other enterprises do not yet run open-source models in production, constrained mainly by risk approval, procurement habits and internal deployment capability

  • The Singapore-based enterprise AI transformation consultant (cross-client) says none of the roughly 50–60 large enterprise and government clients he works with has formally deployed an open-source model. Enterprises still lack confidence in vendor stability, model updates, long-term service capability, and their own deployment capability; model capability and price are no longer the main obstacle. At some Asian enterprises, management also weighs whether the model vendor will still be in business, and whether export restrictions could apply.
‘I work with maybe 50, 60 different big companies. So not one of them wants to deploy an open source model at this stage... CEOs are not buying on model capability anymore; they’re buying on reputation and risk.’
  • Offering open-source models through AWS and other Hyperscalers reduces some procurement concerns but does not fully resolve vendor stability and policy risk. The expert estimates Hyperscaler hosting removes only around 20% of the concerns.
‘If it’s deployed with the hyperscalers, plus people like AWS, it helps... It probably relieves it by about 20% and does make it easier if there’s an integration with pre-existing partners like AWS. But again, it doesn’t resolve all the issues about whether DeepSeek is still going to be in business or still going to be offered to companies outside of China.’
  • For insurance, retail and large manufacturing companies, compliance and security approval remain the more direct constraint. Large European insurance company A has not approved any open-source model for production. Some teams have already deployed DeepSeek or built connectivity in the test environment and could move it into production quickly if approved, but it ultimately did not clear regulatory, governance and security review. The expert notes that insurers have close to zero risk appetite and are therefore cautious about model stability, version updates and ongoing maintenance. Large Southeast Asian retail company A and the local R&D subsidiary of very large Southeast Asian manufacturing company A are likewise still at individual testing and POC stage, and have not entered enterprise production.
‘Because of regulatory and governance issues, they are not given access to those, though some of the teams... had already deployed the DeepSeek model or built connectivity to it in the test environment... Insurers have a very, very low risk appetite. Their appetite for risk is almost zero... If there’s an error or bug, we’ll have to wait for the open source community to fix it.’
  • Others still prefer SaaS and closed models already integrated with their existing cloud platforms. Large Canadian consumer goods company A currently buys SaaS Cloud Solutions first on group instruction, mainly because those products can be used directly, whereas open-source models still require additional data governance, legal frameworks, internal development, and Fine-tuning capability. Its procurement function estimates that an open-source approach may be only ~15%–20% cheaper than the current cloud approach on a fully loaded basis, which is not enough to justify migrating now.
‘There’s a lack of knowledge... We need data governance. We need to have legal... open source is something that is not yet really understood by the leading team.’
  • Large UK public sector organization A is in a similar position. Calling models through Azure Foundry or Snowflake lets it reuse existing procurement and security frameworks; introducing a new open-source model or vendor would require procurement, deployment and security assessment to be redone, and the organization currently lacks the capability to assess model security independently. Large US pharma company A has tested Nemotron locally, but token generation on a Mac is slow, and the cost advantage of Local AI PC and self-built inference has yet to be confirmed.

Most enterprises will keep testing open-source models, but migration is slow and will start with small-scale POCs and cloud-based calls

  • Most samples remain positive on the long-run direction for open-source models but gave no explicit usage share. The Singapore-based enterprise AI transformation consultant expects that over the next 18 months to 2 years, as enterprises get more comfortable with Agents and model deployment, some clients may shift towards cheaper open-source models with more control. Large Canadian consumer goods company A expects to start considering open-source models in about a year, as it hires AI talent and builds out internal governance; large US pharma company A likewise expects to add Open-weight models as a formal option within the next year.
  • Large UK public sector organization A is slower still, expecting a more systematic evaluation of open-source models only in 2–3 years. The main drivers include potential increases in closed-model token prices, growing public finance pressure, and rising European attention to data sovereignty and non-US models.
‘I think over the next two to three years, we might be exploring more the open source area... there’ll be two dimensions: one is geopolitical and the other one would be financial.’
  • Some enterprises remain cautious. Large European insurance company A wants to increase AI usage and value in its business units first, and will only then consider open-source models as a cost lever. Large Southeast Asian retail company A will adopt only after a model clears CISO and data security certification; the local R&D subsidiary of very large Southeast Asian manufacturing company A believes models such as Kimi K3 are already close to closed frontier models on capability, but expects compliance constraints to keep them out of formal products.
  • On deployment channels, pilots will likely run first through Hyperscalers, existing cloud platforms, or small-scale local devices, rather than building a large self-hosted inference environment in the near term. Large US pharma company A is comparing Local AI PC against cloud options such as AWS; large UK public sector organization A may test models on its own cloud GPUs; and large European pharma professional services company A, which already uses Open-weight models, also calls them mainly through AWS and Azure.
  • Whether self-hosting saves money depends mainly on volume and GPU utilization. The expert at large European pharma professional services company A estimates that if GPU utilization on a self-built inference cluster falls below roughly 80%, the cost advantage of local deployment largely disappears. Even as open-source usage rises, therefore, incremental calls are still likely to run through cloud platforms in the near term.
‘If all hundred GPUs are in use less than 80% of the time, you’re already getting rid of the price advantage that you have by hosting it yourself.’

3. AI Penetration

As in the first round, paid AI tool coverage keeps widening, but there is still a clear gap between having a license and using it heavily and deeply

  • Large UK public sector organization A has extended paid Microsoft Copilot licenses to roughly 3,000 employees, taking coverage from ~10% to close to 100%. Part of the rapid expansion came from a steep discount offered by Microsoft; the main uses today are still searching internal material, summarising documents and handling routine office tasks.
  • Large Southeast Asian retail company A has roughly 1,200 white-collar employees, ~30%–40% of whom have enterprise AI tools such as Gemini Enterprise and Claude. Of those with licenses, ~60%–70% use them weekly; adoption was only ~15%–20% at the start of the rollout, and after training and an AI Literacy program, usage in some functions on some tools has reached 90%–95%. The company plans to keep expanding in batches of 200–300 people rather than opening access to all employees at once.
‘Initially, our adoption rate was only 15% to 20%... We’ve been doing a lot of training... some of the functions are already at 90%, 95% in some tools.’
  • Large European pharma professional services company A has opened Microsoft Copilot to everyone in its roughly 6,000–7,000 person service delivery team, but only ~20% actually use the Coding Agent or more specialised Copilot Agents; newer tools such as Claude Cowork are still in a small pilot of roughly 100 people. All ~60,000 employees at large European insurance company A can also use Copilot inside Microsoft Office, but most still use it for simple tasks such as document summarisation and email replies, with deeper usage concentrated among developers and a small number of business teams.
‘Everybody in the organization has Copilot... but there are a majority of users... using it just to summarize this document, summarize the email or draft an automatic reply.’
  • An internal ChatGPT deployment at large Canadian consumer goods company A is already open to Canadian employees, but Copilot and other paid AI Solutions cover only ~100 people. The company first selects employees already comfortable with AI to act as “AI Champions,” who then train and pull along other business teams; the user count on existing business AI Solutions is planned to rise to ~150 by June 2027, with a longer-term goal of ~1,000 Canadian employees.

Usage depth remains concentrated in engineering, the AI team and a selected group of business power users, and non-technical usage remains shallow

  • From the second quarter, the local R&D subsidiary of a very large Southeast Asian manufacturing company A has provided external AI tools to roughly 500 employees on a uniform basis, with a $400 per person per month allowance. Most employees still use them, but usage intensity has clearly diverged by function: most engineers consider the allowance insufficient, and ~10% have applied to raise it to $600, while finance, HR, and similar functions use much less and many employees do not exhaust their allowance. Engineering uses it mainly for unit testing, CI/CD documentation, code analysis and Agent and Workflow work, while other functions remain largely in chat and simple Q&A.
‘From what we have seen, AI is heavily used in our engineering department... the usage is very small for other departments.’
  • Roughly 40% of employees at large US pharma company A have paid AI tool access, equivalent to ~600 people across legal, marketing, commercial, R&D and software development. Productivity gains are concentrated in a small number of mature teams, however: the expert’s own AI team has lifted development throughput by ~52%–120%, while other teams may be up only ~4%–10%.
‘My AI team... I’m seeing anywhere between 52% to 120% of velocity gain... But then other teams, they’re at 10%, 4%.’
  • Adoption speed also varies widely across non-technical functions. Usage at large Southeast Asian retail company A is deepest in marketing operations, merchandising and supply chain; finance has some tool access but, because models still make errors on numbers and tables, only a small number of employees actually use them. Large Canadian consumer goods company A is rolling out mainly through “AI Champions” in sales and marketing, starting with employees already familiar with AI who then train the other business teams.

Enterprises are starting to embed AI into Agents and business workflows, but most samples are still at pilot, phased-expansion or partial-production stage

  • The Singapore-based enterprise AI transformation consultant (cross-client) notes that individual productivity tools are already widespread in large enterprises, but Agentic AI that actually connects to internal systems and executes tasks continuously is still rare. The real constraint is that enterprises remain unwilling to open SharePoint, internal files, databases, and business applications to Agents because of data leakage, misoperation, and system risk; whether employees can use ChatGPT or Copilot is no longer the question.
‘Almost every company in Singapore has allowed employees to use AI for personal productivity... The problem is very few companies are using agentic AI, and that’s where tokens are consumed and internal systems are touched.’
  • A large European insurance company is testing a claims quality assurance Agent. A manual review of one claim typically takes half a day, so coverage is limited today; after introducing multiple Agents, test accuracy has risen from ~70% to above 85%, and the company is preparing the Business Case to take the pilot into production. The company’s model spend of more than $250k/mo also goes mainly to this kind of API and business workflow work, with employee Copilot licenses a small share.
  • Large UK public sector organization A is likewise extending AI from individual office tools into specific business processes. A pilot may start with 10–20 employees before rolling out to a business unit of roughly 600. The expert expects token usage to roughly double over the next six months, with the increase coming mainly from product development, pilot expansion, and embedding AI in live workflows.
‘When a pilot, we might be trying with 10 people or 20 staff; there might be 600 in that unit that we’ll roll that out to.’
  • A small number of samples have built reasonably stable production workflows. Large European pharma professional services company A has deployed custom analytics Agents for clients that handle multiple structured data sources and are used daily by tens to hundreds of client employees; the company has also consolidated repetitive tasks such as transcription, translation and downstream analysis into complete workflows. Of the 6 AI Solutions currently running at large Canadian consumer goods company A, 2 already connect to internal knowledge bases or sales data via API, with 2 more pilots planned over the next six months. Overall, what has moved fastest into production this round is still workflows with clear boundaries, repeatable execution, and business outcomes that are easy to measure.

4. Tokenmaxxing or Budgeting

As in the previous round, the enterprises here broadly run Budgeting, though the mechanics range from project approval to individual allowances to consumption-based chargeback

  • None of the 7 enterprise and institutional samples in this round describe themselves as leaning toward Tokenmaxxing. Three have set individual or team allowances, and two manage mainly by specific use case, project budget and ROI; of the remaining two, large European insurance company A charges back on actual usage and large UK public sector organization A is still mainly monitoring consumption. The cross-client enterprise AI transformation consultant separately notes that some enterprises have set Token allowances or Token budgets, but that limits are generally loose at this stage.
  • Large Canadian consumer goods company A has no Token allowance and manages budget by specific use case. The company is still in trial and validation, but every new AI project must submit a business case requiring ~$2.5 of return per $1 invested, which can come from cost optimization, lower external service spend, or sales growth.
‘On every dollar that I’m spending for an AI solution... I need to justify a $2.5 return on the investment.’
  • Large European insurance company A sets no fixed Token budget and charges each team internally on actual usage. Token consumption has to be forecast before go-live, and a usage dashboard provides daily consumption and cost alerts; if a project forecasts 1m tokens a month and actually uses 5m, the company will ask why. This approach does not cut off service when a team hits a uniform cap, but it does make the consuming team carry the cost.
‘We have not put in a token budget policy as such... We don’t want to restrict you. But it is very clear to them that they will get billed based on what they use.’
  • Large UK public sector organization A has no formal Token cap today, monitors usage by project, and decides case by case whether to continue. The expert expects that once a project moves into production, the spend will be brought into the formal budget on an actual opex and run-rate basis. Some of the enterprises the cross-client consultant works with are at the same stage: Token budgets exist, but limits are loose, the focus is still tool rollout and employee training, and accountability for eventual ROI is not yet clear.

Enterprises generally do not set the same hard cap for every employee; high-value projects and heavy users can still apply for more

  • None of the 3 samples that gave explicit allowances runs a fully fixed one. Enterprises typically set a default allowance for ordinary employees and then allow some users to apply for more based on role, team, and specific project; high-cost models may also sit behind separate access controls.
  • The default Coding Agent allowance at large European pharma professional services company A is $70/user/month, and employees can apply for more; each team sets the final cap, with an average maximum of ~$200/user/month, according to the expert. Teams must define and enforce the allowance up front, and can also do more work within the existing budget by using free or discounted models.
‘Users have a fixed — by default, they have a fixed budget of $70 per month, but they can request an increase... On average, the maximum is around $200 per month.’
  • The local R&D subsidiary of very large Southeast Asian manufacturing company A gives employees a default AI allowance of ~$400/month. Starting in August, employees working on a specific AI project can apply to raise it to $600; the expert says ~10% of applications were approved that month, and approvals continue monthly. API budget is requested separately by project, and employees who do not use AI tools for three consecutive months can lose access.
  • Large US pharma company A has ~600 paid AI users at ~$120/month each, of whom ~120 engineers are on $200/month 20x accounts and can request more Tokens through an enterprise approval if the existing allowance is insufficient. Fable 5 sits behind a separate access control and is open to only a small number of employees today.

Adoption of a unified model Router is relatively low in this round’s samples; model tiering today relies mainly on manual selection, training and workflow optimization

  • None of the 4 samples that answered the Router question clearly says it selects models automatically through a unified Router. Large European pharma professional services company A is planning one, large UK public sector organization A is exploring routing and compression middleware, large US pharma company A explicitly has no centralized Router, and large European insurance company A has built a unified model entry point but still lets each use case team pick the model.
‘We have something called an LLM Lounge, wherein you have API access to all of these different models. As a business user or a business unit, you can go and subscribe to the API and then use any one of the models that you want... We said global pricing for us... everybody is going through the same lounge or the same provisioned LLM.’
  • Large European pharma professional services company A plans to build its own open-weight model inference stack and add a unified Router across models. Each team and user still selects coding models today.
‘For now, for coding, we will leave this to each team... There is a more enterprise-wide initiative to host open-weight models on our own stack and to develop this kind of router between different models.’
  • Large UK public sector organization A is exploring routing and compression middleware such as Portkey. The expert says this is preparation for higher token usage ahead, with the aim of cutting some cost through model routing and context compression. Models are currently called mainly through Azure AI Foundry and Snowflake, and model choice still depends on the platform and the project team.
‘We are looking at routing and managing... compression through services like Portkey... when my token utilization starts to scale up exponentially, I need some middleware.’
  • Large US pharma company A guides model selection through a model-usage coaching Skill built inside Claude and has no centralized Router of its own. The company mainly uses Opus and Sonnet and plans to move toward a three-tier structure: Opus for more complex orchestration, and Sonnet and Haiku for simpler tasks such as summarisation, document Q&A, and proposal generation. The expert notes that developers still use over-powered models for ordinary tasks, so the company wants to push more usage onto Haiku.
‘We have created a coaching skill in Anthropic... but we don’t have a centralized router... We have Opus for more orchestration, while Sonnet and Haiku for simpler tasks.’
  • The local R&D subsidiary of a very large Southeast Asian manufacturing company A also relies mainly on training to drive model tiering. The expert estimates Fable accounted for ~70%–90% of model token traffic shortly after launch; as the company trained employees to distinguish between models and monitor allowance consumption, Fable's share has fallen to ~10%–20%. Employees are asked to use Sonnet first and escalate to Opus or Fable only where the task cannot be completed.
‘At the first month... it was around 70% to 90%... Now... Fable is really only 10% to 20%... We always try to use Sonnet first. And then when it fails... we can start calling Opus.’
  • Cost control at large Southeast Asian retail company A happens more at the workflow layer. The company cuts unnecessary context, caches repeated requests, runs simple steps on small models, and reduces token consumption through batching, better prompts, and fewer wasteful agent loops. The expert says this optimization runs continuously once a project is in production, aiming to control cost without affecting the user experience.
‘One is reducing the unnecessary context... Other areas are caching those repeated requests, using smaller models for simple steps... batching workloads... improving prompts and limiting unnecessary agent loops.’

5. Headroom in AI Cost Optimization

Experts in this round generally did not quantify how much further the current AI bill could fall, but several enterprises are already controlling API / Token cost by changing models, caching, compressing context, capping output, and redesigning workflows. As in the previous round, the unit cost of an equivalent task keeps falling; however, because usage and production use cases are still growing at most enterprises, the optimization mostly shows up as supporting more work within the same budget or slowing spend growth.

Headroom in the current AI bill is still concentrated in APIs / tokens and production workflows

  • Large Southeast Asian retail company A continuously cuts unnecessary context in production use cases, caches repeated requests, moves simple steps to small models, and controls cost through batching, better prompts, and fewer wasteful agent loops. The expert says the focus of the optimization is cutting spend that does not improve the user experience; compressing AI usage itself is not the objective.
‘Once an AI use case goes into production, there are quite a few opportunities to reduce cost without basically affecting the user experience... reducing unnecessary context... caching repeated requests, using smaller models for simple steps, batching workloads... improving prompts and limiting unnecessary agent loops.’
  • Large European insurance company A already uses caching and has started compressing over-long context and capping both output length and the token usage of a single request. If a use case only needs roughly 100 words or a single page, for example, the company limits the output scope directly instead of letting the model generate a full-length report.
  • The first AI product at a large UK public sector organization previously used Gemini; after the product was reworked, the associated Credit consumption fell slightly. The organization is also exploring model routing and context compression tools to tighten costs before token usage scales up.
‘Actually, we reduced the cost a little bit... The first product we launched, we used Gemini... and then the credits actually dropped a little bit for us.’

Model downgrading and open-weight models can lower unit cost, but the size of the saving varies widely by task and by company

  • Because a unified model router is still not widespread, model downgrading today relies more on employee training, manual selection, and workflow evaluation. At the local R&D subsidiary of a very large Southeast Asian manufacturing company A, the expert estimates ~70%–90% of model token traffic went to Fable shortly after launch; after training and demonstrations, employees came to accept that ordinary tasks do not need the most capable model, and the Fable share is now down to ~10%–20%. The company generally asks employees to use Sonnet first and escalate to Opus or Fable only where the task cannot be completed.
‘After the initial months, a lot of people... realized that we don’t need the most capable models... We always try to use Sonnet first. And then when it fails... we can start calling Opus.’
  • For more repetitive agentic workflows, large European pharma professional services company A runs the same task through 5–6 different models and compares output quality against cost. For workflows with simple steps and relatively deterministic outputs, open-weight models can deliver the same task at roughly 3–10x lower cost; complex tasks still run mainly on closed models.
‘For the same type of output... at least between three, four to up to ten times cheaper, depending on what it is.’
  • Other enterprises see materially smaller potential savings. Large Canadian consumer goods company A relays an initial estimate from its legal and procurement functions that an open-weight approach may be only ~15%–20% cheaper than the current commercial cloud approach; the company is still constrained by data ownership, internal technical capability and governance requirements, and has not formally adopted it.
  • Large US pharma company A is also costing out local deployment of open-weight models, but its AI lead’s initial view is that the fully loaded advantage after hardware, hosting and operations may not be significant, and the company is still working through the numbers.
‘He didn’t think there’s going to be that much savings... [on a] total cost of ownership [basis], because the hosting and things like that.’

Once unit cost per task falls, the saving tends to be reabsorbed by more usage and new use cases

  • Large European pharma professional services company A expects the existing AI tool budget to be broadly flat over the next six months even as employees consume more tokens. The incremental usage will be absorbed by cheaper models and a lower cost per task; when assessing model cost, the company also looks more at the actual cost of completing one task, and simply comparing price per million tokens is no longer the main basis.
‘We don’t expect to increase the budget. What we expect is that people will use more tokens, but they will use cheaper models... cheaper per accomplishing the task.’
  • Large Southeast Asian retail company A is in a similar position. Even as the company keeps optimizing context, models and agentic workflows, the expert still expects total AI spend to grow as production use cases increase. Cost optimization therefore mainly cuts waste and slows the rate of growth, and the total AI budget will not be reduced as a result.
‘The goal... is not necessarily to minimize AI spending... it is to eliminate where we are wasteful... While we are continuing to optimize the cost, there will be an increase in the overall AI spend.’
  • Token consumption at large UK public sector organization A has already come down, but as more projects move from development and pilot into rollout, the expert expects token usage to roughly double over the next six months. Falling cost per workflow and rising total token spend can happen at the same time.
  • The Singapore-based enterprise AI transformation consultant (cross-client) is one of the few experts to quantify where the savings go. Assuming model API prices fell 50%, he estimates roughly half the savings would convert into direct cost savings and the other half would be reinvested in AI, in more calls or new workflows.
‘I’d say at least half of that would be put down as savings and half of it will be redeployed back into AI.’
  • Optimization in this round’s samples therefore shows up mainly as lower marginal cost for incremental usage. As long as enterprises keep widening AI coverage, adding production use cases and deploying more agents, falling cost per task and rising total AI spend can continue to happen together.

6. ROI

Overall, and as in the previous round, ROI in API-driven use cases is generally clearer than on enterprise subscriptions and (company-funded) individual productivity tools. Business use cases that have gone into production are now assessed on sales, gross margin, cost savings and reallocatable capacity, and a small number of projects have explicit return requirements or forecasts; employee subscriptions, however, are still measured mainly on adoption and hours saved, and whether that converts into actual financial return remains unclear in several samples. As AI spend and its share of IT spend scale, some enterprises are demanding a clearer ROI and using it to decide whether a project can go into production, whether a business unit gets incremental budget, and whether an employee can request a larger Token allowance.

ROI on production use cases is generally clearer than on enterprise subscriptions and individual productivity tools

  • Large Southeast Asian retail company A runs relatively complete ROI management on production use cases. Every project first establishes a sales, gross margin or cost baseline, then compares actual results through POC and pilot; only projects that hit expectations move into production, with the metrics tracked continuously by the finance team. Roughly 80% of its use cases are judged mainly on sales and gross margin impact, and the remaining ~20% on automation and productivity.
  • By contrast, ROI on employee subscriptions such as ChatGPT and Gemini is harder to calculate directly, and the company relies mainly on employee surveys to estimate hours saved.
‘Every use case has a clear ROI. We track it along with our finance team... For subscriptions... sometimes it is hard to calculate the ROI... We ask teams how much time they were able to save.’
  • The Singapore-based enterprise AI transformation consultant (cross-client) sees the same: after handing out enterprise AI licenses at scale, most companies and organizations still mainly count logins, daily and weekly actives, and estimated hours saved. The expert sees these as adoption indicators that do not yet demonstrate financial return; rolling out employee tools is more of a precondition for the workflow redesign that follows.
‘The current metrics that they’re using are very... vanity metrics. What they’re counting is how many people use it — the raw number of logins, what kind of daily usage there is... estimated cost or time savings, but that would be it.’‘What is the ROI of 10,000 people having ChatGPT? What that really means is... it’s just a stepping stone to eventually figuring it out.’

A small number of enterprises have set explicit return requirements, but most projects are still early, and whether the actual return meets the approval hurdle remains to be seen

  • Large Canadian consumer goods company A requires the business case for a new AI project to show a return of ~2.5x, that is ~$2.5 of expected benefit per $1 invested. It can be met through direct cost savings, lower external agency fees, reallocated employee time, higher volumes, or better service levels. A pricing tool, for example, can justify itself through a ~10%–15% volume uplift, while a sales forecasting project can be measured by moving service levels from 90% to 95%.
‘On every dollar that I’m spending for an AI solution... I need to justify a $2.5 return on the investment.’(This is an approval and monitoring requirement, and does not mean the company’s existing AI projects have delivered 2.5x in aggregate.)
  • Roughly 70% of new AI use cases at a large European insurance company target cost optimization and ~30% revenue growth. Its claims quality assurance project is forecast to deliver ~5–6x: the company can currently review only ~5% of claims manually, and finds previously unreported reinsurance recoveries in the process; if AI lifts review coverage to ~50%, the business unit expects a marked increase in reinsurance cost recovery.
‘They are projecting that they will get at least five to six times return on the investment in the whole project.’ That return is still a project expectation derived by the business unit from wider review coverage. The company plans to observe actual performance before deciding whether to invest further.

Hours saved do not equal an actual cost reduction; what matters is whether reallocatable capacity is released

  • Large European pharma professional services company A explains this most clearly. Employees saving 5%–10% of their hours through AI does not usually lower project cost, because that scattered time is not enough for them to take on another piece of work. For its professional services teams, only when the effort on a class of work falls by roughly 30% or more can the freed time be reallocated to other projects and create real economic value.
‘A reduction of effort of 5% or 10% doesn’t do anything... We need to have a solution that decreases the current effort by at least 30%, so that we can free this person 30% of the time and then redeploy this resource into another project.’ At the same time, clients have started asking for ~10%–20% price reductions on projects, while the expert estimates the company can actually support only around 5% today. Part of the AI efficiency gain therefore goes first to meeting client price demands or protecting existing revenue, rather than converting fully into incremental profit.
  • Large UK public sector organization A also measures project returns mainly on expected hours saved. The expert’s example is that one or two use cases consuming ~£100k of Credit are expected to generate ~£2m of time value; the organization has not yet decided whether that time ultimately shows up as a direct cost saving or as higher output from the existing team. The number is therefore closer to an expected efficiency value than a realized cash return.
  • Large US pharma company A measures the effect of AI on engineering teams mainly through agile delivery metrics, including story points completed, commit volume, and the share of auto-generated code and automated tests. The expert sees development throughput up ~52%–120% in his AI team but only ~4%–10% in some others, which shows how differently the same toolset performs across teams and how hard it is to roll up into a single financial ROI.

ROI is starting to feed directly into project approval and budget allocation

  • After AI spend rose rapidly over the past six months, the CEO and CFO at a large US pharma company A have begun demanding clearer ROI metrics and direct budget savings. The company sees a reasonably clear effect in a few areas such as agentic coding, but productivity gains vary widely by team, so further investment requires more financial support from the relevant functions.
‘My CEO and CFO are asking for better ROI metrics and direct budget reductions because of AI.’
  • The local R&D subsidiary of a very large Southeast Asian manufacturing company A has also started tying incremental Token allowances to specific projects. Requests justified only by “needing stronger analytical capability” are not approved; today only engineers working on AI projects, agents, and workflows can apply for more. The company has also found that many employees do not exhaust their existing allowance, and that non-engineering functions have yet to show a clear productivity gain.
  • At enterprises with established project governance, ROI accountability is also moving to business units. Roughly 80% of AI projects at large European insurance company A require the business unit proposing the use case to set out the expected benefit and resolve the funding, with the remaining ~20% supported by central budget. Large Southeast Asian retail company A likewise makes the business unit responsible for the final business metric, with the AI team building the solution and helping measure the impact; without business ownership of adoption and outcomes, expected ROI is hard to realize even after the technical solution goes live.
‘The business function owns the business outcome. AI builds the solution and helps measure the impact. But if the business isn’t really accountable for the adoption and the KPI, it becomes very difficult to realize the ROI.’
  • Across this round’s samples, ROI measurement is moving from adoption and hours saved towards sales, gross margin, direct cost, and reallocatable capacity. Projects that can already be connected to a specific business outcome tend to find it easier to secure continued investment, while the value of employee tools and broad subscriptions is still hard to realize directly, and some enterprises have started tightening the budget review around them.

7. New Use Cases and Budget Growth

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论