Cheaper Models, Bigger AI Bills: Why the Agentic Cost Paradox Is Now a Board-Level ROI Problem

Category: AI Insights | Author: Colter Mahlum | Published: 2026-09-23

Weekly executive AI trends brief: September 23, 2026 The cost of AI intelligence has collapsed since GPT-4 launched in 2023. Enterprise AI bills are still rising. That is the agentic cost paradox now…

<p></p> <p><strong>Weekly executive AI trends brief: September 23, 2026</strong></p> <p>The cost of AI intelligence has collapsed since GPT-4 launched in 2023. Enterprise AI bills are still rising.</p> <p>That is the agentic cost paradox now confronting CFOs, CIOs, COOs, and boards: cheaper models reduce the price of each unit of capability, but autonomous agents consume more reasoning, more tools, more context, and more infrastructure to complete increasingly complex work.</p> <p>The relevant question is no longer <strong>“What does a token cost?”</strong> It is <strong>“What does a completed task cost, and does the surrounding workflow capture the value?”</strong></p> <h2><strong>Key Findings</strong></h2> <ul> <li>The same agentic task can cost <strong>up to 30x more</strong> from one run to another because agents take different paths to the same result.</li> <li>Approximately <strong>37% of organizations</strong> report measurable EBIT impact from AI, while only <strong>6%</strong> report significant EBIT impact, according to McKinsey’s 2026 State of AI research.</li> <li>Roughly <strong>60% of organizations</strong> plan to increase AI spending, even as approximately <strong>20%</strong> report that AI-related operating costs are constraining adoption.</li> <li>McKinsey’s rule of thumb suggests that a task taking a human <strong>1 hour</strong> may create value with an agent that succeeds more than <strong>10%</strong> of the time if its output can be verified in approximately <strong>6 minutes</strong>.</li> <li>Hyperscaler AI infrastructure spending is projected in the range of <strong>$660 billion to $900 billion in 2026</strong>, with some estimates placing capital expenditure growth at approximately <strong>36% year over year</strong>.</li> <li>Mahlum Innovations’ AI Strategy engagements target an average <strong>3.5x ROI</strong>, while Cloud AI implementation can enable deployment up to <strong>60% faster</strong> when architecture, governance, and operating requirements are defined early.</li> </ul> <h2><strong>What Changed This Week</strong></h2> <p>During McKinsey’s September 22, 2026, McKinsey Live session, <strong>“Improving the Economics of Agentic AI,”</strong> senior partners <strong>Tanguy Catlin</strong> and <strong>Lari Hämäläinen</strong> addressed the widening gap between falling model prices and rising enterprise consumption.</p> <p>As reported by <a href="https://fortune.com/2026/09/23/mckinsey-cheaper-ai-models-bigger-ai-bills-cfo/">Fortune CFO Daily</a>, models delivering performance comparable to GPT-4 now operate at a fraction of the 2023 cost. Yet organizations are asking AI systems to perform substantially more work: planning, coding, searching, calling tools, validating outputs, and repeating failed steps.</p> <p>Vendors capture some efficiency gains through higher margins. Enterprises capture the rest only when they engineer systems that avoid unnecessary token usage and connect AI activity to measurable business outcomes.</p> <p>Traditional software generally produced predictable costs per transaction. Agentic systems do not. An agent may complete a task with one model call, or it may invoke multiple agents, retrieve excessive context, call external tools repeatedly, and revise its answer several times. The result may be similar. The cost may be <strong>30x higher</strong>.</p> <p>That volatility makes agent economics a board-level issue rather than a procurement detail.</p> <p><img src="https://cdn.marblism.com/Ca1htL3oruW.webp" alt="Vector illustration of task-level AI economics with cost, success rate, and verification metrics" style="max-width: 100%; height: auto;"></p> <h2><strong>The New ROI Metric: Cost per Completed Task</strong></h2> <p>Cost per token is an input metric. It is not a business case.</p> <p>Executives should evaluate each AI workflow using at least three measures:</p> <ol> <li><strong>Cost per completed task</strong>: the full inference, infrastructure, integration, and human-review cost.</li> <li><strong>Agent success rate</strong>: the percentage of tasks completed to an acceptable standard without rework.</li> <li><strong>Human verification time</strong>: the time required to approve, correct, or escalate the output.</li> </ol> <p>Consider a claims-review workflow that takes a trained employee <strong>60 minutes</strong>. If an agent produces a usable first draft in <strong>6 minutes</strong>, the agent does not need a 99% success rate to create value. A success rate above approximately <strong>10%</strong> can begin to justify deployment, provided the successful outputs generate meaningful capacity and the failures are inexpensive to identify.</p> <p>The economic model changes when verification time rises from <strong>6 minutes to 30 minutes</strong>. At that point, the agent may be adding review work rather than removing it.</p> <p>This is why AI implementation services must include workflow measurement, not only model selection. The economic unit is the completed workflow outcome.</p> <h2><strong>McKinsey’s Three Cost Levers</strong></h2> <p>Catlin’s central message was direct: <strong>there is no single cost lever</strong>. Effective governance requires three.</p> <h3><strong>1. Create visibility into consumption</strong></h3> <p>Finance and technology leaders need reporting by:</p> <ul> <li>Use case</li> <li>Business unit</li> <li>Agent</li> <li>Model</li> <li>User</li> <li>Workflow</li> <li>Cost per completed task</li> <li>Human verification time</li> </ul> <p>A monthly cloud invoice cannot explain whether $50,000 in AI spend generated $150,000 in labor capacity, $500,000 in incremental revenue, or no measurable benefit.</p> <p><a href="https://mahluminnovations.com/ai-systems">Data Analytics</a> turns raw usage data into an operating view executives can act on. The objective is not simply to reduce consumption. It is to identify where consumption produces the highest return.</p> <h3><strong>2. Optimize the workflow and architecture</strong></h3> <p>System design directly determines task economics.</p> <p>Leaders should require:</p> <ul> <li>Model routing based on task complexity</li> <li>Smaller models for classification, extraction, and routine transformations</li> <li>Larger models only for high-complexity reasoning</li> <li>Cached reusable context</li> <li>Explicit limits on tool calls</li> <li>Maximum agent-loop thresholds</li> <li>Evaluation gates between multi-agent steps</li> <li>Human review at defined risk boundaries</li> </ul> <p>Single-agent and multi-agent architectures can produce dramatically different costs. A multi-agent system may improve reliability for complex research or software tasks, but unnecessary orchestration can multiply inference costs without improving the outcome.</p> <p>Mahlum’s <a href="https://mahluminnovations.com/case-studies/axiom-ai-orchestration">Axiom AI orchestration case study</a> illustrates the production discipline required for multi-agent systems: model routing, evaluation loops, integrations, monitoring, and support for multiple cloud and local models.</p> <h3><strong>3. Apply sourcing discipline</strong></h3> <p>Cheaper models do not compensate for unmanaged vendor sprawl.</p> <p>Executives should review:</p> <ul> <li>Unused licenses</li> <li>Per-user quotas</li> <li>Model and API commitments</li> <li>Enterprise discounts</li> <li>Data-processing terms</li> <li>Portability requirements</li> <li>Failover options</li> <li>Single-vendor concentration risk</li> </ul> <p>The goal is not to force every workflow onto the lowest-cost model. The goal is to match the cost structure to the value of the task while preserving resilience and negotiating leverage.</p> <h2><strong>The Workflow Redesign Imperative</strong></h2> <p>The hardest part of agent deployment is not connecting a model to a process. It is redesigning the process around the capacity the agent creates.</p> <p>If an agent saves <strong>20 hours per week</strong> but the organization has no plan for those hours, the financial result may remain zero. If an agent accelerates forecasting but inventory, procurement, and staffing decisions remain unchanged, predictive accuracy does not become margin.</p> <p>This is where <strong>Digital Transformation with AI</strong> becomes materially different from tool adoption. The workflow must define:</p> <ul> <li>Which decisions move faster</li> <li>Which roles absorb freed capacity</li> <li>Which approvals remain human-owned</li> <li>Which service levels improve</li> <li>Which revenue or cost metric changes</li> <li>How performance is monitored after launch</li> </ul> <p><a href="https://mahluminnovations.com/ai-systems">Machine learning consulting</a> is valuable when predictive models, classification systems, or anomaly detection are integrated into the decisions that control revenue, risk, and operating cost. For example, <a href="https://mahluminnovations.com/ai-systems">Predictive Analytics for business</a> can support demand forecasts, churn scoring, equipment-failure alerts, and risk prioritization, with forecast accuracy reaching <strong>up to 95%</strong> in appropriately scoped applications.</p> <h2><strong>Why Q4 AI Contracts Face More Scrutiny</strong></h2> <p>The macro environment is tightening.</p> <p>Analysts estimate that hyperscaler capital expenditure may reach approximately <strong>$660 billion to $900 billion in 2026</strong>. J.P. Morgan projections indicate that capital spending could consume <strong>90% to 100% of operating cash flow</strong>, compared with approximately <strong>60% in 2025</strong>, while projected cash shortfalls for Microsoft, Alphabet, Amazon, and Meta could reach approximately <strong>$600 billion through 2028</strong>. <a href="https://broadbandbreakfast.com/capex-to-consume-nearly-all-hyperscaler-operating-cash-flow-in-2026/">Industry reporting on the capex outlook</a> describes the scale of the infrastructure cycle.</p> <p>At the same time, cloud capex growth is expected to slow from approximately <strong>54% in 2025 to 19% in 2026</strong>. The implication for buyers is clear: capital discipline is increasing, and vendors will compete for fewer dollars subject to stricter ROI thresholds.</p> <p>Google Cloud and Accenture responded to this implementation gap with the September 8 launch of the <a href="https://newsroom.accenture.com/news/2026/accenture-and-google-cloud-deepen-partnership-with-formation-of-new-accenture-gemini-enterprise-business-group">Accenture Gemini Enterprise Business Group</a>. The initiative plans to establish a workforce of <strong>1,000 forward-deployed engineers</strong> to help customers integrate AI, move beyond pilots, and prove measurable value.</p> <p>The FDE model reflects the market’s central problem: enterprises are willing to invest, but they cannot justify systems that fail to integrate, scale, or produce auditable returns.</p> <p><img src="https://cdn.marblism.com/I07Ltt5LHgw.webp" alt="Minimal vector illustration of redesigned enterprise workflows, model routing, monitoring, and ROI" style="max-width: 100%; height: auto;"></p> <h2><strong>Questions to Ask Before Signing a Q4 AI Contract</strong></h2> <p>Whether you are evaluating an AI platform, cloud provider, systems integrator, or implementation partner, require written answers to these questions:</p> <ol> <li><strong>What is the expected cost per completed task at current and projected volumes?</strong></li> <li><strong>What is the agent success rate on our real data, not a vendor benchmark?</strong></li> <li><strong>How many minutes of human verification are expected per task?</strong></li> <li><strong>What happens when the agent fails, loops, exceeds a quota, or calls an unavailable tool?</strong></li> <li><strong>Can the system route work across multiple models and providers?</strong></li> <li><strong>What percentage of the workflow is genuinely automated versus merely accelerated?</strong></li> <li><strong>Which business KPI will change within 90, 180, and 365 days?</strong></li> <li><strong>Who owns the prompts, evaluation data, workflow logic, monitoring, and deployment artifacts?</strong></li> <li><strong>Can we export our data, logs, configurations, and model evaluations if we change vendors?</strong></li> <li><strong>What is the implementation plan for the freed capacity?</strong></li> <li><strong>What costs are excluded from the proposal, including integration, review, security, and change management?</strong></li> <li><strong>What production evidence demonstrates that similar customers achieved measurable ROI?</strong></li> </ol> <p>If a vendor cannot answer these questions, the contract is not yet investment-ready.</p> <h2><strong>Executive Action List</strong></h2> <p>For the next 30 days, executives should:</p> <ul> <li>Establish a baseline cost for the human workflow being considered for automation.</li> <li>Measure cost per task, success rate, and verification time, not only tokens.</li> <li>Identify the top five AI use cases by expected economic value.</li> <li>Audit licenses, quotas, unused models, and vendor concentration.</li> <li>Require model-routing and agent-loop controls in all production designs.</li> <li>Redesign the surrounding workflow before scaling agent volume.</li> <li>Set a documented payback period and executive owner for every AI deployment.</li> </ul> <p>Mahlum Innovations applies the <strong>RAPID Framework</strong> as a light execution discipline: define the economic target, instrument spend and outcomes, and move production systems forward with auditable decisions. The result is a practical path from <a href="https://mahluminnovations.com/ai-consulting-montana">AI strategy consulting</a> to <a href="https://mahluminnovations.com/ai-systems">AI automation services</a>, Cloud AI, machine learning, predictive analytics, and production deployment.</p> <p>The winners of the agentic era will not be the companies buying the cheapest models. They will be the companies engineering the lowest-cost path to a valuable completed task.</p> <p>Start with Mahlum Innovations’ <a href="https://mahluminnovations.com/ai-readiness-assessment">free AI Readiness Assessment</a> to identify where AI can produce measurable value, or <a href="https://mahluminnovations.com/contact">contact Mahlum Innovations</a> to discuss an accountable AI implementation plan.</p> <p><img src="https://cdn.marblism.com/raL8eNN_WbN.webp" alt="Vector illustration of executive AI governance, vendor selection, budget guardrails, and production deployment" style="max-width: 100%; height: auto;"></p>

About The Author's Firm

Colter Mahlum, Founder & CEO of Mahlum Innovations
Colter Mahlum — Founder & CEO, Mahlum Innovations, Bigfork, Montana

Colter wrote this article and personally leads every engagement at Mahlum Innovations. Mechanical engineer turned AI builder, based in Bigfork and Kalispell, Montana. 11 production AI systems shipped across healthcare, wellness, legal, wealth management, fitness, manufacturing, and consumer apps. Read full bio · LinkedIn.

Related Articles

Hire an AI Employee

Reading about AI? Put it to work. AxiomAI offers pre-trained AI employees for every business function — subscription-based, deployable in minutes.

Browse all AI employees →

← Back to Blog | Discuss this topic with us →