According to the external review, the large-scale testing of AI by an enterprise over the past two years was a reasonable choice, but in 2026 management was more interested in moving from “deployment of more AI” to “how to get real results while controlling costs”. According to the article, reliance on cheaper token or models alone does not solve the core problems of the Enterprise AI.
Five common waste points.
The author points out that many companies currently use AI in a way that looks like a leak pipe. From the construction of the hint to the data call, to the model selection and re-use of the results, there are multiple links that continuously consume unnecessary token and budget. Even if the manufacturers OpenAI, Anthropic, Google and others have recently reduced their prices to theken and introduced features such as hint caches, this will only relieve some of the pressure.
The article summarizes the inefficiency of Enterprise AI in five categories. The first is the expansion of the context. Overloading of information into the hint not only increases costs but may also divert model outputs from focus. In the author ' s view, it would be more efficient to organize the bottom data first so that the system can directly call on credible information, rather than ad hocly collating large amounts of original content into prompt.
The second is the lack of governance of data access. Without a clear catalogue of data, permissions and quality tags, AI agents search the data warehouse repeatedly and verify the results ex post, waste token and can easily lead to inconsistent implementation results. The third is a model mismatch. According to the article, all tasks are assigned to the strongest and most expensive front-line models, the most common and expensive habits of enterprises.
Memory and cache determine long-term costs
The fourth problem is the lack of sustained memory. In each case, the non-state agent reloads the background, reruns historical information and re-identifies the exceptions that have been processed. The article argues that if the system retains user preferences, established decision-making and to-do matters, it can reduce double counting and the single cost of interaction.
Fifthly, repeated requests were not repeated. Many business agents, when faced with high frequency, similar tasks, reread the context, reretrace and recreate the answers as they did for the first time. The authors suggest that the combination of hint caches, vector embedding and expected results make repeat requests cheaper and faster.
The competition points shift to governance capacity
The article cited, for example, the thousands of interactions that could occur every day when large enterprises used AI agents to sell or serve their customers. Token consumption may be 5-10 times higher if raw unorganized data are sent directly to the large model hint each time. Costs are rapidly accumulated with the number of agents at 10 to 15 per million token.
According to the article, the next stage of competition for AI lies not in who buys more token, nor in who uses more models, but in who can turn the AI process into a manageable and measurable business capability. The real advantages come from a cleaner data base, clearer access systems and more sophisticated modeling.
Additional information:The author is the Vice-President of Salesforce, who also refers to the Informatica Master Data Management and Salesforce Data 360 integration case to illustrate its judgement on the enterprise data base.
