Strategy

June 2026

·

18 min read

The Build Trap: Why Enterprises That Try to Build Their Own AI Will Pay Dearly for It

And why the companies winning with AI stopped playing software company.

JV

Joe Valeri

Co-Founder & Chief Revenue Officer, Surfaice

· MBA, MS

Originally on Substack

This article was first published on Joe Valeri's Substack. Read and share the original post.

The current AI cycle is the fastest I have ever seen. ChatGPT became the fastest-growing consumer application in history within months of launch. Within two years, it became a reference point in every enterprise software conversation on the planet.

That is the power of horizontal software at scale.

ChatGPT, Gemini, Microsoft Copilot, Claude — these are extraordinarily capable platforms designed for broad, general-purpose use. Because they are so capable, they are now being stretched into nearly every industry to solve highly specific, deeply contextual business problems.

At first glance, this appears to work. A retail real estate analyst can ask ChatGPT to summarize a lease and get a reasonable response. A construction manager can use Copilot to draft an RFI. The outputs are useful. The promise feels real.

But this is the same illusion that made early adopters believe horizontal ERP could manage retail leases. The outputs are plausible because the AI is broadly capable. They are insufficient because the AI does not understand the business context — the specific way percentage rent is calculated for this landlord, the approval chain required for this change order, the operational implications of this dispatch decision.

If you have not noticed yet, the same names are showing up on the cap tables of the horizontal AI players. OpenAI with Microsoft. Gemini at Google. Einstein GPT at Salesforce. Watson at IBM. Anthropic backed by Google and Amazon. Microsoft, Google, Salesforce, IBM — the same horizontal incumbents who lost vertical after vertical in the last software cycle are now building the horizontal AI of this one. They will win the broad market. They will lose the verticals. Again.

The numbers already show it. In 2025, enterprise companies spent $37 billion on generative AI — up from $11.5 billion in 2024. Horizontal AI copilots still dominate that spend. But vertical AI investment grew nearly three times in a single year, from $1.2 billion to $3.5 billion, and the trajectory is accelerating. Gartner projects that by 2028 more than half of all enterprise generative AI models will be domain-specific, up from less than 5% in 2023.


There is a moment that plays out in boardrooms and operations reviews across retail, real estate, and commercial property every quarter right now. Someone on the technology team runs a live demo: a conversational AI agent that can answer lease questions, pull deal data, or summarize a site report — built in-house, in a matter of weeks, using one of the powerful foundation models available via API. The room is impressed. A vice president asks the obvious question: “Why would we pay a vendor for something we can apparently build ourselves?”

It is a reasonable question. And it is almost always the wrong one.

What follows is a data-driven reckoning with what actually happens when enterprises that are not software companies decide to become software companies — specifically in the domain of agentic AI.

The Demo Is Not the Product

Modern foundation models are genuinely extraordinary. Within hours of API access, a technically curious team can build an agent that answers domain questions, retrieves documents, and holds a coherent conversation. This is not a trick — it reflects genuine progress in AI accessibility. The problem is that a compelling demo and a production-grade enterprise solution are separated by a chasm most organizations dramatically underestimate.

Gartner put a number on it: over 40% of agentic AI projects will be canceled by end of 2027 due to escalating costs, unclear business value, or inadequate risk controls. MIT's 2025 GenAI Divide report found that roughly 95% of generative AI pilots delivered zero measurable financial return. And according to recent industry analysis drawing on data across hundreds of enterprise deployments, only one in nine organizations that have AI agents “running in some capacity” actually has those agents operating in production at scale.

The Standish Group's long-running CHAOS research puts a harder edge on this: among large enterprises specifically, fewer than 9% of internally-developed technology projects are completed successfully — on time, on budget, and delivering the promised functionality. AI projects, which introduce unique complexity around data readiness, model drift, and governance, fail at roughly twice the rate of conventional IT initiatives according to RAND Corporation research.

For every 33 AI proofs-of-concept an enterprise starts, four reach production. Four.

The Talent Problem Is Not Solvable at Your Timeline

The first instinct when committing to an in-house AI build is to hire. This is where the math breaks badly.

The global AI talent market in 2025 and 2026 is not a buyer's market. ManpowerGroup's 2026 Global Talent Shortage Survey — covering 39,063 employers across 41 countries — found that AI skills are the hardest to hire for in the world, for the first time surpassing engineering, trades, and IT. Demand exceeds supply by a 3.2-to-1 ratio across key roles.

Average base salaries for senior AI engineers reached $206,000 in 2025, representing a $50,000 year-over-year increase. Specialists in large language model fine-tuning — exactly what you need to build a domain-specific agent — command 40–60% premiums above that baseline. For companies offering below a $200,000 base salary floor, the average time to fill a senior AI role is 114 days.

That means your in-house AI team is, at minimum, 12 to 18 months away from being assembled and productive — and that assumes your offers are competitive enough to attract talent that every technology company, healthcare system, and financial institution is simultaneously chasing. A specialist vertical AI vendor, by contrast, can typically begin delivery within four to eight weeks.

This is not a gap you can hire your way out of at the speed AI is moving.

The True Cost of Ownership Is Not in the Demo Budget

The cost conversations about in-house AI builds tend to focus on the initial development. This is the wrong number to watch.

A comprehensive 2025 analysis of enterprise AI implementation found that 85% of organizations misestimate AI project costs by more than 10% — with the most common failure being severe underestimation. Data preparation alone — the process of cleaning, labeling, and structuring your operational data into AI-ready formats — typically runs between $100,000 and $380,000, and can exceed the total cost of the models and tools themselves. Gartner projects that through 2026, organizations will abandon 60% of AI projects specifically because of a lack of AI-ready data.

Beyond the build, ongoing operational costs for enterprise AI include a maintenance run-rate layer — engineering upkeep, infrastructure, model monitoring, retraining as performance drifts, prompt engineering, and governance overhead. Industry benchmarks put that layer at roughly $1.0M per year for one production-grade agentic system. Xenoss and similar TCO analyses cite 15–30% of initial build cost annually for that ops layer. When organizations fail to budget for ongoing cost, budget overruns of 30–40% appear within the first year.

Our 3-year build model uses that flat $1.0M/yr operate figure for Years 2 and 3 — the same large-enterprise benchmark as our Buy vs. Build slide. It does not climb with retailer size in the early years.

There is also the LLM infrastructure cost that rarely surfaces in early business cases: token consumption at scale. An agent that handles thousands of queries per day across a large portfolio generates token costs that compound quickly and unpredictably.

The hidden cost no one puts on a spreadsheet, however, is this: every engineering hour spent building, debugging, and maintaining an AI system is an engineering hour not spent on your actual business.

Buy vs. Build

You could build it yourself. Here's what it actually costs.

A foundation-model API and a sharp engineer can produce an impressive retail-AI demo in a weekend. A production-grade agentic system — one that handles your lease clauses, permit sequences, construction timelines, and system integrations reliably, at scale, and stays current as models change — is a different animal entirely.

Build in-house

Your own agentic AI team

3–4 specialists + data, infrastructure, and ongoing operations

$3.9M

estimated total cost of ownership over 3 years

Year 1 — AI/ML talent (3–4, fully loaded)

1

$1,150,000

Year 1 — Data preparation

2

$380,000

Year 1 — Infrastructure & MLOps setup

$150,000

Year 1 — LLM token consumption

$120,000

Year 1 — Recruiting & onboarding

3

$100,000

Year 1 subtotal — Build

$1,900,000

Year 2 — Operate & maintain

4

$1,000,000

Year 3 — Operate & maintain

4

$1,000,000

3-year TCO

$3,900,000

Years 2–3 use a flat $1.0M/yr operate & maintain run-rate for a production-grade agentic system — the same large-enterprise benchmark as our Buy vs. Build slide. It does not scale up with retailer size in the early years.

1 Senior AI engineer base salaries reached ~$206K in 2025 (LLM fine-tuning specialists command a 40–60% premium); 12–18 months to assemble a productive team — Signify Technology; ManpowerGroup 2026 Global Talent Shortage Survey. 2 Data preparation typically runs $100K–$380K — Xenoss / Riseup Labs. 3 114-day average time-to-fill for senior AI roles below a $200K base — Signify Technology. 4 Years 2–3: flat $1,000,000/yr operate & maintain run-rate (engineering upkeep, cloud/GPU, tokens, integrations, security, MLOps, QA). Xenoss cites 15–30% of build cost annually for the ops layer; this $1.0M figure is the illustrative large-enterprise maintenance benchmark and does not grow with company size in early years — BenchmarkIT · DataRobot · EY · McKinsey practitioner data.

Buy Surfaice

Purpose-built for retail, live in weeks

A predictable annual subscription — no team to hire, no infrastructure to run

Weeks

to production — not the 12–18 months a build requires1

  • 150+ production-tested skills, ready on day one
  • A knowledge base built by practitioners with 100+ years of retail store-lifecycle experience
  • Token economics optimized across the whole retail market — not one cold-start enterprise
  • Connectors that already know Lucernex, MRI, CoStar, Procore & JLL
  • Improves with every deployment across the industry — your build only learns from you
  • Zero ML headcount, zero infrastructure babysitting, zero model-drift firefighting

The compounding gap. A home-grown build accumulates knowledge only from your own operations — and walks out the door when your best people do. Surfaice gets smarter for every client, every deployment.

2.3× ROI

40%+

of agentic AI projects will be canceled by end of 2027⁵

<9%

of large-enterprise in-house tech projects finish on time, on budget, in scope⁶

95%

of generative-AI pilots deliver zero measurable financial return⁷

76%

of enterprises now buy AI rather than build — up from 47% a year earlier⁸

See Surfaice pricing

Book a demo

The Domain Expertise Gap Is Enormous and It Does Not Close

There is a specific belief driving many in-house AI builds in operational industries: “Our use cases are so specific that only we can understand them well enough to build the right solution.” This sounds like a competitive insight. The data says it is a rationalization.

Domain-specific AI solutions trained on years of industry data, workflows, and terminology outperform general-purpose models by a measurable margin — and the gap is significant. In benchmark studies comparing vertical AI agents to general-purpose LLMs on industry-specific workflows, generic outputs perform at roughly half the accuracy of vertical agents.

What takes a vertical AI vendor years to build is not primarily the model. It is the corpus. The training data. The edge cases and exceptions baked in from thousands of real operational decisions made across the industry. You can pick up a foundation model API this afternoon. You cannot pick up that expertise, and you cannot manufacture it quickly.

The same pattern played out with ERP. The vendors who knew the domain won the vertical, every time.

The Market Has Already Voted

According to Menlo Ventures' 2025 State of Generative AI in the Enterprise report — based on a survey of 495 U.S. enterprise AI decision-makers — the shift has been dramatic: in 2024, 47% of AI solutions were built internally; by 2025, 76% are purchased rather than built. That is a near-inversion in a single year.

Enterprise investment in vertical, industry-specific AI solutions reached $3.5 billion in 2025, nearly three times the $1.2 billion invested in 2024. Companies deploying vertical AI solutions report 2.3× higher average ROI than those deploying only general-purpose LLMs.

The enterprises that tried to build have, in large numbers, tried and revised that decision. The organizations winning with AI are not the ones who became software companies. They are the ones who found the right software partner and focused their energy on using the technology to do their actual jobs better.

What This Means for Operations Leaders

If you lead a real estate, retail, or operational function, the question in front of you is not whether AI will transform your industry. It will. The question is whether you will capture that transformation or spend the next eighteen months — and several million dollars — discovering why most enterprises don't.

The honest framing of the in-house build argument is this: your team can build something that demonstrates the concept. They probably already have. What they cannot build, at any reasonable cost or timeline, is a production-grade solution that handles the full complexity of your domain, stays current as models and APIs evolve, operates reliably at scale, and continues to improve with each new deployment across your industry.

That last part matters most. Every time a vertical AI vendor deploys a solution, they learn something that makes the system better for every client. Your in-house build accumulates knowledge only from your own operations. The compounding curve bends against you over time, not in your favor.

What This Looks Like in Practice: The Retail Store Lifecycle

To make this concrete, consider one of the most complex operational domains in retail: store development. A single store opening touches real estate site qualification, lease negotiation, permitting, construction scheduling, fixture procurement, IT provisioning, and grand opening sequencing — across dozens of interdependencies, jurisdictions, and counterparties.

Surfaice was built specifically for this problem — and the way it was built illustrates everything the previous sections describe.

The platform launched with a knowledge base trained by industry practitioners representing, collectively, over 100 years of retail store lifecycle experience. Surfaice's library of over 150 trained skills with refined, production-tested prompts is the operational output of that investment.

Three architectural decisions in Surfaice's design directly address the failure modes that sink in-house builds:

  • The collective knowledge problem. Surfaice's architecture inverts the single-tenant trap: every client inherits the collective knowledge base of the platform on day one, then builds their own layer on top.
  • The token economics problem. Surfaice minimizes token use through stored intelligence and benefits from cost efficiencies of serving the entire retail market rather than a single enterprise.
  • The connector intelligence problem. Connectors learn from experience across Lucernex, MRI, CoStar, Procore, JLL platforms, and the specific combinations retail customers actually run.

The result is a platform that a retail operations team can deploy and use meaningfully within weeks — without a team of ML engineers, without an 18-month data preparation project, and without diverting their most experienced people away from running stores to babysit an AI infrastructure build.

Further reading

These arguments are explored in greater depth in the forthcoming book The Startup Consigliere: From SaaS to Agentic AI: A stage-by-stage playbook for the Founder who wants to Build the Next Category.

Sources and further reading: Menlo Ventures 2025 State of Generative AI in the Enterprise; Gartner agentic AI and domain-specific model research; MIT GenAI Divide report; Standish Group CHAOS research; Xenoss and Riseup Labs TCO analyses; ManpowerGroup / Signify Technology talent market data.