MIT researchers who tracked 300 real AI deployments found 95% show zero measurable financial return, and the gap wasn't about model quality or regulation. **Every case study that did deliver a hard number automated one specific, well-defined process instead of trying to run a whole department — and most bought that system from a specialized partner rather than building it in-house** (MIT NANDA, 2025). That pattern holds from a jewelry brand with thousands of employees down to a two-person firm in Brooklyn.
Why Do 95% of AI Projects Show Zero Return?
MIT Project NANDA's "State of AI in Business 2025" report is based on a systematic review of over 300 publicly disclosed AI initiatives, 52 structured interviews, and 153 survey responses from senior leaders collected across four industry conferences, fielded January through June 2025 (MIT NANDA, 2025). Despite an estimated $30-40 billion in enterprise GenAI investment, the researchers found 95% of organizations saw zero measurable P&L impact. Just 5% of integrated pilots were extracting millions in value. The report calls this the "GenAI Divide," and it explicitly rules out the two explanations most people reach for first: "this divide does not seem to be driven by model quality or regulation, but seems to be determined by approach" (MIT NANDA, 2025).
Is It Better to Buy an AI System From a Vendor or Build It Yourself?
The same report found external partnerships with learning-capable, customized tools reached deployment about 67% of the time, compared to about 33% for tools built internally (MIT NANDA, 2025). The researchers are upfront that this is a self-reported figure from a 52-organization interview sample, not a controlled study — companies that choose an outside partner may already have more procurement sophistication or risk tolerance than ones that build in-house, so the gap may partly reflect who chooses which path, not just the path itself. Even with that caveat, the direction was consistent enough across interviewees that the report treats it as a real signal, not noise: a narrow, specific tool built by people who've built it before beats a general one built by a team doing it for the first time, most of the time.
What Do Real, Named AI Case Studies Actually Show?
Salesforce's own Agentic Enterprise Index, published August 7, 2026 from aggregated Agentforce usage data, and its customer-story archive both put names and numbers behind the MIT pattern — and it's worth saying plainly that a vendor publishing case studies about its own product is not a neutral source, the same way any of the numbers below should be read as the vendor's most favorable examples, not an average outcome (Salesforce, 2026). What's useful is the shape they share, not the specific numbers.
| Business | Size | What got automated | Published result |
|---|---|---|---|
| SkySync | 2-person firm, Brooklyn | Lead capture from inbound inquiries | 3 days to 5 minutes; leads grew 24x |
| 360 MMS | Small mortgage brokerage, 4 brands | Marketing campaign builds + commission tracking | 16x faster campaign build time |
| Pandora | Global jewelry brand | Routine customer inquiries during peak seasons | 60% of routine requests handled; NPS up 10% |
| Siemens | Enterprise, 18,000 sellers | Triage of 2,800 unqualified inbound leads/week | Capability described; no headline metric published |
| PenFed | Federal credit union | Account balances, loan status, fund transfers | Capability described; no headline metric published |
Every single one is scoped to one job: capture this lead, answer this routine question, triage this queue, look up this balance. None of them describe an AI system that runs a department end to end, even at the enterprise end of the list (Salesforce customer stories, 2026).
Why Don't All AI Case Studies Publish a Real Number?
Notice the gap in that table: Pandora, SkySync, and 360 MMS each publish a specific before/after metric. Siemens and PenFed describe what the AI does in real detail but stop short of a headline result — no percentage, no dollar figure, no time saved (Salesforce customer stories, 2026). That's not automatically a red flag; PenFed operates under federal credit-union compliance rules that make some metrics harder to disclose publicly. But it's a genuinely useful filter for reading any AI case study, from any vendor: a mechanism description ("the AI evaluates X and does Y") is not the same claim as a measured result ("X went from A to B"), and a page that only offers the first one hasn't actually told you whether the project worked.
When Does Building an AI Tool Yourself Make More Sense Than Buying One?
The 67%-versus-33% gap isn't a rule against ever building internally — it's an average across a mixed sample, and MIT's own researchers flag that it may partly reflect who chooses each path. Building in-house is the better call when the workflow is genuinely unique to the business and not something a specialist vendor has built a dozen times before, when there's already technical staff who'll own and maintain it, or when the project is a low-stakes internal experiment where a wrong answer costs an afternoon, not a customer. None of the five case studies above fit that description — they're all customer-facing or revenue-facing processes handled by a purpose-built partner, which is exactly where the report's own data says the odds favor buying, not building.
What Does This Mean for a Small Montana Business Picking Its First AI Project?
A five-person shop in Kalispell doesn't need Siemens' 18,000-seller sales org or PenFed's compliance team to apply the same lesson. SkySync is the closer comparison: a two-person firm that picked one process — lead capture — and got a specific, measured result from it, rather than trying to wire AI into every part of the business at once. The temptation for a small operator moving fast is to buy one tool for the phones, another for scheduling, another for follow-up texts, and hope they add up. The pattern in the data says the opposite works better: pick the single process bleeding the most value right now — for most Northwest service businesses, that's the phone ringing with nobody free to answer it — and get that one thing built right by someone who's built it before.