Stop funding AI pilots. Start running an AI portfolio. Here is the number that should end the debate: last year, forty-two percent of organizations abandoned the majority of their artificial-intelligence initiatives before they ever reached production — up from seventeen percent the year before. The abandonment rate more than doubled in twelve months. And it is about to repeat: Gartner projects more than forty percent of agentic-AI projects will be cancelled by twenty twenty-seven, for the same reasons.
This article lays out the discipline that breaks the cycle — the AI Use-Case Portfolio — and the operating rhythm that runs it: the Portfolio Test.
Pilot purgatory is a capital-allocation failure
Most banks fail here not because the models are weak, but because they fund AI the wrong way. They scatter dozens of disconnected pilots, each a little technology project on a capital-expenditure budget, with no shared thesis, no graduation criteria, and no rule for when to stop. The result has a name: pilot purgatory — capital trapped in proofs-of-concept that never ship. By some estimates more than ninety percent of pilots languish in this state rather than reaching production.
The waste is concrete. In one documented case, a pilot delivered about seven hundred and fifty dollars of value against a five hundred thousand dollar build cost. The same tool in production could have returned many times its cost — but nobody bridged the gap between sandbox and production, so the money became a total sunk loss. Multiply that across an industry and the untapped value runs into the trillions. And the time-in-pilot penalty compounds it: sixty-eight percent of technology chiefs name legacy systems as the top obstacle, with integration delays of twelve to eighteen months before a model even functions.
The point is not that banks run too many pilots, or too few. It is that they fund them without portfolio discipline. A bank would never fund a loan book without underwriting, a thesis for each loan, and a rule for writing off the bad ones. Yet that is exactly how most banks fund AI. The fix is to treat AI as an investment portfolio — with a profit-and-loss thesis at the entrance and a kill rule at the exit.
The portfolio: two axes, four verdicts
Score every AI use case as an investment, on two axes.
The first axis is Value — but only value that reaches the profit-and-loss statement. Vanity metrics are banned. "Hours saved" counts for nothing unless it maps to an audited line: lower cost, an avoided loss, or new revenue. If a tool saves ten thousand hours but headcount and vendor spend never move, the realized value is zero. Every use case is anchored to one of the three value zones — cost-out, loss-avoidance, or revenue.
The second axis is Feasibility — and here is the trap. Feasibility is almost never about how clever the model is. It is about your internal friction: whether the data is unified and decision-grade, whether the operating model can absorb the output, and whether you can buy the capability or must build it from scratch. The algorithm is rarely the constraint; the technical debt is.
Portfolio verdict = Value (to the P&L) × Feasibility (your internal friction)
Cross the two axes and every candidate lands in one of 4 quadrants.
Scale — high value, high feasibility
Fund it without constraint and push it enterprise-wide. This is the small set of use cases that actually changes the bank; the discipline of the portfolio exists to find them and feed them the capital freed from the failures.
Gated Pilot — high value, low feasibility or high risk
Release capital in tranches, only as the blockers clear. The prize is real but the friction — unified data, regulatory load, integration debt — is not yet paid down, so you buy down the risk one gate at a time rather than committing the full build.
Quick Win — low value, high feasibility
Deploy it fast on a vendor tool, build momentum, and accept the modest upside. These earn their place as proof and morale, not as the return that moves return on equity — so do not let them absorb the budget the Scale use cases need.
Kill / Park — low value, low feasibility
Decommission it before it bleeds more capital. Parking a foundational-but-early capability is legitimate; funding a dead pilot because no one wants to book the loss is not.
Three things the simple grid forgets. Payback now runs two to four years, not the old seven to twelve months, so penalize long horizons. Some low-value foundations — unified data, an identity platform — unlock everything later, so do not kill them on near-term profit alone. And heavy regulation is a feasibility penalty all its own: the moment a use case touches automated credit decisioning, its feasibility drops regardless of its theoretical value.
The Portfolio Test
Knowing the quadrants is not enough. You need an operating rhythm — the Portfolio Test, run like private-equity stage-gating.
It starts with the entry ticket: a profit-and-loss thesis. No use case gets funded without naming the exact line it will move and by how much. That single rule kills most of pilot purgatory at the door, because most pilots cannot name the line.
Then every initiative passes four gates. Gate one tests feasibility and architecture before serious money is committed. Gate two validates the controls and the evidence — including shadow-mode results, where the model runs in parallel on live data, proving itself without execution risk. Gate three confirms a safe launch. And gate four is the one that matters most: value realization. Are the financial outcomes actually showing up in production?
If they are not, the kill rule fires — automatically. The decision must be mechanical, because left to humans it never happens. Three forces keep dead pilots alive: sunk cost (after five hundred thousand dollars is spent, no one wants to realize the loss), diffuse ownership (the technology team blames the data, the business blames the technology, and inertia funds the pilot), and vanity metrics (a comfortable number that was never tied to the P&L). A mechanical rule strips all three out. Regulators are now forcing the issue: in one market the central bank wants a literal kill switch built into every automated system, so the off-ramp exists from day one.
And the portfolio is never static. Reprice it quarterly, the way you reprice credit risk. A use case that was a Gated Pilot last quarter can become a Quick Win the moment a vendor ships the capability that removes its blocker. When a pilot misses its gate, leaders should kill it without ceremony and recycle the capital into the use cases that are clearing theirs.
The proof, and the regulation that rewards it
The banks that win are not the ones with the most pilots. They are the ones with the most discipline.
JPMorgan treats AI as a managed portfolio — over four hundred fifty use cases in production, scaled because the bank invested in the connective tissue, not just the models. DBS unified its data first, then scaled more than fifteen hundred models across hundreds of use cases, cutting time-to-market from fifteen months to under three, and projecting around a billion Singapore dollars of value. A Gated Pilot looks like the banks that ran new compliance logic in shadow mode first — one validated its agent against a hundred and twenty-seven thousand transactions and reached ninety-four percent agreement with human officers before a single live decision. And the Kill verdict looks like a firm outside banking that spent three years on an AI ordering system before finally decommissioning it — three years too late, because it had no Gate Four. (No equivalent public banking kill-case is disclosed; this cross-industry example stands in.)
The regulation now rewards exactly this discipline. The revised United States model-risk guidance is materiality-based and traces data lineage — so an out-of-scope generative agent that feeds a regulated credit model pollutes the regulated space, and examiners will follow the trail to its origin. Europe classifies credit scoring as high-risk, loading it with obligations that lower its feasibility. Vietnam restricts where financial data can travel and demands human oversight and fast incident reporting. Read together, the rules are a portfolio gate: any use case touching a material, regulated decision must clear continuous oversight to qualify.
Where the series comes together
This is the capstone of the whole framework. The value axis is the Boardroom Equation and the three Value Zones — what reaches the P&L, and where. The feasibility axis is the Operating-Model Multiplier and the Build-Buy-Partner choice — whether the organization can absorb the output and source it well. And when a use case earns a Scale verdict, it deploys through the Agent Army operating model, governed by the four Stone Guardians. The portfolio is the engine that runs all of it.
So here is the Portfolio Test to run on Monday. Demand a profit-and-loss thesis for every use case. Run the four gates. Enforce a mechanical kill rule. Rebalance every quarter.
Stop funding pilots. Start managing the portfolio — and the capital you free from the failures will fund the few that change the bank.
Sources & note. Reported industry estimates, hedged as in the article: 42% of organizations abandoned most AI initiatives in 2025, up from 17% (reported, varies across sources); more than 90% of pilots reportedly never reach production; Gartner — more than 40% of agentic-AI projects projected cancelled by 2027; the ~$750-value-vs-$500,000-cost pilot case; 68% of technology chiefs naming legacy systems as the top obstacle, with 12–18-month integration delays. Discipline exemplars: JPMorgan's 450+ use cases in production; DBS's 1,500+ models with time-to-market cut from ~15 months to under 3 and ~SGD 1bn projected value; a shadow-mode compliance agent validated against 127,000 transactions at ~94% agreement with human officers; a cross-industry three-year AI-ordering kill-case (no equivalent public banking kill-case is disclosed). Regulatory references: U.S. materiality-based, data-lineage model-risk guidance; EU AI Act — credit scoring classified high-risk; Vietnam — financial-data-localization limits, human oversight, and fast incident reporting. FACT discipline: all figures are reported estimates and vary across sources.