AI Proof of Concept: What It Should Prove

Building a working AI demo is now the easy part.

Asghar MirzaieSeptember 29, 2026
A polished AI demo recedes behind an evaluation spreadsheet where real test cases trigger a pre-set no-go decision.

This guide is for operators, product owners and teams deciding whether an AI idea deserves real investment. It covers what a PoC should prove, how to set pass and fail criteria, where these projects usually break, and what a sensible timeline and budget look like.

Why a Working AI Demo Proves Less Than It Seems

A demo proves the model can do the task once, on inputs someone picked. It does not show the system is reliable, affordable, safe, or something your team will actually use. Those questions sink most projects, and a curated walkthrough answers none of them.

One practitioner on Reddit described taking over a PoC that management had already called a success after six months. When the receiving team tested it, the system sent nearly 50 repetitive queries per request and took over 75 seconds to respond, against 2 to 5 seconds for the existing system. It produced wrong answers in user testing. Operating it would have cost an estimated $1.2 million a year. The phrase from that thread sums up the problem well: a proof of concept is not a proof of value.

The AppJet guide to production AI apps makes the same point from the engineering side. A working demo tells you nothing about malformed inputs, changing model behavior, failed providers or slow retrieval. We see this gap constantly at Refact, and we cover it in more depth in our piece on avoiding the demo-to-production trap. A useful AI PoC tests value and operational readiness. It asks whether one specific workflow can run reliably, safely and economically under real conditions.

The Real AI PoC Failure Rate Depends on What You Count

You have probably seen the claim that 95% of AI pilots fail. It comes from MIT’s NANDA report, and it is widely misread. That 5% figure refers to integrated pilots of task-specific enterprise tools that produced sustained, substantial financial value. It does not mean 95% of proofs of concept failed technically. The report itself calls its findings directional.

The numbers change with the definition of success:

  • Reached production at all: about 40 to 50%. Gartner reported 48%. S&P Global’s 46% scrapped rate suggests roughly half go ahead.
  • Reached wide-scale deployment: much lower. IDC research with Lenovo found only 4 of every 33 AI PoCs got there, roughly 12%.
  • Created substantial value: lower again. BCG put it at 4% of companies in 2024.

Gartner also forecast that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. BizTech Magazine’s coverage of that forecast lists the causes: poor data quality, weak risk controls, rising costs and unclear business value. Notice that “the model couldn’t do it” is not on the list. Every serious source puts the bottleneck around the model. It sits in data, workflow, ownership, cost and verification.

PoC, Pilot, Prototype or MVP: Pick the Question You Need Answered

Teams mix these stages up, and it costs them twice: once in wasted work and again in lost time. Each stage answers a different question. Passing one does not mean you have passed the next.

Stage Question it answers What passing looks like
Proof of concept Can this approach do the defined task to an agreed quality, cost and speed? Meets thresholds you set in advance, on representative data
Prototype What should this look and feel like for users? Users understand the flow and can finish the core task
Pilot Does it hold up with real users in a controlled live setting? Adoption, override rates and outcomes hold over weeks
MVP Will people use or pay for it in the market? Real usage or revenue from real customers

The MindInventory guide to AI PoCs draws similar lines between PoC, prototype, pilot and MVP. For product-side examples of what each early test proved and missed, see our breakdown of minimum viable product examples. The practical rule is simple. Don’t build a full product to answer a yes-or-no question, and don’t treat a yes to the PoC question as a yes to the pilot question.

Start From a Costly Workflow, Not From the Tool

The most common origin story for a failed PoC goes like this: an executive sees a demo, and the organization scrambles to find a use for it. One Reddit commenter saw this exact pattern at three companies. On X, a practitioner described a $40 million professional services firm that ran four AI pilots in 18 months. Each one worked in the demo. None reached daily use. The diagnosis was that every pilot started from the tool, not from the work.

The successful cases in the research share a shape. Ally Financial built call summarization with Microsoft and reported a 30% cut in post-call work, with staff reviewing every summary. C.H. Robinson reported that emailed price quotes dropped from hours to an average of 32 seconds. These are company and vendor-reported figures, so treat them as examples, not benchmarks. But the pattern holds up. The work was narrow, repetitive and high-volume. It had a measurable baseline. The AI sat inside existing work, with a human still accountable.

Our own work shows the same thing. For a daily newsletter publisher, the curator was manually checking more than 30 websites several times a day to find stories. That gave us a clear, costly task with an obvious baseline before anyone chose a tool, and it became the automated news pipeline we built. With Workform, an AI assistant for project managers, the first brief was an assistant that helped with “everything.” Our blueprint process narrowed it to one focused job: an assistant that understands projects by pulling together data from Slack, email, Asana and meetings. That narrowing was the most valuable decision in the project.

Before any build, write down three things: the task, what it costs today (minutes per case, error rate, cost per transaction), and the number that would justify scaling. If you can’t describe the use case in one sentence, the PoC isn’t ready. Our guide to AI workflow automation covers how to map the process before choosing a model.

Check the Data Before You Choose a Model

MindInventory’s guide flags data readiness as the highest-risk technical failure mode, above model selection. The research supports that. Gartner warned in February 2025 that a lack of AI-ready data puts projects at risk. The most useful evidence comes from IBM’s published report on an enterprise RAG system, the kind of assistant that answers questions from your own documents.

When IBM’s team first looked at real user questions, 40% of valid questions had no relevant content in the knowledge base. After they filled gaps, relevant content existed for 75% of questions, but search still failed to find it 47% of the time. Some wrong answers were fixed with a small edit to the source content, not a prompt change. Complex tables, long procedures and differences between how documents were written and how users asked also hurt results. This is one deployment, not a rule for all RAG systems. The lesson carries over anyway: “we have the documents” does not mean the answers are there or findable.

RAG pipeline diagram for an AI proof of concept showing retrieval and generation stages
Tracing the pipeline from document indexing to vector retrieval and language model prompting reveals how each distinct phase introduces its own potential point of failure. · Source: www.leewayhertz.com

For a RAG-style PoC, track three failure points separately:

  1. Existence: Does a source that answers this question exist at all?
  2. Retrieval: Did search find it?
  3. Generation: Did the model use it correctly?

Each one has a different fix. The first is a content problem, the second is a search and information architecture problem, and only the third is a model or prompt problem. Teams that lump them together end up tuning prompts to fix missing documents.

Also check access, privacy, freshness and whether you have ground truth to grade answers against. Use real inputs, including messy ones. One practitioner on X described a clean test set that never included a sideways scanned PDF, which exposed the system’s assumptions the moment real files arrived.

Set Pass, Fail and Kill Criteria Before Anyone Builds

A PoC without thresholds set in advance can’t fail. It just keeps going. Decide what “good enough” means before the team starts building or tuning, while nobody has a stake in the result.

You need two kinds of metrics. The MLflow guide to AI evaluation metrics separates task-quality measures like accuracy and grounding from operational ones like p95 latency, throughput, cost per request and error rates. A model that scores well offline can still fail in production because it is too slow or too expensive. For an AI PoC, cover at least:

  • Task success rate on a representative test set
  • Tolerance for serious errors, which is different from the average error rate
  • Latency at p95, not just the median
  • Cost per successful task, including retries and failures
  • Escalation behavior: does it hand off when it should?
  • The business outcome against your baseline
  • For anything that reaches users: adoption and how often people override the output

There is no universal accuracy threshold. 85% might be fine for draft summaries a person reviews and unacceptable for refund decisions made without one. The right bar depends on what an error costs and whether a human catches it.

Klarna shows why business metrics alone mislead. In early 2024 the company reported that its AI assistant handled 2.3 million conversations in its first month, about two-thirds of chats, and did the work of roughly 700 agents, cutting resolution time from 11 minutes to under 2. By May 2025, CEO Sebastian Siemiatkowski said cost had been too dominant a measure, quality had suffered, and customers should always be able to reach a human. Pair every efficiency metric with a quality metric such as repeat contacts, correct escalations or customer satisfaction.

Finally, write kill criteria. For example: “If serious errors exceed 2% on the adversarial set after two iterations, we stop or change approach.” Without these, failure turns into endless tweaking instead of a decision.

Test the Whole System, Not the Model

OpenAI’s evaluation guidance warns against “vibe-based” testing, where a team judges quality from a handful of impressive answers. NIST’s Generative AI Profile warns against drawing conclusions from narrow tests. What you need is a task-specific evaluation set drawn from real traffic, graders checked against human review, and a re-run after every meaningful change.

Test what will actually ship: prompts, retrieval, tools, permissions, integrations and fallbacks together, not a model in a playground. Include ambiguous requests, stale data, tool failures and attempts to make the system misbehave. A fluent answer is not a correct one, and your eval set should be built to catch the difference.

LLM evaluation dashboard used to test an AI proof of concept against a task-specific dataset
Granular, prompt-by-prompt metrics replace subjective impressions with measurable evaluation across custom test cases. · Source: mlflow.org

Tool design changes accuracy and cost

When Anthropic built its agent for the SWE-bench coding benchmark, the team spent more effort on its tools than on its prompts. One recurring filepath error went away once the tool required absolute paths. A community benchmark on Dev.to found something similar. When a tool returned a count, every model tested answered correctly at a flat cost. When the tool returned raw rows, models that reasoned through the list got it right but used 6 to 26 times the tokens, and models that answered straight away got as few as 0 of 21 correct. It is one unreviewed experiment, but the mechanism is easy to believe. How the model’s tools are designed decides accuracy and cost as much as which model you pick.

Prefer a workflow over an agent until the agent earns its place

Anthropic’s guidance on building agents recommends starting with the simplest approach that works. Use fixed workflows for predictable tasks, and add agents only when the flexibility is worth the extra cost, latency and risk. Research from METR explains the caution. In 2025, frontier agents succeeded almost every time on tasks that take a person under about four minutes, and less than 10% of the time on tasks over about four hours.

For a PoC, run a fixed workflow and an agent against the same eval set and compare cost and latency per successful task. If you do need an agent, the AletheionAGI agentic implementation guide covers the decisions to lock in early: how much autonomy the agent has, where state lives, and when it must stop. Our guide on building an AI agent that works goes further into scoping.

Put the Guardrails Outside the Model

“It’s only a PoC” is how teams talk themselves out of controls. The public incidents show why that’s a mistake.

In July 2025, SaaStr’s Jason Lemkin reported that a Replit coding agent deleted a production database during a code freeze. Replit’s CEO called it unacceptable and said the company was rolling out automatic separation of development and production databases. The lesson: telling a model not to do something is not access control. Give dev, staging and production separate credentials, block destructive operations by default, require approval for high-impact actions, and test that you can actually recover.

Cost needs the same hard limits. In a story recounted from a Databricks panel, a team validated an LLM process on 10,000 records, then ran it on 4 million over a weekend with no budget cap. The bill came to about $150,000. The account is secondhand, but the failure is common. Set token and spending caps from day one, limit retries, and watch the slowest and most expensive 5% of requests, not the average. As one practitioner put it, the median cost hides the p95 tail that makes the pilot unpriceable.

API usage limits dashboard setting budget caps for an AI proof of concept
Configuring explicit monthly spending caps directly within the platform’s billing dashboard prevents runaway API costs before an AI prototype ever touches production. · Source: community.openai.com

Accountability doesn’t disappear because a model wrote the answer. In February 2024, a Canadian tribunal ordered Air Canada to pay C$812.02 after its chatbot gave a customer wrong bereavement-fare guidance. The airline was held to what its bot said. In 2025, Deloitte Australia agreed to refund part of an AU$440,000 government contract after a 237-page report turned out to contain made-up citations and a fabricated court quote. If output carries legal, financial or reputational weight, a person verifies it before it goes out.

Australia’s guidance on AI proof of concept to scale puts privacy, security, oversight and performance across user groups in scope at the PoC stage, not later. Averages can hide a subgroup the system fails badly. NIST’s framework adds red-teaming, ongoing monitoring and incident response. A small, isolated experiment can justify lighter controls. It cannot justify skipping data and privacy obligations.

What a Realistic AI PoC Timeline and Budget Look Like

A focused PoC is short. HSO’s guide to AI proofs of concept describes a typical run of 4 to 8 weeks ending in a clear go or no-go decision, often using sample or synthetic data. That length works if the scope is one workflow. The caution is with the data. Synthetic and hand-picked samples can show feasibility while hiding the long tail of real inputs. Use real, representative cases wherever privacy allows, and include historical edge cases on purpose.

The PoC is also not the whole journey. Gartner’s eight-month average from prototype to production is a better guide for planning total effort. One commenter on r/LLMDevs said a week-long PoC gets you “around 70%” of the way, and the last 30% (edge cases, integration, monitoring, maintenance) takes far more effort than the first 70. Budget and staff for that from the start.

No credible, typical price range for an AI PoC exists in the research, and you should be wary of anyone quoting one before seeing your data. A realistic budget covers model usage plus data preparation, integration, evaluation, security review, and a projection of cost per request at expected volume. That projection matters most. The PoC bill is not the operating bill, and the $1.2 million estimate above was found after the PoC was declared a success.

Standard sprint estimates also mislead for AI work, because experiments can fail. Time-box the PoC around what you need to learn, not a feature list, and make the stop-or-go decision part of the plan. Price discovery separately from the full build, and be skeptical of any partner who guarantees accuracy or savings before seeing your data. Our AI integration services playbook covers how to scope that first workflow and when a good partner should tell you to stop.

What Happens After the PoC Says Yes, or No

A passing PoC earns you the right to plan a pilot. It doesn’t mean you’re ready for production. The next stage needs integration with real systems, funding, monitoring, security review and a plan for adoption. The gaps between a demo and live systems are covered in our guide to AI systems integration.

Name one owner during the PoC, not after it. Practitioners keep describing pilots that “belong to everyone and no one” and quietly stop. That owner should sit between the people building the system and the people whose work it changes. Bring the team that will run it into testing before the demo ends, so it doesn’t get a system it can’t support.

Roll out carefully, because trust is hard to win back. Early bad outputs teach users not to rely on the tool, and that wariness stays after the fixes land. Start with work where a mistake is easy to catch and cheap to fix. Ship only the parts that met your thresholds, and keep uncertain or high-impact cases with a person. Keep evaluating after launch. One r/MachineLearning practitioner reported that after 2.5 years in production, only 8% of a model’s original training data was still in use. The rest had been collected during deployment. That is an older, pre-generative AI example, but the point about upkeep still applies.

A “no” is a good result when it comes early and cheaply. Some practitioners’ PoCs showed that machine learning wasn’t needed at all and a simpler rule or statistical baseline did the job. Others showed the requirements couldn’t be met. Either finding is worth far more than an eight-month build that never ships.

An AI Proof of Concept Planning Checklist

  1. Name the job and its limits. Write what the AI may do, what it must not do, and when it hands off to a person.
  2. Measure the baseline. Record current time per case, error rate or cost per transaction.
  3. Audit the data. Check availability, access, privacy, freshness and ground truth, using real inputs.
  4. Set thresholds and kill criteria. Cover task success, serious-error tolerance, p95 latency, cost per successful task and escalation.
  5. Start simple. Test a fixed workflow before an agent, and a simple baseline before a model if that could work.
  6. Test the real system. Use representative, messy and adversarial cases, and re-run evals after every change.
  7. Enforce controls outside the model. Separate environments, set least-privilege access, add spending caps and approval gates.
  8. Project production economics. Estimate cost per request at real volume and at peak.
  9. Name the owner and next stage. Decide who runs it, who can halt it, and what the pilot needs.
  10. Decide. Go, fix or stop, based on the criteria you wrote in step 4.

Everything above comes down to one habit: decide what would count as proof before you build anything that could impress you. The model is rarely the hard part now. The work is picking the right workflow, being honest about the data, and setting a bar the demo can actually fail. If you want help pressure-testing an AI idea before committing budget to it, Refact’s AI development team starts with that scoping and discovery work, and our product discovery process is built to reach a clear go or no-go before development starts.

Building a product and unsure what to scope first? Let’s talk. Free 30-minute call, no pitch.

Share
Written by
Asghar Mirzaie
Asghar Mirzaie

Asghar Mirzaei is a backend developer at Refact, focused on the APIs, integrations, and infrastructure that power the studio’s products. His work spans data pipelines, third-party services, backend architecture, and deployment systems, helping ensure that products are stable, scalable, and ready for real-world use. Asghar works closely with the team to connect product requirements with reliable technical foundations, especially in systems where performance, automation, and integration quality matter. At Refact, he contributes to the engineering work behind the interfaces, making sure the products the studio builds can run smoothly and dependably

More from Asghar Mirzaie
Questions

Common questions.

What clients usually want to know before starting a project.

Still have a question? Ask us
Related Insights

More on Digital Product

See all Digital Product articles

Membership Platform: How to Choose One

A member’s card declines on the third of the month. Someone now has to decide whether she loses access that afternoon, after a grace week, or not at all. Most teams shopping for a membership platform haven’t made that call yet, and it shapes the product more than anything on a pricing page. A membership […]

Product Strategy Framework: The Decisions That Matter

Maya has three customer interviews, a landing page that gets some interest, and a Notion document full of feature ideas. A developer has quoted $30,000 to build version one. She still can’t answer the question behind the quote: who is this product for first, and what would prove it deserves the money? Most people in […]

Minimum Viable Product Examples, Explained

Basis spent about seven months building its software. After release, the team wrote in its January 2025 shutdown post, not one person tried it. Tract, a UK property-tech startup, raised £744,000, heard plenty of praise from design partners, and closed in March 2025 with zero revenue. Both teams knew the famous minimum viable product examples: […]

Your next step

Let's build something that lasts.

Tell us what you are trying to build, improve, or migrate. We will help you identify the right next step and what it will take to move forward.