CURIOUSRUBIK
Let’s talk about your next move ↗View complete sitemap
Back to the blog

Why AI Projects Fail Without Process and Data Readiness

An AI project can produce plausible predictions and still fail operationally when the business has no stable definition of the task, reliable evidence of outcomes, or capacity to handle exceptions. In that situation, improving the model addresses only one part of the problem. The organization may be asking software to learn a process that different teams perform and record differently.

For a customer operations director considering AI request routing, the investment decision is whether the operation is ready for a meaningful pilot. Readiness means the team can define a correct destination, supply the information available at decision time, observe what happens afterward, and act when the system cannot decide. A large historical dataset does not establish those conditions.

The practical deliverable is a readiness backlog tied to one operating outcome. This article uses a hypothetical packaging supplier to show how ambiguous case states and inconsistent historical labels can undermine an otherwise credible routing project. The recommendations concern operating readiness, not generated software or a particular model product.

Define the task before collecting examples

State exactly what the system will decide. Routing an incoming message to a department differs from identifying the owner of an entire customer case. One message may contain a delivery discrepancy, a damaged-item report, and a specification question. The business must decide whether these become separate tasks or one coordinated case.

Specify the point at which the decision occurs. A model trained on a complete case history may use information that does not exist when the first message arrives. A successful historical test then says little about the proposed live workflow. Reconstruct the information actually available at intake.

Agree on the destination’s responsibility. Does receiving the case mean acknowledging it, investigating it, authorizing a remedy, or coordinating several specialists? A routing label is useful only when the receiving team understands and accepts the associated work.

Define abstention as a valid outcome. Some messages lack enough information or span several responsibilities. The workflow needs a triage route for those cases, with an owner and service expectation. Forcing every input into a department can make the model appear decisive while transferring ambiguity to employees.

Examine the process that produced the historical data

Historical records reflect operating practice, including workarounds. A case may have reached a team because someone knew a specialist personally, because another queue was understaffed, or because a product category changed. Those destinations may be poor examples of the policy the business now wants.

Read representative cases with staff from each receiving team. Trace the original request, successive assignments, decisions, and final outcome. Ask why the case moved. Separate a correct initial route from a route that eventually found the right owner after several transfers.

Check the meaning of closure. One department may close after sending an acknowledgment; another may close only when the customer accepts a remedy. Training on a shared closed flag would mix different outcomes. The project needs a common definition or deliberately separate labels.

Preserve known uncertainty in the records. If nobody can determine whether a case was correctly routed, mark it as unresolved for evaluation rather than inventing a clean label. Excluding such records can be reasonable for a particular test, but report what population remains outside its scope.

A hypothetical packaging supplier

Imagine a hypothetical packaging supplier serving manufacturers across two regions. Its customer operations team wants to route incoming requests among delivery support, quality review, and technical specification support. The initial dataset contains several years of tickets with department and closed-status fields.

A readiness review discovers that Region North historically closes delivery tickets once a carrier investigation begins. Region South closes them when the outcome is communicated and the agreed next step is completed. The identical status label therefore describes different points in the process.

The review also finds that technical questions were routed to quality during a staffing shortage. Those records may accurately describe what happened, but they do not represent the intended future routing policy. Copying historical department labels would teach the model a temporary workaround.

One sample request reports both a late delivery and damage to several cartons. A department classifier can choose only one destination, yet the operating process requires a coordinating owner and two related tasks. The missing capability is case decomposition and ownership, not simply a more accurate department prediction.

The proposed pilot is narrowed accordingly. It handles single-purpose delivery inquiries with a verified order reference. Mixed or incomplete requests enter a staffed triage queue. The team creates a current routing guide, relabels an independently reviewed evaluation sample, and records when a receiving owner accepts responsibility.

The supplier has not proved that AI will save time. It has established a task that can be tested. If the revised routing guide and structured intake already remove most of the burden, simpler automation may be sufficient. If significant language variation remains, the narrower AI pilot has a credible operating foundation.

Hypothetical packaging-service readiness review links distinct problems to repairs. Inconsistent North/South closure meanings require current definitions and ownership. Temporary historical routing workarounds require independently reviewed evaluation labels for the intended policy. Mixed-purpose requests require bounded intake and staffed triage rather than one department prediction for two tasks. The pilot covers single-purpose delivery inquiries with a verified order reference; mixed or incomplete cases go to staffed triage. Measure whether the receiving owner accepts responsibility. These repairs make routing testable, without proving AI benefit.
Hypothetical readiness sequence. Process definitions and usable outcome evidence determine whether a routing pilot can be evaluated meaningfully.
Open full-size diagram

Check whether the data represents the intended work

Inventory the data needed for the task, including identifiers, timestamps, source channels, product context, and outcome evidence. Record what is missing and whether it can be supplied at the decision point. An enrichment available only days later cannot support an immediate intake decision without changing the workflow.

Assess coverage across relevant conditions. Include message types, languages, regions, products, and ordinary exceptions. A dataset dominated by easy inquiries may be large but poorly suited to evaluating the work that causes most operational difficulty.

Datasheets for Datasets proposes documenting why a dataset exists, how it was collected, what it contains, and its intended uses. That documentation helps assess suitability; it does not certify that the data is appropriate for a specific routing task. Gebru et al., Datasheets for Datasets, 2021 version

Identify duplicates and related cases before splitting training and evaluation data. Near-identical messages from the same incident can make a test look easier than future use. Choose a split that reflects the deployment question, and preserve a test set that is not repeatedly used to tune the system.

Establish a trustworthy evaluation answer

Write labeling instructions that a qualified reviewer can apply. Define each destination, the evidence needed, and the treatment of mixed or incomplete requests. Include examples that distinguish neighboring categories, not only obvious cases.

Have more than one reviewer assess a useful sample and investigate disagreements. The goal is not to force perfect agreement by hiding ambiguity. It is to identify where the task definition, available evidence, or reviewer training needs improvement.

Keep evaluation outcomes independent of the model’s own suggestions. If reviewers see a proposed label first, it may influence their judgment. Where practical, establish reference decisions before exposing the model output, and document how disputed cases are resolved.

Measure the consequence of error. Sending a routine inquiry to a neighboring team differs from losing a time-sensitive contractual complaint. The business owner should define which mistakes block the pilot and which can be managed through an approved correction route. A single overall accuracy number cannot make that decision.

Prepare the receiving operation

Routing is useful only if the receiving team can take ownership. Confirm queue coverage, escalation, reassignment, and the response expected when a case is misrouted. AI that sends work faster into an unstaffed queue has not improved the customer process.

Estimate the residual workload. A narrow pilot may deliberately send many cases to triage. That is acceptable if triage capacity is planned and the benefit is measured after including it. A target automation percentage should not encourage unsafe expansion just to reduce the visible exception count.

Give employees a correction path that improves the record. Reassignment should capture the relevant reason and final owner without making staff enter an essay for every ordinary transfer. Validate corrections before treating them as new training labels; an employee may reroute work for a temporary staffing reason.

Sculley and colleagues describe machine-learning system debt arising from data dependencies, feedback loops, and changing environments. For this project, that supports treating labels and receiving processes as maintained dependencies rather than one-time preparation. Hidden Technical Debt in Machine Learning Systems, 2015

Turn readiness findings into funded work

For each gap, assign an owner, a corrective action, evidence of completion, and the pilot capability it blocks. A missing routing policy belongs to operations. Inconsistent customer identifiers may belong to data stewardship. An unstaffed exception queue requires a service-management decision.

Avoid demanding perfect enterprise data before any experiment. A bounded pilot can proceed when its eligible population is clear and the relevant gaps are controlled. The important question is whether the proposed use has sufficient evidence and support, not whether every system has been modernized.

Distinguish a blocker from an improvement. The absence of a correct outcome definition blocks meaningful evaluation. A convenient additional customer attribute may improve performance but need not block a limited test. Make that judgment explicitly so the readiness program does not become an indefinite cleanup project.

At the readiness review, ask the receiving teams to process a small set of representative cases using the proposed definitions and fallback. Confirm that the organization can establish a correct result and observe it afterward. Fund the AI pilot when those conditions hold. Where they do not, repair the operating foundation first, with a concrete backlog rather than another model demonstration.

Further Reading

What’s on your mind?

A little context is all it takes to begin.

Please leave out passwords, payment details and confidential account data.