CURIOUSRUBIK
Let’s talk about your next move ↗View complete sitemap
Back to the blog

Diagnose a Broken Process Before Automating It

A return request waits in three queues before anybody decides whether the goods should come back. Automating the routing can make the request move faster through each queue while preserving the reason it was delayed: nobody has authority to resolve an ambiguous warranty claim.

The title captures a real danger, but it needs a qualification. Automation can improve a flawed process when it removes a specific defect, strengthens a control, or makes failure visible. The mistake is assuming that faster execution of the existing sequence will produce a better outcome.

Before approving automation, leaders should identify what is broken: the purpose, decision rules, information, authority, capacity, or execution. Only then can they decide whether to simplify the process, redesign a decision, repair a dependency, or automate a stable part. That diagnosis protects the investment from becoming an efficient way to reproduce an expensive problem.

Start with the failed outcome

A process is not broken merely because employees dislike its steps. Some delays reflect necessary verification, legitimate customer choice, or scarce specialist capacity. Conversely, a process can appear fast while producing incorrect results that somebody corrects later.

Define the outcome and the unacceptable failure. For returns, the outcome may be a timely, authorized resolution that protects customer service and inventory integrity. Failures could include unnecessary transport, unapproved refunds, goods received without identification, or customers waiting without a clear next step.

Establish where the process begins and ends. A returns team may measure time from receipt of a complete form, excluding the customer’s earlier attempts to obtain the form. Finance may measure refund processing after approval, excluding the unresolved claim. Both measures can be accurate and still omit most of the customer experience.

Choose boundaries that match the decision being evaluated. Preserve sub-process measures for diagnosis, but do not mistake a faster local step for an improved end-to-end result.

Reconstruct cases instead of accepting the process map

A documented workflow describes intended behavior. Select recent cases and trace what actually happened, including messages, returned requests, missing information, and manual corrections.

Include successful cases, routine failures, and unusual but consequential exceptions. Ask employees why they took each detour. A workaround may conceal poor discipline, but it may also be the only available way to handle a legitimate case the official process excludes.

Record active work separately from elapsed waiting. Identify who held the case, what they were waiting for, and whether they had authority to move it forward. “Awaiting approval” is not a diagnosis if the approver was waiting for a technical decision owned elsewhere.

GAO’s reengineering guidance considers customer needs, performance problems, organizational change, risks, and process implementation together. Its relevance here is the breadth of the inquiry; it does not prove that any particular automation method succeeds. GAO, Business Process Reengineering Assessment Guide.

Distinguish six causes of failure

A practical working heuristic is to classify the dominant defect as purpose, rule, information, authority, capacity, or execution. This is a diagnostic aid proposed here, not a formal standard.

A purpose defect occurs when an activity no longer serves a necessary outcome. A report copied to several managers may have survived after the decision it supported disappeared. Automating its production preserves waste.

A rule defect occurs when policy is ambiguous, contradictory, or badly matched to the situation. If warranty eligibility is undefined for repaired equipment, a workflow engine cannot infer the commercial answer responsibly.

An information defect occurs when the decision needs facts that arrive late, disagree, or cannot be linked to the case. An automated request for missing data can help, but only if somebody can obtain and verify it.

An authority defect occurs when responsibility and decision rights diverge. A coordinator may own turnaround time but lack permission to accept a return or obtain a specialist decision.

A capacity defect occurs when the relevant role cannot handle the arriving work within the required time. Faster intake can make the queue grow sooner.

An execution defect occurs when a sound, adequately supported process is performed inconsistently. This is often a strong automation candidate, provided exception and recovery paths remain workable.

A case may involve several defects. Identify the sequence: missing serial information may trigger extra review, which overloads a specialist, which encourages unauthorized bypasses. Treating the final bypass as the only problem leaves the earlier causes intact.

From a failed process outcome, investigate six possible causes and match the response: purpose, remove obsolete work; rule, resolve policy ambiguity; information, repair evidence; authority, assign decision rights; capacity, address the constraint; execution, automate stable work. Test outcomes and harms after the chosen intervention.
Working diagnostic heuristic: a slow process can have several causes, and only some are primarily execution problems.
Open full-size diagram

A repair business tests the reason for returns delays

Consider a hypothetical company servicing industrial pumps. Customers request warranty returns through email. A coordinator enters a case, a service manager reviews it, a finance manager approves a credit where applicable, and the warehouse receives the equipment.

The proposed automation creates the case automatically and sends sequential approvals. A case review shows that the service manager cannot determine eligibility because the equipment’s repair history is stored under a different identifier. Finance then holds the request because nobody has established whether the remedy is repair, replacement, or credit.

The business changes the process before automating it. Intake distinguishes product serial number, prior repair reference, reported fault, and requested remedy. Missing identity information enters a verification route rather than a generic approval queue. A technical owner determines eligibility within defined policy; the commercial owner approves exceptions beyond that policy.

The warehouse receives a return authorization only after the required disposition is established, except for a separately approved diagnostic-return route. Returned goods retain an identity and status that prevents them being treated as saleable inventory without the necessary checks.

Automation now has a narrower job: create the case, validate identifiers, retrieve relevant history, route unresolved identity issues, and notify the appropriate decision owner. It does not decide an ambiguous warranty obligation simply because a form has been completed.

The pilot tests a valid routine return, a missing serial number, a repaired unit with conflicting history, and equipment already shipped back without authorization. It records customer elapsed time, unnecessary approval touches, incorrect dispositions, and warehouse identification problems. The example is hypothetical; the process must demonstrate improvement rather than inherit an assumed saving from faster routing.

Remove work only after understanding its protective function

A duplicated check may be unnecessary, or it may compensate for unreliable upstream evidence. Removing it without repairing the cause can expose the business to errors that were previously caught informally.

Ask what risk each control addresses, which evidence it uses, who acts on failure, and whether another control already covers the same risk. Distinguish prevention from detection. A check after shipment may reveal an error but cannot prevent the original dispatch.

NIST’s risk-assessment guidance distinguishes threats, vulnerabilities, likelihood, and impact in an information-security context. The general discipline of identifying the actual failure mechanism and consequence is useful when reviewing automation controls; the business-process application here is an analogy, not a claim that NIST prescribes the workflow. NIST SP 800-30 Revision 1.

In the pump example, finance approval may be unnecessary for a standard repair already covered by policy, but essential for an exceptional refund. That conclusion requires the appropriate business and control owners, not a developer deciding that an approval slows the workflow.

Test the repaired process before scaling the software

A manual or lightly configured pilot can establish whether the new decision rules work before investing in broad automation. Use enough real variation to expose the important exceptions, with safeguards appropriate to the process.

Define the intervention and the comparison. If intake information, approval policy, staffing, and software all change together, the overall result may improve, but attribution to automation will remain uncertain. That may be acceptable for a business transformation; it should be disclosed in the business case.

Select measures that could reveal harm. Faster resolution should be assessed alongside incorrect outcomes, reopened cases, unapproved exceptions, and work displaced to another team. Do not declare success merely because the automated system reports fewer manual steps.

Record the reasons for failure. A small pilot can reveal a missing rule or misunderstood customer need even when it cannot support a precise benefit estimate. Use that evidence to revise the design rather than forcing a premature enterprise-wide return calculation.

Preserve a route for cases the design did not anticipate

Even a well-designed process encounters new circumstances. Automation should recognize when its assumptions are not met and refer the case with useful context.

Specify the conditions for continuing, pausing, or escalating. Give the exception owner the evidence needed to decide and authority appropriate to the consequence. A queue without a responsible decision maker merely relocates the original failure.

Include technical uncertainty. If a return authorization was created but the acknowledgment was lost, a retry should not issue a second authorization. Operators need to distinguish confirmed success, confirmed failure, and an unknown outcome that requires reconciliation.

Provide a safe way to stop the automation while preserving pending work. If the eligibility rule proves wrong, the business should be able to contain the affected cases without losing their histories or disabling unrelated services.

Recognize when a limited automation is still worthwhile

An organization may lack authority or time to redesign the entire process. A bounded intervention can still help: validating a reference, consolidating evidence, or highlighting aging cases can improve work while a wider policy issue remains unresolved.

Be explicit about the limitation. Do not claim that an automated dashboard has repaired the process if staff still cannot resolve the underlying dispute. Record the remaining dependency and the owner responsible for it.

Urgency can justify a temporary bridge, particularly when a failing legacy tool must be replaced. Give the bridge a review trigger and avoid embedding unresolved policy in hard-to-change code. A temporary solution becomes expensive when nobody remembers which compromises were temporary.

The decision before the automation decision

Ask the sponsor to produce a small number of traced cases, a clear failure diagnosis, and a proposed change that addresses the cause. Ask the process owner which rules and controls will change. Ask the technical team how exceptions, retries, and stopping will work.

A credible supplier should be willing to recommend less automation when the process needs a policy decision, better information, or additional capacity first. Buyers should value that diagnosis rather than treating every unautomated step as a missed opportunity.

The useful question is not simply whether the sequence can run faster. It is whether the organization now knows what a correct outcome requires and can reliably recognize when those conditions are absent. Automation becomes valuable when that understanding is designed into the work.

Further reading

What’s on your mind?

A little context is all it takes to begin.

Please leave out passwords, payment details and confidential account data.