NetSuite Insights & Guides | CuriousRubik

Management by Exception Requires Evidence and Response Capacity

Written by Chaitanya Tej | Aug 3, 2023, 1:00:00 PM

An organization operating by exception does not ask people to inspect every transaction. It allows defined routine work to proceed and directs attention to cases that need judgment, correction, or intervention.

The attraction is clear: scarce expertise can focus on consequential work. The risk is equally important. If “normal” is poorly defined, the organization may stop looking at cases that are wrong but produce no alert. If exceptions are badly designed, the promised reduction in work becomes an overloaded queue of ambiguous warnings.

Exception-based operations require a complete operating design: eligibility for routine processing, reliable detection, meaningful routing, enough response capacity, and independent checks on what passes silently. The decision for a COO is therefore not how many alerts to create. It is which work can safely proceed without individual review and how the business will know when that assumption no longer holds.

Define normal as an eligible operating condition

“Nothing went wrong” is too weak a definition of normal. A transaction may have missing information, an unrecognized identity, or a new business context that the rules never considered.

Define the conditions for routine processing: known work type, valid evidence, authorized action, acceptable data freshness, and a tested path to completion. The exact conditions depend on the process and consequence.

Treat unknowns separately from confirmed acceptable results. If a supplier status cannot be retrieved, the absence of a negative flag should not be interpreted as clearance. If a service event lacks an identifier, it should not disappear from monitoring because it cannot be joined to the main record.

Write the routine-processing boundary in language the business owner can approve and the technical team can test. A vague promise to handle “standard cases” leaves the important decisions to implementation assumptions.

Design exceptions around a required response

An exception should tell someone what needs attention and why. It should include the affected case, relevant evidence, consequence or deadline, and an owner who can act.

Distinguish business exceptions from technical failures and information-quality issues. A policy conflict requires a decision; an unavailable interface requires recovery; a questionable identifier requires verification. They may affect the same transaction but need different skills and response paths.

Avoid making every unusual event urgent. A warning that does not require timely action can be recorded for analysis instead of interrupting an operator. Reserve immediate notifications for cases where speed changes the outcome.

NIST’s continuous-monitoring guidance for information security emphasizes maintaining awareness to support risk decisions. Applied by analogy, exception operations should connect detection to a defined management response rather than equating more monitoring data with better control. NIST SP 800-137.

Use a complete exception loop

Working operating loop: low exception volume is not evidence of success unless the silent path remains trustworthy. Open full-size diagram

A practical working heuristic has six parts: qualify, detect, route, decide, reconcile, and learn. This is an author-proposed operating model, not an industry standard.

Qualification determines whether a case belongs in the routine path. Detection identifies a departure from that path or a failed assumption. Routing assigns the case to a capable owner with appropriate urgency.

Decision records what will happen and under which authority. Reconciliation confirms that the intended effect occurred and that related records agree. Learning examines whether repeated exceptions indicate a poor rule, an upstream defect, or a changed business need.

The loop should include an independent sample of routine cases. Without that check, the business measures only what the detection system already knows how to recognize.

A contract caterer moves from universal checking to targeted review

Consider a hypothetical company supplying meals to workplace locations. Each day, planners review recurring orders, delivery changes, and customer requests. The proposed automation would process recurring orders that meet approved conditions and refer deviations.

The business first defines routine eligibility: a current customer agreement, a recognized service location, an approved menu arrangement, a valid delivery date, and quantities within the authorized operating policy. Any food-safety or dietary requirements remain under the qualified owners’ established procedures; the automation design does not invent or relax them.

Exceptions are separated by cause. A customer-requested quantity change goes to the planning role. A delivery-location mismatch goes to customer operations. An unavailable source record creates a data-verification case. A request outside an approved service arrangement requires an authorized commercial decision.

The queue includes the decision needed and the relevant cutoff. An unresolved change before production planning has different urgency from a historical reporting discrepancy. Operators can see whether work is safely held, continuing under an approved fallback, or awaiting a decision.

The pilot includes routine recurring orders, a location change, an amendment received after the planning cutoff, and a stale agreement record. It also samples orders that passed without intervention to check whether eligibility rules missed a problem.

The company measures exception volume by cause, age, resolution effort, repeat occurrence, and errors found in the routine sample. A declining exception rate is not automatically success; it could mean better inputs, weaker detection, or missing events. The hypothetical example establishes what to test rather than claiming an operational result.

Plan response capacity before widening the routine path

An exception rate looks small until it is multiplied by volume and handling effort. A process producing many cases can overwhelm a specialist team even when most transactions pass automatically.

For a hypothetical illustration, 2,000 daily transactions with a 3 percent referral rate create 60 referrals. At an average of eight minutes of active review, that is eight hours of review work before supervision, breaks, complex cases, or absence cover. These assumptions are illustrative, not a staffing recommendation.

Variation matters. A source-system outage can create a burst far above the ordinary rate. Define surge handling, prioritization, and the conditions under which routine processing must pause because the team can no longer manage exceptions responsibly.

Do not lower detection sensitivity merely to fit an understaffed queue. Investigate avoidable referrals, improve upstream evidence, or change capacity and scope. A quieter queue is not useful if it conceals more consequential failures.

Separate thresholds from specifications

An operational threshold can express a policy limit, a capacity boundary, or a statistical signal. Those are different concepts and should not share an unexplained red line.

NIST’s statistical-process-control guidance uses control charts to assess process behavior over time. Statistical control limits are not the same as customer specifications or business approval limits. A process can be statistically stable while producing unacceptable outcomes, and an unusual signal does not by itself establish the cause. NIST/SEMATECH, Control Charts and control limits versus specifications.

For exception operations, document what each threshold means, how it was chosen, and who can change it. A limit based on a commercial policy needs policy authority. A statistical signal needs appropriate data and interpretation.

Review thresholds when the population changes. New products, customer segments, or seasonal demand can make an old rule generate too many irrelevant referrals or miss emerging problems. Preserve versions so changes can be evaluated against the right baseline.

Make resolution an accountable state change

Closing an exception should mean that a defined issue has been resolved, accepted within authority, or transferred to a responsible process. It should not simply mean that an operator has dismissed the alert.

Record the decision, evidence, actor, relevant authority, and resulting action. Keep the record proportionate to consequence. Low-risk corrections may need a concise reason; significant exceptions may require formal approval evidence.

Verify the downstream effect. If an operator corrects a customer location, confirm that the order, dispatch instruction, and relevant copies are consistent. Otherwise, the queue can appear clear while the original operational problem remains.

Handle uncertain outcomes explicitly. If an action may have succeeded before a connection failed, do not automatically repeat it. Reconcile the result or route it to someone who can establish what happened.

Sample the silent path

Exception-based management creates a visibility bias: staff see the difficult cases and may know little about the work that proceeds normally. Build routine-case assurance into the operating model.

Choose samples according to risk and learning needs. Include new work types, recently changed rules, and cases near relevant boundaries. Where statistical assurance is required, use an appropriate sampling design and qualified analysis rather than treating an informal spot check as proof.

Compare sampled outcomes with the system’s classification. Investigate false negatives, not only noisy alerts. A routine path that misses material issues needs redesign even if the exception team is performing well.

Protect privacy and access when sampling. Reviewers should see the information needed for the check, and results should be used for process improvement rather than unsupported judgments about individual employees.

Turn recurring exceptions into process improvement

A repeated exception is evidence about the design. It may indicate a missing parameter, unclear instruction, unstable source, or policy that no longer fits the work.

Give recurring causes an owner outside the daily resolution queue. Operators need to close today’s cases, while process owners address why similar cases keep arriving. Without that distinction, the organization becomes efficient at treating symptoms.

Do not automatically convert a common exception into a routine rule. Frequency does not establish acceptability. Review the consequences, authority, evidence quality, and recovery behavior before expanding the automated boundary.

Track whether a change reduces total work and preserves outcomes. A rule that removes referrals but increases downstream correction has shifted the burden rather than solved it.

What buyers should demand

Ask a supplier to demonstrate an unknown case, a burst of exceptions, a missed detection found by sampling, and a resolution that fails downstream. Request clear ownership, aging, escalation, and safe-pause behavior.

Evaluate the cost of operating the queue, not only the percentage of transactions processed automatically. The business case depends on the complexity of the remaining work and the capability required to resolve it.

Operating by exception is a shift in attention, not an abandonment of oversight. It works when routine eligibility is trustworthy, exceptions lead to useful decisions, and the organization continues to test the cases that pass silently. That combination makes reduced manual review a controlled operating choice rather than an act of faith.

Further reading