CURIOUSRUBIK
Let’s talk about your next move ↗View complete sitemap
Back to the blog

NetSuite Scheduled Script and Map Reduce Failure Triage

Triage a failed NetSuite background script by separating execution status from business completion. Preserve the task and deployment evidence, identify which inputs produced durable results and determine the safe recovery action for each remaining item. Restarting the whole job before reconciling its effects can create duplicates or repeat an external action.

Scheduled scripts and Map/Reduce scripts have different recovery behavior. Map/Reduce can restart work after certain interruptions, but that does not automatically undo records the script already created. SuiteScript 2.x scheduled scripts do not provide the legacy recovery-point and yielding APIs. The runbook must match the actual script type and version.

Capture the run before resubmitting it

Record the account, script, deployment, task or execution reference, submission time and timezone. Preserve the version of the code and important non-secret parameters. A deployment can run several different populations over time, so its name alone does not identify the failed work.

Collect the current status, execution errors, stage information and last useful application log. For a Map/Reduce run, inspect the input, map, reduce and summary evidence relevant to the failure. If the script catches errors internally, confirm whether those errors reach the summary or only a custom log.

Record the business deadline and what is blocked. A failed overnight classification update may be deferrable; a job that partially created billing documents may require immediate containment and finance reconciliation. The same technical error can have very different consequences.

Prevent an automatic resubmission from overlapping with manual recovery where the approved operational design allows it. Do not stop unrelated processing merely because it shares the same broad application name.

Locate the failure boundary

Ask whether the run failed before selecting input, while processing a particular item, after saving a result or during final reporting. Each boundary changes what must be checked before replay.

A getInputData error can leave the run without the intended input population. A map or reduce error may affect one key while other keys continue. A summary-stage problem may leave business work completed even though the completion notification failed.

Separate platform interruption from an uncaught application error. Map/Reduce responds differently according to stage and configuration. For example, an uncaught input-stage error proceeds to summary rather than simply rerunning input, while map and reduce retry behavior depends on the configured options.

Avoid interpreting a final task state as a complete transaction ledger. A job can finish with individual failures, and a failed notification does not mean all record writes failed. Use application-level evidence to classify the affected inputs.

Understand what restart protection does and does not cover

The Map/Reduce framework manages its internal key-value processing during restarts. Durable business records and external side effects require protection in the script's own design. An order created before an interruption may still exist even if the processing key is retried.

The isRestarted flag indicates that a stage or invocation is part of restarted work. It does not prove that this particular business effect already happened. The job may have been interrupted before reaching that operation, so a blanket skip can lose work just as a blanket retry can duplicate it.

Use a stable business identity and a verifiable result state. Recovery should determine whether the intended output already exists and is correct, whether it is absent, or whether the result is uncertain. The details depend on the operation and any external system involved.

Review Buffer Size carefully. Increasing it can enlarge the group of keys exposed to reprocessing after an interruption. Keep performance tuning separate from the incident unless evidence shows that a setting is causal and the proposed change is tested.

Classify each input before choosing an action

Create a recovery ledger at the grain of the business operation. It should contain the input identifier, intended effect, observed target reference, evidence, classification and approved next action.

Useful classifications are:

  • Completed and verified: retain the result and do not repeat its effect
  • Failed before the effect: correct the cause and consider a bounded retry
  • Effect uncertain: investigate target and external-system state before replay
  • Completed incorrectly: obtain approval for correction rather than recreating blindly
  • Not yet attempted: process through the agreed remaining-work route

Do not use a record count alone to classify completion. Two outputs for one input and no output for another can produce the expected total while still being wrong. Reconcile identifiers and relevant amounts or quantities.

If an external request timed out, check the receiving system using the original event identity. A lost response is not evidence that the receiver did nothing.

Hypothetical example of an interrupted billing batch

A fictional batch has 100 approved input requests. Investigation confirms 72 correct invoices, eight inputs that failed validation before creation and 15 not yet attempted. Five inputs have uncertain outcomes. The populations reconcile because 72 + 8 + 15 + 5 = 100.

Target-state checks show that three of the five uncertain inputs already created correct invoices. Two have no invoice and no pending external operation. The ledger now contains 75 verified completions and 25 remaining inputs: eight needing correction, 15 unattempted and two cleared for retry.

The team repairs the eight input issues and processes only the approved 25-item recovery population. It verifies one correct invoice per input and reconciles financial totals with the business owner. Replaying all 100 would have risked repeating the 75 effects already completed.

This example describes a recovery design, not a claim that every billing process supports the same identity lookup or replay mechanism. Where evidence remains uncertain, the corresponding input stays on hold.

Handle scheduled scripts on their own terms

For a SuiteScript 2.x scheduled script, do not prescribe a nonexistent direct yield or recovery-point method. Review whether the implementation uses a deliberate durable checkpoint and supported resubmission design, or whether it should be redesigned as Map/Reduce for suitably separable work.

A checkpoint must correspond to verified completion. Advancing it before the business effect is durable can skip work after failure. Updating it only after the effect creates an interruption window that also needs duplicate protection. The design must handle both sides of that boundary.

For legacy SuiteScript 1.0 code, inspect its actual recovery logic rather than assuming the 2.x behavior applies. A migration or rewrite is a separate engineering change requiring testing, not an automatic incident response.

Do not increase a batch size or repeatedly submit additional deployments to overcome an unexplained backlog. That can create overlap, contention or more uncertain outcomes without addressing the original failure.

Prove recovery with business evidence

After the recovery run, reconcile the original input population with completed results and explicitly excluded items. Record any corrections, unresolved exceptions and external acknowledgments. A summary message saying “processed 25” is useful only when the identifiers and effects support it.

Check downstream consequences. An invoice may exist but remain unapproved; a fulfillment may have been created without its intended notification; a custom record may have been updated while a related export failed. Define completion at the business boundary the script was meant to serve.

Review notifications and operational ownership. The person able to diagnose the job should receive actionable errors with the relevant run identifier. Avoid placing sensitive payloads or credentials in broadly accessible execution logs.

Improve the runbook from the incident

Document the confirmed cause separately from contributing conditions. Add the exact diagnostic evidence and the approved recovery population definition. Retain the failure case as a regression scenario, including interruption after a durable effect and before completion tracking.

Consider stronger per-item result logging, smaller coherent work units or clearer summary reporting where the investigation exposed gaps. Those improvements should reduce ambiguity rather than create a second unreliable ledger.

For a difficult recovery, CuriousRubik's NetSuite support services can help connect execution evidence with business reconciliation. The immediate goal is a known final state for every input, followed by a tested design that handles the same interruption safely next time.

Frequently asked questions

Does a failed Map/Reduce task mean no records were saved?

No. Some business effects may already be durable. Inspect per-input target state and external acknowledgments before retrying any operation.

Does isRestarted mean the current item should always be skipped?

No. It indicates restarted work, not certainty that the item's business effect completed. Use stable identity and result checks to determine whether to continue, retry or investigate.

Can a SuiteScript 2.x scheduled script use legacy yield APIs?

No. It lacks the legacy direct yielding and recovery-point methods. Use an appropriate tested checkpoint/resubmission design or consider Map/Reduce when the work can be separated suitably.

Should every error be retried automatically?

No. Validation errors need correction, uncertain effects need reconciliation and incorrect completed effects may need approved remediation. Retry only the population whose recovery conditions are understood.

What proves a recovery run succeeded?

Reconcile every original input to its intended durable result or an explicit unresolved exception. Verify identifiers and business amounts or quantities, including downstream steps that define completion.

What’s on your mind?

A little context is all it takes to begin.

Please leave out passwords, payment details and confidential account data.