NetSuite Insights & Guides | CuriousRubik

Outgrowing Business Software: Diagnose the Real Constraint

Written by Krishna | Jun 7, 2023, 1:00:00 PM

A business can double revenue without outgrowing its software. It can also add one new sales channel and discover that a previously adequate stack no longer supports its operating model. Revenue is an unreliable replacement trigger because the harder problem is often a change in complexity, coordination or control.

Growth exposes assumptions embedded in earlier technology decisions. A system may assume one warehouse, one currency, one person approving payments or a product sold only as a single item. Those assumptions can remain invisible while the business is small. When the model changes, employees compensate with extra files, manual checks and delayed updates.

The decision for leadership is not whether the stack looks old. It is whether the company can support its next stage of growth reliably and economically, and whether targeted improvement can restore that fit. Replacing everything too early can consume the resources needed for growth. Waiting until workarounds become critical infrastructure can make replacement harder.

Separate four different ceilings

The first ceiling is transaction capacity. Systems or integrations cannot process the required work within an acceptable time. Symptoms include growing queues, slow batch jobs and failures during peak periods. Infrastructure tuning, scheduling changes or a better integration design may solve the problem without replacing the enterprise platform.

The second ceiling is business-model fit. The software cannot represent a new commercial relationship or operating structure coherently. Subscription amendments, project billing, shared inventory and multiple legal entities create different demands. Adding more processing power cannot resolve a data model that lacks the required concept.

The third ceiling is coordination. Individual applications perform adequately, but information crosses their boundaries too slowly or ambiguously. Teams cannot establish whether an order is approved, available, shipped or billable without asking several people. The remedy may involve ownership and integration rather than new core applications.

The fourth ceiling is control. Access, approvals, audit trails or recovery arrangements depend on personal knowledge and informal supervision. Growth increases the number of people, exceptions and places where something can go wrong. A familiar workaround may no longer provide acceptable oversight.

These categories are a working diagnostic heuristic. They are not a maturity certification, and a company can encounter several simultaneously. Their value is preventing a technical capacity complaint from being used to justify an unrelated platform purchase.

Working diagnostic heuristic. Several constraints may coexist; growth alone does not determine the intervention. Open full-size diagram

Measure the constraint in its own units

For transaction capacity, measure backlog age, completion time, failure rates and resource consumption under realistic demand. Average daily volume can conceal a short peak that determines customer experience. A monthly batch that finishes overnight may become unacceptable when management needs hourly availability information.

Google’s Site Reliability Engineering chapter on handling overload explains why request counts alone can be a poor capacity measure: different requests consume different resources, and retries can amplify load. Those observations come from large-scale serving systems. The practical inference for enterprise applications is to test realistic work mixes and failure behavior rather than extrapolate from a simple record count.

For business-model fit, count the critical scenarios that require data to be represented outside the system. Examine the consequences rather than the number of spreadsheets. A planning spreadsheet can be a deliberate analytical tool. A spreadsheet that alone determines contractual billing may indicate a much more consequential gap.

For coordination, measure the time between a business event and the next team’s ability to act on it. For control, test whether an unfamiliar but authorized employee can execute the process correctly and whether someone else can reconstruct what happened. A process that works only when its longest-serving employee is available has a people dependency that growth can magnify.

A hypothetical wholesaler identifies the wrong bottleneck

Consider a fictional wholesaler preparing to add an online channel and a third warehouse. Management believes its core application is too slow and proposes a full replacement. The following numbers are illustrative.

The company receives an average of 900 orders per day, but a promotional campaign can generate 600 orders in one hour. A connector submits orders to the core application in five-minute batches. Each batch triggers individual stock queries, and failed requests are retried without a controlled limit. During a promotion, the connector’s queue grows even when the core application’s ordinary users see acceptable response times.

A test separates new orders from retries and measures processing time by order type. It shows that orders with many lines consume considerably more resources. A revised connector combines appropriate lookups, limits concurrent requests and applies a bounded retry policy. Stable order identifiers allow the receiver to recognize repeated submissions without creating duplicate orders; when a timeout leaves the outcome unknown, the team checks the recorded result before resubmitting unsafely. The team also defines what customers see while an order awaits confirmation. This addresses a transaction-capacity problem without first replacing the core system.

The same investigation uncovers a different limitation. The current system records stock by warehouse but cannot apply the company’s proposed allocation policy across channels. Sales staff reserve goods in a spreadsheet, while the online shop presents a separate availability figure. Better connector performance could make both channels consume the same stock more quickly without resolving the conflict.

Leadership therefore separates the decisions. It funds the connector repair immediately, while testing whether supported allocation functionality can represent the intended policy. The commercial team defines which orders have priority, how long reservations last and what happens when a customer changes an order. If the platform cannot support that model, replacement or a specialized allocation service becomes a legitimate option.

The two findings produce different investments and different acceptance tests. One concerns throughput under peak demand. The other concerns the correctness of promises made to customers. Calling both “the system is slow” would have hidden the distinction.

Illustrative wholesaler scenario, not a benchmark. A throughput repair does not resolve business-model fit. Open full-size diagram

Look for the rising cost of adaptation

A stack may remain operational while becoming increasingly expensive to change. A small pricing adjustment requires coordinated edits across several tools. Adding a warehouse creates another set of spreadsheets. A release is postponed because only one person understands the integrations.

Track a few representative changes over time. How long does it take to introduce a product, add a reporting entity or change an approval rule? How much testing is required? Which teams must coordinate? What breaks after the change? Use comparable examples; a complex acquisition should not be compared with a simple screen adjustment.

Also examine the proportion of change effort spent protecting workarounds rather than improving capability. There is no universal percentage at which replacement becomes mandatory. The signal is a deteriorating relationship between business ambition and the effort required to deliver it.

Include the cost of delay. If a new channel waits for months because its transactions cannot be represented safely, the technology decision has a commercial consequence. Quantify that consequence with the business owner, using ranges and explicit assumptions. Avoid attributing all projected channel revenue to a system change when demand, pricing and operational execution remain uncertain.

Test resilience before adding more dependence

Growth often concentrates more activity on components that were never designed as critical services. An integration server initially used for a daily export may eventually coordinate orders, inventory and billing. Its operational importance has changed even if its technical architecture has not.

Ask what happens when each critical dependency is unavailable. Can orders be accepted safely? Can staff see which transactions are pending? Can processing resume without duplication? How much information can the business afford to lose, and how long can it tolerate interruption?

NIST’s Contingency Planning Guide for Federal Information Systems, issued in 2010, links contingency planning to operational priorities and organizational resilience. Although written for federal systems, it provides a useful basis for asking business-led recovery questions. It does not establish a universal recovery target for a commercial application.

Recovery requirements should influence the upgrade path. A platform with sufficient transaction capacity can still be unfit for growth if restoring it depends on an untested backup or a single employee. Conversely, improving recovery and support may extend its useful life more effectively than buying additional features.

Compare three credible paths

The first path is to stabilize the existing stack. Remove avoidable bottlenecks, clarify ownership, improve monitoring and address critical recovery gaps. This is appropriate when the underlying business model fits and the main problems are implementation or operating discipline.

The second is to extend selectively. Add a specialized capability where the business needs a distinct function and the boundary can be kept clear. Evaluate the cost of integration, support and data consistency. A narrow extension can preserve a sound core; a web of extensions can postpone a necessary redesign while increasing its eventual difficulty.

The third is to replace a major platform or redesign the architecture. This becomes more persuasive when essential business concepts cannot be represented, control gaps cannot be closed credibly, or the cumulative cost of adaptation exceeds the value of preserving the current arrangement.

Compare all three against the same growth scenarios and time horizon. Include transition risk, parallel operation, retraining and the cost of retaining historical information. The replacement proposal should explain which constraints it removes and which organizational problems will still need separate attention.

Build a trigger-based plan

A roadmap tied only to a distant implementation date can drift. A trigger-based plan defines observable conditions that require action. Examples include a critical queue exceeding its business tolerance, a new entity requiring unsupported accounting structures, or a recovery test failing to meet an agreed operational requirement.

For each trigger, assign an owner and an available response. Some trigger a capacity adjustment; others trigger a selection process or a restriction on further growth. Set thresholds using business consequences and measured evidence rather than adopting arbitrary industry numbers.

Maintain a short decision file containing current constraints, attempted fixes, test results and unresolved risks. This prevents every budget cycle from restarting the same debate. Review it when the business model changes, not merely when a renewal invoice arrives.

There are limits to this approach. Forecasts can be wrong, and a platform selected for one expansion may become inappropriate after an acquisition or strategy change. Preserve options where possible: maintain exportable data, documented interfaces and a credible exit path. Flexibility has value, but it should be weighed against the extra complexity it can introduce.

Outgrowing a stack is ultimately a loss of fit between the way the business operates and the way its systems represent, coordinate and control that work. Diagnose that loss precisely. The right response may be a modest repair, a carefully bounded extension or a substantial replacement. Growth alone does not choose among them.

Further reading