NetSuite Insights & Guides | CuriousRubik

A Resilient Single Source of Truth Needs More Than Replication

Written by Akshay | Oct 13, 2023, 1:00:00 PM

A single source of truth should establish authoritative meaning and controlled versions, not require every reader to depend on one running database or one specialist. The design challenge is to preserve a coherent answer while distributing the physical means of serving and recovering it.

For a data-platform owner, the immediate decision is how an approved operational dataset remains usable when the primary service fails or a defective transformation is released. Those are different failure modes. Replication can help with infrastructure availability while rapidly copying an incorrect result. Version control and recovery evidence are needed to address the second problem.

Begin by defining what is authoritative: the underlying facts, the transformation rules, the published dataset version, or an approved reporting interpretation. These may be controlled by different components. Calling one platform “the truth” without identifying those responsibilities makes both change management and incident recovery harder.

Separate authority from physical location

A dataset can have one approved definition and several read copies. The copies remain consistent in meaning when they identify the version, source boundary, transformation, and publication status they represent. The fact that data is stored in several places does not itself create several competing definitions.

Conversely, a central database can contain competing interpretations. Two teams may calculate “open order” differently inside the same warehouse. Physical consolidation has not resolved the disagreement; it has merely placed both calculations on shared infrastructure.

Create a publication contract for consequential datasets. It should identify the owning team, business definition, included population, effective cutoff, transformation version, quality checks, and permitted uses. Consumers should be able to distinguish the latest arrived data from the latest approved publication.

W3C’s Data on the Web Best Practices addresses provenance, version information, and persistent identification. These are useful design references for identifying shared enterprise datasets, although the organization’s approval and authority model is a separate responsibility. W3C Data on the Web Best Practices, Sections 8.4–8.7

Design serving copies around consumer needs

Different consumers may need different availability and freshness. A near-term dispatch decision can require a more current view than a weekly capacity discussion. A previously approved period report may need stable reproducibility rather than constant refresh.

Use read replicas, materialized views, cached snapshots, or separately hosted publications where they fit those needs. Each adds synchronization, access, and operating obligations. Do not create copies without an owner and a way to detect whether they are outdated or incomplete.

Preserve the source version in every served result. A cache should expose when its business data was effective, not merely when the page was refreshed. If a consumer combines datasets, the combined result needs a policy for compatible versions and cutoffs. Two individually valid snapshots can still produce a misleading comparison when their boundaries differ.

Limit the authority of serving copies. A read-only analytical view should not independently authorize a transaction that requires current state from the operational owner. Availability of a cached answer does not establish permission to use it for every action.

A hypothetical outage with two valid responses

Suppose a hypothetical order-backlog publication is produced at 09:45 and copied to a secondary serving environment. A new publication is created at 10:00, but its transfer is delayed. The primary service fails at 10:03. The secondary still holds the approved 09:45 version, which is eighteen minutes old.

Assume the organization has established a five-minute freshness limit for one operational use and a thirty-minute limit for a planning use. These limits are invented for the example and are not recommended service targets. The same secondary copy is too stale for the operational use but still within the planning use’s permitted window.

The correct fallback therefore differs by consumer. The operational screen should show that its freshness requirement cannot be met and route the decision through an approved alternative or pause it. The planning view can show the 09:45 publication with a clear age and limitation, provided its other acceptance criteria remain satisfied.

At 10:15, that publication reaches the planning view’s thirty-minute boundary. The design needs an explicit response if no newer approved version has arrived. A fallback that was acceptable at the beginning of an outage can become unacceptable as time passes.

This example separates service availability from fitness for use. The secondary can answer requests successfully throughout while failing the operational freshness requirement. A 100 percent HTTP-success metric would miss that distinction.

It also separates stale display data from permanent data loss. The primary source records might remain durable and replayable even while the serving layer is unavailable. Recovery objectives should distinguish rebuilding the latest publication from recovering source transactions that could otherwise be lost.

Hypothetical consumer limits. A successful response does not establish freshness, and acceptable fallback can expire during an outage. Open full-size diagram

Make recovery objectives specific

Define the acceptable interruption and data-loss exposure for each important capability, with the business owners who experience the consequences. NIST’s contingency-planning guidance discusses business impact analysis and recovery requirements in a federal information-system context. It provides a planning reference, not a universal target for enterprise analytics. NIST SP 800-34 Revision 1, Business Impact Analysis

A recovery-time objective concerns the intended restoration window. A recovery-point objective concerns the point to which data must be restored. Neither should be confused with a consumer’s acceptable age for a cached dataset. The three can be related, but they answer different questions.

Map the dependencies required to restore a usable publication: source records, transformation code, schemas, configuration, access, keys, scheduling, and quality rules. A backed-up table is not a complete recovery plan if the team cannot reproduce the transformation or obtain the permissions needed to serve it.

Test restoration into an isolated environment. Verify record populations, versions, definitions, and representative outputs before treating the recovered dataset as authoritative. The objective is a usable, explainable result, not merely a database process that starts successfully.

Understand the replication tradeoff

Replication behavior depends on the selected technology and configuration. As one concrete example, PostgreSQL 15 documentation distinguishes asynchronous and synchronous replication and their effects on waiting and potential loss around failover. These product-specific mechanisms illustrate the tradeoff; they do not establish the behavior of every managed data service. PostgreSQL 15, Log-Shipping Standby Servers

A more synchronous design can reduce some data-loss exposure while making commits depend on additional components. An asynchronous design can reduce that waiting while allowing a standby to lag. The business should understand the consequence before the architecture team selects settings.

Do not infer end-to-end consistency from database replication alone. An analytical publication may also depend on files, transformation jobs, model definitions, and external reference data. Those components need compatible recovery and version handling.

Watch shared dependencies. Two database replicas in different environments may still depend on one identity service, one metadata store, one network path, or one person with recovery knowledge. Independence must be assessed across the path that produces and serves the result.

Protect against a wrong answer that is highly available

Replication helps serve a version; validation and reproducible version bundles address whether the version is correct. Open full-size diagram

A defective transformation can publish incorrect data to every serving copy. Infrastructure redundancy will not identify the semantic error. Add quality gates and controlled publication so suspicious results can be held before they become the approved version.

Maintain a last-known approved publication where the use case permits it. If a new version fails validation, consumers should not silently receive a mixture of old data and new definitions. The fallback needs a coherent version bundle and an explicit freshness policy.

Test rollback of both data and meaning. Restoring yesterday’s rows while retaining today’s changed calculation can produce an answer that never existed in either approved version. Retain the transformation version and relevant reference-data versions needed to reproduce the publication.

Some corrections require withdrawal rather than fallback. If the prior publication is also known to be wrong, labeling it “last known good” would be misleading. The appropriate state may be unavailable pending repair, with the affected decision owners informed through the established process.

Distribute knowledge and decision rights

A platform can have redundant infrastructure and still depend on one analyst who understands its calculations. Maintain readable definitions, runbooks, test cases, and a trained alternate owner. Exercise the process with the alternate rather than assuming documentation alone establishes readiness.

Define who can approve a recovered publication, accept a temporary limitation, or suspend a consumer’s use. These decisions need business and technical input, but they should not require an improvised committee during every incident.

Test access behavior during failover as well as data availability. A secondary copy should not expose a dataset to people whose access has been revoked merely because its permission snapshot is old. Define whether authorization is checked independently, replicated with a suitable freshness rule, or causes the fallback to stop when it cannot be established. The right design depends on sensitivity and the access model, but uninterrupted service is not a reason to silently weaken permissions.

There are costs to resilience. Additional environments, retained versions, testing, and operating capacity consume resources. A low-consequence monthly dataset may not need the same architecture as a time-sensitive operational view. Select the design from the consequence of interruption and incorrectness, and document where a manual fallback is sufficient.

Begin with one widely used dataset. List its authoritative definition, publication evidence, serving copies, freshness limits, and recovery dependencies. Then simulate both a primary outage and a bad release. A single source of truth becomes resilient when every consumer can tell which coherent answer it has, whether that answer is fit for its purpose, and what happens when no suitable answer is available.

Further Reading