CURIOUSRUBIK
Let’s talk about your next move ↗View complete sitemap
Back to the blog

The Future of Enterprise Application Development with AI

AI-assisted development changes the cost of producing a plausible implementation. It does not remove the need to decide what the application should do, prove that it behaves correctly or maintain it after release. For enterprise leaders, the important question is how the development process must change when draft code becomes easier to produce than reliable evidence about that code.

The opportunity is broader than typing assistance. Teams can explore alternatives, prepare explanations and propose tests with AI support. The risk is that a convincing draft moves through a workflow whose review capacity and acceptance criteria were designed for a slower rate of production.

A useful strategy therefore treats AI as a contributor within an accountable engineering system. The organization’s advantage will depend on the quality of its specifications, verification and operating feedback as much as on the generation tool it chooses.

Separate a generated artifact from a delivered capability

A code suggestion is an intermediate artifact. A delivered capability includes the behavior users need, integration with the existing system, security controls, tests, documentation, deployment and support. Evaluating only the time taken to produce code leaves most of that responsibility outside the measurement.

Start with a bounded development activity. Examples include drafting a well-specified transformation, explaining an unfamiliar module, proposing test cases or preparing a routine interface. Specify what a qualified person must check before the output can be used.

Avoid treating these activities as equally suitable for automation. A suggestion that helps a developer understand code has a different consequence from a change to a data-retention job. The latter can affect records irreversibly and needs stronger evidence, permissions and review.

The scope should reflect both consequence and verifiability. A small change with ambiguous expected behavior may be harder to supervise than a larger change governed by precise, independently checkable examples.

Improve specifications before accelerating implementation

AI does not resolve a requirement that the business has left contradictory. If two teams disagree about when an account may be closed, faster code generation can produce a more polished expression of the wrong interpretation.

Define the actors, permitted actions, relevant state and observable outcomes before accepting implementation. Include the exceptions that change the result, not just the straightforward example. Record where a business decision remains unresolved rather than letting the implementation invent a default.

Make context selective and explicit. A developer or tool needs the relevant interfaces, conventions and constraints, but not indiscriminate access to all organizational information. Outdated documentation should not quietly outrank the current contract or approved rule.

Version the acceptance examples with the requirement. When the rule changes, the team should be able to identify which implementation and tests depended on the old interpretation. Otherwise, generated code can accumulate assumptions that nobody knows to revisit.

A report export illustrates the assurance problem

Consider a hypothetical internal application where regional managers download service-performance reports. A team asks an AI assistant to draft a new export that includes a date range, regional filter and spreadsheet output.

A demonstration can look successful when the file downloads and contains the expected columns. Yet several questions remain: does the server restrict the manager to authorized regions? Does the export include records created during the job consistently? What happens when the selected period is empty? Can a large request exhaust the worker? Does the downloaded file contain only information permitted for that role?

The request’s visible shape does not communicate those requirements automatically. The team needs approved rules and test data that distinguish a correct export from a merely well-formed file.

For instance, the test environment can contain permitted-region and prohibited-region records with similar dates. An expected result prepared from the approved access rule should include only the permitted records. A test that simply asserts a file exists would miss the important failure.

Hypothetical AI-assisted regional report export: approved access and report rules inform both a generated implementation proposal and independently established expected-record fixtures. Expected results include permitted-region records and exclude prohibited-region records. Review authorization, record consistency, limits and failure behavior before controlled release; operating feedback follows release. A downloadable file alone is only a demonstration, and generated code does not authorize its own release.
The generated export is a proposal. Acceptance depends on independently established business and engineering conditions; a successful download alone is insufficient evidence.
Open full-size diagram

This example is not an argument that AI-generated code is necessarily worse than human-written code. It shows why the source of an implementation cannot replace evidence about its behavior.

Keep the expected answer independent

AI can propose useful tests, but tests generated from the same mistaken interpretation can reinforce the implementation’s error. The team should establish important expected outcomes independently from the code being checked.

Business owners can approve representative rules and examples. Engineers can add boundary, failure and security cases. Reviewers should understand why a test’s expected result is correct rather than relying on the fact that the test passes.

For the report export, this includes prohibited access, revoked membership, invalid ranges, empty results, interrupted execution and retry behavior. It also includes the meaning of totals and dates. A technically valid file can still summarize the wrong population.

Use multiple forms of evidence where the consequence warrants it: focused code review, automated tests, dependency checks and controlled operational observation. No single check should be described as proving the whole application safe.

NIST’s Secure Software Development Framework 1.1 describes practices that can be integrated into a software development lifecycle to reduce and respond to vulnerability risk. It does not make the origin of a code fragment an exemption from the development process. Applying those practices to AI-assisted work is an engineering interpretation, not a claim that the 2022 framework validates a particular AI tool. NIST, SSDF 1.1, February 2022.

Protect the development environment and its information

A development assistant may receive source code, documentation, logs or sample data. The organization needs a clear rule for what may be provided to which service, under what terms and with what retention and access conditions. Public availability of a tool does not authorize sending it proprietary material.

Use suitably minimized or synthetic examples where possible. Remove credentials and secrets from prompts, repositories and diagnostic material. If an accidental disclosure occurs, follow the organization’s incident process rather than assuming deletion from the visible conversation resolves the exposure.

Tool access also needs boundaries. The ability to propose a patch is different from permission to install dependencies, modify infrastructure or deploy to production. Keep consequential actions subject to the appropriate controls, even when the assistant can technically perform them.

Treat instructions found in repository content, documents or external pages as untrusted input to the workflow. They should not expand the assistant’s authority or override the task’s approved constraints. Enforce important boundaries outside the discretionary reasoning of the model.

These controls should support useful work with an approved path. If the only practical option is an unofficial account and copied production data, the organization has not created a workable adoption model.

Measure the complete development cycle

A pilot should compare equivalent work and include specification, generation, review, correction, testing and integration. Track what happens after acceptance as well: escaped defects, support demand and the effort of later changes.

Suppose a hypothetical task takes eight engineering hours without AI assistance, including implementation and verification. An assisted version takes two hours to draft, four to review and correct, and two to test and integrate. The modeled total is still eight hours. Faster drafting has not established an overall effort reduction.

If review reveals a requirement problem earlier, that can still be valuable. Describe the benefit accurately rather than forcing it into a time-saving claim. Similarly, a reduction in engineering hours is capacity, not automatically a cash saving or a shorter delivery date.

Compare tasks by complexity and context. A team familiar with an application may evaluate suggestions differently from a new hire. Repetitive well-specified work may behave differently from cross-system design. One successful demonstration should not become a universal productivity multiplier.

Watch queue effects. Producing more proposed changes can overload reviewers and delay integration. The organization should improve verification capacity and limit work in progress rather than assume that more generated output means more delivered value.

Preserve maintainability and ownership

The organization owns the application it releases, including code proposed by an assistant. Someone must be able to explain the relevant design decisions, diagnose failures and make the next change.

Ask reviewers to examine unnecessary abstractions, duplicated logic, unfamiliar dependencies and code that is harder to maintain than the problem requires. An implementation can pass current tests while creating avoidable future obligations.

Record significant AI-assisted decisions at an appropriate level. The objective is traceability of requirements, accepted changes and evidence, not an indiscriminate archive of every prompt containing sensitive information. Maintain the normal version history and review record.

Consider what happens when the tool changes or becomes unavailable. Developers should retain access to the source, tests and build process, with documentation sufficient to continue work. A generated application that can only be understood through one external assistant creates a new dependency for the business.

Build capability in stages

Begin with a limited set of approved tasks, trained reviewers and clear acceptance conditions. Observe where assistance helps and where it creates correction work. Expand scope only when the evidence supports the next use case and the required controls exist.

NIST’s AI Risk Management Framework 1.0 organizes risk management around governing, mapping, measuring and managing AI risk. Its lifecycle emphasis supports evaluating the actual use context rather than treating a tool purchase as the end of governance. It is not a productivity forecast or a guarantee of trustworthy output. NIST, AI RMF 1.0, January 2023.

Train people to challenge plausible output, identify uncertainty and escalate unresolved requirements. Review quality depends on competence and available time. Adding a human approval step without giving the reviewer either does little to establish control.

Keep the rollout reversible at the workflow level. If a use case produces excessive correction or weak evidence, narrow it while improving the supporting process. Adoption should be driven by demonstrated value and acceptable risk rather than pressure to use AI everywhere.

Prepare for cheaper drafts and more important judgment

The future pace and capability of AI development tools remain uncertain. A useful strategy does not require a confident prediction of autonomous software delivery. It requires a development system that can benefit from better assistance without losing accountability for what reaches users.

Invest in clear requirements, reusable test evidence, secure environments and observable operation. Those capabilities make current assistance more useful and provide a foundation for evaluating future tools. The lasting objective is a reliable application that the business can understand, change and support, regardless of how its first draft was produced.

Further reading

What’s on your mind?

A little context is all it takes to begin.

Please leave out passwords, payment details and confidential account data.