New Directions in Software Technology (NDIST)

Prevent,Contain,Prove

A Framework for Urgently Securing Critical Digital Systems

Strategic Guidance for an Era of AI-Accelerated Offense

The full report will be available here on Tuesday, September 23.

In brief

For decades, the cost of expert labor set a practical limit on software attack: finding an important vulnerability took weeks or months of specialist work, and turning it into a reliable exploit took more. Frontier AI is removing that limit.

The shift is measurable. In 2025, autonomous systems in DARPA's AI Cyber Challenge found 86 percent of the vulnerabilities planted in critical-infrastructure software.⁠[1] By spring 2026, frontier models were surfacing thousands of previously unknown vulnerabilities across every major operating system and browser — more than ten thousand high- or critical-severity findings from some fifty organizations in a single month, with independent security firms confirming over 90 percent of a sampled set as real.⁠[2] Today's models are better at finding flaws than at exploiting them end to end, but their makers caution that the gap may narrow.⁠[3] In July 2026, those capabilities appeared in the wild.⁠[4] Capabilities that once required a nation-state are reaching criminal groups and individuals, and on the current trajectory, defensive adoption will fall well behind offensive automation: the window is defined by lead time, not by any present limit on what attackers can do.⁠[5]

The answer is not simply to patch faster.⁠[6] This report argues for changing how consequential software is built, through high-assurance engineering: memory-safe languages and hardware, compartmentalized architectures, and formal methods — a body of practice, adopted in proportion to consequence, that prevents important categories of failure by construction, contains them by architecture, or rules them out by rigorous analysis. None of this is theoretical: Apple, Amazon, Google, Intel, and Microsoft run verification programs today;⁠[7] after Android moved new code to memory-safe languages, the share of its vulnerabilities attributable to memory safety fell from 76 percent to under 20;⁠[8] and a growing cohort of venture-backed startups is driving the frontier of automated mathematical reasoning.⁠[9]

The challenge is a limited pool of experts who can lead industrial-scale high-assurance engineering. The deployments above were mostly built by unusually expert teams using tools that remain difficult by industry standards; they show that the guarantees are real, and leave open whether any competent team could reproduce them today. Meanwhile, the long tail of critical infrastructure — community hospitals, water systems, regional utilities — is run by organizations with no verification team, no way to evaluate a vendor's claims, and no procurement rule or insurer asking for evidence. Given the track record of these techniques in the best-resourced parts of industry, the question is not whether they work, but how quickly their adoption can be made ordinary and brought to scale.

Cities did not fight fires by hiring more firefighters alone; they wrote building codes — requiring fire-resistant materials, mandating firewalls between units, and inspecting the result — and applied them first and most strictly to the buildings where failure would cost the most.

From the report

For decision makers, the techniques serve three functions, which together define the strategy:

  • PreventMake important defect classes impossible by construction.
  • ContainLimit what a successful compromise can accomplish.
  • ProveEstablish critical properties through rigorous, machine-checkable analysis.

If memory-safety vulnerabilities cannot exist in the code, no amount of automated probing will find one. If verified compartment boundaries contain a breach, a foothold does not become a takeover.

From the report

The report lays out a phased adoption arc — from practices any competent engineering team can adopt today to machine-checked proofs for the most critical components — and asks organizations to back their assurance claims with inspectable artifacts: evidence, not checklists. Its actions for decision makers are deliberately limited and initial, each justified whether AI capability keeps advancing at its current rate or plateaus:⁠[10] inventory the software whose failure would be intolerable; make memory-safe implementation and explicit compartmentalization the default for new critical code; fund bounded pilots on a few high-consequence components; direct AI productivity gains toward higher-assurance code; assign executive ownership and decide how procurement and insurance will recognize adoption; and build the enabling ecosystem of tools, workforce, open-source assurance, shared services for under-resourced operators, and independent evaluation. For legacy systems, it points to compartmentalization as the most immediately deployable protection and to AI-assisted specification lifting, checked by experts, as the first step. These steps can begin now, while defenders can still act deliberately rather than in response to catastrophe.

This guidance is offered by industry and the technical community as a first step toward consensus. Detailed criteria are expected to emerge from rapid, iterative convening among practitioners and leaders across industry, academia, civil society, and government — cycles of weeks rather than years, each producing working guidance meant to be superseded by the next⁠[11] — and the committee intends to follow this document promptly with further work along those lines.

Formally verified code is often more performant than the unverified code it replaces.

Byron Cook, AWS Security Blog (2024)⁠[12]

The full report includes the actions for decision makers, the production evidence, the phased adoption arc with the assurance artifacts a buyer can request at each stage, the emerging ecosystem of independent evaluation, a candid account of costs and limitations, and an appendix of techniques and representative tools.

The full report (PDF) will be available on this page on Tuesday, September 23, 2026.

Cite as: NDIST Strategic Guidance Committee (2026). Prevent, Contain, Prove: A Framework for Urgently Securing Critical Digital Systems. New Directions in Software Technology.

About NDIST

New Directions in Software Technology (NDIST) is an annual invitation-only workshop hosted by Kestrel Institute, a nonprofit computer science research center in Palo Alto, California. For over two decades, NDIST has convened researchers, practitioners, and decision makers from academia, industry, and government to examine emerging challenges and opportunities in software technology. Each year's workshop focuses on a different theme; past topics have included synthetic biology, 3D printing, and artificial intelligence.

The Strategic Guidance Committee was formed at NDIST to develop this document through an open review process, drawing on the workshop's longstanding role as a forum for cross-disciplinary dialogue on the future of software. The report reflects extensive expert input, multiple rounds of open review, and sustained technical discussion among committee members and a wider community of stakeholders.

Committee co-chairs
Gopal Sarma (RAND), Brad Martin (Galois), Bryan Loyall (Charles River Analytics)
Committee members
Mike Dodds (Oath Technologies), David Hardin (Collins Aerospace), Robert Laddaga, Pat Lincoln (SRI), Paul Robertson (DOLL Labs), Howard Shrobe (MIT CSAIL), Eric Smith (Kestrel Institute), János Sztipánovits (Vanderbilt University), Stephen Westfold (Kestrel Institute)