What the Model Says About Evidence and Review
The CMMC 20X model separates evidence handling from review capacity and shows why improving either one alone leaves a different constraint in place.
Current for schema-v3 model commit bf91444e3fcc and assumption set graduated-verification-2026-08-13.v1. This is conditional systems analysis; it is neither an observed program result nor a measurement of software performance.
If evidence moves faster, does the review queue disappear?
No. Better evidence handling and greater review capacity act on different parts of the system. Either can improve the path. Neither automatically removes the constraint addressed by the other.
That distinction is central to the CMMC 20X program model. The model tracks an aggregate Level 2 supplier population through readiness, evidence preparation, verification routes, corrective work, expiration, and renewal. It also tracks finite queues for independent assessment, assisted human review, and Government sampling. Evidence methods affect how much repeated handling occurs and how quickly work can be prepared or reviewed. Capacity determines how much qualified review can finish during a period.
The model is a design instrument. Its results are conditional on published assumptions and named policy portfolios becoming effective and resourced. It does not report observed contractor behavior, actual AI accuracy, realized reviewer productivity, or a forecast of Department action.
Two resources that are easy to collapse into one
Evidence is the support attached to a security claim: its source, scope, collection time, coverage, responsibility, conflicts, and review history. A reviewer uses that record to decide whether an implementation claim withstands examination.
Review capacity is the available throughput of people and institutions authorized or qualified to perform a particular review. It depends on staffing, training, authorization, case complexity, productivity, quality controls, and the time each case consumes.
The two interact. A well-structured package may reduce reconciliation work. A larger reviewer pool may clear more cases. Yet an immaculate package still needs the review required by the program, and a larger reviewer pool still loses time when every case arrives with disconnected files, uncertain scope, or missing provider facts.
| Program lever | Direct effect represented in the model | Constraint left behind |
|---|---|---|
| Evidence modernization | Reduces some repeated preparation and handling; supports repeatable checks | The policy-required review route and its finite capacity |
| Independent capacity expansion | Raises the throughput of full third-party assessment over time | Supplier readiness work and package friction |
| Assisted human capacity | Creates a finite queue for qualified exception review | Need for tested task boundaries, escalation, and sampling |
| Graduated verification | Allocates different review depth by mission and data consequence | Need to calibrate routes and govern every route |
| Coordinated reform | Sequences scope, readiness, capacity, evidence, and routes | Implementation risk and uncertain empirical parameters |
This table describes model mechanisms. It makes no claim that a particular technology has achieved the assumed labor or exception rates.
What the comparison shows
The published analysis uses 100 paired primary runs for each of seven policy options. Each paired run shares the same exogenous seed so that policy options face the same modeled world shocks. The primary estimand asks how the system behaves once the named intervention portfolio is effective. A separate implementation-risk branch represents first-attempt policy and legal process through 2035.
In the median evidence-modernization path, final current assurance reaches 97 percent, supplier loss reaches 13 percent, and the peak review backlog reaches 23,000 organizations. The path assumes that a common evidence structure and assisted handling reduce some repeated work. It retains the full-assessment route for the same population. The queue therefore remains large.
In the median capacity-expansion path, final current assurance also reaches 97 percent, supplier loss reaches 18 percent, and peak backlog reaches 27,000. Capacity grows after training and authorization delays. Every supplier still faces the modeled readiness and evidence work associated with that route.
The coordinated path changes several mechanisms in sequence. It corrects scope, phases demand, supports readiness, expands independent capacity, introduces structured evidence, and activates three review routes. Its median run ends at 91 percent current assurance, 1 percent supplier loss, and a peak backlog of 7,400. Those figures belong to this assumption set. They should travel with the model version, seed design, estimand, and limitations.
The finding is structural: improving the inputs to a queue differs from changing the queue’s service capacity, and changing service capacity differs from deciding which work enters it. A complete operating design has to address all three.
Why the final coverage number can mislead
A route can finish with high current assurance after a difficult transition. That endpoint may conceal a large earlier backlog or substantial supplier exit. When organizations leave the modeled active population, the denominator for coverage also changes. A high final percentage can therefore coexist with a weaker industrial-base outcome.
For that reason the analysis presents current assurance beside peak backlog, supplier loss, high-risk coverage, contractor burden, and a control-exposure index. The exposure index is a comparative measure of modeled safeguard implementation. It is not a breach probability or loss estimate.
Review capacity also has more than one form. C3PAO capacity, assisted-review capacity, and Government sampling capacity are distinct stocks in the model. Moving work away from one queue without resourcing the receiving queue simply moves the constraint. A self-review route without credible sampling can create coverage on paper while weakening the check. An assisted route without enough qualified people can accumulate its own backlog.
What the model cannot settle
Several important values remain decision hypotheses:
- The reference 40/35/25 route mix has not been calibrated against Government mission and data-risk cohorts.
- Assisted-review labor, exception, and failure rates are assumptions used for sensitivity analysis. They are not measured Deep Fathom results.
- The model is aggregate. It does not create a synthetic firm population with company-specific architectures, contracts, providers, or behavior.
- Cost layers are kept separate and expressed as ranges. Official estimates and broader small-business cost claims are not blended into a single point.
- Legal and process failure is represented separately from the primary conditional comparison.
These limits prevent a common analytical error: reading a simulated portfolio as proof that its operational components already work at scale.
An inspectable test of the mechanism
Government could test the disputed link between evidence and capacity without changing any contractor’s status.
First, freeze a representative set of synthetic or de-identified cases. Include complete, stale, conflicting, provider-dependent, sparse, and adversarial packages. Have qualified reviewers establish reference findings before seeing software output.
Second, randomize the same cases across current document handling and a structured, source-linked workflow. Keep the required review depth constant. Measure preparation labor, reviewer labor, elapsed time, conflicts found, unsupported conclusions, corrections, disagreement, and rework.
Third, use the observed distributions to replace the model’s evidence-labor and exception assumptions. Rerun capacity sensitivities while preserving the published seeds and scenarios. Publish the cases, protocol, versioned system, errors, corrections, and revised parameter ranges where security permits.
Finally, test allocation separately. Do not infer that faster evidence handling justifies a different verification route. Route depth should remain a policy decision based on mission consequence, data sensitivity, threat exposure, supplier criticality, and material change. Evidence weakness should trigger correction, escalation, or sampling.
The policy implication
Evidence modernization deserves evaluation because repeated reconstruction can consume scarce contractor and reviewer time. Capacity expansion deserves action because independent review remains finite and takes time to grow. Graduated verification deserves controlled testing because scarce review should be allocated deliberately.
CMMC 20X proposes that these mechanisms be designed together: one safeguard baseline, source-linked evidence, qualified human judgment, distinct review routes, Government sampling, and refresh after material change. This is a public proposal. It is not part of the current Government standard unless and until an authorized process adopts specific elements.
The model’s contribution is narrower and useful: it makes the queue visible. It shows which assumption changes preparation, which changes throughput, and which changes allocation. Policymakers can then ask for evidence about each mechanism instead of treating “automation” or “more assessors” as a complete program.
Inspect the complete methodology and compare the seven CMMC policy options.
Sources and revision history.
Primary and governing sources
- 01CMMC Program final rule, 89 FR 83092U.S. Government Publishing Office
- 02Cybersecurity Maturity Model Certification Program, 32 CFR part 170Electronic Code of Federal Regulations
- 03Defense Industrial Base Cybersecurity Strategy Implementation Plan, FY 2024–2027U.S. Department of Defense
- 04Artificial Intelligence Risk Management Framework (AI RMF 1.0)National Institute of Standards and Technology
- 05Open Security Controls Assessment LanguageNational Institute of Standards and Technology
- 06CMMC Reform Analysis MethodologyDeep Fathom
Corrections and material revisions
No material corrections recorded.
See an error or a source we missed? Send a correction. Material changes are recorded here rather than silently overwritten.
Inspect the model behind the finding.
Compare all seven CMMC policy options, then read the assumptions and limits that determine what the results can and cannot support.