CMMC20X
Government pilot proposal

Built and ready
for a controlled pilot.

Deep Fathom has completed the core R&D and built the working system. A Government sponsor can now put it against selected cases, qualified reviewers, explicit thresholds, and real operating constraints—then decide where and how it should scale.

Download the 5-page Government pilot charter Discussion draft v0.1 · August 2026 · Tagged PDF/UA-1 · 47 KB
Working systemComplete Level 2 modelRunnable synthetic benchmarkEvaluation charter ready
Two ways to take part

Sponsor the decision—or strengthen the test.

Sponsor a Government pilot

A Government sponsor names the decision, selects the cases and reviewers, approves the guardrails, and controls how results may be used.

Discuss sponsorship

Contribute to the evaluation

Contractors, providers, advisors, assessors, primes, and researchers can contribute controlled cases, field constraints, reference review, or interoperable workflows without sponsoring the pilot.

See how to contribute
What Deep Fathom brings

The pilot starts with working software.

Deep Fathom can demonstrate the complete workflow now, supply the first versioned case, and instrument the comparison. Government does not have to fund the invention of the core system before testing it.

  • Complete Level 2 model

    All 110 requirements and 320 assessment objectives connected to implementation, ownership, evidence, findings, and open work.

  • Working evidence graph

    Source, collection time, scope, coverage, limits, integrity, responsibility, conflicts, and human findings remain traceable.

  • Repeatable and AI-assisted review

    Rules and AI suggestions identify missing, stale, conflicting, and out-of-scope support while preserving the reviewer’s authorship.

  • Runnable synthetic case

    A versioned MFA case, open JSON export, and inspectable human-readable record are ready for the first benchmark run.

What Government controls

Government sets the question and the rules.

The sponsor defines the decision the pilot will inform. Government designates the reviewers, controls the cases and environment, and sets every threshold before scoring begins.

The pilot should answer one operational question: can software-assisted review surface unsupported or conflicting security claims sooner, reduce repeated work, and focus qualified reviewers on the cases where failure would matter most?

Decision
Name the Government decision this study will inform—and the decisions it cannot make.
Sponsor
Name the Government problem owner, method owner, data owner, and security owner.
Cases
Select versioned synthetic cases first, then approve any representative inputs, environments, access, retention, and use.
Comparison
Run ordinary review and AI-assisted review on the same frozen cases where feasible.
Reference findings
Qualified reviewers work independently of vendor output, cite sources, and record how disagreements were resolved.
Pass rule
Burden improves while every minimum quality and safety result set before testing holds for every relevant case group.
Stop rule
Stop for a predeclared dangerous error, data-handling failure, integrity failure, or protocol violation.
Path to scale

Earn the next stage with measured results.

  1. 1. Demonstrate

    Run the working system live against the published synthetic case. Inspect every source, rule result, AI suggestion, reviewer action, and export.

    Stage outputShared understanding of the system and its boundaries
  2. 2. Benchmark

    Freeze Government-selected cases, methods, versions, measures, success floors, and stop rules. Qualified reviewers establish the reference findings independently.

    Stage outputBenchmark cases, reference findings, signed protocol
  3. 3. Compare

    Give ordinary and AI-assisted reviewers the same frozen cases. Record quality, errors, labor, elapsed time, disagreement, corrections, and provenance.

    Stage outputMeasured results tied to each person, method, and version
  4. 4. Expand under control

    If the benchmark clears its thresholds, add representative cases and operating constraints under Government-approved data, access, retention, and security rules.

    Stage outputResults across technical, procedural, inherited, and difficult cases
  5. 5. Decide how to scale

    Convert measured results and observed failures into performance requirements, human controls, integration requirements, and a deployment decision that names permitted uses, controls, and stop conditions.

    Stage outputDecision memo, requirements, controls, and next-stage plan
What the pilot measures

Faster review only matters if quality and safety hold.

Compare ordinary and AI-assisted review on the same frozen cases. Publish the errors and corrections alongside accuracy and time results.

Burden

Supplier and reviewer labor, duplicated handling, elapsed time, and costs separated into security, evidence, and verification work.

Quality

Coverage, source traceability, supported findings, missed conflicts, agreement, disagreement, and reproducibility.

Safety

Unsupported conclusions, false negatives, unsafe reliance, resistance to misleading inputs, changes in results over time, correction behavior, and results by supplier and operating context.

Test cases

Include the cases most likely to expose a bad method.

Synthetic cases come first. Report each case, evidence type, supplier context, and operating environment separately—not only as an average.

  • Technical

    Configuration, identity, logging, vulnerability, network, endpoint, and source-system evidence.

  • Procedural

    Policies, interviews, demonstrations, judgment-dependent objectives, and implementation gaps.

  • Inherited

    MSPs, ESPs, cloud services, shared responsibility, customer duties, and provider limits.

  • Constrained

    Operational technology, legacy systems, small suppliers, compensating conditions, and sparse telemetry.

  • Degraded

    Incomplete, stale, contradictory, duplicated, mislabeled, and internally inconsistent packages.

  • Adversarial

    Misleading artifacts, hidden instructions, poisoned context, suspicious provenance, and crafted omissions.

  • Changed

    Material change after review, altered scope, expired evidence, and dependent findings.

  • Ambiguous

    Reasonable reviewer disagreement, uncertain responsibility, insufficient support, and appeal.

Pilot guardrail

During evaluation, experimental output does not change certification, SPRS, award, affirmation, enforcement, supplier status, or procurement preference.

What Government receives

An answer it can act on.

The pilot produces more than a vendor score. It leaves behind reusable cases, measured performance, requirements other methods can meet, and a documented decision about the next stage.

  • Reusable benchmark

    Cases + reference

    Versioned cases, cited reference findings, resolved disagreements, and a protocol that another team can rerun.

  • Measured performance

    Burden + quality + safety

    Results by case group, method, reviewer, and version—not one average that hides a dangerous failure.

  • Deployment requirements

    Technical + human controls

    Performance floors, stop rules, audit records, integration needs, correction paths, and explicit decision boundaries.

  • Scale decision

    Proceed, revise, narrow, or stop

    A documented decision about the work the method may support, the cases where it may be used, and the next stage of deployment.