Sponsor a Government pilot
A Government sponsor names the decision, selects the cases and reviewers, approves the guardrails, and controls how results may be used.
Discuss sponsorshipDeep Fathom has completed the core R&D and built the working system. A Government sponsor can now put it against selected cases, qualified reviewers, explicit thresholds, and real operating constraints—then decide where and how it should scale.
Download the 5-page Government pilot charterDeep Fathom can demonstrate the complete workflow now, supply the first versioned case, and instrument the comparison. Government does not have to fund the invention of the core system before testing it.
All 110 requirements and 320 assessment objectives connected to implementation, ownership, evidence, findings, and open work.
Source, collection time, scope, coverage, limits, integrity, responsibility, conflicts, and human findings remain traceable.
Rules and AI suggestions identify missing, stale, conflicting, and out-of-scope support while preserving the reviewer’s authorship.
A versioned MFA case, open JSON export, and inspectable human-readable record are ready for the first benchmark run.
The sponsor defines the decision the pilot will inform. Government designates the reviewers, controls the cases and environment, and sets every threshold before scoring begins.
The pilot should answer one operational question: can software-assisted review surface unsupported or conflicting security claims sooner, reduce repeated work, and focus qualified reviewers on the cases where failure would matter most?
Run the working system live against the published synthetic case. Inspect every source, rule result, AI suggestion, reviewer action, and export.
Freeze Government-selected cases, methods, versions, measures, success floors, and stop rules. Qualified reviewers establish the reference findings independently.
Give ordinary and AI-assisted reviewers the same frozen cases. Record quality, errors, labor, elapsed time, disagreement, corrections, and provenance.
If the benchmark clears its thresholds, add representative cases and operating constraints under Government-approved data, access, retention, and security rules.
Convert measured results and observed failures into performance requirements, human controls, integration requirements, and a deployment decision that names permitted uses, controls, and stop conditions.
Compare ordinary and AI-assisted review on the same frozen cases. Publish the errors and corrections alongside accuracy and time results.
Supplier and reviewer labor, duplicated handling, elapsed time, and costs separated into security, evidence, and verification work.
Coverage, source traceability, supported findings, missed conflicts, agreement, disagreement, and reproducibility.
Unsupported conclusions, false negatives, unsafe reliance, resistance to misleading inputs, changes in results over time, correction behavior, and results by supplier and operating context.
Synthetic cases come first. Report each case, evidence type, supplier context, and operating environment separately—not only as an average.
Configuration, identity, logging, vulnerability, network, endpoint, and source-system evidence.
Policies, interviews, demonstrations, judgment-dependent objectives, and implementation gaps.
MSPs, ESPs, cloud services, shared responsibility, customer duties, and provider limits.
Operational technology, legacy systems, small suppliers, compensating conditions, and sparse telemetry.
Incomplete, stale, contradictory, duplicated, mislabeled, and internally inconsistent packages.
Misleading artifacts, hidden instructions, poisoned context, suspicious provenance, and crafted omissions.
Material change after review, altered scope, expired evidence, and dependent findings.
Reasonable reviewer disagreement, uncertain responsibility, insufficient support, and appeal.
During evaluation, experimental output does not change certification, SPRS, award, affirmation, enforcement, supplier status, or procurement preference.
The pilot produces more than a vendor score. It leaves behind reusable cases, measured performance, requirements other methods can meet, and a documented decision about the next stage.
Versioned cases, cited reference findings, resolved disagreements, and a protocol that another team can rerun.
Results by case group, method, reviewer, and version—not one average that hides a dangerous failure.
Performance floors, stop rules, audit records, integration needs, correction paths, and explicit decision boundaries.
A documented decision about the work the method may support, the cases where it may be used, and the next stage of deployment.