AI in GRC for federal agencies is shifting compliance from manual evidence collection to continuous, verifiable control monitoring.

Artificial Intelligence · Governance, Risk and Compliance · By KSG Team · September 2026

Governance, Risk and Compliance work in federal agencies has historically been a documentation exercise. A control is assessed, a screenshot is captured, a narrative is written, a package is assembled, and an Authority to Operate (ATO) is granted against a point in time. By the time an auditor examines the evidence, the environment it describes has often changed.

That model is being replaced. The introduction of AI in GRC for federal agencies is part of a broader shift toward continuous compliance, in which control evidence is generated by systems as they operate rather than reconstructed by people after the fact.

This article covers three things in order. It describes what is driving the shift, it identifies where AI genuinely contributes and where it does not, and it addresses a governance question that most coverage of this subject omits entirely: an AI system deployed to reduce compliance burden may itself carry a compliance obligation.

The Cost of Manual Evidence Collection

The difficulty with the traditional model is not that compliance teams work slowly. It is that the evidence they produce describes a moment that has already passed.

Four characteristics make manual evidence collection expensive for an agency.

  1. Evidence ages immediately. A screenshot of an access control configuration is accurate on the day it is captured. Assessment cycles run annually. The gap between the evidence and the environment widens for the remainder of the year.
  2. The same proof is collected repeatedly. An agency subject to FISMA, holding cloud services under FedRAMP, and carrying framework obligations across multiple systems will often maintain separate evidence repositories for overlapping controls. A single access management control is evidenced several times, in several formats, for several audiences.
  3. Assembly consumes the expertise. Experienced Information System Security Officers spend a substantial share of their time collecting, formatting and cross-referencing artifacts rather than analyzing risk. That is the most expensive possible use of a scarce skill set.
  4. Documentation drifts from reality. A System Security Plan describes an intended architecture. Systems change continuously. Without automated reconciliation, the plan and the environment diverge quietly, and the divergence is usually discovered during an assessment rather than before one.

Continuous compliance addresses these by changing where evidence originates. Instead of being assembled by a person at assessment time, evidence is produced by the environment itself, continuously, as a structured output.

What Is Driving the Shift

FedRAMP 20x and continuous validation

The most visible federal example is FedRAMP 20x, the General Services Administration's modernization of federal cloud authorization, announced in March 2025. It moves the program away from static documentation toward continuous, automated validation of security outcomes.

Three elements of that model matter beyond cloud authorization. The first is Key Security Indicators, a defined set of security outcomes validated automatically rather than described narratively. The second is machine-readable evidence, submitted as structured data generated from the running environment through application programming interfaces and configuration state rather than transcribed into documents. The third is a changed role for the assessor, whose work shifts from examining static documents toward confirming that the validation pipeline itself works correctly.

As of mid-2026 the program is in its third phase, following completed pilots at the Low and Moderate baselines. Consolidated rules and a public submission pipeline are expected within the current fiscal year. Timelines published by the program are described as estimated goals rather than firm commitments, so agencies should track them rather than plan against them.

OSCAL and machine-readable evidence

The enabling standard is the Open Security Controls Assessment Language (OSCAL), developed by the National Institute of Standards and Technology. OSCAL expresses security controls, System Security Plans, assessment results and Plans of Action and Milestones as structured, version-controlled data rather than as prose documents.

The practical consequence is that a control assessment becomes reproducible. A finding carries its source tool, the time the evidence was gathered, the control identifier it maps to, and the assessment objective it relates to. Re-running the pipeline against the same input produces the same result. That property is difficult to achieve with a document and a screenshot, and it is the foundation on which any useful automation is built.

It is worth being precise about sequence here. OSCAL and evidence pipelines are not AI. They are structured data and automation. AI becomes useful on top of that foundation, and considerably less useful without it.

Continuous ATO

Continuous ATO applies the same principle to authorization itself. Rather than a periodic reauthorization exercise, the authorizing official maintains visibility into control state continuously, and risk decisions are made against current information. This depends on the same prerequisites: instrumented systems, structured evidence, and a defined set of indicators that can be validated without human transcription.

Where AI Contributes, and Where It Does Not

Discussion of AI in compliance tends toward two unhelpful positions. The first treats AI as a replacement for the compliance function. The second dismisses it as unsuitable for regulated work. Neither is accurate, and the useful distinction is narrower than either.

AI performs well on tasks where the source material exists, the output is checkable, and a qualified person reviews the result before it carries any weight.

Task Why AI is suited to it The condition attached
Control mapping across frameworks Overlapping control sets across NIST SP 800-53, NIST SP 800-171, FedRAMP baselines and agency policy are a language problem. AI handles the first pass of identifying where one control satisfies another. A human validates each mapping. An incorrect mapping produces a control gap that appears closed.
Evidence summarization and gap identification Reading across large volumes of scan output, configuration state and log data to identify which controls lack supporting evidence is well suited to automation. The system must cite the underlying artifact for every statement. A summary without provenance is not evidence.
Draft narrative generation Producing first drafts of control implementation narratives from structured configuration data removes a substantial amount of transcription work. A draft is a draft. The System Security Plan is a legal representation of the environment and requires human authorship and accountability.
POA&M triage and trend analysis Clustering related findings, identifying systemic root causes across many items, and surfacing trends that are difficult to see in a spreadsheet. Risk acceptance and prioritization remain judgments made by accountable officials.
Policy and document question answering Retrieval across policy libraries, framework text and prior assessment records, so that an analyst finds the governing language rather than recalling it. Retrieval must be restricted to authoritative sources and must return citations.

Where AI should not be given authority

The limitations are as important as the capabilities, and they are underdiscussed because the commercial incentive in this market runs in the other direction.

  • AI should not determine whether a control is satisfied. That determination has legal and authorization consequences. It belongs to an accountable official working from evidence, not to a model working from inference.
  • AI should not generate evidence. It should organize, summarize and cite evidence that already exists. A model that produces a plausible description of a control implementation without an underlying artifact has produced a liability, not an assessment.
  • AI output without provenance has no assessment value. An assessor cannot accept a conclusion that does not trace to a source. Any system used in this context must return the artifact, the tool, the timestamp and the control identifier alongside the conclusion.
  • AI should not be the sole reviewer of its own output. Automated review of automated output removes the independent judgment that makes the evidence credible.

Expressed as a single principle: AI reduces the labor of compliance. It does not transfer the accountability.

The Obligation That Most Coverage Omits

This is the part of the subject that receives the least attention, and it has become time-sensitive.

An agency that deploys AI to reduce compliance burden has deployed an AI system. That system falls within the federal AI governance requirements the agency is already subject to, and those requirements carry their own inventory, assessment and reporting obligations.

What OMB M-25-21 requires

OMB Memorandum M-25-21, issued on April 3rd, 2025, is the governing guidance for federal agency use of AI. It rescinded and replaced M-24-10 and sets binding requirements across the executive branch. It requires each agency to designate a Chief AI Officer, to convene an AI governance body, to maintain an AI use case inventory, and to apply minimum risk management practices to systems classified as High-Impact AI.

The memo defines High-Impact AI as AI with an output that serves as a principal basis for decisions or actions with legal, material, binding or significant effect on a defined set of areas. One of those areas is strategic assets or resources, including high-value property and information marked as sensitive by the federal government.

A deadline that has now passed

M-25-21 required agencies to report their minimum risk management practices for High-Impact AI to OMB by September 22nd, 2026. The memo states that any use case not compliant with those minimum practices must be discontinued.

Agencies that completed the classification exercise should confirm that GRC automation was considered within it. Agencies that have not completed it are now past the reporting date.

Is a GRC automation system High-Impact AI?

The honest answer is that it depends on how the system is designed, and the determination is not automatic.

The list of presumed High-Impact use cases in the memo does not include compliance automation. The general definition, however, turns on whether the AI output serves as a principal basis for a decision with material effect. An AI system whose output is the principal basis for accepting a control as implemented, or for an authorization decision concerning a system that holds sensitive federal information, is closer to that definition than most agencies assume.

The determining factor is architectural rather than technical. It is whether the system is designed to decide or to advise.

Design choice What the AI produces Governance consequence
Advisory Draft narratives, candidate control mappings, evidence summaries with citations, and flagged gaps. An accountable official reviews the output and makes every determination. The AI output is an input to a human decision rather than the principal basis for it. This design also produces better assessment outcomes, because the reviewing step is where errors are caught.
Decisional Control satisfaction determinations, automated risk acceptance, or authorization inputs applied without substantive human review. The system is considerably closer to the High-Impact AI definition, and the agency should expect to carry the associated minimum practices, including pre-deployment testing, an AI impact assessment, ongoing monitoring, human oversight and appeal mechanisms.

Two points follow from this, and both should be settled before deployment rather than after.

  1. Human-in-the-loop design is a governance decision, not only an engineering preference. Keeping the accountable official as the decision-maker is what keeps a GRC automation system on the advisory side of the definition. That choice has the additional benefit of matching how assessment evidence is actually evaluated.
  2. The classification should be documented, whichever way it goes. A determination that a system is not High-Impact AI is a determination. It should be recorded with its reasoning, reviewed by the Chief AI Officer's process, and reflected in the agency's AI use case inventory. An undocumented conclusion is indistinguishable from an unexamined one.

A scoping note on applicability

M-25-21 exempts the Department of War and National Security Systems from its provisions. The requirements described above therefore apply to civilian federal agencies. Defense organizations operate under separate AI governance direction and should confirm their obligations through their own channels rather than applying this memo by analogy.

Contractors supporting civilian agencies should note that the obligation sits with the agency, and that AI capabilities delivered under contract will be evaluated within the agency's governance process. Understanding that process before delivery is substantially easier than retrofitting to it afterward.

The supporting framework

The NIST AI Risk Management Framework, published as NIST AI 100-1, provides the structure most agencies use to operationalize these requirements. Its four functions, Govern, Map, Measure and Manage, supply the common vocabulary for AI risk management across the federal government. The framework is voluntary guidance, but OMB requirements have made it the practical reference standard for federal AI risk work.

For agencies already operating mature Governance, Risk and Compliance functions, this is a recognizable pattern rather than a new discipline. The same practices used for system authorization apply to AI systems: inventory what exists, classify it by impact, test before deployment, monitor after deployment, and document the decisions.

How KSG Approaches AI in GRC

Kaizen Solutions Group has secured federal and state government systems since 2016. We are an SBA 8(a)-certified small disadvantaged business, ISO 27001:2022, ISO 9001:2015 and ISO 20000-1:2018 certified, and we hold GSA HACS designations across all five categories: Risk and Vulnerability Assessment, High Value Asset assessment, Penetration Testing, Incident Response, and Cyber Hunt.

We do not treat AI as a standalone tool. We integrate AI into governed business and security processes with access control, auditability and risk management already in place, so that an agency gains the operational benefit without weakening its security posture or its compliance position.

Our view of this work is that the sequence matters. An agency that applies AI to a compliance process it has not yet structured will automate an unstructured process. Evidence pipelines, control mappings and defined indicators come first. AI makes that foundation faster to work with, and it is of limited value without it.

Where we typically begin

  • Governance, Continuous ATO and Continuous Monitoring - We establish the control baselines, evidence sources and monitoring cadence that continuous compliance depends on, including FISMA program support and Assessment and Authorization services.
  • AI for Governance, Risk and Compliance - We apply AI to control mapping, evidence summarization, gap identification and draft narrative generation, with citation to source artifacts and an accountable reviewer at every determination point.
  • AI governance and responsible AI - We support AI use case inventory, impact classification, pre-deployment testing and ongoing monitoring aligned to the NIST AI Risk Management Framework and OMB direction.
  • Agentic AI and workflow automation - We design automation that keeps accountable officials as decision-makers, with auditable records of what the system did and what a person approved.
  • Risk Management and POA&M analytics - We deliver continuous monitoring, vulnerability and patch management, and Key Risk Indicator and Key Performance Indicator dashboards in Tableau and Power BI, so that control state and remediation progress are visible to leadership.
  • Security operations and engineering - We support Security Operations Center monitoring, SIEM and SOAR, Identity and Access Management, and DevSecOps, which are the systems that generate control evidence in the first place.

A useful place to begin

For most agencies, the first question is not which AI capability to adopt. It is which control evidence is already produced automatically, which is still assembled by hand, and which AI systems already in use have been classified within the agency's AI inventory.

Those three answers determine whether an agency is positioned for continuous compliance or is preparing to automate a manual process without restructuring it.

Talk to our team · Download our capability statement · Explore our Artificial Intelligence capabilities

Frequently Asked Questions

Is an AI system used for GRC automation considered High-Impact AI?

It depends on the design, and the determination is not automatic. OMB M-25-21 defines High-Impact AI as AI whose output serves as a principal basis for decisions with legal, material, binding or significant effect on defined areas, including strategic assets and information marked as sensitive by the federal government. Compliance automation is not on the memo's list of presumed High-Impact use cases. A system that drafts narratives and summarizes evidence for human review sits on the advisory side. A system whose output determines control satisfaction or drives risk acceptance without substantive human review is considerably closer to the definition. The classification should be documented either way.

Does OMB M-25-21 apply to the Department of War?

No. The memo explicitly exempts the Department of War and National Security Systems from its provisions. It governs civilian executive branch agencies. Defense organizations operate under separate AI governance direction and should confirm their obligations through their own channels.

What was the September 22nd, 2026 deadline?

M-25-21 required agencies to report their minimum risk management practices for High-Impact AI to OMB by September 22nd, 2026. The memo states that any use case not compliant with those minimum practices must be discontinued. Earlier milestones in the same memo required a designated Chief AI Officer by June 30th, 2025, an AI governance body by August 12th, 2025, an agency AI strategy and a compliance plan by December 26th, 2025, and updated internal policies on IT infrastructure, data, cybersecurity and privacy by May 6th, 2026.

What is OSCAL, and why does it matter for continuous compliance?

The Open Security Controls Assessment Language is a NIST standard that expresses security controls, System Security Plans, assessment results and Plans of Action and Milestones as structured, machine-readable data rather than as prose documents. It matters because it makes control assessment reproducible: a finding carries its source tool, collection time, control identifier and assessment objective. That structure is the foundation continuous compliance and any useful automation are built on.

Can AI write a System Security Plan?

AI can produce a first draft from structured configuration data, which removes a significant amount of transcription work. It should not be the author of record. A System Security Plan is a representation of an environment that carries authorization and legal consequences, and it requires human authorship and an accountable reviewer. Any AI-generated content should cite the underlying artifact it was derived from.

What are Key Security Indicators under FedRAMP 20x?

Key Security Indicators are a defined set of security outcomes that can be validated automatically rather than described in narrative form. They were the central proof of concept in the FedRAMP 20x Phase One pilot at the Low baseline and were extended in the Phase Two Moderate pilot. They represent the shift from describing a control to continuously demonstrating an outcome.

Does continuous compliance eliminate the need for an ISSO?

No. It changes what the role spends time on. Automating evidence collection removes assembly, formatting and cross-referencing work, which is the least valuable use of the expertise. Determinations about control satisfaction, risk acceptance and authorization remain human responsibilities, and they become more central as the routine work is automated.

Where should an agency start?

With an inventory rather than a tool. Three questions establish the position: which control evidence is already produced automatically by existing systems, which is still assembled manually, and which AI systems already in use have been classified within the agency's AI use case inventory. Applying AI to a compliance process that has not been structured produces an automated version of an unstructured process.