Artificial intelligence (AI) isn’t entering the control environment as simply another application. It’s entering as a participant, a virtual employee described by some, that can create information, interpret policy, recommend judgments, communicate with stakeholders, and increasingly act autonomously across systems.
That distinction should change how management thinks about internal controls. Every management team should be asking: How should the control environment be redesigned when judgment and authority are being shared with systems whose outputs are probabilistic, dynamic, and not always fully observable?
Fortunately, 2026 guidance from the Committee of Sponsoring Organizations of the Treadway Commission (COSO) report, “Achieving Effective Internal Control Over Generative AI (GenAI),” offers a practical starting point for addressing this question.
According to COSO, management doesn’t need a separate framework for every new AI model. Instead, management can apply the five components of its Internal Control-Integrated Framework to a changed operating environment: control environment, risk assessment, control activities, information and communication, and monitoring activities.
Traditional technology governance begins with access. Specifically, it aims to answer who can enter a system and what data they can see. However, AI requires management to govern something broader by finding out how much judgment and authority a system may exercise.
Every material AI use should have a named business owner accountable for its purpose, risks, outputs, and controls. IT may operate the platform, but it shouldn’t own the business consequence. Therefore, finance should remain accountable for an AI-assisted reconciliation; human resources should remain accountable for an AI-supported employment decision; and the compliance team should remain accountable for an AI-generated regulatory assessment.
Management should also establish authority tiers. A tool that drafts an internal email shouldn’t be governed like an agent that recommends a journal entry, changes access rights, approves a claim, or communicates externally. The more AI moves from assistance toward execution, the more approval, competence, documentation, and oversight management should increase.
Additionally, tone at the top matters. Leaders can’t demand rapid AI adoption and then be surprised when employees treat validation as bureaucracy. Performance expectations should reward not only adoption and efficiency but accuracy, escalation, transparency, and adherence to controls. COSO similarly emphasizes clear ownership, ethical boundaries, role-appropriate competence, oversight, and accountability for AI outcomes.
The practice of maintaining an approved vendor listing as support for systems inventory shouldn’t be mistaken as an AI risk assessment. Instead of asking questions about completeness of an inventory listing, the better questions to ask are: What can the system do? What can it influence? What can it reach?
To answer these, COSO recommends assessing AI by its capability, such as ingesting data, transforming information, processing transactions, orchestrating workflows, generating judgments, monitoring activities, retrieving knowledge, or interacting with people. After all, the same product may be relatively low risk when summarizing internal notes but high risk when influencing financial reporting or acting through connected systems.
For each use case, management should identify the data consumed, output produced, decisions affected, systems accessed, and speed at which an error could spread. The assessment should also consider hallucinations, bias, privacy loss, model drifts, vendor changes, prompt injections, deepfakes, synthetic records, excessive agent authority, and insider or outsider fraud. COSO specifically notes that AI agents can create authorization risks and operate through insecure interfaces.
Additionally, management should address another but less familiar possibility: An advanced system may not faithfully describe its own conduct.
Recent disclosures from OpenAI and Anthropic have revealed that autonomous AI models have gone rogue by exploiting loopholes, concealing actions, or appearing compliant while pursuing another objective. Both incidents reportedly occurred within controlled research environments; however, researchers haven’t characterized this as evidence that current systems are conscious or poised to seize control. Regardless, the implication for management teams is clear: An autonomous system should never be allowed to operate without a human in the loop.
Given that models, prompts, integrations, and vendor settings can quickly change, AI risk assessment shouldn’t be merely an annual exercise. Risk assessment must become a living process with defined triggers for reassessment.
I often hear and use the phrase “human in the loop.” The phrase signals that AI, even autonomous AI, isn’t acting without oversight and, if necessary, intervention.
However, I want to be clear: Having a human in the loop doesn’t automatically mean that a control is effective. What matters is whether the human reviews the output before irreversible harm occurs, has sufficient evidence to challenge it, and possesses the competence and time to do so.
Management should match controls to consequences. Low-risk drafting may permit review after generation. Material accounting judgments, legal conclusions, customer decisions, or external communications should require qualified approval before use. AI agents that execute transactions should operate with the least amount of privilege access, value and volume limits, dual authorization for sensitive actions, and shutdown mechanisms they can’t override.
To support that risk-based approach, COSO recommends implementing access restrictions, segregation of duties, documented approvals, independent validation, rollback plans, and records sufficient to reconstruct what the system received, produced, and did.
Overall, human review must be substantive rather than ceremonial. A reviewer who accepts a polished AI output without corroboration isn’t performing a control. Management should define what evidence is required, what must be independently verified, which exceptions require escalation, and how the quality of the review will be tested.
AI can make information easier to produce while making its origin harder to establish. That creates risk wherever management, auditors, regulators, customers, or investors depend on reliable evidence.
The higher the risk, the more evidence management should preserve to show how an AI output was produced, reviewed, and approved, including source data, model and configuration, significant prompts, resulting output, reviewer approval, and known limitations. Users should also know when content is AI-generated, when confidence is uncertain, and when a model or vendor change may have affected performance.
To support this, COSO recommends reporting output-quality indicators, such as hallucinations, citation coverage, and bias, along with control performance measures. COSO also calls for defined protocols governing who must be informed of incidents, limitations, and material changes.
A key nuance of AI controls is that they may weaken because of an update, changed source data, or a new prompt. Therefore, monitoring must determine not only whether a control exists but whether it still works.
Monitoring must also remain sufficiently independent from the AI system itself. A system shouldn’t be able to alter the evidence used to evaluate it, suppress an alert, change its own monitoring thresholds, or approve its own exception.
COSO calls for continuous observation combined with periodic deeper reviews and prompt evaluation and communication of deficiencies. The National Institute of Standards and Technology’s AI Risk Management Framework similarly emphasizes incorporating trustworthiness throughout the design, use, and evaluation of AI systems.
Overall, the lesson across these five components isn’t that AI makes COSO’s framework obsolete, but that it must be applied with greater precision and care.
This means that management must evolve from governing access to software toward governing delegated judgment and authority. Risk assessment must follow changing capabilities, not static vendor names. Control activities must validate probabilistic outputs and constrain autonomous action. Information systems must preserve sources and disclose limitations. Monitoring must test actual behavior rather than merely confirm that a policy exists.
Ultimately, durable advantage won’t necessarily belong to organizations that adopt AI the fastest but to those that govern it with the discipline, evidence, and oversight needed to scale responsibly.