SOC 2 and AI coding agents: what auditors ask, and how to evidence it
AI coding agents act on instructions stored in your repositories. Here is how that maps to SOC 2, ISO 27001 and ISO 42001 controls, and what evidence to collect.
Updated 2026-10-06
Auditors increasingly ask how an organisation governs AI coding agents. The honest answer at most companies is a policy document and good intentions. But agents do not read policies: they read AGENTS.md, CLAUDE.md, Cursor rules, Copilot instructions and MCP configuration, files that live in every repository and are changed like code, often without review.
Where agent instructions touch your controls
| Risk | Typical control | Evidence that answers it |
|---|---|---|
| Nobody knows which agents and tools are in use | SOC 2 CC6.1 · ISO 27001 A.5.9 · NIST AI RMF GOVERN 1.6 | An inventory of instruction files and MCP servers per repository |
| Secrets written into agent context or MCP config | SOC 2 CC6.1 · ISO 27001 A.5.17, A.8.12 | A scan for credentials across every agent-readable file |
| Injected or hidden instructions | SOC 2 CC6.8 · ISO 27001 A.8.7, A.8.28 | Checks for invisible Unicode, encoded payloads and download-and-run instructions |
| Agents told to skip tests or hide changes | SOC 2 CC8.1 · ISO 27001 A.8.29 | A scan for bypass instructions |
| Unpinned third-party MCP servers and skills | SOC 2 CC9.2 · ISO 27001 A.5.21 · ISO 42001 A.10.3 | Version and transport checks for every server |
| Unowned, unreviewed changes to instructions | SOC 2 CC8.1 · ISO 27001 A.8.32 | CODEOWNERS coverage of instruction files, plus PR checks |
| No ongoing monitoring | SOC 2 CC4.1, CC7.1 · ISO 27001 A.8.16 | Repeated audits compared over time |
Collecting the evidence
- Inventory every repository, not a sample: agent instructions are spread across teams.
- Test the instructions themselves, with deterministic checks whose false-positive rates are known.
- Keep the raw data and a hash of it, so the evidence can be tied to what was tested.
- Repeat on a schedule and keep the comparison, which is what demonstrates monitoring.
- Gate changes in pull requests so new problems cannot merge unnoticed.
threadctx’s org audit does this from your machine: it reads your repositories with your own GitHub token, tests eight controls, and writes an evidence pack with a framework cross-reference, an inventory and CSVs. Nothing is uploaded.
npx threadctx audit --org your-orgFramework mappings indicate where evidence is commonly relevant. They are not an audit opinion; your auditor decides whether evidence is sufficient.