SOC 2 for AI Startups: Scope Models, Data, and Providers
A practical SOC 2 scope plan for AI startups: trace customer data through model calls, logs, training, and providers before writing controls.

SOC 2 for AI startups starts with the product’s real data path. Name the service customers use, then trace their information through your application, retrieval store, model calls, logs, feedback, and any training or evaluation process. Decide which people and providers operate each step, what you promise customers, and where you can find dated evidence. That gives management a candidate scope to discuss with a CPA firm before it writes controls or sets a report date.
What changes when the product uses AI?
The SOC 2 framework does not become a new framework because the service calls a model. The AICPA’s Trust Services Criteria still provide the criteria for the selected categories, while the Description Criteria address management’s description of the actual service and system. The AICPA also published technical questions and answers on AI in SOC examinations in September 2026. Ask the CPA firm how its current guidance applies to your architecture.
For an AI product, the usual scoping mistake is stopping at the web app and cloud account. A production request may send customer text to another provider, add retrieved documents, write a trace to logs, and save feedback for later evaluation. Each step may have a different owner, retention rule, and evidence source. Those facts change the system description and the controls management needs to operate.
This article is about an AI product going through SOC 2. If you mean an AI agent helping your team maintain compliance records, use the separate AI agent for SOC 2 workflow.
Map one request before choosing controls
Take a representative production request and write down every place its data goes. Do the same for batch jobs and administrative workflows that handle customer information. A useful first map looks like this:
| Step | Question to answer | Source to inspect |
|---|---|---|
| Input | What do users submit, and can it contain confidential or personal information? | API schema, upload path, customer terms |
| Retrieval | Which documents or embeddings join the request, and how is tenant access enforced? | Store configuration, access tests |
| Model call | Which provider or hosted model receives the assembled input, in which environment? | Application configuration, contract, request traces without secrets |
| Output | Where does the response go, and who can see or change it? | Application permissions, delivery logs |
| Logs and feedback | Are prompts, outputs, metadata, or ratings retained, and for how long? | Logging settings, retention configuration |
| Training and evaluation | Does any customer material enter datasets, fine-tuning, or testing? | Dataset inventory, approval records, pipeline configuration |
Do not copy this example into a system description as fact. Check the running service and the contracts. A setting that prevents retention in one environment does not prove the same rule applies to every integration or log sink. If a developer tool also receives production information, map that path too.
NIST’s AI Risk Management Framework recommends mapping components and third-party software and data. That is useful discipline here, but the NIST framework does not decide your SOC 2 engagement scope. The general SOC 2 scope guide shows how to turn the map into a bounded service, information types, components, and exclusions.
Decide which AI dependencies matter to the service
Use the map to test each proposed inclusion or exclusion. Start with the customer-facing inference path, then add supporting work that can change its security or operation.
| Dependency | Include in management’s scope analysis when… | Check before claiming control |
|---|---|---|
| External model API | Production service calls it or sends it scoped data | Contract, data handling, access, configuration, and provider oversight |
| Self-hosted model | Your team deploys or runs it for the service | Deployment review, access, monitoring, changes, and recovery |
| Retrieval store | It holds or selects customer context for a model call | Tenant boundaries, access, retention, backup, and deletion |
| Prompt or output logs | They can contain customer data or support a control | Actual fields, access, retention, and sample retrieval |
| Training or evaluation set | Scoped customer data enters it or it can change the service | Approval, lineage, access, version, and deletion process |
| Internal AI assistant | It can access production data, code, credentials, or control work | Allowed use, access, data flows, and human review |
This is an analysis list, not a claim that every item must appear in every report. A provider’s role also needs care: a relevant vendor is not automatically a subservice organization. Discuss provider treatment and any carve-out or inclusive method with the CPA firm. The vendor review guide explains how to keep the provider decision and follow-up work reviewable.
Turn the map into controls and evidence
For each path that matters, write an owner, a rule, an operating source, and a way to retrieve a dated record. A small team can start with five decisions:
- Define what customer information may enter prompts, retrieval, logs, and training. Check the product behavior against the customer terms.
- Restrict who can change model routing, prompt templates, retrieval access, provider settings, and production credentials. Keep the change record in the source systems.
- Review model and data providers against the service’s commitments. Record the contract, relevant assurance material, open questions, and owner.
- Decide how to detect and handle failures, misuse, information exposure, and bad changes. Keep incident and change records tied to the actual systems.
- Test that the evidence can be retrieved without exposing raw customer prompts or secrets to people who do not need them.
The exact controls depend on your service, risks, selected categories, and customer commitments. For example, a product promise about output handling may raise different questions from a promise about uptime. Do not infer that model accuracy is always tested under SOC 2, or that a Security-only scope answers every buyer question about AI behavior. Set those expectations with the report user and CPA firm.
Keep the program reviewable in Git
For this workflow, filegrc can hold the proposed System boundary, Components, information types, vendors, risks, controls, obligations, and evidence references in one Git-native GRC workspace. JSON holds structured records, Markdown holds long-form work, and Git supplies the change history. Starter records are proposals for management to review, not compliance claims.
Keep operational proof in the systems that produce it. Your application, identity provider, cloud account, model provider, monitoring tools, and other source systems still operate the controls. FileGRC can organize fixed evidence or references, but it does not operate those systems, collect their data automatically, perform the examination, or decide whether evidence is enough.
Start with a single reviewed data-flow map, then use the SOC 2 system description guide to turn the agreed boundary into management’s narrative. When the model path, provider, retention rule, or customer promise changes, revisit the map, risks, controls, and evidence sources before carrying old claims into a new report.
Run your SOC 2 program as files in Git.
Keep policies, controls, work, and evidence indexes in a repository your team and agents can inspect. Add optional hosted email and Slack reminders to keep work moving.
Frequently asked questions
Can an AI startup get a SOC 2 report?
Yes. An AI startup can pursue a SOC 2 examination for a defined service organization system. Management describes the service, its AI components, data, providers, commitments, and controls; an independent CPA firm agrees on the engagement scope and examines the controls.
Does a third-party model provider belong in SOC 2 scope?
A model provider is relevant when the scoped service depends on it or sends it customer information. Record the actual data flow, contract and configuration, provider responsibilities, and the controls your team operates. Management and the CPA firm decide how that provider is treated in the engagement; do not assume every vendor is a subservice organization.
Do prompts and model outputs count as customer data?
They can. Classify prompts, retrieved context, outputs, logs, and feedback by their actual contents and customer commitments. A prompt can carry customer information even when the application does not store the original document.
Does SOC 2 require an AI startup to prove model accuracy?
There is no universal SOC 2 model-accuracy test. The relevant controls depend on the scoped service, selected Trust Services Criteria, commitments, and risks. If the company promises a particular output process or quality measure, discuss its treatment with the CPA firm rather than assuming Security alone covers it.
Can FileGRC collect model-provider logs automatically?
No. FileGRC organizes structured SOC 2 records as JSON, long-form work as Markdown, and changes in Git. Your application, model provider, monitoring, and other source systems produce the operational records; your team collects or references fixed evidence from them.