Insights · Buyer Guides
AI Agent Data Privacy: What to Verify Before You Share Customer Records
By the Augex team · 13 min read · 2026-08-20
AI Agent Data Privacy: What to Verify Before You Share Customer Records
Every AI agent that does useful work needs your data. Support triage needs your tickets. Lead scoring needs your CRM. Invoice extraction needs documents with customer names, addresses and bank details on them.
The moment you paste a customer list into an agent, you have made a legal decision, not just a technical one. Under the GDPR you remain the controller, and the vendor becomes your processor. That relationship carries obligations you cannot delegate away, and they attach whether or not anyone in your company thought about it first.
This article is the due diligence pass to run before that first upload. It is written for a business without a legal team or a privacy officer, because that is who usually ends up making this call. It covers what to ask, what a good answer sounds like, and the three answers that should end the evaluation on the spot.
One boundary worth stating up front: this is a practical checklist, not legal advice. If you operate in health, finance or children's services, the sector rules on top of this are real and you need a lawyer, not a blog post.
Why this is your problem and not the vendor's
The instinct is to assume that a vendor with a polished security page has taken care of compliance. That is not how the law allocates responsibility.
Under GDPR Article 28, a controller may only use processors that provide sufficient guarantees of appropriate technical and organisational measures, and the arrangement has to be governed by a contract that specifies subject matter, duration, nature and purpose of processing, the type of personal data, and the categories of data subjects (GDPR Article 28). The duty to check sits with you. If the vendor turns out to be careless, "they told us it was fine" is not a defence, because the article requires you to have verified the guarantees before processing began.
Article 32 adds the security dimension: both controller and processor must implement measures appropriate to the risk, taking into account the state of the art and the severity of risk to data subjects (GDPR Article 32).
So the practical translation is blunt. You cannot outsource the decision. You can only do the checking properly or skip it and hope.
The nine checks
Work through these in order. The first four are contractual and can be settled by reading documents. The last five need a direct answer from a human at the vendor, which is itself a useful test of how they operate.
1. Is there a data processing agreement, and does it name what it must name?
Ask for the DPA before the trial, not after. A real one specifies the categories of personal data, the purposes, the duration, and the obligations on the processor to act only on documented instructions.
What a good answer looks like: a linked, versioned DPA you can read without asking, plus a named contact for signing.
What should worry you: a DPA that exists only as a paragraph inside the general terms of service, or a vendor who says a DPA is unnecessary because the data "is not really personal." Names, email addresses and account identifiers are personal data. That answer tells you they have not done this before.
2. Who are the subprocessors, and do you get notice of changes?
An AI agent is almost never one company. There is the agent vendor, the model provider behind it, probably a cloud host, possibly a vector database, sometimes an observability tool that logs prompts and responses. Each of those is a subprocessor touching your customer data.
Article 28 requires the processor not to engage another processor without prior specific or general written authorisation, and where general authorisation applies, the processor must inform the controller of intended changes and give you a chance to object (GDPR Article 28).
What a good answer looks like: a public subprocessor list with company names, roles and locations, plus an email subscription for change notices with a notice period stated in days.
What should worry you: "we use industry standard providers" with no names. You cannot assess a risk you are not allowed to see.
3. Is your data used to train models, and is that the default or an option?
This is the single question buyers most often forget, and the one with the least reversible consequence. Data that has been absorbed into a training set cannot be meaningfully recalled by a deletion request.
Ask it in three parts, because vendors answer different parts and let you assume the rest:
- Is customer data used to train or fine tune the vendor's own models?
- Is it passed to the underlying model provider in a way that permits training on it?
- Is human review of inputs and outputs performed for quality purposes, and by whom?
What a good answer looks like: training is off by default for business accounts, the model provider is on a no training API tier, human review either does not happen or happens only on data you explicitly flag, and all three points are stated in the contract rather than in a support article.
What should worry you: an answer that covers only the vendor's own models and goes quiet about the model provider behind them.
4. What is the retention period, and what triggers deletion?
Prompts, outputs and uploaded files are usually retained somewhere for debugging and abuse monitoring. That is legitimate. Indefinite retention with no stated period is not.
What a good answer looks like: a specific number of days for prompt and output logs, a separate number for uploaded files, a documented deletion process on account termination, and confirmation of whether deletion covers backups and on what timeline.
What should worry you: "we retain data as long as necessary to provide the service." That phrase is doing no work at all.
5. Where is the data processed, and what covers any transfer?
If your customers are in the EU or UK and the agent runs on infrastructure elsewhere, you need a lawful transfer mechanism. Chapter V of the GDPR governs this, and transfers to a third country require either an adequacy decision or appropriate safeguards (GDPR Article 44). In practice most vendors rely on Standard Contractual Clauses, which the European Commission publishes and maintains (European Commission on SCCs).
What a good answer looks like: the vendor names the processing regions, states whether regional pinning is available on your plan, and points to executed SCCs in the DPA.
What should worry you: a vendor who has heard of SCCs but cannot say whether they have signed any, or a vendor who offers EU residency for storage while inference still runs elsewhere. Storage residency and processing residency are different claims, and the second one is what matters when the model runs.
6. Can you get your data out, and can you get it deleted on request?
You will need this twice: when a customer exercises their rights, and when you leave the vendor.
What a good answer looks like: a self serve export in a documented format, a deletion endpoint or a support process with a stated turnaround, and clarity on whether deletion of an individual record is possible or only wholesale account deletion. If a customer asks you to erase their data and your agent vendor can only wipe everything, you have a problem you want to discover now.
7. Does the agent make decisions about people without a human in the loop?
This one changes the compliance picture significantly and is easy to trip over accidentally. GDPR Article 22 gives data subjects the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects concerning them or similarly significantly affects them (GDPR Article 22).
Screening job applicants, scoring creditworthiness, deciding whether to close an account: if an agent does any of these with no meaningful human review, you are in the scope of that article. Note the word "meaningful." A human who rubber stamps every agent output at a rate of four seconds each is not providing review, and a regulator will read it that way.
What a good answer looks like: you can describe exactly where the human decision point sits and show that the human has the information and the authority to disagree.
8. What is the incident notification commitment, in hours?
If a vendor suffers a breach involving your customer data, your own notification clock starts, and you cannot start it if nobody told you.
What a good answer looks like: a stated notification window in hours, a named channel, and a commitment to provide the detail you need to assess the risk rather than a vague heads up.
What should worry you: "we will notify you without undue delay" and nothing more. That phrase is lifted from the regulation and, on its own, commits the vendor to nothing you can measure.
9. What can you actually verify rather than take on trust?
Everything above is an assertion until something backs it. Three artifacts carry real weight, in ascending order of effort for the vendor to obtain:
- A completed security questionnaire, which at least forces specific answers.
- A SOC 2 Type II report, which reports on the operating effectiveness of controls over a period rather than at one moment. The distinction between Type I and Type II is the reason to ask which one they hold (Secureframe on SOC 2, AICPA on SOC reporting).
- A penetration test summary from a named third party, dated within the last twelve months.
For a small vendor, having none of these is not automatically disqualifying. Being unable to describe what they do instead is.
A worked example: support triage on a real ticket queue
Take a ten person company evaluating an agent to categorise inbound support email and draft replies. The queue contains names, email addresses, order numbers, occasionally a partial card number a customer pasted in without being asked.
Running the nine checks produces a specific picture rather than a general feeling:
- DPA: published, versioned, signable. Pass.
- Subprocessors: model provider and cloud host named, observability vendor also named. Pass, and the observability vendor is the one to look at, because that is where prompt text tends to sit in plain form.
- Training: off by default, model provider on a no training tier, human review only on flagged items. Pass.
- Retention: 30 days for prompts and outputs, 7 days for attachments, backups purged within 35 days. Pass, and short enough that a breach exposes a bounded window.
- Residency: EU processing available on the business plan, inference included, SCCs executed. Pass.
- Export and deletion: JSON export self serve, per record deletion by API. Pass.
- Automated decisions: agent drafts, human sends. No Article 22 exposure as configured, and worth writing down that the exposure appears the moment someone enables auto send.
- Incident notification: 72 hours, named security contact. Acceptable, though 24 hours would be better and is worth asking for.
- Evidence: SOC 2 Type II from last year, pen test summary from eight months ago. Pass.
That queue also surfaces a problem the vendor cannot solve: customers occasionally paste card numbers into support email. That is your data hygiene issue, not theirs, and the answer is redaction before the agent sees the ticket. Due diligence on a vendor often exposes work on your own side, and skipping the exercise means skipping that discovery too.
The three answers that should end the evaluation
Most weak answers are negotiable. These three are not:
"Your data may be used to improve our models, and there is no way to opt out." For a business account handling customer records, this is disqualifying on its own. The consequence is irreversible.
"We cannot tell you which subprocessors we use." You cannot fulfil Article 28 obligations you are not permitted to see, and a vendor confident in their stack has no reason to hide it.
"We do not sign DPAs." For a vendor processing personal data on your instructions, this is not a policy position. It is a signal that the compliance work has not been done, and everything else they tell you inherits that doubt.
Where this sits alongside security and regulation
Data privacy overlaps with two adjacent areas, and it is worth being clear about the boundary.
Security is about whether the agent can be made to do something it should not. The OWASP GenAI project maintains a catalogue of agentic threats and mitigations covering prompt injection through retrieved content, unintended tool invocation and privilege escalation through chained actions (OWASP GenAI Security Project, OWASP Top 10 for LLM Applications). If an agent can be talked into exfiltrating the records you gave it, your DPA does not help.
Regulation is about obligations attaching to the AI system itself. The EU AI Act introduces transparency duties for certain systems, including disclosure when people interact with an AI system (EU AI Act Article 50, European Commission AI framework). If your agent talks to customers directly, that is a live question for you and not only for the vendor.
For structuring the wider risk picture, the NIST AI Risk Management Framework is the reference most teams reach for, with more specific control guidance in NIST AI 100-2 (NIST AI RMF, NIST AI 100-2e2025).
FAQ
How long should this take? About two hours for a small team, most of it spent reading a DPA and waiting for email replies. That is a reasonable cost against granting software access to your entire customer base.
We are a small company. Does anyone really check? Enforcement attention scales with harm, not with company size, and the practical exposure arrives sooner than a regulator does. It arrives when an enterprise prospect sends you their vendor questionnaire and you cannot answer the subprocessor question about your own stack.
Does anonymising the data first solve this? Sometimes, and it is genuinely worth doing where the agent's job permits it. Be careful with the word though: removing names from records that still contain order numbers, timestamps and postcodes is pseudonymisation, and pseudonymised data remains personal data under the GDPR. True anonymisation, where re-identification is not reasonably possible, is harder than it looks.
What if the vendor is a solo builder on a marketplace rather than a company? The same nine questions apply, and the answers will be thinner. What matters is whether they can describe their stack precisely and tell you which subprocessors sit behind it. Precision from a solo builder is worth more than a polished trust page from a company that cannot name its own model provider.
Should any of this be in the contract rather than in email? Yes, for training use, retention, residency and incident notification. An email from a support agent does not survive that person leaving, and these four are the ones you would actually need to prove later.
Sources
- GDPR Article 28: Processor
- GDPR Article 32: Security of processing
- GDPR Article 22: Automated individual decision making
- GDPR Article 44: General principle for transfers
- European Commission: Standard Contractual Clauses
- ICO: Data sharing code of practice
- EU AI Act Article 50: Transparency obligations
- European Commission: Regulatory framework for AI
- OWASP GenAI Security Project: Agentic AI Threats and Mitigations
- OWASP Top 10 for Large Language Model Applications
- NIST AI Risk Management Framework
- NIST AI 100-2e2025: Control overlays for securing AI systems
- Secureframe: What is SOC 2
- AICPA: SOC reporting guidance
Which specialist task does your team keep pushing to 11pm? Start there.
Related reading