How to Use AI Safely with Beneficiary and Donor Data
June 29, 2026
Last year a programme coordinator we work with sent a spreadsheet of her programme's beneficiaries to a public AI chat assistant and asked it to summarise the age distribution by district. It is the kind of request that feels harmless. The answer came back in seconds and was useful.
What had actually happened is that a file containing the names, ages, caste, household composition and village of roughly nine hundred children had been transmitted to a service she had no contractual relationship with, stored wherever that service stores things, retained according to a policy she has never read, and reviewed by humans at that company under terms she never agreed to. She could not delete it. She cannot tell you whether it was used to train anything.
This is not a story about a careless person. She is competent and careful. The tools gave no warning, the interface felt like a text box, and there was no policy telling her not to. That gap is the actual problem, and it is solvable this week.
What "the AI might keep it" means in practice
The honest answer, without going into anyone's internals, is that once data leaves your control you have given up control of it. Specifically:
- Transmission. Whatever you paste leaves your device immediately, over a connection you do not control.
- Retention. Different services retain conversation history for different periods, and some retain it by default so that a user can resume a conversation.
- Human review. Some providers reserve the right to review flagged content, and their definition of a flag is not written down precisely.
- Deletion is hard. Deleting a chat from your history does not necessarily delete it from their systems. A genuine deletion request under data protection law may be honoured for data you control and may not cover an account created by someone else.
- Terms change. The provider's terms are the ones in effect when you use the tool, and they may differ from what you agreed to when you signed up.
None of that means you should not use AI. It means the decision about what goes in has to be made before the paste, not after.
The three-tier classification that actually works
Data protection regulations like India's Digital Personal Data Protection Act 2023 distinguish between personal data and non-personal data, and between notice, consent and lawful purpose. Useful for a lawyer. Not useful at nine in the morning for someone who needs a summary.
For an operating rule, three tiers are enough.
| Tier | What it covers | Into a consumer AI tool? | Instead |
|---|---|---|---|
| Public | Published reports, your website copy, public programme pages, published statistics, your own marketing material | Yes, freely | Nothing needed |
| Internal | Draft reports, budgets, staff information, board papers, meeting notes, internal plans, funder templates | Only after checking the provider's data terms and settings, and with training opt-out enabled | An enterprise or business tier with an appropriate data agreement |
| Restricted | Anything identifying a specific beneficiary or child, case records, donor KYC and bank details, medical or disability information, photographs of beneficiaries, anything received under a signed MoU or grant contract | No | Anonymise first, or do the task manually |
The rule that follows is one sentence: public is fine, internal needs a check, restricted never goes in. That is short enough to write on a wall.
Restricted data is the strictest because it is data about people who did not choose to deal with you and cannot consent meaningfully. Children are the clearest case. A beneficiary's bank details for a small cash transfer are the second. A donor's PAN and bank account are the third, and they show up more often than anyone expects because the natural instinct is to ask an assistant to reconcile a donation spreadsheet.
Ten habits that keep you safe
- Anonymise before you analyse. Replace names with IDs before uploading anything. Keep the mapping on your own machine. "Bene_0001, district, age band, household size" is as useful for a summary as the names are, and carries none of the risk.
- Strip addresses and phone numbers. They are rarely needed for the analysis and frequently are the whole problem.
- Never paste a document a funder gave you. Grant agreements, MoUs and budgets usually contain a confidentiality clause. Your obligation to the funder survives whatever tool you use.
- Minimise before you upload. If the question is about age distribution, upload age bands, not ages. If it is about district, upload district totals, not row-level records.
- Check the settings for a training opt-out. Consumer tiers often have this off by default. Turn it on. It does not remove every retention issue, but it removes one large class.
- Use a business or enterprise tier for anything internal. These carry contractual terms about your data that consumer tiers do not.
- Keep a note of what went where. One line per project: which tool, which dataset, which date. If something does go wrong later, this is the difference between knowing and guessing.
- Understand that deletion is not guaranteed. Assume anything you send may be permanent and behave accordingly.
- Train the team rather than policing them. A rule people understand gets followed. A rule announced as prohibition gets worked around and hidden. Staff who tell you they used a tool are doing the right thing; staff who do not are the problem.
- Get the professional advice. This post is a practical operating rule, not legal advice. Your obligations depend on your registration status, your data volume, and your funder contracts. Talk to a lawyer about yours.
What anonymising looks like in practice
The coordinator at the start of this post had a simple fix available and did not know it. Her sheet had name, age, district, village, caste, and household composition. The question was about age distribution by district.
She deleted the name column. She converted age to a band. She dropped village because district was enough. What remained could go anywhere, and it answered her question.
The general version: before uploading, ask what the question actually requires, then send only that. Most of the time the requirement is far less than the file you are holding, and the gap between those two things is where the risk lives.
Where you genuinely need row-level analysis, do it in a tool you control, on a machine you control, and use the AI tool for interpreting a result you have already de-identified. That order matters. De-identify first, ask second.
A policy your team will actually follow
One page. Written by someone who uses AI themselves, which is what makes the difference between a policy and a prohibition.
- The three tiers, defined in terms your team recognises, not legal categories.
- Two or three named examples of each tier using your actual work.
- The one-sentence rule: public fine, internal after a check, restricted never.
- The anonymise-first habit, with the instruction that it happens before any upload.
- A note to keep a record of what was sent where.
- A named person to ask when it is genuinely unclear.
That is enough. The organisations that do well on this have policies that are read, understood and followed. The ones that struggle have policies that are twenty pages long and read once by a board that does not use the tools.
Want your team to build these habits rather than work around them? Book a digiSarathi AI workshop for your organisation.