Skip to content
hypris.ai
AI and your data

Where your company data goes when you send it to AI

What happens to a prompt sent to a model: who processes it, in which region, and how long they keep it. Plus the questions to put to a vendor in writing.

6 min read

The question "where does our data actually go" tends to arrive late — months after someone started pasting quotes into a chat window, when a larger customer sends a security questionnaire with it on the third line.

The answer is less dramatic than it sounds and less simple than you would like. A prompt goes not to one place but through several, each with its own rules. Below: that route, and the questions to put to a vendor in writing.

The route a prompt takes: three stops

When an employee types a question into an AI tool, the data crosses three layers. Even where you see one logo, three separate companies usually sit behind it.

The application you work in. The work platform, the CRM, the messenger. It collects the question and attaches context: fragments of records, attached documents, earlier turns of the conversation. So the model rarely gets the question alone, and that added scope is invisible from the interface.

The model vendor. The application sends that package to the company that runs the model — very often not the company you pay. Your contract is with one party; the contents of the prompt reach another.

The infrastructure the model runs on. The model executes in a specific data centre, in a specific country, frequently at a third cloud provider. Nobody raises this layer in a sales call: formally it is not your counterparty.

Every question about data therefore has to cover the whole chain, not the first link. "We don't do that" can be true and worthless at the same time, if it covers only layer one.

Training a model versus processing a prompt

This is where most conversations about AI security fall apart: three separate things get treated as one.

Processing a prompt. The content goes into the model, the model computes an answer, and it leaves working memory. Nothing remains inside the model — its knowledge does not change because you sent it something.

Training on your data. Prompt content joins the set used to train the next version of the model. Only then can a fragment of your information surface in an answer given to someone else — the fear people mean when they ask whether company data will "leak into the AI".

Retaining prompts. The most commonly overlooked one. Vendors typically hold prompt content for a period: diagnostics, abuse detection, billing. Not training and no change to the model, but a copy of your data outside the company, on a disk you cannot inspect.

So "we don't train on your data" does not answer "does our data stay with you". Both need asking; the risk profile differs.

Retention: how long, and in whose hands

The most concrete part of the topic: it comes down to a number of days.

How many days, and what for. "For as long as necessary" is not an answer. You want a number and the reason behind it.

Whether it can be reduced to zero. Some business plans allow no retention of prompt content, but you have to request it; it is not the default.

Who can access what is retained. Usually a narrow group of vendor staff, in defined circumstances. Worth knowing which.

Two clocks, not one. The application deletes conversation history on its own schedule while the model vendor keeps its copy on another. Two independent periods, both worth asking about.

Processing region and subprocessors

The region is where the servers computing your answer physically stand. It does not follow from where the company selling you the tool is registered, nor from the language you write your prompts in.

Guaranteed or merely default. Guaranteed means written into the contract, stable under load. Default means it is that way today and may be different tomorrow.

Who else is in the chain. The subprocessor list should be public, current, and covered by an obligation to notify you of changes. If a vendor cannot say where it is published, that is an answer in itself.

The questions to ask in writing

In writing, because a reassurance given in a sales call is worth nothing a year later when an auditor asks. An email is enough.

  1. Is our data used to train models — yours or your subprocessors'? The question has to cover the whole chain.
  2. How long do you retain the content of prompts and responses, and why? A number of days, not a policy description.
  3. In which region are prompts processed, and is that contractually guaranteed?
  4. Who are the subprocessors involved, and where do I find the current list?
  5. What does deletion of our data cover? Ask explicitly about backups, logs, and the timeframe.

Keep the answers in one file next to the contract. When a customer's security questionnaire arrives, you fill it in on the spot instead of starting a round of emails to vendors.

What you can do on your side

Whatever the vendor answers, part of the control stays with you.

Establish what the application attaches to a question. How much context travels with a prompt is often a setting, not a fixed property of the tool. If it can be narrowed to one area instead of the whole database, that is the cheapest change on this list.

Name the categories that never leave. Payroll data, documents under confidentiality, correspondence with a lawyer. A short list of what goes into no chat window works better than a general call for caution.

Refresh the answers at renewal. Vendor policies shift along with plans and product versions. What you were given in writing a year ago describes the position a year ago.

Write down who can enable new integrations. One named person approving each new connection is usually the whole procedure a small company needs.

Summary

A prompt sent to AI passes through the application, the model vendor, and a specific data centre — each layer with its own rules. Training a model, processing a prompt, and retaining its content are three different things, asked about separately. Five questions, emailed and filed in one place, settle the matter for years.

This is one piece of a larger picture — the rest, including whose permissions the model inherits, is in the guide on connecting AI to company data.

Frequently asked questions

Is data sent to an AI model used to train it?
That depends on the plan and the specific vendor, so the only answer worth having is one you get in writing. Bear in mind that training a model and retaining prompts are two different things — a vendor can decline to train on your data and still hold a copy of it long after the answer was returned.
Where is a prompt sent to AI physically processed?
In the model vendor's data centre, which need not sit in the same country as the company selling you the application. Processing region is often a business-plan setting rather than a default, so ask about it directly and get the answer written into the contract.
What should you ask an AI vendor before signing?
Five things: whether your data trains models, how long prompts are retained and why, the processing region, the list of subprocessors, and the deletion procedure. Ask all five about the whole chain of suppliers, not only the company whose contract you are signing.

Read next

Hypris

See what this looks like in practice

Hypris is a work platform with a built-in AI agent that acts only after you approve it. Your whole company, its data and its agents in one place — field crews included.

No strings attached. We show a working product, not a slide deck.