Comparison of where AI PDF uploads go

Where Do Your PDF Uploads Go? AI Document Privacy Compared

A PDF upload can be stored, used for training, or sent to another lab. Here is how ChatGPT, Claude, Venice, and API routers handle the file.

Venice.aiVenice.ai

You drop a PDF into a chat box because retyping the contract is worse. The file can hold names, prices, medical details, or unreleased numbers. The icon does not tell you whether the company stores that file, trains on it, or forwards it to the lab that runs the model. Those are three different destinations, and which one applies depends on the product and the privacy mode you selected.

Short answer

A PDF upload goes to the company operating the chat, and often to the company running the model if those are not the same. For example, free ChatGPT can train on conversations unless you opt out, so a file you discuss there is part of that risk. Claude keeps consumer chats 30 days if you opt out of training and 5 years if you opt in.

On Venice, chat uploads accept PDF among other formats, up to 25MB per file and 50MB per message, and Private mode does not store the prompt or response. Pro TEE still supports file uploads. Pro E2EE is text only, so paste the text rather than attaching the file.

Why do PDF uploads leave your computer?

The viewer on your laptop renders the file, but the model cannot, so something has to send the text or the file itself to the machines running inference. From there, three things can happen, and products differ on each.

The company can keep a copy so the chat can reload tomorrow. For example, if the transcript appears when you sign in on a phone, the server has it. Venice stores chat history in the browser instead, and the default model Kimi K2.5 does not log conversations, which is a different storage design from an account archive.

The company can use the conversation, including what you asked about the file, to train models. OpenAI's data controls FAQ says this is the free-tier ChatGPT default, with an opt-out in settings.

Anthropic's consumer terms update requires Claude Free, Pro, and Max users to make that choice, and the retention that follows is 5 years if you opt in and 30 days if you opt out. Safety-flagged Claude chats may still be used for training. The Claude API sits outside that consumer training policy by default, though Anthropic still receives what you send the API.

A third company can receive the file because they host the weights. Venice Anonymous mode proxies GPT, Claude, and Gemini with identity stripped, and the provider still receives the content and stores it under its own policy. For example, if you upload a PDF in a Claude thread on Venice, Anthropic receives the conversation content.

OpenRouter does not store prompts or completions by default, per its data collection guide, though the upstream provider still receives the prompt. Poe shares chat contents with the underlying model providers.

Venice's own upload limits are 25MB per file and 50MB per message. The formats include PDF, DOCX, XLSX, CSV, and plain text, plus common images. Private mode, the default, stores no prompt or response and does not train on inputs. The agents page says the same to developers: you can send files, and Venice does not keep them as training data.

Pro TEE is the hardware-isolated option, and Venice's launch notes say file uploads still work there. Pro E2EE does not accept file uploads at all, since it runs text models only, with no web search and no memory, so copy the passage into the message box.

Character documents have a smaller, separate cap: a character can take an uploaded document up to 5MB, which is not the chat attachment limit. Character chats are still stored on your device, per the Characters page.

None of this is a SOC 2 or HIPAA program. If the PDF is regulated health or payment data and your counsel wants a named certification, Venice does not advertise SOC 2, HIPAA, ISO 27001, PCI, or FedRAMP. What you can check for yourself is storage, training, and who receives the file. You can compare the products built for that analysis work in private AI document analysis.

How to keep a PDF more private when you upload it to AI

1. Upload an excerpt, or a PDF you already redacted

Delete signature blocks, account numbers, and customer names in the PDF itself before you upload it, or paste the one clause instead. The model can explain a redacted section without the appendix of email addresses, and this works in every product, including the one you are not ready to leave.

  • Best for: A single question about a long document.

2. Check the training control before the upload finishes

On free ChatGPT, opt out of training first. On Claude, make the training choice knowing the retention that comes with it, either 30 days or 5 years. On Grok, consumer prompts can be used for training unless you opt out, and Private Chat is excluded where it is available, per xAI's FAQ. Turning training off does not prove the file was never stored, but it does stop the default "use this to improve the model" path those pages describe.

  • Best for: Accounts you will keep using for low-sensitivity PDFs.

3. Match the model to the privacy mode, then upload

On Venice, confirm you are on a Private model such as Kimi K2.5 before you attach the PDF, because selecting Claude or GPT instead sends an Anonymous request and the lab receives the content. Venice chat history stays in the browser, and the privacy page describes what each of the four modes does.

If you are a developer sending extracted text through the API, the same split applies, and the Venice API page is the place to create a key. Default retention is zero.

  • Best for: Uploads where the file is the sensitive object, not just the question around it.

4. Use Pro TEE when the file should stay in an enclave

TEE runs inference in a hardware-isolated enclave where GPU providers cannot access the prompts, and you can check remote attestation yourself. Venice's launch description says this mode still supports file uploads, web search, memory, and auto-routing. Replies can be slower, and fewer models are available. It is a Pro feature at $18/month to start, and it is the upload path when you still need to attach a file and Private mode's contractual no-storage rule is not enough.

  • Best for: PDFs you want isolated in hardware, with an attestation report you can look at.

5. For E2EE, paste text instead of attaching the PDF

E2EE encrypts the prompt on your device so Venice cannot read it, and decryption happens only inside a verified TEE. The mode does not include the file tools, web search, or memory that TEE keeps, so copy the relevant pages into the message. If the document is 25MB of scans, E2EE is the wrong mode, and you want TEE or a redacted Private upload.

  • Best for: A passage of text, not a full scanned binder.

The best AI tools for analyzing PDFs

1. Venice

You can attach a PDF in chat, up to 25MB, on a Private model that does not store the prompt or response and is not used for training. Use Pro TEE when you want file upload inside an enclave, and use E2EE for pasted text. The free plan is 10 text prompts a day, and there is no phone know your customer (KYC) check.

  • What it does differently: The upload limit, the storage rule, and the exception for third-party models are all stated, rather than folded into one "private PDF" claim.
  • Pricing: Free, Pro $18/month, Pro Plus $68/month, Max $200/month.
  • Best for: Confidential PDFs when you can stay on a Venice-hosted private or TEE model.

2. Anthropic API

The consumer Claude retention rules are not the API rules, and commercial Claude, including the API, is excluded from that training policy by default. The PDF's text still goes to Anthropic, so this is the right pick when a contract requires Claude with no consumer training, and the wrong one when your requirement is that the operator never receives the file.

  • What it does differently: A commercial training exclusion, with the lab still receiving the content.
  • Pricing: Consumer Pro is $20/month, and API tokens are billed separately. On Venice, sending the same work to Claude goes through Anonymous mode rather than the direct API contract.
  • Best for: Teams required to use Claude under Anthropic's commercial terms.

3. OpenRouter

Developers who extract text from a PDF and send it as a prompt get OpenRouter's default: the prompt is not stored, metadata is kept, and the upstream model provider receives the text. The data-collection guide describes no separate handling for files, so the file's contents become the prompt and follow the prompt rules.

  • What it does differently: One key, many models, with logging of the prompt text turned off unless you opt in.
  • Pricing: 50 free requests a day, then provider rates with no markup. Card credit fees are 5.5%.
  • Best for: Pipelines where you control the extractor and you will read the upstream policy for that model ID.

OpenRouter does not replace Venice TEE or E2EE, because the lab can read the extracted text.

Is it safe to upload a PDF to ChatGPT?

It is a poor default for a confidential file, since free-tier conversations can be used for training unless you opt out and the account holds the chat. If you stay, opt out, redact the PDF, and send a short excerpt. For a file that should not be stored or used for training, use Venice Private or Pro TEE. If you are building this into code, the developer picks are ranked in best LLM APIs for private document processing.

Does Claude store uploaded files?

Consumer Claude retention follows your training choice: 30 days if you opt out, 5 years if you opt in. Those windows are the published retention for those chats rather than a separate rule for uploaded files. The API is excluded from the consumer training policy by default, and Anthropic still receives API inputs. Safety-flagged consumer chats may still be used for training.

Where does a PDF go on Venice?

It goes to the model you selected, under that model's privacy mode. On Private, the prompt and response are not stored, Venice does not train on inputs, and chat history stays in the browser. On Anonymous third-party models, the provider receives the content. Chat attachments allow PDF up to 25MB per file, Pro TEE still supports file uploads, and Pro E2EE is text only.

What file types can I upload for private analysis?

Venice's documented formats include PNG, JPEG, WebP, GIF, PDF, DOCX, XLSX, CSV, JSON, XML, YAML, Markdown, and plain text. The caps are 25MB per file and 50MB per message. A character document is a separate 5MB upload. Those caps are the product limits, and if you are sending files from code, check the API docs for the limits there.

Will an API aggregator keep my PDF?

OpenRouter keeps metadata and, by default, not the prompt text. The provider running the model still receives the prompt, which is where your extracted PDF text went, and Poe shares chat contents with the underlying providers. Venice's API default is zero data retention, with the same Anonymous exception when you call a frontier model. You can see how the main options compare in best LLM API aggregators.

Can end-to-end encryption cover a PDF upload?

Not on Venice's E2EE mode, which runs text models only with no web search and no memory. TEE is the Pro mode that still supports file uploads, inside a hardware enclave and with remote attestation. Use TEE when you need the PDF attached and you need hardware isolation, and use E2EE, with the text pasted in, when Venice should not be able to read the words.

For the next confidential PDF, redact it, select a Private model, and attach it in Classic Chat. If the file should sit in an enclave, use Pro TEE, and if you only need a paragraph checked, paste it into E2EE.

Back to all posts