Third-Party Providers
Every Chat-lab.ai tool delegates its heavy lifting to an external service. If you work with unpublished manuscripts, confidential recordings, student work, client materials, or any other potentially sensitive data, where that data goes matters as much as what the tool does. This page summarizes each provider's data processing practices, including free-tier vs. paid-tier distinctions that are easy to miss, and how the apps are designed to minimize what providers ever see. Given the current level of site usage, Chat-lab.ai apps rely on free tiers from all these providers.
Policies verified against official sources in July 2026.1 Providers change terms; the linked pages are authoritative. When in doubt, check with your institution's data-governance office.
How the apps minimize data exposure
- Deterministic sources first. Bibliobot resolves papers through CrossRef and arXiv, non-profit scholarly registries that receive only an identifier (a DOI or arXiv ID), never your document. An LLM sees your text only when no identifier can be found.
- Minimal excerpts, never full documents. When an LLM is needed, Bibliobot sends only the first ~4,000 characters — enough for title-page metadata, not the paper. OCR submits only the first 5 pages of scanned files.
- Bring your own key to control the tier. Several providers treat free-tier and paid-tier data very differently (see the table). Entering your own paid-tier key in a tool's settings routes your content under your account's policy. Keys live only in your browser session and are never stored.
- Nothing persists here. The apps themselves keep nothing. See Privacy & Data Retention.
At a glance: model training on your data
| Provider | Free tier | Paid tier | Default retention1 |
|---|---|---|---|
| Mistral | Trains by default (this app's workspace has opted out) | Not trained (opted out by default) | 30 days (abuse monitoring) |
| AssemblyAI | Trains — no opt-out on free accounts | Trains by default, self-serve opt-out | Audio ≤48h; transcripts 72h default |
| Groq | Not trained, on any tier | Not trained | None by default (≤30-day troubleshooting logs) |
| OpenAI | (API is bring-your-own-key) | Not trained by default | ≤30 days (abuse monitoring) |
| Google Gemini | Trains + possible human review (why it's last-resort only) | Not used for improvement | Per Google's terms |
| Jina AI | Not trained; stated explicitly | Not trained | ~5-min cache; logs ≤7 days |
| Firecrawl | Policy does not address training | Policy does not address training | Crawled pages cached (~2-day default) |
| CrossRef | Sees only DOIs; nothing to train on | — | — |
| arXiv | Sees only arXiv IDs | — | — |
AI & speech providers
Mistral AI
Used by: Bibliobot (default LLM extraction).
Receives: the first ~4,000 characters of papers with no detectable identifier.
A French company with an EU-first hosting posture, often the friendliest choice for GDPR-minded institutions. The tier distinction is important: on the free tier, inputs and outputs are used for model training by default, with an opt-out available in the account console's Privacy settings; on the paid (Scale) plan, training is off by default. API data is otherwise retained for 30 rolling days for abuse monitoring.
This app has opted out
The Chat-lab.ai workspace has disabled model training in Mistral's privacy settings. Text these apps send to Mistral for LLM extraction is not used to train Mistral's models.
Privacy Policy · Data Processing Addendum · Training opt-out · Usage tiers · Trust Center
Groq
Used by: Audio Insights (default summary generation, serving open-weight models).
Receives: your transcript text plus any context notes you add.
Not to be confused with Elon Musk's Grok, Groq is an inference provider hosting various open-source models. Because Groq does not create its own models, it has the strongest default privacy posture of the LLM providers here: Groq does not train on customer data on any tier, and by default retains nothing from inference requests (narrow exception: up to 30 days of troubleshooting/abuse logs).
Your data in GroqCloud · Services Agreement · Rate limits · Trust Center
OpenAI
Used by: Audio Insights and Bibliobot, only when you supply your own OpenAI API key.
Receives: transcript text (Audio Insights) or paper excerpts (Bibliobot).
OpenAI's API has no free tier, so these apps only ever send data to OpenAI when you supply your own key. API data is not used for training by default regardless of spend, which is why an OpenAI key is the "bring your own key" recommendation in these apps. Inputs/outputs are retained up to 30 days for abuse monitoring; zero-retention and regional (including EU) data residency exist for qualifying use cases.
gpt-oss ≠ OpenAI's servers
Audio Insights' free default summarizer is gpt-oss-120b, an open-weight model published by OpenAI but hosted by Groq. That data is governed by Groq's no-training, minimal-retention policies; OpenAI never receives it.
API data privacy · Privacy Policy · Your data (developer docs) · Rate limits & usage tiers
AssemblyAI
Used by: Audio Insights (transcription and speaker diarization).
Receives: your full audio/video file.
The provider academics should read most carefully. AssemblyAI's terms allow customer data to be used to train their models: free accounts cannot opt out; paid accounts train by default but can opt out self-serve in the dashboard (not retroactive). Retention is short: uploaded audio is deleted within 24–48 hours, and transcripts default to a 72-hour lifetime (configurable down to 1 hour, or deletable immediately via API).1 An EU (Dublin) endpoint exists for data-residency needs.
This app deletes transcripts immediately
Audio Insights deletes each transcript from AssemblyAI via their API as soon as it has been retrieved, so the remote copy doesn't sit out the 72-hour default window. Your transcript then exists only in your browser session. (If you use your own AssemblyAI key, a toggle in the app lets you choose: deletion is on by default, or turn it off to keep transcripts in your own account.)
Sensitive recordings
For interviews under IRB protocols, confidential meetings, or anything with identifiable participants, use your own paid AssemblyAI key with the training opt-out enabled (new accounts include free credits), or don't upload the recording.
Privacy Policy · Terms of Service · Training opt-out · Data retention · Security
Google Gemini
Used by: Bibliobot, only as the OCR fallback for scanned PDFs, and only when you haven't supplied your own OpenAI key (which takes over OCR entirely).
Receives: the first 5 pages of a PDF, and only when the file has no readable text layer (or every other extraction path came up empty).
Gemini is deliberately the last resort in the pipeline, for a simple reason: Google's free-tier terms are the least agreeable of any provider here. For unpaid API usage, Google states that it "uses the content you submit to the Services and any generated responses to provide, improve, and develop Google products and services and machine learning technologies," that "human reviewers may read, annotate, and process your API input and output," and it warns directly: "Do not submit sensitive, confidential, or personal information to the Unpaid Services." (Paid-tier Gemini API usage is not used for improvement.) That's why this app sends Gemini as little as possible, as rarely as possible: most papers never reach it (they resolve through registries, native text extraction, or Mistral first), and the ones that do are truncated to the first 5 pages.
Scanned + sensitive = don't upload
If a scanned document is confidential or contains sensitive information (e.g., an unpublished manuscript, a student record), don't upload it. Scanned text with no identifiers (e.g., DOI numbers) is exactly the case that falls through to Gemini's free tier, where training on your data and human review is possible.
Gemini API Terms · Rate limits · Pricing/free tier
Extraction & crawling providers
Jina AI (Reader)
Used by: PDF-MD, URL-MD, Sitemap Crawler.
Receives: the URLs you convert, and full PDFs you upload for extraction.
Jina (acquired by Elastic in October 2025) states explicitly that it never trains on API inputs or outputs. Identical requests are served from a ~5-minute cache, and server logs are deleted within seven days. The apps authenticate with their own key, which raises Jina's rate limit substantially and is why batches move quickly.
Reader & rate limits · Legal / privacy
Firecrawl
Used by: Sitemap Crawler (site mapping only; page content is fetched via Jina).
Receives: the site URL you map; the resulting link list.
Firecrawl does not create or host AI models, and its policies do not address model training either way. Crawled pages are cached on their infrastructure by default (~2 days), though this app only uses the mapping endpoint on URLs that are public by definition. The free plan's monthly allowance is modest; academics can get a much larger one through the student program.
Privacy Policy · Terms · Pricing · Student program
Scholarly registries (no content ever sent)
CrossRef
Used by: Bibliobot (DOI and SSRN metadata lookups).
Receives: only a DOI.
A not-for-profit membership organization running core scholarly infrastructure. It never sees your document — a DOI in, public bibliographic metadata out. The apps identify themselves to CrossRef's "polite pool" per their etiquette guidelines.
arXiv
Used by: Bibliobot (arXiv metadata lookups).
Receives: only an arXiv ID.
An independent non-profit (formerly hosted at Cornell) with no commercial data interest; API logs are used operationally, never sold. The apps follow arXiv's API etiquette by pacing their requests.
API Terms of Use · Privacy Policy · About
-
Specific figures on this page (retention periods, cache windows, and similar) are reported as of July 2026 and were collected by Claude Code from each provider's published documentation. While I have reviewed the contents of this page and found it to be generally consistent with my own manual research, inaccuracies are possible.
Moreover, model providers can (and do) change their terms of service often. To that end, I use AI to maintain this page updated, but you should independently verify the information before submitting sensitive data to any of these services. ↩↩↩