Limited time: Save up to 33% on every planView pricing
Voibe Logovoibe Resources

AI Tool Privacy Tracker

What every major AI tool actually does with your data. Training behavior, retention, and on-device support — verified against primary sources, with a separate row for consumer and business tiers because the answer is different.

Last updated September 10, 202657 tools trackedNext review October 10, 2026

Recent Changes

Dated policy shifts that changed what a tool does with your data. Each entry is linked to a primary source.

  1. Tracker expansion (23 new rows, new Clinical Documentation table)

    The tracker grew from 38 to 61 tools. New Clinical Documentation table for ambient AI scribes: Abridge, Freed, Heidi Health, Nabla, Suki, and DeepScribe. New meeting notetakers: Fathom, tl;dv, Read AI, Notta, and Gemini in Google Meet. New coding agents: OpenAI Codex, Gemini CLI, and Kiro. New dictation tools: Handy, OpenWhispr, FluidVoice, VoiceDash, Paraspeech, Voicy, Blip AI, WisprType, and Dictaflow. Every new cell was verified against the vendor's own policy, terms, security page, or docket on September 10, 2026; cells a vendor does not publish are marked as not published rather than guessed, and cells that rest on secondary reporting (where a vendor page blocked automated retrieval — Blip AI, parts of Notta and Read AI) say so in the cell text.

    Source: This page (methodology)
  2. tl;dv

    tl;dv published an updated privacy policy dated September 9, 2026 (Tldx Solutions GmbH, Aachen). It states Customer Content is not used "to train, fine-tune, or improve foundation models, large language models, or other generative AI models" by tl;dv or its AI providers, names Anthropic and Google Vertex AI as AI processors and AssemblyAI and ElevenLabs for transcription, and keeps the 3-month auto-delete for free-plan recordings.

    Source: tl;dv privacy policy
  3. Granola

    Granola published a new privacy policy effective September 8, 2026. The consumer training default is unchanged — "We only use de-identified data to train AI models, which you can opt-out of within your Granola account settings" — while Enterprise Workspace products are described as having admin-enforced settings that are opted out by default. On the Chamberlain docket (Judge Edward M. Chen), Granola was served August 5 and the parties stipulated on August 20 to extend the response deadline; no motion to dismiss has been filed.

    Source: docs.granola.ai (privacy policy)
  4. Voibe (speech-to-text API + MCP server)

    Voibe launched a batch speech-to-text REST API at api.getvoibe.com and a hosted MCP server (api.getvoibe.com/mcp) for agents. Both run on the same zero-retention pipeline as the dictation apps: audio is transcribed by Whisper large-v3-turbo (open-source) on zero-retention providers, speaker separation uses pyannote community-1, and summaries use GPT-OSS 120B (open-weight). Per the API page, the recording is deleted the moment the transcript exists, is never used to train models, and every read is scoped to the account that created the job; transcripts stay readable for 24 hours after completion and are then deleted. Pricing is prepaid ($10 / 2,000 min through $100 / 24,000 min), billed per second only when a job completes, with 15 free minutes on every new account. This is a new product surface, not a change to the dictation apps' posture, so Voibe's matrix scores are unchanged; the business-tier row now notes the API.

    Source: getvoibe.com (speech-to-text API)
  5. Cursor

    Cursor removed two sentences from its Data Use & Privacy Overview that had described how codebase embeddings and metadata (hashes, file names) are handled and stored during indexing; the page was re-stamped September 3, 2026 with no replacement disclosure. Privacy Mode semantics are otherwise unchanged. The same page now notes that models requiring provider-side retention (Claude Fable 5.1 and Fable 5) are blocked for Privacy Mode and Enterprise workspaces until an admin approves them. Separately, the August 13, 2026 Terms of Service replaced arbitration with mandatory litigation in Texas. Cursor's Track record scores were lowered by one point on both tiers for the transparency removal.

    Source: conductatlas.com (Data Use diff)
  6. Abridge

    Abridge published a revised privacy policy (Last Updated: August 27, 2026). The archived 2022 text — "We use de-identified data for research and development of new products or tools, to refine our algorithms and machine learning applications" — is gone; the new policy says Customer Data "is governed exclusively by the terms of our agreements with those customers, including our Business Associate Agreements (BAAs)." The training permission moved from a public document into the contract.

    Source: Abridge privacy policy
  7. Microsoft Copilot (consumer)

    Microsoft shipped an updated Microsoft Copilot app for web, desktop, and mobile. Microsoft Support now splits its consumer privacy documentation into pages for the older app and the updated app, and the training controls are exposed as two separate toggles under Profile → Privacy: "Training on conversation activity" and "Training on voice conversations." Opting out excludes future conversation activity from model training but, per Microsoft, conversations may still be used for other product improvements, advertising, digital safety, security, and compliance purposes. Personalization can stay on with training off. The default-on training posture for signed-in consumer users is unchanged, as is Microsoft 365 Copilot's contractual no-training commitment.

    Source: support.microsoft.com (Copilot privacy controls)
  8. Google (Gemini in Gmail, Chat and Meet)

    Plaintiffs filed a Second Amended Class Action Complaint in Thele v. Google LLC, No. 5:25-cv-09704 (N.D. Cal., Judge Noël Wise), adding Edward Goldstein as a named plaintiff, after the court dismissed the first amended complaint on July 7, 2026 for failure to allege concrete harm, with leave to amend. The suit alleges Google switched Gemini "smart features" on by default for Gmail, Chat and Meet on or about October 10, 2025. Google stipulated on August 20 to extend its response deadline; no merits ruling.

    Source: Thele v. Google LLC docket (PacerMonitor)
  9. Atlassian (Jira / Confluence) — watchlist / industry signal

    Watchlist item — not a scored matrix row (Atlassian is a SaaS suite, outside this tracker's assistant / coding / voice / meeting / local-LLM scope). Logged here as an industry signal: effective August 17, 2026, Atlassian began using customer data from Jira, Confluence, Jira Service Management, and other Cloud products to train AI offerings including Rovo and Rovo Dev, on an opt-out basis. Affects roughly 300,000 customers. Org admins can turn off in-app data contribution under Atlassian Administration → Security → Data contribution; opting out removes in-app data within 30 days and content attributes within 90 days per Atlassian. Reporting indicates metadata collection cannot be disabled on Free, Standard, and Premium plans, only on Enterprise.

    Source: us.seibert.group
  10. Fathom

    Fathom updated its privacy policy (Last Updated August 16, 2026). Training language is unchanged in substance: de-identified Meeting Content Information may be used to train Fathom's in-house models depending on account settings, with an opt-out in account settings, and third parties such as OpenAI, Anthropic and Google are not authorized to train on it.

    Source: Fathom privacy policy
  11. DictaFlow

    DictaFlow published an updated consumer privacy policy (last updated August 14, 2026) under operator SmartBids.ai Corp. The provider list now reads "OpenAI, Deepgram, and Groq for transcription and related AI processing"; NVIDIA, named in the version we reviewed in July 2026, no longer appears. Retention language: "We do not permanently retain audio recordings on our servers after processing in the ordinary course," with transcription text retainable for features, device sync, result recovery, support, or legal compliance. The standard service remains excluded from PHI use, with healthcare customers directed to DictaFlow Medical and a BAA.

    Source: dictaflow.io (privacy policy)
  12. Otter.ai

    Judge Eumi K. Lee ruled on the motion to dismiss in In re Otter.AI Privacy Litigation, No. 5:25-cv-06911-EKL (N.D. Cal.), granting it in part and denying it in part. The federal Wiretap Act (ECPA) claim, the California Invasion of Privacy Act § 631 claim, both Illinois BIPA voiceprint claims, the unjust-enrichment claim, and the UCL claim survive: the court held plaintiffs plausibly allege Otter "independently collects, retains, and uses communications for its own commercial purposes" — including training its models on recordings — which makes Otter a potential third-party eavesdropper rather than merely the host's recording tool. The two Computer Fraud and Abuse Act counts, the California CDAFA claim, the Washington Privacy Act claim, and most intrusion-upon-seclusion claims were dismissed with leave to amend; plaintiffs voluntarily dropped conversion, trespass to chattels, and some CIPA provisions. Plaintiffs had until about August 27, 2026 to amend, with Otter's response due 21 days later; the case proceeds toward discovery and class certification. Otter's Track record scores were lowered on both tiers.

    Source: recordinglaw.com (ruling explainer)
  13. Freed

    Freed updated its Platform Terms of Use (August 13, 2026; privacy policy updated August 11). The Terms state Freed may use De-identified Data to train its AI/ML models "with your consent provided via Platform settings," will link de-identified data to the customer ID to personalize the model, and reserve the right "in our sole discretion to use De-identified Data and to disclose such De-identified Data to third parties." The default state of the consent setting is not published.

    Source: Freed Platform Terms of Use
  14. OpenWhispr

    OpenWhispr updated its privacy policy to describe three processing modes (local, BYOK, OpenWhispr Cloud), to state that cloud audio is "processed in real time, and discarded," and to disclose that "API keys you enter for third-party providers are currently stored in plaintext in your local app data directory." The Terms of Service effective July 23, 2026 carry a matching no-training clause and identify the operator as Gizmo Labs Inc., a Delaware corporation.

    Source: openwhispr.com/privacy
  15. tl;dv

    Security researcher bobdahacker publicly disclosed that a missing tenant-isolation rule on tl;dv's Firestore "meetings" collection allowed any authenticated account to enumerate 181,874 meeting records across 84,312 users (creator email, conference ID, provider, recording status, timestamps), including roughly 1,000 joinable live calls, after first reporting it on January 28, 2026. tl;dv told GIGAZINE on August 13 that these were "two separate vulnerabilities," that no audio, transcripts, notes, passwords or billing data were reachable, that the second vector was patched within 24 hours of discovery, and that Firebase has been removed from its infrastructure.

    Source: bobdahacker disclosure
  16. Granola

    Chamberlain v. Granola, Inc., No. 3:26-cv-07926 (N.D. Cal., filed July 30, 2026) — a putative class action against Granola, Inc. and UK affiliate Granola Labs Ltd. alleging Granola's AI notetaker intercepts and records virtual-meeting participants without their knowledge or consent and, by default, uses the contents to train Granola's AI models. Because Granola captures system audio instead of joining meetings as a visible bot, non-user participants get no notice — the complaint quotes Granola's own marketing that "other people in the room won't know" — and the training opt-out exists only in Granola users' account settings, leaving other participants no opt-out path. Claims: federal Wiretap Act (ECPA, 18 U.S.C. § 2510 et seq.), California Invasion of Privacy Act §§ 631 and 632, California CDAFA (Penal Code § 502), common-law intrusion upon seclusion, California UCL, and unjust enrichment, on behalf of a nationwide class and a California subclass. Granola's Track record scores were lowered on both tiers.

    Source: courtlistener.com (Chamberlain docket)
  17. DictaFlow Medical

    DictaFlow Medical's privacy policy was updated July 24, 2026. It states the product "is designed not to permanently store raw audio or transcript text on the backend unless a specific support, security, legal, or customer-requested workflow requires retention," keeps audit and disclosure metadata "without intentionally storing raw PHI in those audit records," and requires covered entities to "execute a Business Associate Agreement before using DictaFlow Medical with PHI." The medical subprocessor list (Deepgram, OpenAI, Groq, Railway/Firebase, Resend/Postmark, Stripe) was last updated May 19, 2026.

    Source: dictaflow.io/medical (privacy policy)
  18. Heidi Health

    Heidi revised its support article "How Heidi protects your data" (July 23, 2026) to state: "Heidi does not use consultation data, session transcriptions, or generated notes to train its AI model in any way," and to list ISO 27001, ISO 42001, and SOC 2. The privacy policy page still displays an October 2024 date and retains a de-identified/aggregate-use clause.

    Source: How Heidi protects your data
  19. Paraspeech

    Paraspeech's privacy policy (controller Burlis Management GmbH, Germany) was last updated July 23, 2026. It names Deepgram as the cloud-transcription processor and Groq and Cerebras as the processors behind the Cloud Cleanup rewrite feature, which the docs describe as available "only when cloud processing is allowed." The policy makes no statement about AI training and publishes no audio or transcript retention window for cloud mode.

    Source: Paraspeech privacy policy (iubenda)
  20. Voibe

    Voibe moved from an on-device-only architecture to a hybrid model with two user-selectable modes. On-device mode still transcribes entirely on the Mac (Apple Silicon, M1 or later), with nothing leaving the device. A new zero-retention cloud mode runs only open-source / open-weight models on zero-retention providers — Whisper Large Turbo (open-source) on Groq or Cloudflare for transcription, and GPT-OSS 120B (open-weight) on Cerebras for text refinement — with audio encrypted in transit, deleted the moment transcription completes, and neither voice nor text stored or used to train any AI model. No third-party AI lab is in the audio path. Cloud mode lets Voibe run on all Macs (Intel and Apple Silicon) and on Windows. Users choose the mode in Settings. Voibe's own matrix row was updated to reflect both modes.

    Source: getvoibe.com (AI & privacy)
  21. Google (Gemini in Gmail, Chat and Meet)

    Judge Noël Wise (N.D. Cal.) granted Google's motion to dismiss the first amended complaint in Thele v. Google LLC, No. 5:25-cv-09704, finding plaintiffs had not alleged "concrete harm sufficient to establish their standing to sue in federal court," and gave leave to amend. The claims — intrusion upon seclusion, California constitutional privacy, Stored Communications Act, CIPA and CDAFA — concern the October 2025 default-on switch for Gemini smart features across Gmail, Chat and Meet.

    Source: Bloomberg Law (July 8 2026)
  22. Perplexity

    Perplexity replaced its privacy policy with a restructured Privacy Notice (last revised July 8, 2026). Three written commitments were removed from the document: the AI-training opt-out, the 30-day account-deletion timeline (replaced by retention "as long as reasonably necessary"), and the promise to notify users of material changes (replaced by posting updates online). The in-product controls still exist — the AI Data Retention toggle still disables training and account deletion is still available — but they are now product features rather than commitments in the policy text. The rewrite also expanded the stated collection scope to cover the Comet browser, voice input, and health and genetic data, and added Global Privacy Control recognition and a commitment not to re-identify de-identified data. Logged on this pass because Perplexity's row had not been re-verified since April.

    Source: venpo.com (policy diff)
  23. Google Meet (Take notes for me)

    Google opened Gemini's "Take notes for me" in Google Meet to personal Google AI Pro and AI Ultra subscribers (previously Workspace-only). Notes are saved to a Google Doc in the organizer's Drive and all participants are notified. Google's Workspace no-training commitment applies to work and school accounts; no equivalent statement was published for personal-plan Meet notes.

    Source: Google blog announcement
  24. Read AI

    Read AI updated its privacy policy (Last Updated June 29, 2026). Model improvement remains off by default via an opt-in Customer Experience Program; the policy states Read does not, and does not permit third-party AI tools to, use Google Workspace API data to train generalized models, and caps audio and video storage at 2 years.

    Source: Read AI privacy policy
  25. Gemini CLI

    Google retired the consumer sign-in path for Gemini CLI. Per the May 19, 2026 announcement, "Gemini CLI and Gemini Code Assist IDE extensions will stop serving requests for Google AI Pro and Ultra, as well as those using it free of charge using Gemini Code Assist for individuals," with users pointed to Antigravity CLI, whose terms use Interactions "to evaluate, develop, and improve Google and Alphabet research, products, services and machine learning technologies" unless changed in settings. Standard / Enterprise licenses and paid API keys were unaffected. The former individuals privacy notice, and its training opt-out checkbox, now serve only a deprecation page. Individuals still using Gemini CLI on a free Gemini API key are governed by the Unpaid Services terms, which permit product improvement and human review with no toggle.

    Source: Google Developers Blog (Gemini CLI to Antigravity CLI)
  26. Google (Search Services History)

    Google introduced a "Search Services History" setting, with a "Save Media" subsetting, that lets Google use images, files, audio, and video from Search services — Google Lens images, Search Live and Translate speaking-practice recordings, uploaded content, and voice searches — to develop and improve its AI models. The setting is migrated on for accounts that had Web & App Activity enabled, which is the default for most users; Computerworld characterizes it as on by default. It does not cover Gemini Live or other Gemini Apps media, which remain governed by Gemini Apps Activity. Opt out under Google Account → Web & App Activity → Search Services History → Save Media.

    Source: support.google.com (Save Media)
  27. Abridge

    Abridge and Nvidia announced a clinical-conversation AI model trained on de-identified doctor–patient conversations from Abridge's platform (about 100 health systems, including Kaiser Permanente, Mayo Clinic, Johns Hopkins, and Yale New Haven Health). The model is to run only inside Abridge's platform on Abridge-controlled hardware. This is the clearest public confirmation that Abridge trains on de-identified encounter data.

    Source: PYMNTS
  28. Apple Intelligence (Private Cloud Compute)

    Apple announced it is running new Apple Intelligence workloads on Google Cloud — NVIDIA GPUs with confidential computing, Intel TDX CPUs, and Google's Titan chip — extending Private Cloud Compute's stateless, verifiable-transparency guarantees to third-party data centers for the first time, with all binaries published for inspection. Apple's third-generation foundation models were built in collaboration with Google; Apple restates that it does not use users' private personal data or interactions when training them. Siri AI is withheld in the EU on the 27-series OS releases under the DMA.

    Source: security.apple.com (Expanding PCC)
  29. Windsurf → Devin Desktop (Cognition)

    Cognition rebranded Windsurf as Devin Desktop; windsurf.com security, privacy, and terms pages now redirect to Cognition / Devin documentation. Under the current docs, training is on by default and the opt-out — which also enables Zero Data Retention with model providers — is available on paid plans, exercised by an administrator on Teams; Enterprise customers are not trained on without express written consent. The self-hosted deployment is in maintenance mode in favor of a hybrid deployment. The matrix row was renamed and its sources replaced.

    Source: docs.devin.ai (security)
  30. VoiceDash

    VoiceDash published a Data Processing Addendum effective May 26, 2026, identifying the operator as Jumpstart Ventures FZE (Dubai Silicon Oasis) and listing nine subprocessors — OpenAI, Groq, Cloudflare, Hetzner, Sentry, Metabase, CustomerIO, RevenueCat, and Apple. It states transcription data is "processed transiently, structurally excluded from AI model training loops" but sets no fixed retention window and contains no HIPAA or BAA language.

    Source: voicedash.ai/dpa
  31. WisprType

    WisprType's first privacy policy took effect April 29, 2026, one day after the domain was registered. It states that "all speech-to-text processing happens entirely on your device using WhisperKit" by default, that audio is "not retained after transcription completes," that cloud providers (OpenAI, Groq, Deepgram) are used only when the user enters an API key and selects cloud mode, and that PostHog telemetry is "disabled by default." The policy contains no training clause and names no legal entity.

    Source: wisprtype.com (privacy policy)
  32. GitHub Copilot

    GitHub began using Free / Pro / Pro+ user interaction data, including code snippets, to train AI models by default. Existing opt-outs are honored. Business and Enterprise are unaffected.

    Source: github.blog/news-insights
  33. Abridge (Sutter Health / MemorialCare deployments)

    Three California patients filed Washington et al. v. Sutter Health et al. in the U.S. District Court for the Northern District of California against Sutter Health, Memorial Health Services, and MemorialCare Medical Foundation, alleging their Abridge ambient-scribe deployments recorded and transmitted clinical conversations without informed consent, in violation of CIPA, CMIA, the UCL, and the federal Wiretap Act, plus intrusion upon seclusion. Abridge is not a defendant. Sutter said its technology is "implemented in accordance with applicable laws and regulations."

    Source: HIPAA Journal
  34. Granola

    The Verge reported Granola makes meeting notes accessible to anyone with the share link by default and opts users into internal AI training in default settings, contradicting the app's "private by default" marketing. Enterprise tier has training off by default; consumer tier requires manual opt-out in account settings.

    Source: themeridiem.com
  35. Wispr Flow

    Wispr's prior SOC 2 Type II (Accorp Partners) and ISO 27001 (Gradient) certifications were proactively invalidated over platform-integrity concerns at the original auditor. Wispr re-engaged A-LIGN: a clean, unqualified SOC 2 Type I was completed in April 2026 (Type II in progress), and ISO 27001:2022 Stage 1 was completed in April 2026 with Stage 2 scheduled for June 2026. The Trust Center moved to trust.wispr.ai.

    Source: docs.wisprflow.ai (compliance FAQ)
  36. Fireflies.ai (row correction)

    Fireflies' privacy policy dated March 6, 2026 states: "We do not use personal information for AI model training and we contractually prohibit our vendors from using this information for their own model training," with a Zero Data Retention section covering third-party processors. Earlier versions of this tracker recorded Fireflies as training on consumer transcripts by default; on the September 2026 pass the row was corrected to "No (stated policy)" on both tiers and the Training scores raised accordingly. The BIPA litigation (Fricker and Martinez, now consolidated before Judge Seeger) concerns voiceprint capture and consent, not model training.

    Source: fireflies.ai (privacy policy)
  37. Zoom

    Zoom Privacy Statement update expanded the definition of "Customer Content" to cover content created using Zoom products beyond meetings, webinars, and messages, and clarified host/participant recording and sharing. The no-AI-training commitment on customer content (audio, video, chat, screen sharing, attachments, transcripts) remains in force as a stated contractual prohibition rather than an opt-out.

    Source: zoom.com/trust/privacy
  38. X / xAI (Grok)

    X's January 2026 Terms of Service update explicitly classified Grok prompts and outputs as user "Content" available for AI training and fine-tuning. Combined with the platform's prior default-on training of public X posts, this codifies the consumer-tier Grok experience as opt-in-by-default: opt-out paths exist in three separate settings surfaces (Grok on X / Grok mobile / grok.com), each requiring a separate toggle.

    Source: privacy.x.com
  39. OpenAI Codex

    OpenAI revised its Enterprise privacy page (the document that governs Codex under API-key, Business, Enterprise, and Edu sign-ins). The commitments as of that revision: "We do not train our models on your data by default" across ChatGPT Business, Enterprise, Healthcare, Edu, Teachers, and the API Platform; admin-controlled retention for Enterprise, Healthcare, and Edu with deleted conversations removed within 30 days; SOC 2 Type 2. The page does not itemize what changed in this revision, and the consumer-side Codex posture (training on by default, two separate toggles) is unaffected.

    Source: Enterprise privacy at OpenAI
  40. Microsoft 365 Copilot

    Microsoft enabled Anthropic as a default subprocessor for Microsoft 365 Copilot — used in Researcher, Copilot Studio, Word / Excel / PowerPoint agents, and Agent Mode in Excel. Processing occurs outside the EU Data Boundary; EU and UK tenants have it disabled by default. The shift is opt-out for commercial tenants in unaffected regions. Microsoft 365 Copilot's no-training contractual posture is unchanged — the data still isn't used to train foundation models — but the surface area for who processes customer content expanded.

    Source: learn.microsoft.com
  41. Fireflies.ai

    Cruz v. Fireflies.AI Corp., No. 3:25-cv-03399, filed in the Central District of Illinois under the Illinois Biometric Information Privacy Act (BIPA), alleging Fireflies captured the voiceprints of all meeting participants without written consent. The "Fireflies.ai Notetaker" bot joins meetings as a visible participant, creating consent complications in two-party-consent jurisdictions. The case was terminated in March 2026 and refiled in the Northern District of Illinois (Fricker v. Fireflies.AI Corp., No. 1:26-cv-02675; Martinez v. Fireflies.AI Corp., No. 1:26-cv-03512, both before Judge Steven C. Seeger).

    Source: courtlistener.com (Cruz docket)
  42. Meta AI

    Meta began using Meta AI conversations for ad personalization in addition to model training across US and most non-EU regions. EU, UK, and South Korea are carved out pending GDPR/regional clearance. Sensitive categories (health, religion, politics) are excluded from ad targeting. First major consumer AI platform to feed active chat content into ad-targeting pipelines at scale; EPIC requested an FTC suspension in response.

    Source: snopes.com (fact-check Nov 2025)
  43. OpenAI / ChatGPT

    OpenAI's obligation to indefinitely retain consumer ChatGPT and API content, imposed under the NYT litigation order, ended. Standard 30-day retention practices resumed. Limited April–September 2025 data is still preserved under the order.

    Source: openai.com
  44. Anthropic / Claude (incl. Claude Code on Pro/Max)

    Anthropic shifted consumer Claude (Free, Pro, Max) from "not used for training" to a user-choice model. Users who opt in have data retained up to 5 years (vs. previous 30 days). Existing-user choice deadline: October 8, 2025. The update explicitly applies to Claude Code sessions running on Pro/Max accounts — a developer using Claude Code via a Pro/Max account is on consumer terms (default-on training unless toggled off), while Claude Code via API key / Console / Enterprise remains on Commercial Terms (no training).

    Source: anthropic.com/news
  45. Grok / xAI

    Approximately 370,000+ Grok shared chat conversations were exposed publicly on Google Search because Grok's share feature generated unauthenticated URLs without `noindex` directives. The exposed prompts included medical information, business details, passwords, and requests for illegal activity. Disclosed first by Forbes / Fortune / TechCrunch; xAI patched after media intervention. Combined with the July 2025 "MechaHitler" content failure and Turkey's criminal probe + content block, this anchors the consumer-tier track-record floor.

    Source: techcrunch.com
  46. Otter.ai

    Brewer v. Otter.ai filed in the Northern District of California alleging Otter "deceptively and surreptitiously" recorded private conversations and used meeting data to train AI models without explicit permission from all participants. Complaint alleges violations of the Electronic Communications Privacy Act (ECPA), Computer Fraud and Abuse Act (CFAA), and California Invasion of Privacy Act (CIPA). The suit was later consolidated with three related cases as In re Otter.AI Privacy Litigation, No. 5:25-cv-06911-EKL, on October 22, 2025; the motion-to-dismiss hearing was reset to July 15, 2026.

    Source: courtlistener.com (docket)
  47. DeepSeek (cumulative bans Jan 2025 – Apr 2026)

    Italy's Garante imposed a 72-hour ban (January 30, 2025) followed by investigations in 13 European jurisdictions and the formation of a dedicated EDPB AI Enforcement Task Force. Government device bans followed in Australia, Taiwan, South Korea, Czech Republic, the Netherlands, Germany, multiple US federal agencies (Pentagon, NASA, US Navy), and several US states. The bipartisan "No DeepSeek on Government Devices Act" is pending in the US Senate. DeepSeek's privacy policy explicitly stores personal data — including keystroke patterns, IP addresses, and uploaded files — on servers in the People's Republic of China.

    Source: iapp.org

How this tracker is maintained

  • Each cell is verified against the vendor's own privacy policy, terms of service, or technical documentation. We don't paraphrase what the policy "probably" says — only what it actually states.
  • We separate consumer and business / API tiers because they operate under different contracts. The tier filter on each table reflects this; conflating them is the most common error in third-party comparisons.
  • "Trains on your data?" answers what happens by default at sign-up. The note under each chip describes the toggle, if one exists.
  • Each tier carries four independent scores on a 0–25 scale: Training, Retention, On-device, and Track record. The four axes sum to a 0–100 composite. We publish per-axis scores because the dimensions trade off differently for different use cases, and the composite for readers who need a single answer.
  • Track record reflects documented incidents, breaches, and unfavorable policy changes — sourced from the vendor's own announcements, court filings, and reporting that we link from this page or from related Voibe Resources articles. We don't penalize a vendor for a past mistake they have visibly fixed; we do penalize opaque reverses and silent expansions of data collection.
  • The matrix is reviewed on a roughly monthly cadence and updated immediately whenever a vendor announces a policy change. Each row carries its own Last verified date.
  • Errors get fixed fast: email hi@getvoibe.com with a primary-source link and we'll update the row.

Spotted an error? hi@getvoibe.com. Include the cell and a primary-source link and we’ll update on the next pass.

Legend:Yes (default)Trains by default with no opt-out pathYes (opt-out)Trains by default; user can disable in settingsUser choiceUser must actively choose during signup or in settingsNoDoes not train on user data, period

Score legend — per axis (0–25)

22–25StrongArchitectural or contractual guarantee
17–21SolidDefault-off training, short retention, or partial on-device
11–16MixedUser must take action, or policy has caveats
6–10WeakUnfavorable default, opt-out only
1–5PoorNo opt-out path, or indefinite retention
0UnclearPolicy does not address this dimension

Total score (0–100, sum of four axes)

85–100Excellent
70–84Strong
55–69Adequate
40–54Weak
0–39Poor

Privacy scoreboard

One row per tool, showing its best-scoring tier in the selected tier group. Each axis is scored 0–25; the total adds the four axes for a 0–100 composite. Sorted by total by default. Tools with multiple sub-tiers in the same group (e.g., Superwhisper’s on-device vs cloud modes) are detailed in the matrix tables below.

Sort
#ToolCategoryPlan tierTrainingRetentionOn-deviceTrack recordTotal
1Local LLM RuntimesOpen-source (MIT) — local install (optional Ollama Cloud models, Pro $20/mo)2525252398/100Excellent
2Local LLM RuntimesOpen-source (Apache 2.0) — local install2525252398/100Excellent
3Voice & DictationOn-device mode (Apple Silicon — M1 or later)2525252297/100Excellent
4Voice & DictationFree build (GPL v3) / Solo $25 / Personal $39 / Extended $49 (one-time)2525252297/100Excellent
5Voice & DictationFree (MIT open-source) — macOS / Windows / Linux2525252196/100Excellent
6Voice & DictationLocal mode (Free — MIT open-source client)2525251994/100Excellent
7Voice & DictationOn-device modes (Fast / Nano / Standard / Parakeet / Cohere Transcribe — Free + Pro)2523252093/100Excellent
8Local LLM RuntimesFree desktop app — Mac / Windows / Linux2523252093/100Excellent
9Voice & DictationFree (GPLv3 open-source) — macOS 15+2523251992/100Excellent
10Voice & DictationLocal Only Mode (Free)2525251489/100Excellent
11Voice & DictationLocal mode on Apple Silicon ($14.99/mo, $99/yr, or Lifetime local-only)2423251789/100Excellent
12Voice & DictationPro (Gumroad) / Whisper Transcription (App Store)2320232288/100Excellent
13AI AssistantsmacOS / iOS / iPadOS / visionOS (Apple Silicon, supported regions)2322222087/100Excellent
14Coding ToolsOpen-source extension (BYOK)2325132384/100Strong
15Voice & DictationDragon Professional v16 — Windows desktop, fully offline (no Mac version since 2018; Dragon Home discontinued 2023)2222231683/100Strong
16Voice & DictationLocal mode (default; free during early access)2223241079/100Strong
17Coding ToolsCode Assistant Platform ($39/user/mo, annual; no free or individual tier)232432272/100Strong
18Voice & DictationmacOS / iOS (Apple Silicon, supported languages)2013181869/100Adequate
19Coding ToolsClaude Code via Pro / Max account (Aug 2025 consumer terms apply)1515131962/100Adequate
20Clinical DocumentationSelf-serve clinician plan (prices not published on nabla.com)202131862/100Adequate
21Voice & DictationFree (2,000 words/mo) / Pro ($7/mo or $69/yr; App Store IAP $7.99/mo, $79.99/yr)1515141761/100Adequate
22Voice & DictationFree (1,000 words/mo) / Pro ($15/mo or $12/mo billed yearly)221731658/100Adequate
23Meeting TranscriptionFree / Pro ($19.75/mo, $15/mo annual)211332057/100Adequate
24Clinical DocumentationFree (unlimited transcription) / Clinician (14-day trial; price not shown on pricing page)191532057/100Adequate
25Voice & DictationFree trial (30 min) / Pro ($8.49/mo annual) / Lifetime ($260)162131555/100Adequate
26Voice & DictationFree (1,000 words + 10 notes) / Pro ($15/mo or $144/yr)181513854/100Weak
27Clinical DocumentationIndividual clinician — Starter $39/mo (40 notes), Core $79/mo, Premier $104/mo annual ($119 monthly); 7-day trial131832054/100Weak
28AI AssistantsFree / Pro / Max151831551/100Weak
29Voice & DictationFree Trial / Individual ($12/mo annual = $144/yr, Private Mode default)201351351/100Weak
30Meeting TranscriptionFree / Pro (prices unverified — pricing page returned 404)22153949/100Weak
31Clinical DocumentationNo self-serve plan published — practice contract (pricing page not found)91731948/100Weak
32Meeting TranscriptionAI Companion (included on paid Zoom plans; on by default, admin-controllable)201231247/100Weak
33AI AssistantsFree / Plus / Pro101831546/100Weak
34Voice & DictationFree / Pro15233546/100Weak
35Coding ToolsCodex via ChatGPT Free / Go / Plus / Pro sign-in87131644/100Weak
36Meeting TranscriptionFree / Premium ($20/mo, $16/mo annual)9932041/100Weak
37AI AssistantsFree / Microsoft 365 Premium $19.99/mo (Copilot Pro retired August 1, 2026) / Windows Copilot101281040/100Weak
38Voice & DictationFree / Pro ($9.99 App Store IAP) / AppSumo lifetime ($59 Tier 1)101431239/100Poor
39Coding ToolsIndividual (Gemini API key, unpaid tier; Google-account consumer login retired June 18, 2026)65131337/100Poor
40Coding ToolsIndividual default (Privacy Mode OFF)83131236/100Poor
41Voice & DictationFree (data sharing ON / Zero Data Retention off, default — formerly called Privacy Mode)20331036/100Poor
42Clinical DocumentationAbridge for Clinicians app — no self-serve plan; requires a health-system deployment91331136/100Poor
43Meeting TranscriptionPersonal account on Google AI Pro / AI Ultra101031235/100Poor
44Clinical DocumentationNo self-serve plan — "Request a demo" only7831735/100Poor
45Coding ToolsFree / Pro / Pro+ / Pro Max / Power via GitHub, Google, or AWS Builder ID sign-in9681134/100Poor
46AI AssistantsFree / Gemini Advanced101031033/100Poor
47Meeting TranscriptionFree (800 min/mo) / Pro2053533/100Poor
48Meeting TranscriptionFree (120 min/mo) / Pro ($13.61/mo, $8.17/mo annual) / Business ($27.78/$16.67)8731533/100Poor
49Coding ToolsIndividual (Free / paid; training on by default)10531331/100Poor
50Meeting TranscriptionFree / Pro10133531/100Poor
51Coding ToolsFree / Pro / Pro+1053826/100Poor
52Voice & DictationStarter (free) / Pro ($8/mo annual) / Max ($24/mo annual, Realtime Mode)8031526/100Poor
53Meeting TranscriptionFree / Pro883625/100Poor
54AI AssistantsFree / Pro / Max833721/100Poor
55AI AssistantsMeta AI in WhatsApp / Instagram / Facebook / Messenger / Threads / Ray-Ban Meta462517/100Poor
56AI AssistantsGrok on X (Free / Premium $8/mo / Premium+ $40/mo) + grok.com + mobile533314/100Poor
57AI AssistantsFree / Pro / Premium (consumer chatbot and mobile app)513110/100Poor

Use case fit: which score is right for what?

A composite score is only useful if you know what threshold to look for. The table below maps four common buyer scenarios to the minimum score we recommend, the reasoning, and the tools that currently meet that bar.

Healthcare, legal, financial — regulated work

85+

Patient records, attorney-client privileged content, financial data — anywhere a leak triggers regulator notification or contract liability.

Why this floor: Below 85, the vendor either retains data longer than 30 days, lacks a contractual training exclusion, or has a track-record incident in the past 24 months. Regulated work has no margin for any of those.

Watch out for: BAA / DPA availability — score doesn't reflect HIPAA contracts. A tool can score 85+ and still be unusable for PHI without a signed BAA. Always verify the contract before using a tool for regulated content.

Tools that meet the bar (business tier)

  • OllamaLocal LLM Runtimes · Open-source (MIT) — local install (same posture; optional cloud models)98/100Excellent
  • JanLocal LLM Runtimes · Open-source (Apache 2.0) — local install (same posture)98/100Excellent
  • VoiceInkVoice & Dictation · Same posture (open-source local install)97/100Excellent
  • HandyVoice & Dictation · Same posture (open-source local install; no vendor tier)96/100Excellent
  • LM StudioLocal LLM Runtimes · Free desktop app (same posture)93/100Excellent
  • FluidVoiceVoice & Dictation · Same posture (open-source local install; no vendor tier)92/100Excellent
  • MacWhisperVoice & Dictation · Same posture (no separate enterprise tier)88/100Excellent
  • Apple Intelligence / SiriAI Assistants · macOS / iOS (same posture)87/100Excellent

Proprietary code, internal docs, M&A drafts

70+

Sensitive but not strictly regulated. Strategy docs, source code, contract drafts, customer data without compliance overlay.

Why this floor: 70+ means contractual ZDR or short retention with a vendor that has a clean recent track record — you trust them not to retain or train, even if you'd still avoid pasting raw secrets.

Watch out for: Consumer-tier mistakes. Most consumer tiers score below 50 — make sure your team is on the business plan, not pasting internal docs into ChatGPT Free. The tier filter on each table makes this visible.

Tools that meet the bar (business tier)

  • OllamaLocal LLM Runtimes · Open-source (MIT) — local install (same posture; optional cloud models)98/100Excellent
  • JanLocal LLM Runtimes · Open-source (Apache 2.0) — local install (same posture)98/100Excellent
  • VoiceInkVoice & Dictation · Same posture (open-source local install)97/100Excellent
  • HandyVoice & Dictation · Same posture (open-source local install; no vendor tier)96/100Excellent
  • LM StudioLocal LLM Runtimes · Free desktop app (same posture)93/100Excellent
  • FluidVoiceVoice & Dictation · Same posture (open-source local install; no vendor tier)92/100Excellent
  • MacWhisperVoice & Dictation · Same posture (no separate enterprise tier)88/100Excellent
  • Apple Intelligence / SiriAI Assistants · macOS / iOS (same posture)87/100Excellent

Day-to-day drafting, research, light coding

55+

Public-facing content, general knowledge work, code that isn't a trade secret, prompts you wouldn't mind appearing in a leak.

Why this floor: 55+ means the tool is reasonably well-behaved by default or has a clear, working opt-out path that most users will actually flip.

Watch out for: Default settings. Many tools score 55+ only after the privacy toggle is on. Verify each user has actually flipped it — onboarding teams to a private-by-default workflow is more reliable than chasing settings.

Tools that meet the bar

  • OllamaLocal LLM Runtimes · Open-source (MIT) — local install (optional Ollama Cloud models, Pro $20/mo)98/100Excellent
  • JanLocal LLM Runtimes · Open-source (Apache 2.0) — local install98/100Excellent
  • VoibeVoice & Dictation · On-device mode (Apple Silicon — M1 or later)97/100Excellent
  • VoiceInkVoice & Dictation · Free build (GPL v3) / Solo $25 / Personal $39 / Extended $49 (one-time)97/100Excellent
  • HandyVoice & Dictation · Free (MIT open-source) — macOS / Windows / Linux96/100Excellent
  • OpenWhisprVoice & Dictation · Local mode (Free — MIT open-source client)94/100Excellent
  • SuperwhisperVoice & Dictation · On-device modes (Fast / Nano / Standard / Parakeet / Cohere Transcribe — Free + Pro)93/100Excellent
  • LM StudioLocal LLM Runtimes · Free desktop app — Mac / Windows / Linux93/100Excellent

Personal, low-stakes use

Any

Notes to self, brainstorming, creative writing — content you wouldn't mind seeing in a leak.

Why this floor: Any tool works if you understand the tradeoff. Default consumer tiers of major assistants land in the 30–50 range; that's fine for non-sensitive prompts.

Watch out for: Voice input. Audio is uniquely sensitive — even casual dictation may capture identity-revealing details, ambient conversations, or addresses. The On-device score matters more for voice than for text.

Tools that meet the bar

  • OllamaLocal LLM Runtimes · Open-source (MIT) — local install (optional Ollama Cloud models, Pro $20/mo)98/100Excellent
  • JanLocal LLM Runtimes · Open-source (Apache 2.0) — local install98/100Excellent
  • VoibeVoice & Dictation · On-device mode (Apple Silicon — M1 or later)97/100Excellent
  • VoiceInkVoice & Dictation · Free build (GPL v3) / Solo $25 / Personal $39 / Extended $49 (one-time)97/100Excellent
  • HandyVoice & Dictation · Free (MIT open-source) — macOS / Windows / Linux96/100Excellent
  • OpenWhisprVoice & Dictation · Local mode (Free — MIT open-source client)94/100Excellent
  • SuperwhisperVoice & Dictation · On-device modes (Fast / Nano / Standard / Parakeet / Cohere Transcribe — Free + Pro)93/100Excellent
  • LM StudioLocal LLM Runtimes · Free desktop app — Mac / Windows / Linux93/100Excellent

AI Assistants

Chatbots and search assistants. Consumer tiers vary the most — check whether your account is logged-in vs. logged-out, and whether you've reviewed your data settings since the last policy change.

ToolPlan tierData collectedTrains on your data?RetentionOn-deviceTrack recordTotalLast verifiedSource
Free / Plus / ProPrompts, outputs, uploaded files, usage, IP, device info, account info. Opt-in Computer History (macOS, August 13, 2026) logs clicks, typing, and app switches locally in unencrypted files kept up to 48 hours; OpenAI states the event files are processed server-side but not retained or used for training, while memories derived from them follow your data-controls setting.Yes (opt-out)

Off via Settings → Data Controls → "Improve the model for everyone." Temporary Chat is never used for training. OpenAI's Rest-of-World Privacy Policy was updated February 6, 2026. Audio and video from Advanced Voice chats are not used for training by default; opting in requires the separate "Include your audio recordings" and "Include your video recordings" sub-toggles nested under the same "Improve the model for everyone" control. Free and Go plans show ads (US test since February 9, 2026; 31 European markets since August 24, 2026); OpenAI states conversations are not shared with advertisers, and ad personalization is a separate control under Settings → Ad Controls.

Training10/25
30 days after deletion. April–September 2025 data preserved due to NYT order; standard practice resumed Sept 26, 2025.
Retention18/25
No
On-device3/25

March 2023 chat-history bug; April–September 2025 NYT-mandated indefinite retention.

Track record15/25
46/100WeakSep 10, 2026
Free / Pro / MaxChats, coding sessions (when using Claude Code with consumer accounts), feedback (thumbs)User choice

Active choice required during signup or in Privacy Settings ("You can help improve Claude"). Off by default for users who decline. Policy changed August 2025. In January 2026 Anthropic added a separate Consumer Health Data Privacy Policy (effective January 12, 2026) covering consumer/personal use by residents of Washington and other US states with consumer-health-data laws; it applies only where Anthropic acts as a data controller and does not cover commercial or enterprise plans. A further consumer Privacy Policy update announced June 8, 2026 and effective July 8, 2026 added sections on data received through connected apps and multi-step tasks, verification data for age / identity checks, and research-study participation; Anthropic reaffirmed that it does not sell data, Claude stays ad-free, and the training choice remains user-controlled. Team, Enterprise, and API are unaffected.

Training15/25
30 days if declined. 5 years if enabled. Flagged conversations: 2–7 years for trust & safety.
Retention18/25
No
On-device3/25

August 2025 reversal: consumer Claude moved from 'never used for training' to user choice.

Track record15/25
51/100WeakSep 10, 2026
Free / Gemini AdvancedChats, files, photos, videos, screen content, account info, IP, device infoYes (opt-out)

Off via the "Keep Activity" setting (the renamed "Gemini Apps Activity" toggle; on by default). Even when off, future chats are kept for 72 hours so Gemini can respond and process feedback, and Temporary Chats are never used for training. Separately, a June 2026 "Search Services History" setting (with a "Save Media" subsetting) lets Google use images, files, audio, and video from Search services — Google Lens images, Search Live and Translate speaking-practice recordings, uploaded content, and voice searches — to develop and improve Google's AI models. Google says the setting is "rolling out gradually over the next few months"; until it appears, Search-services media remain governed by Web & App Activity, which is on by default for most accounts. Per Google's own pages this does not cover Gemini Live or other Gemini Apps media, which remain governed by Keep Activity.

Training10/25
18 months default (adjustable to 3 months / 36 months / never). Human-reviewed chats retained up to 3 years (disconnected from account).
Retention10/25
No
On-device3/25

Keep Activity 'off' still keeps chats 72h; human-reviewed conversations retained up to 3 years.

Track record10/25
33/100PoorSep 10, 2026
Free / Pro / MaxQueries, prompts, AI responses, usage, device infoYes (opt-out)

Off via Account Settings → Preferences → "AI Data Retention." Logged-out users are trained on by default with no opt-out path. On July 2, 2026 Perplexity replaced its privacy policy with a restructured Privacy Notice (revised July 8, 2026) that no longer describes the training opt-out, the 30-day deletion timeline, or a commitment to notify users of material changes; the toggle still exists in the product but is no longer a written policy commitment. The notice also expanded stated collection to the Comet browser, voice input, and health and genetic data.

Training8/25
Threads kept until manually deleted. The July 2026 Privacy Notice replaced the 30-day account-deletion timeline with retention "as long as reasonably necessary."
Retention3/25
No (Comet browser stores some data locally — separate policy)
On-device3/25

2024 reporting documented robots.txt evasion via undisclosed user-agent; logged-out users still trained on. July 2026 Privacy Notice rewrite removed the written training opt-out, 30-day deletion timeline, and change-notification commitments while widening stated collection (Comet, voice, health and genetic data). A March 31, 2026 class action alleges embedded trackers sent chat content to Google and Meta.

Track record7/25
21/100PoorSep 10, 2026
Free / Pro / Premium (consumer chatbot and mobile app)Prompts, outputs, uploaded files, IP, device info, account info. (Keystroke patterns were listed in the 2025 policy versions; the February 10, 2026 text no longer names them.)Yes (default)

DeepSeek's privacy policy (last updated February 10, 2026) states: "we directly collect, process and store your Personal Data in People's Republic of China." There is no in-app training toggle; the policy grants "the right to opt-out of using your Personal Data for training our models or optimizing our technologies," exercisable by emailing privacy@deepseek.com. Note: open-source MIT-licensed weights can be self-hosted, which removes this concern entirely — but that is not the consumer chatbot row.

Training5/25
Retained "for as long as necessary to provide our Services" on servers in the People's Republic of China — no fixed window published. The policy's jurisdiction-specific supplemental clause repeats that personal data is processed and stored in China.
Retention1/25
No
On-device3/25

Banned by Italy (Garante, Jan 2025); 13-jurisdiction EU probes; US gov device bans (Pentagon, NASA, US Navy); Jan 2025 breach exposed 1M+ records; Feroot Security found code linking to China Mobile authentication registry.

Track record1/25
10/100PoorSep 10, 2026
macOS / iOS / iPadOS / visionOS (Apple Silicon, supported regions)Prompts, optional personal context (only when feature invoked), request size and duration metadata. ChatGPT integration is opt-in and routes through OpenAI's enterprise terms with IP obfuscation.No

Apple's published policy states Apple does not use users' private personal data or user interactions when training its foundation models. ChatGPT integration is a separate, opt-in boundary not covered by Apple's guarantees.

Training23/25
Stateless. Private Cloud Compute processes data only to fulfill the request and returns results to device; data is not stored or made accessible to Apple. Apple collects only request metadata (size, feature, duration) — not content.
Retention22/25
Yes (primary). The on-device foundation model handles most tasks. Private Cloud Compute runs on Apple Silicon servers and, since June 8, 2026, also on NVIDIA GPU / Intel TDX hardware in Google Cloud under the same stateless, verifiable-transparency guarantees (published binaries, no SSH or admin access); Apple states data is not accessible to Apple or Google.
On-device22/25

Verifiable transparency model; signed binaries publicly inspectable; no documented incidents in the Apple Intelligence era. June 2026: PCC extended to third-party (Google Cloud) data centers for the first time, and third-generation foundation models were built in collaboration with Google — Apple still states user data is not used for training. Siri AI is withheld in the EU on the 27-series OS releases under the DMA. ChatGPT integration introduces an opt-in external boundary not covered by PCC guarantees.

Track record20/25
87/100ExcellentSep 10, 2026
Free / Microsoft 365 Premium $19.99/mo (Copilot Pro retired August 1, 2026) / Windows CopilotPrompts, conversation activity, voice inputs, uploaded images and files, Bing / MSN browsing signals, ad interactions, identifiers tied to Microsoft Account when signed in. Verbatim from Microsoft: "your voice and conversation activity with Copilot, including the images or files you upload."Yes (opt-out)

"Except for certain categories of users or users who have opted out, Microsoft uses data from Bing, MSN, Copilot, and interactions with ads on Microsoft for AI training." The updated Copilot app (August 18, 2026) exposes two toggles under Profile → Privacy: "Training on conversation activity" and "Training on voice conversations"; opting out "will exclude your future conversation activities" from model training, while conversations may still be used for other product improvements, advertising, digital safety, security, and compliance. Signed-out users, and users of Copilot inside Microsoft 365 apps on Personal / Family / Premium subscriptions, are excluded; chat at copilot.com on those plans is not. Excluded regions per Microsoft: Brazil, China (excluding Hong Kong), Israel, Nigeria, South Korea, Vietnam.

Training10/25
18 months default for conversation activity per Microsoft Support. User can delete individual items or full history; opt-out from training is separate from history deletion.
Retention12/25
Partial — Recall and Click to Do run locally on Copilot+ PCs (NPU ≥40 TOPS); snapshots stored on-device only. Most Copilot chat (text / voice / image generation) runs in the cloud. Recall is opt-in (off by default) post-2024 backlash, with encrypted local database.
On-device8/25

Recall controversy (2024 — Kevin Beaumont / DoublePulsar plaintext-database findings) is the adjacent track-record drag; Copilot itself has no comparable public breach. Migliaccio & Rathod LLP announced an investigation into Copilot's default-on consumer training on August 26, 2026. Microsoft's general advertising / data history weighs on the consumer composite.

Track record10/25
40/100WeakSep 10, 2026
Meta AI in WhatsApp / Instagram / Facebook / Messenger / Threads / Ray-Ban MetaPublic Facebook and Instagram posts, captions, and comments from adults (18+); all interactions with Meta AI (prompts, queries, responses); public profile data; images shared publicly or in which the user is tagged; voice recordings from Ray-Ban / Oakley Meta glasses (stored by default since April 2025, retained up to 1 year); device and location data; ad and interaction signals; public Vibes (AI-generated content shared to feed).Yes (default — regional opt-out)

Opt-in by default globally on public posts and Meta AI conversations. EU / EEA users have opt-out via the "Right to Object" form in Privacy Center → "How Meta uses information for generative AI models and features" — Meta resumed EU training May 27 2025 under GDPR Article 6(1)(f) "legitimate interest" after a one-year pause. UK opt-out form available after ICO concessions. US / Australia / most non-GDPR jurisdictions have no general opt-out — Meta confirmed in Australian Senate hearings (Sept 2024) it scrapes AU public posts back to 2007 and will only offer an opt-out "if governments force it to." WhatsApp 1:1 and group messages remain end-to-end encrypted and are not used for training UNLESS the user explicitly invokes @Meta AI in chat. As of December 16 2025, Meta AI conversations also feed ad personalization.

Training4/25
Indefinite unless user manually deletes (Settings → Data and privacy → Manage your information → Delete all chats and media). No published auto-purge window. Ray-Ban Meta voice recordings: up to 1 year by default. Memory is managed separately via Accounts Center.
Retention6/25
No for the Meta AI assistant. WhatsApp "Private Processing" (2025) is a server-side TEE-based confidential compute layer with hardware attestation, not on-device.
On-device2/25

June 2024 EU training pause forced by Irish DPC + noyb pressure; May 2025 noyb cease-and-desist threatening class action; Sept 2024 admission to AU Senate that AU public posts scraped back to 2007 with no general opt-out; April 2025 Ray-Ban Meta voice-recording-by-default expansion (removed prior user control); Dec 2025 ad-targeting expansion → EPIC FTC complaint. Pattern of expand-first, retreat-only-under-regulator.

Track record5/25
17/100PoorSep 10, 2026
Grok on X (Free / Premium $8/mo / Premium+ $40/mo) + grok.com + mobilePublic X posts, likes, follows, replies, reposts; Grok prompts, inputs, outputs, voice prompts, images uploaded for context; device data (IP, browser, OS, advertising ID, carrier, language, memory, apps installed, battery level); approximate location (IP-derived) plus GPS if granted; account info, contact list if shared; DM contents, recipients, timestamps; Grok memory stores conversation context across sessions.Yes (default)

Opt-in by default for X posts AND Grok chats / outputs. Three separate opt-out paths across surfaces: (1) Grok on X — Settings → Privacy and safety → Data sharing and personalization → Grok & xAI → uncheck "Allow your posts as well as your interactions, inputs, and results with Grok and xAI to be used for training and fine-tuning." (2) Grok mobile app — Settings → Data Controls → uncheck "Improve the model." (3) Grok web (grok.com) — Settings → Data → uncheck "Improve the model." EU / EEA users excluded from X-post training under the Sept 4 2024 DPC undertaking; DPC opened a fresh statutory inquiry into XIUC on April 11 2025. The Jan 15 2026 X ToS update explicitly classifies Grok prompts and outputs as user "Content" available for AI training. Private Chat mode on grok.com auto-opts-out of training for that session.

Training5/25
Indefinite unless user-deleted; deletion processed within 30 days. No tier-based retention differences across Free / Premium / Premium+.
Retention3/25
No — cloud-only via xAI's Memphis Colossus / Colossus 2 supercomputer.
On-device3/25

Aug–Sept 2024 Irish DPC High Court action over Grok EU-post training; April 2025 DPC fresh statutory inquiry into XIUC; July 8–12 2025 "MechaHitler" content failure → EU Commission technical meeting + Turkey criminal probe + court-ordered content block + Poland DSA referral; August 20 2025 — ~370,000+ Grok shared-chat conversations indexed on Google because share feature generated unauthenticated URLs without `noindex` directives; March 2026 Lieff Cabraser class action over Grok-generated CSAM / non-consensual sexual deepfakes. February 17, 2026: the Irish DPC opened a further inquiry into XIUC over Grok-generated non-consensual intimate images of EU / EEA residents, including minors.

Track record3/25
14/100PoorMay 12, 2026

AI Coding Tools

IDE assistants and agents. Consumer defaults shifted in April 2026 (GitHub Copilot now trains on consumer interaction data by default). Most tools offer a Privacy / Zero-Data-Retention mode that flips the answer; check whether yours is on.

ToolPlan tierData collectedTrains on your data?RetentionOn-deviceTrack recordTotalLast verifiedSource
Individual default (Privacy Mode OFF)Codebase data, code, prompts, editor actions, code snippetsYes (default)

Default for individual accounts is Privacy Mode OFF (Cursor no longer labels this "Share Data"). With it off, per the Data Use page (last updated September 3, 2026): "we may use and store codebase data, prompts, editor actions, code snippets, and other code data and actions to improve our AI features and train our models." Toggle Privacy Mode ON to opt out — code is not trained on and cached plaintext is discarded after each request. Requests go through Cursor's backend even with your own API key.

Training8/25
Privacy Mode OFF: Cursor "may use and store" code, prompts, and editor actions with no published retention period; inference providers may temporarily store inputs and outputs, "deleted after use." Privacy Mode ON: cached files are temporary, encrypted with client-generated keys that exist on Cursor servers only for the request.
Retention3/25
No (configurable to use local Ollama / LM Studio models, which bypass Privacy Mode entirely)
On-device13/25

Privacy Mode off by default for individuals; no documented incidents. The August 29, 2026 Data Use update removed the disclosure of how indexing embeddings and metadata are retained, and the August 13, 2026 Terms replaced arbitration with mandatory Texas litigation.

Track record12/25
36/100PoorSep 10, 2026
Free / Pro / Pro+Inputs, outputs, code snippets, associated contextYes (opt-out)

Policy changed April 24, 2026: GitHub now trains on consumer interaction data by default. Existing opt-outs honored. Toggle in Settings → Privacy.

Training10/25
User Engagement Data: 2 years. Coding-agent session logs persist as account records until deleted (no published retention period). Private repo code at rest is NOT used for training; in-flight interaction data IS.
Retention5/25
No
On-device3/25

April 2026 reversal trains consumer interaction data by default (Business / Enterprise unaffected); no documented incidents.

Track record8/25
26/100PoorSep 10, 2026
Individual (Free / paid; training on by default)Per Cognition's privacy policy (March 9, 2026), User Content is used "to train, fine tune and improve the models that power our Services" unless you opt out.Yes (opt-out)

Windsurf was rebranded Devin Desktop by Cognition on June 2, 2026; all windsurf.com legal pages now redirect to Cognition / Devin pages. Training is on by default. Per docs.devin.ai: "If you're on a paid plan, you can opt out at any time on the Data Controls settings page. After you opt out, your data will not be used for training and Zero Data Retention will be enabled with our model providers." No opt-out is documented for the free tier.

Training10/25
With the paid-plan opt-out: ZDR with model providers — Customer Data "not saved to disk or otherwise persistently retained" and "deleted upon generation of the relevant Output." Otherwise Cognition "only retains data processed through Devin for the duration of the relationship with a given Customer."
Retention5/25
No
On-device3/25

2024 Codeium → Windsurf rebrand; 2025 Cognition acquisition; June 2, 2026 rebrand to Devin Desktop. Training opt-out limited to paid plans; no documented incidents.

Track record13/25
31/100PoorSep 10, 2026
Open-source extension (BYOK)Cline operates no model server. With your own API keys, code goes only to your configured provider (Anthropic, OpenAI, Bedrock, Gemini, etc.) and Cline "does not receive or store your input tokens, output tokens, underlying code, or other User Content." If you use Cline-provided API keys, prompts and code snippets transit Cline's infrastructure to the provider.No (by Cline)

Cline's stated principle: "Code never leaves your machine" toward Cline servers. Anonymous telemetry (features used, task completion) is opt-out via the "Cline Telemetry" setting. Code, file contents, command arguments, and conversation content are not collected by telemetry.

Training23/25
Cline retains nothing about your code on the bring-your-own-key path; no retention period is published for the Cline-key path. Provider retention applies (e.g., Anthropic API ZDR, OpenAI API 30 days), and per Cline's ToS providers "may use User Content for training unless you have explicitly opted out where such an option is offered."
Retention25/25
Partial — extension runs locally; inference happens at your chosen provider, or fully on-device if you configure Ollama / LM Studio.
On-device13/25

Open-source; no model server; no documented incidents.

Track record23/25
84/100StrongSep 10, 2026
Code Assistant Platform ($39/user/mo, annual; no free or individual tier)Code context (open files, signatures) sent for inference, then discarded. No source code stored. Operational metrics/logs (no code, no PII) retained ~1 week for support.No

Tabnine's proprietary completion models are trained only on permissively licensed open-source code (MIT, Apache 2.0). Customer code is never used to train models and is never shared with third parties. Stated verbatim on the code-privacy page: "We never use your code to train any of our models."

Training23/25
Zero code retention — requests are processed in-memory and immediately discarded after the server returns the suggestion. No source code logged. Operational metrics/logs (no code, no PII) kept ~1 week for support.
Retention24/25
No — cloud inference for the SaaS tier.
On-device3/25

SOC 2 Type II, ISO 27001, GDPR. Long-running vendor with no documented privacy incidents. No free or individual plan; entry tier is Code Assistant Platform at $39/user/mo (Agentic Platform $59). Acquired by Tricentis, announced July 30, 2026; the August 28, 2026 privacy-policy and terms edits were cosmetic with no change to data handling.

Track record22/25
72/100StrongSep 10, 2026
Claude Code via Pro / Max account (Aug 2025 consumer terms apply)All user prompts and model outputs, encrypted via TLS. Includes file contents the model reads, tool-use results, and Bash command outputs that are returned to the model. Additionally: telemetry (latency / reliability / usage metrics — no code or file paths per the docs), error reports to a third-party error-tracking service (on only for Pro / Max sign-ins on v2.1.198+ connecting directly to the Claude API), and optional `/feedback` submissions (full conversation history including code).User choice (default-on)

Critical disambiguation: Claude Code run from a Claude.ai Pro or Max account inherits the August 28 2025 consumer-terms update — default-on training, user choice required by Oct 8 2025. Anthropic's announcement explicitly states the update applies "when they use Claude Code from accounts associated with those plans." Opt-out at claude.ai/settings/data-privacy-controls.

Training15/25
5 years if model-improvement enabled; 30 days if disabled. Local session transcripts cached in plaintext under ~/.claude/projects/ for 30 days (tunable via cleanupPeriodDays). /feedback transcripts retained 5 years. Optional session-transcript uploads retained 6 months.
Retention15/25
Partial — Claude Code binary executes locally (file edits, shell execution, MCP servers, local session log). Inference is cloud-only via Anthropic API. Code never leaves machine unless sent as part of an inference request, telemetry fires (no code in telemetry per docs), an error report fires, or the user runs /feedback, /bug, or /share.
On-device13/25

No documented Claude Code data-handling incidents to date; newer surface than Cursor (Claude Code GA'd 2025). Slight deduction for the Aug 2025 consumer-terms reversal scope catching some users by surprise (developers assumed API-tier terms).

Track record19/25
62/100AdequateSep 10, 2026
Codex via ChatGPT Free / Go / Plus / Pro sign-inPrompts, file contents and tool outputs the agent reads, model outputs, and (for cloud tasks) the repository checked out into an OpenAI-hosted container. Per the help center, ChatGPT sign-in means the ChatGPT Terms of Use and Privacy Policy "apply to data shared between Codex and ChatGPT." Per the Codex docs, "Local execution does not mean offline or device-only model inference."Yes (opt-out, two separate toggles)

OpenAI's help center (updated August 2026) names Codex directly: "When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models." The ChatGPT Settings → Data Controls → "Improve the model for everyone" toggle covers "your ChatGPT conversations and Codex tasks." A second control exists: "Codex has separate controls for allowing training on full environments, which you can manage in the Codex Settings. Note that adjusting your settings in the ChatGPT interface or privacy portal will not affect these full-environment Codex settings." The default state of that environment control is not published (unverified).

Training8/25
No Codex-specific retention window is published. The consumer privacy policy applies: deleted content is removed "within 30 days" unless retained for legal, abuse, or security reasons, or already de-identified for training. Cloud task containers are cached "for up to 12 hours."
Retention7/25
Partial — the CLI and IDE extension run locally (file edits, shell, MCP servers); inference is cloud-only. Cloud tasks run in "VM-backed sandboxes" on OpenAI-managed infrastructure.
On-device13/25

No documented data-handling incident. Command-injection flaw that exposed GitHub user access tokens (ChatGPT web, Codex CLI, SDK, IDE extension) reported December 16, 2025 and patched February 5, 2026 with no exploitation reported; CVE-2025-61260 (auto-loaded MCP config) fixed in v0.23.0, August 2025. Deduction for the second, separately managed full-environment training toggle that ChatGPT's main opt-out does not touch.

Track record16/25
44/100WeakSep 10, 2026
Individual (Gemini API key, unpaid tier; Google-account consumer login retired June 18, 2026)Prompts, file contents and tool outputs the agent reads, and model responses, governed by the Gemini API Terms (effective March 23, 2026). Google retired the consumer sign-in path: "Starting June 18, 2026, Gemini Code Assist IDE extensions stopped serving requests for the Gemini Code Assist for individuals, Google AI Pro, and Google AI Ultra tiers" — a change Google says also applies to Gemini CLI. Separately, anonymous usage statistics (tool names, model, durations, session config) are on by default; per the config docs "We do not log the content of your prompts or the responses."Yes on the unpaid API tier (no toggle; switch to paid billing to stop)

Gemini API Terms, Unpaid Services: "Google uses the content you submit to the Services and any generated responses to provide, improve, and develop Google products" and "human reviewers may read, annotate, and process your API input and output." No opt-out setting exists on the unpaid tier; the Paid Services section is the off-switch. The former "Gemini Code Assist for individuals" privacy notice — which offered an opt-out checkbox — has been replaced by a deprecation page (last updated September 2, 2026), so its text could not be re-verified this pass. Usage statistics can be disabled with privacy.usageStatisticsEnabled: false.

Training6/25
Not published as a number. Unpaid tier: content is used for product improvement after "disconnecting this data from your Google Account, API key, and Cloud project." Paid tier: logs kept "for a limited period of time, solely for detecting and preventing violations of the Prohibited Use Policy."
Retention5/25
Partial — the CLI runs locally (file edits, shell, MCP servers); inference is cloud-only. No local-model option is documented.
On-device13/25

Open-source (Apache 2.0). Critical advisory GHSA-wpqr-6v78-jr5g published April 24, 2026 (headless workspace-trust and tool-allowlist bypasses; fixed in 0.39.1); 2025 Tracebit prompt-injection RCE. Consumer login sunset announced May 19, 2026 with 30 days' notice. GitHub issue #20569 (February 27, 2026) on conflicting free-vs-paid training docs was closed as not planned.

Track record13/25
37/100PoorSep 10, 2026
Free / Pro / Pro+ / Pro Max / Power via GitHub, Google, or AWS Builder ID sign-inPer the data-protection page (last updated August 4, 2026): "your questions to Kiro, other inputs you provide, and the responses and code that Kiro generates," plus telemetry (client version, OS, machine ID, request counts, errors, latency, crash reports). "Individual subscribers" are users "that have a paid Kiro subscription and access it through a social login provider (like GitHub or Google) or through AWS Builder ID." Pricing page: Free $0, Pro $20, Pro+ $40, Pro Max $100, Power $200 per user per month.Yes (opt-out)

"Kiro may use certain content from Kiro Free Tier and Kiro individual subscribers for service improvement ... to provide better responses to common questions, fix Kiro operational issues, for de-bugging, or for model training." IDE opt-out: Settings → User → Application → Telemetry and Content, uncheck "Data Sharing and Prompt Logging: Content Collection for Service Improvement" and "Usage Analytics And Performance Metrics." CLI exposes `kiro-cli settings telemetry.enabled`; its default and a CLI-side content-collection switch are not published (unverified). Paying does not opt you out — sign-in method does.

Training9/25
Free Tier: "we may store your inputs for up to 60 days for the purposes of detecting activity that violates the Agreement." "For OpenAI GPT models, classifier-flagged traffic will be retained for up to 30 days." No retention window is published for content collected for service improvement. Encrypted with AWS-owned KMS keys.
Retention6/25
No — the IDE and CLI run locally but every request is inferred in AWS (Bedrock-hosted Claude, OpenAI GPT, and open-weight models per the FAQ); no local-model option is documented.
On-device8/25

No customer data breach reported, but a dense 2026 vulnerability record: CVE-2026-10591 (RCE via file writes to .vscode/tasks.json, bulletin June 2, 2026), CVE-2026-18656/18657 (Windows executable resolution in IDE and CLI, bulletin August 4, 2026), and the Mindgard Kiro Powers prompt-injection exfiltration (fixed in 0.8.140, reported August 27, 2026). GitHub issue #2206 on AWS terms still permitting training for paid Kiro users has been open since August 18, 2025; issue #4024 (content-collection checkbox not disableable for Q Developer Pro logins) was closed as not planned.

Track record11/25
34/100PoorSep 10, 2026

Voice & Dictation

Speech-to-text and dictation tools. Voice input is uniquely sensitive — audio carries identity, biometric data, and ambient context — so the on-device column matters more than for text-only tools.

ToolPlan tierData collectedTrains on your data?RetentionOn-deviceTrack recordTotalLast verifiedSource
On-device mode (Apple Silicon — M1 or later)None — audio is transcribed entirely on your Mac and never transmitted. Account holders: email (account auth) plus non-identifying usage analytics; crash reports exclude dictated content.No

In on-device mode nothing leaves your Mac: your voice is transcribed entirely on the device. Because audio never crosses the network, there is no training to opt out of.

Training25/25
Audio: not transmitted, not retained. Account email: kept while account is active.
Retention25/25
Yes — Whisper models running on the Apple Silicon Neural Engine. Requires an Apple Silicon Mac (M1 or later).
On-device25/25

New entrant; on-device processing removes the surface for retention or training incidents.

Track record22/25
97/100ExcellentSep 10, 2026
Zero-retention cloud mode (open-source models on zero-retention providers — all Macs and Windows)Audio is encrypted in transit and sent to zero-retention providers for transcription, then deleted the moment transcription completes. Formatting is a second hop that sees only the text, never the audio. Account email plus non-identifying usage analytics as above.No

Cloud mode runs only open-source / open-weight models on zero-retention providers — Whisper Large Turbo (open-source) on Groq or Cloudflare for transcription, and GPT-OSS 120B (open-weight) on Cerebras for text formatting. No third-party AI lab (OpenAI, Google, Anthropic, Microsoft) is in the audio path, and users never create an AI-vendor account or paste an API key. Voice is never stored, never sold, and never used to train any model; the same applies to text. Users pick a mode at setup and can switch in Settings. Windows and Intel Macs are cloud-mode only.

Training23/25
Audio is deleted the moment transcription completes. Neither voice nor text is stored.
Retention23/25
No — cloud mode runs on zero-retention providers (works on all Macs — Intel and Apple Silicon — and on Windows, a ground-up native app launched in 2026).
On-device3/25

Hybrid model documented publicly on Voibe's AI & privacy page; open-source models only, zero retention, no training, no third-party AI lab in the audio path.

Track record22/25
71/100StrongSep 10, 2026
Free (data sharing ON / Zero Data Retention off, default — formerly called Privacy Mode)Audio, transcripts, edits, optional Context Awareness (screen content from active app)Yes (opt-in)

After 2024 community backlash, training is now off by default and requires opt-in ("If you opt to share your content with us for model training"). The privacy policy (last updated August 19, 2026) no longer states a fixed retention window for retained dictation data — "only for as long as necessary" — and says content passed to third-party LLM providers "is not used to train their models and is generally deleted within 30 days." Dictation Cloud Storage and Notetaker transcript retention are user-configurable under Settings → Data and Privacy. The August 2026 policy also newly covers Wispr's Notetaker (meeting audio) and Google Workspace data.

Training20/25
No fixed retention window published for retained dictation data ("only for as long as necessary"); about 30 days for third-party LLM passthrough per the August 2026 policy.
Retention3/25
No — transcription always happens in the cloud, even in Privacy Mode (zero-retention cloud, not local).
On-device3/25

2024 community backlash forced opt-in training and Privacy Mode; Free tier still indefinite by default. March 2026: Wispr's prior SOC 2 Type II (Accorp Partners) and ISO 27001 (Gradient) certifications were proactively invalidated over platform-integrity concerns at the original auditor; A-LIGN re-audited, completing a clean SOC 2 Type I in April 2026 and ISO 27001:2022 Stage 1 in April 2026. As of September 10, 2026 Wispr's compliance FAQ still lists the SOC 2 Type II observation period as underway with no report issued, and ISO 27001 Stage 2 as "scheduled June 2026" with no completion announced.

Track record10/25
36/100PoorSep 10, 2026
On-device modes (Fast / Nano / Standard / Parakeet / Cohere Transcribe — Free + Pro)None — audio processed locally and never transmittedNo

"Your data is not retained on Superwhisper servers" and "not used for training AI models or any other machine learning purposes." Audio recordings are saved to local disk by default — opt out in settings.

Training25/25
N/A on servers. Local recordings persist until the user deletes them.
Retention23/25
Yes
On-device25/25

Stable privacy-first stance. Cloud modes were added before the public privacy policy addressed them; the June 2026 Terms now state zero-retention, no-training handling for cloud modes. Trust center (trust.mycroft.io/superwhisper) lists a SOC 2 Type II report dated March 3, 2026 and subprocessors Anthropic, Cloudflare, AWS, Argmax, and OpenAI.

Track record20/25
93/100ExcellentSep 10, 2026
Cloud modes (Ultra transcription / Super Mode LLMs — Pro)Audio sent to Superwhisper's proxy infrastructureNo (per vendor)

Superwhisper says cloud audio is proxied through their infrastructure and providers "see a proxy request, not your account, and nothing is retained for training." The privacy policy (still dated June 19, 2024) does not distinguish cloud modes, but the Terms (updated June 2026) now state that cloud models "process your content transiently on a zero-retention basis and do not store it after processing" and that "your data is not used to train, fine-tune, or improve AI models." Cloud transcription models include Ultra, S1-Voice, Scribe V2, and Nova 3; Super Mode LLMs include Claude, GPT-5-series, Gemini, and Grok models.

Training18/25
Terms (June 2026): cloud models process content "transiently on a zero-retention basis"; the 2024 privacy policy still does not address cloud modes separately.
Retention18/25
No
On-device3/25

Stable privacy-first stance. Cloud modes were added before the public privacy policy addressed them; the June 2026 Terms now state zero-retention, no-training handling for cloud modes. Trust center (trust.mycroft.io/superwhisper) lists a SOC 2 Type II report dated March 3, 2026 and subprocessors Anthropic, Cloudflare, AWS, Argmax, and OpenAI.

Track record20/25
59/100AdequateSep 10, 2026
Pro (Gumroad) / Whisper Transcription (App Store)On-device modes: none transmitted. App Store version discloses "Usage Data" and "Product Interaction" as Data Not Linked to You. Cloud Assistant or BYOK (OpenAI / ElevenLabs) features send audio to those providers under their terms.No (by MacWhisper)

MacWhisper does not train its own models on user audio. Cloud Assistant and BYOK integrations inherit the chosen provider's terms (e.g., OpenAI Whisper API, Anthropic / ElevenLabs).

Training23/25
On-device transcription: not retained. Cloud Assistant / BYOK: per third-party provider's terms.
Retention20/25
Yes (primary mode) — local Whisper models plus Apple Foundation Models for AI features. Cloud Assistant is opt-in for higher-quality transcription.
On-device23/25

Long-running indie tool; on-device by default; no documented incidents.

Track record22/25
88/100ExcellentSep 10, 2026
Starter (free) / Pro ($8/mo annual) / Max ($24/mo annual, Realtime Mode)Audio inputs, technical data (IP, browser, OS, performance metrics), session metadata. With Privacy Mode disabled, "we may securely store transcript data on our servers." Instant Mode uses Aqua's own Avalon models; Realtime Mode streams audio to a third-party realtime speech provider.Yes (opt-out)

Privacy Mode toggle stops transcript storage on Aqua Voice servers; with it enabled, "transcript data is not collected" though session metadata may still be. The privacy policy (last updated July 29, 2026) does not use the word training, but the FAQ now states: "By default, dictation transcripts may be retained to improve the model; switch on Privacy Mode and they aren't." SOC 2 Type II certified by Advantage Partners. Per the FAQ, "Aqua doesn't sign HIPAA BAAs yet."

Training8/25
With Privacy Mode disabled: not specified in policy. With Privacy Mode enabled: transcripts not stored; session metadata (timestamps, device type, performance metrics) may be retained.
Retention0/25
No — cloud transcription
On-device3/25

SOC 2 Type II via Advantage Partners. The privacy policy is silent on training, but the FAQ states default transcript retention is used to improve the model; no HIPAA BAA.

Track record15/25
26/100PoorSep 10, 2026
Free / ProAudio plus limited contextual information, processed on Typeless's cloud servers. Subprocessors include third-party LLM providers, analytics, and cloud infrastructure.No (per vendor)

Privacy policy (last updated August 25, 2026): "Your data is never used to train these services and is configured for zero retention by the providers." The data-controls page adds: "Your voice audio is not stored. Dictation data is not used for model training by us or any third party." Note: the November 2025 reverse-engineering analysis documented in our Typeless privacy issues investigation reported collection beyond what the public policy describes — verify against the current policy and subprocessor list before sensitive use.

Training15/25
Per the privacy policy, audio + contextual information are "processed in real time on our cloud servers and immediately discarded once the transcription result is returned." Optional History cloud sync stores history text on Typeless servers until you turn it off or delete it.
Retention23/25
No — cloud-processed in real time
On-device3/25

Nov 2025 reverse-engineering analysis documented collection (URLs, window-title metadata, broad permissions) beyond what the public policy describes.

Track record5/25
46/100WeakSep 10, 2026
macOS / iOS (Apple Silicon, supported languages)Audio inputs, plus contextual data (contacts, app names, etc.) when sent to serversOpt-in only

"Improve Siri & Dictation" must be enabled. Default at setup is to be asked.

Training20/25
If opted in: audio + transcripts kept under a rotating random ID for up to 6 months, dissociated and kept up to 2 years for improvement; reviewed subset retained beyond 2 years. If opted out: not retained for improvement.
Retention13/25
Yes (partially) — most languages on Apple Silicon process locally for general text fields (Notes, Mail, Messages). Server fallback applies to unsupported languages, search-box dictation, and some third-party Speech Recognition API uses.
On-device18/25

2019 Siri grading scandal led to opt-in for human review; otherwise privacy-forward.

Track record18/25
69/100AdequateSep 10, 2026
Free Trial / Individual ($12/mo annual = $144/yr, Private Mode default)Private Mode (default): basic technical and account-related data only — verbatim: "In private mode, Willow only collects basic technical and account-related data needed to run the app and nothing else. No voice, no dictated text." Opt-In Mode: anonymized text and usage data to improve text-correction models.No (Private Mode default)

Private Mode is the default. Training only occurs if the user explicitly opts into Opt-In Mode — "You can allow Willow to collect minimal usage data to help improve our text correction models." This is the inverse of Wispr Flow Free, which had Privacy Mode off by default until the 2024 community backlash. Privacy policy last updated April 30 2025.

Training20/25
Private Mode: no audio or transcript collected on Willow servers. Opt-In Mode: "We keep anonymized text and usage data only as long as needed to train and improve the app." No specific retention window disclosed for the Opt-In path.
Retention13/25
No — cloud-first transcription. The privacy policy's "this is only stored locally on your device for you to view" refers to local UI storage of past transcripts, not local transcription. No offline transcription mode advertised on the current pricing page.
On-device5/25

Newer entrant (YC X25 cohort). Marketing pricing page advertises HIPAA + SOC 2 + zero data retention at the Enterprise tier, but the privacy policy text itself only references "GDPR and SOC 2" — HIPAA / BAA scope is not documented in the policy. Minor discrepancy between marketing claims and policy substantiation.

Track record13/25
51/100WeakSep 10, 2026
Free (1,000 words + 10 notes) / Pro ($15/mo or $144/yr)Per the rewritten monologue.to data-privacy page: "Monologue does not save audio or transcripts from Dictation on its servers." Notes are different: "When you record a Note, Monologue stores the Note's audio, transcript, and summary in your account so they can sync across your devices," and saved Notes are also reachable through Monologue's MCP, API, and CLI tools. "Deep Context screenshots are deleted immediately after processing." Custom modes and dictionaries now sync across devices.No (per vendor, for Dictation)

The data-privacy page now states: "Monologue's AI providers process Dictation with zero data retention, so your dictation is not retained for model training." The statement covers Dictation and is phrased as provider-side retention rather than an explicit Monologue-side training exclusion; it does not address the stored audio, transcripts, and summaries from Notes. The Notion-hosted privacy policy remains sparse.

Training18/25
Dictation: audio and transcripts not saved on Monologue servers; AI providers process with zero data retention. Notes: audio, transcript, and summary stored in your account for sync. Deep Context screenshots deleted immediately after processing.
Retention15/25
Partial — "Offline transcription on Mac using local models" is listed on the marketing site, but the default architecture is cloud (AI providers process Dictation) and Notes are stored in the cloud.
On-device13/25

Indie product (1-person team). Privacy policy hosted on Notion and sparse — no list of data categories, no subprocessor list, no SOC 2 / HIPAA / BAA. The data-privacy page was rewritten in 2026 to add a no-retention-for-training statement for Dictation while disclosing that Notes audio, transcripts, and summaries are stored in the account.

Track record8/25
54/100WeakSep 10, 2026
Free build (GPL v3) / Solo $25 / Personal $39 / Extended $49 (one-time)None on VoiceInk servers — there is no VoiceInk-operated cloud service. By default audio is transcribed on-device using whisper.cpp. Optional bring-your-own-key cloud features exist: cloud transcription can send audio to the provider you choose, and AI Enhancement sends "only the transcribed text (never audio)" to that provider.No

Privacy policy (last updated April 20, 2026): "No data leaves your computer unless you explicitly choose to enable optional cloud services." Optional BYOK providers include OpenAI, Anthropic, Google Gemini, Groq, Mistral, Deepgram, and ElevenLabs; with them enabled, handling follows that provider's terms. VoiceInk itself operates no server and does not train. Source code is auditable on GitHub under GPL v3.0.

Training25/25
Default: nothing leaves the device. With optional BYOK cloud modes enabled, retention follows the chosen provider.
Retention25/25
Yes (default) — local Whisper models via whisper.cpp on macOS 14.4+. Optional BYOK cloud transcription and AI Enhancement are user-enabled.
On-device25/25

Open-source GPL v3.0; auditable on GitHub at Beingpax/VoiceInk; 6,400+ stars and 900+ forks; latest v2.13 (August 27, 2026). No documented incidents; GitHub issue #537 (February 2026) flagged that the privacy policy did not fully document Context-Aware / clipboard data sent to BYOK providers, and the April 2026 policy revision added the optional-cloud disclosure. Mild deduction for being a single-maintainer project (continuity risk distinct from privacy risk).

Track record22/25
97/100ExcellentSep 10, 2026
Local Only Mode (Free)None during transcription — audio runs through OpenAI Whisper Large-v3 and NVIDIA Parakeet locally on Apple Silicon with no network calls. Analytics per the privacy policy are limited to button clicks and page views, with no personally identifiable information stated.No

In Local Only Mode there is no transmission to train on. The privacy policy (last updated July 7, 2026) states audio recordings are "not stored" on Spokenly's servers; in this mode nothing reaches a server in the first place.

Training25/25
Audio: not transmitted, not retained. Local recordings persist on the Mac until the user deletes them.
Retention25/25
Yes — Whisper Large-v3 and Parakeet on Apple Silicon, no network calls during transcription.
On-device25/25

Indie product by named solo developer Vadim Akhmerov (disclosed via App Store, not the privacy policy); no disclosed corporate entity or jurisdiction; no SOC 2 / HIPAA / ISO 27001 / GDPR / CCPA attestations. App Store 4.4/5 from 56 ratings; v1.12.3 released September 2026. No documented incidents.

Track record14/25
89/100ExcellentSep 10, 2026
BYOK cloud (Free — bring your own API key)Audio routed to whichever provider you configure keys for — OpenAI, Deepgram, Groq, Anthropic, or Google. Spokenly passes audio through; data handling is governed by the chosen provider's terms.Depends on provider

Spokenly itself does not train on audio, but BYOK transfers the training and retention question to the provider whose key you supply. Each provider (OpenAI, Deepgram, Groq, Anthropic, Google) has its own training and retention terms — verify the specific provider before sensitive use.

Training16/25
Spokenly: "not stored" on its servers. Provider-side retention follows whichever API you bring keys for.
Retention14/25
No — audio leaves the device for the configured cloud provider.
On-device3/25

Indie solo-developer product; no disclosed corporate entity; no compliance attestations. App Store 4.4/5 from 56 ratings.

Track record14/25
47/100WeakSep 10, 2026
Pro managed cloud ($9.99/mo)Audio forwarded by Spokenly's backend to third-party transcription services: the July 2026 policy names OpenAI, Deepgram, Soniox, Cartesia, ElevenLabs, Groq, Mistral AI, xAI, OpenRouter, fal.ai, and Cerebras (eleven providers; Fireworks was removed). Each is a separate data-handling entity.Best-effort ("prefer" no-training providers)

The July 7, 2026 policy states Spokenly's backend "does not store these recordings" and "immediately forwards them to third-party transcription services," and that Spokenly "prefer[s] options that do not use customer data for model training where available." That is a hedged preference, not a contractual no-training guarantee across the eleven named providers.

Training15/25
Spokenly: audio "not stored" on its servers. Each of the eleven named providers retains per its own terms — not enumerated in Spokenly's policy.
Retention16/25
No — audio is processed through the managed-cloud subprocessor chain.
On-device3/25

No SOC 2 / HIPAA BAA / ISO 27001; no disclosed corporate entity. Eleven subprocessors named but their handling terms are not enumerated.

Track record14/25
48/100WeakSep 10, 2026
Dragon Professional v16 — Windows desktop, fully offline (no Mac version since 2018; Dragon Home discontinued 2023)None transmitted in local desktop operation. Voice profile (acoustic data, custom vocabulary, dictionary) stored locally on the user's Windows machine.No

Dragon Professional v16 is a locally installed Windows desktop application that processes audio on-device once activated. Audio is not transmitted to Nuance/Microsoft servers in normal desktop operation. Product is Windows-only — Mac version was discontinued in 2018.

Training22/25
Profile and acoustic data stored locally on the user's machine; nothing sent off-device in desktop dictation mode.
Retention22/25
Yes — fully offline/local processing on Windows (x86). Apple Silicon Mac users have no native Dragon option since the 2018 Mac discontinuation.
On-device23/25

2017 NotPetya attack disrupted Nuance's hosted services (including Dragon Medical) for weeks. Development largely stalled since Microsoft's 2022 Nuance acquisition (v17 ≈ v16). Mac version discontinued 2018; Dragon Home (consumer) discontinued 2023. Non-medical product marketing pages now redirect to Microsoft's health-solutions hub.

Track record16/25
83/100StrongSep 10, 2026
Dragon Anywhere (iOS/Android) — cloud; end of sale July 1, 2026 (no new subscriptions or renewals)Audio and transcripts transmitted to Nuance/Microsoft cloud for recognition. Account/device metadata also collected.Yes — product improvement (Nuance Privacy Statement)

As of July 1, 2026 Dragon Anywhere Mobile "is no longer available for sale" and new subscriptions or renewals are no longer possible, per the App Store listing; existing subscribers continue to run a discontinued cloud product. The Nuance Privacy Statement (March 14, 2025) that still governs Dragon states voice data "may be manually reviewed and transcribed into text for Nuance's product operation, tuning, maintenance, and enhancement" and that "to build, train, and improve the accuracy of our automated methods of processing (including AI), we manually review some of the predictions." No Dragon Anywhere-specific opt-out is documented; Microsoft's general Privacy Statement (September 2026) also reserves the right to use data to "develop, train, and fine-tune our AI models."

Training8/25
Nuance Privacy Statement: data retained "for as long as you remain an active customer … and for 3 years afterwards." App Store privacy label lists audio recordings, email, product interaction, and crash data linked to the user.
Retention10/25
No — cloud recognition.
On-device3/25

Same 2017 NotPetya history. Post-acquisition product reorganization removed many Nuance trust-center URLs (now redirect to Microsoft health-solutions), creating a discovery gap that itself is a track-record drag for buyers trying to verify privacy commitments.

Track record14/25
35/100PoorSep 10, 2026
Free (MIT open-source) — macOS / Windows / LinuxNone on any Handy server — there is no Handy-operated cloud service and no account system. Audio runs through local Silero voice-activity detection and a locally downloaded speech model (Whisper, Parakeet, and others). The only default network call is a signed update check against GitHub releases, toggleable in Settings.No

No policy document exists — handy.computer/privacy returns 404 (checked September 10, 2026). The governing text is the homepage and README: "Your voice stays on your computer. Get transcriptions without sending audio to the cloud." Because no vendor server receives audio or text, there is no training path. The README roadmap still lists "Opt-in Analytics" ("Privacy-first approach with clear opt-in") as unshipped as of v0.9.6. An optional bring-your-own-key LLM post-processing path that sends transcript text (never audio) to a provider you configure was found in our July 2026 source audit; it is not documented in the README and was not re-audited this pass.

Training25/25
Nothing leaves the device. Transcript history lives in a local SQLite database under the user's control. No vendor retention surface.
Retention25/25
Yes — all transcription is local on macOS (Intel and Apple Silicon), Windows, and Linux. No cloud transcription endpoint exists in the codebase.
On-device25/25

Open-source MIT at github.com/cjpais/Handy; 31,300+ stars and 2,800+ forks; releases v0.9.4 (July 21), v0.9.5 (August 8), and v0.9.6 (August 24, 2026). No documented incidents, CVEs, or advisories. Deductions: no privacy policy document at all (procurement has nothing to file), solo maintainer CJ Pais with no legal entity behind the project (sponsor-funded: Wordcab, Epicenter, Bolt AI), and a roadmap analytics item to re-check each release.

Track record21/25
96/100ExcellentSep 10, 2026
Local mode (Free — MIT open-source client)None during transcription — Whisper or NVIDIA Parakeet run on-device via whisper.cpp and sherpa-onnx. The privacy policy (last updated August 12, 2026) states that in local processing "no audio is transmitted anywhere." An OpenWhispr account (email, name, hashed password) is only required for cloud features; optional diagnostic metrics cover "audio processing metrics (duration, format)" and timing data, not content.No

Nothing is transmitted in local mode, so there is no training path. The August 12, 2026 privacy policy adds a blanket commitment across all modes: "We never use your transcribed content, notes, or audio for advertising, marketing to third parties, or training our own AI models." The desktop client is MIT-licensed and auditable at github.com/OpenWhispr/openwhispr.

Training25/25
Nothing leaves the device. Local audio files are kept "up to 30 days" only if the user enables local audio retention; the user sets the window.
Retention25/25
Yes — local Whisper / Parakeet on macOS, Windows, and Linux with Metal, CUDA, and Vulkan acceleration.
On-device25/25

MIT client, 7,900+ stars, 975 forks; releases v1.9.0 (August 24), v1.9.1 (August 27), v1.9.2 (August 29, 2026). Legal entity Gizmo Labs Inc. (Delaware). No documented incidents. Deductions: the README's "No data collection, no telemetry" line predates the 2026 hosted product with accounts and sync; the August 2026 policy discloses that BYOK API keys are "stored in plaintext in your local app data directory."

Track record19/25
94/100ExcellentSep 10, 2026
BYOK cloud (Free — bring your own API key)Audio and/or transcript text routed with your own key directly to the provider you configure — the policy says BYOK audio is "sent directly to user's provider" (OpenAI, Groq, and Gemini transcription per the v1.9.1 release notes; LLM formatting can use further providers). OpenWhispr's servers are not in the path.Depends on provider

OpenWhispr does not train, but BYOK moves the training and retention question to the provider whose key you paste in; that provider's API terms govern. Key-handling caveat from the August 12, 2026 policy: "API keys you enter for third-party providers are currently stored in plaintext in your local app data directory." The v1.9.2 release (August 29, 2026) notes custom endpoints now fail safely instead of silently routing to default providers.

Training16/25
Not retained by OpenWhispr (not in the path). Provider-side retention follows the API you bring keys for.
Retention14/25
No — audio leaves the device for the configured provider.
On-device3/25

Plaintext local storage of BYOK keys is disclosed in the policy rather than fixed. No incidents.

Track record18/25
51/100WeakSep 10, 2026
OpenWhispr Cloud — Pro ($6.67/user/mo billed $80/yr)Audio sent to OpenWhispr's API and forwarded to named subprocessors — OpenAI OpCo, LLC (speech-to-text and language models), Parasail, Inc. (AI inference), and Groq, Inc. (inference and transcription failover) per DPA Annex 3 — plus account data in Neon, payments in Stripe, hosting on Vercel, and optional Google OAuth / Calendar. Synced notes and transcripts are stored for the account.No (contractual)

Terms of Service (effective July 23, 2026): "We will not use Customer Content — audio, transcriptions, notes, or AI agent prompts — to train any artificial intelligence or machine learning model." DPA v1.2 (updated August 21, 2026): OpenWhispr "will not use Customer Personal Data — including voice audio, transcriptions, notes, prompts and AI-agent conversations — to train, retrain or improve any artificial intelligence or machine learning model." Homepage adds: "If that ever changes, it will be opt-in and we will ask first."

Training23/25
Audio: "sent to our API, processed in real time, and discarded" (policy) — marketed as "0% Data Retention." Synced notes/transcripts persist for the account and are "deleted within 30 days of account deletion." Subprocessor-side retention at OpenAI / Groq / Parasail is not enumerated in the policy.
Retention21/25
No — audio is processed through the managed-cloud subprocessor chain. The same app can be switched back to local mode.
On-device3/25

SOC 2 Type 2 (report period May 1–July 31, 2026), ISO/IEC 27001:2022 (valid through July 27, 2027), and HIPAA / GDPR attestations listed in the DPA and on the Comp AI-hosted trust center; auditor not named on either page. No documented incidents. Young company (Delaware, product line expanded to hosted accounts in 2026); the README still says "no telemetry" while the policy describes optional diagnostics.

Track record19/25
66/100AdequateSep 10, 2026
Free (GPLv3 open-source) — macOS 15+No audio or text on any vendor server. README: "Your voice, audio, and transcribed text never leave your machine unless you explicitly opt in to a cloud AI provider." What does reach the vendor by default is anonymous analytics: "a random installation ID, activity date, app version, and macOS platform label" plus daily feature/model usage totals, onboarding progress, and model-download outcomes — the README states voice, audio, transcripts, prompts, and clipboard content are not collected.No

No vendor receives audio or transcript text in the default configuration, so there is no training path. Speech models (Whisper, Parakeet, Nemotron, Cohere Transcribe, Apple Speech) are local downloads. The enhancement model Fluid-1 (about 3.5 GB) is closed — "We're keeping Fluid Intelligence private for now so we can sustainably offer the core dictation experience for free" — but runs offline on the Mac. Opt-in cloud enhancement sends transcript text to OpenAI, Groq, or a custom provider under that provider's terms; keys are stored in the macOS Keychain. There is no formal privacy policy document; the README and altic.dev/fluid (last updated February 23, 2026) are the governing text.

Training25/25
Audio and transcripts stay on the Mac. Analytics metadata is sent to the vendor; its retention window is not published.
Retention23/25
Yes — all speech-to-text and default enhancement are local. Apple Silicon required for every model except Whisper (Intel supported from v1.5.1 via Whisper only). "Detailed anonymous analytics" are on by default and can be disabled at Settings → Share Detailed Anonymous Analytics.
On-device25/25

GPLv3 since February 23, 2026 (Apache 2.0 before); 11,400+ stars, 811 forks; releases v1.6.7 (August 5), v1.6.8 (August 11), v1.6.9 (August 18, 2026). No documented incidents. Deductions: default-on analytics in a local-first app, the closed Fluid-1 model cannot be audited, no privacy policy document, and a small team with no published legal entity or compliance attestations.

Track record19/25
92/100ExcellentSep 10, 2026
Free (1,000 words/mo) / Pro ($15/mo or $12/mo billed yearly)Audio and dictated text pass through VoiceDash's pipeline to OpenAI on every dictation — the privacy policy states: "Your voice recordings are transmitted to OpenAI for the sole purpose of real-time transcription and response generation." The DPA (effective May 26, 2026) names nine subprocessors: OpenAI, Groq, Cloudflare, Hetzner, Sentry, Metabase, CustomerIO, RevenueCat, and Apple. Account and billing data are retained; content the user saves to Notes is stored.No (policy + DPA)

Privacy policy (undated): "Your voice data is strictly excluded from our machine learning training sets. Your input does not influence the development of VoiceDash AI models." DPA: data is "processed transiently, structurally excluded from AI model training loops." OpenAI's API data-usage page confirms API inputs are not used for training by default. Groq appears in the DPA subprocessor list but not in the privacy policy.

Training22/25
VoiceDash: "does not store your voice files or transcriptions on our servers" (exception: user-saved Notes). Second perimeter: OpenAI retains API abuse-monitoring logs up to 30 days by default; the founder's AppSumo claim that VoiceDash holds a Zero Data Retention arrangement with OpenAI is unverified — it appears in neither the privacy policy nor the DPA.
Retention17/25
No — cloud-only; no local or offline mode is offered.
On-device3/25

Jumpstart Ventures FZE, Dubai Silicon Oasis (disclosed in the DPA, not the privacy policy); founded February 2025. No documented incidents. AppSumo 4.42/5 from 234 reviews. Deductions: privacy policy carries no effective date; terms of service (October 27, 2025) governing-law clause reads "the laws of us"; no SOC 2, ISO 27001, or HIPAA attestation published; DPA subprocessor list (Groq) is broader than the policy's disclosure.

Track record16/25
58/100AdequateSep 10, 2026
Local mode on Apple Silicon ($14.99/mo, $99/yr, or Lifetime local-only)No audio during transcription — docs: "When the active speech-recognition backend is local, speech audio is transcribed on your Mac." Account, purchase, analytics, diagnostics, download, and update flows still use the network; the privacy policy (last updated July 23, 2026) lists PostHog, Cloudflare Web Analytics, Ahrefs, Simple Analytics, and Dub.co for analytics and Hetzner, Supabase, and Apple for hosting.No (architectural; policy silent)

In local mode nothing reaches a server to train on. The privacy policy makes no statement about AI training either way — an omission worth noting because it means cloud-mode users have no written commitment. Data controller is Burlis Management GmbH, Philippstrasse 27, 52349 Dueren, Germany (an EU entity under GDPR). Local models require Apple Silicon and macOS 14+; "Intel Macs use cloud models with a subscription."

Training24/25
Audio: "Audio input is temporary during transcription. The last audio file is stored locally on the user's device until replaced by a new recording or deleted by the user." Account data held "for as long as required by the purpose"; raw journey events cleared after 90 days.
Retention23/25
Yes on Apple Silicon — "Supported local modes can run offline after setup." Not available on Intel Macs. Cloud Cleanup must stay off.
On-device25/25

Named EU GmbH controller and named subprocessors are creditable. No documented incidents. Deductions: no SOC 2 / ISO 27001 / HIPAA / BAA at any tier; policy silent on training; iOS Voice Keyboard privacy label (v1.0.4) lists Audio Data under "Data Linked to You"; Cloud Cleanup shipped in 2026 after a local-first launch, so the default is local but the product is no longer local-only.

Track record17/25
89/100ExcellentSep 10, 2026
Cloud mode — Deepgram transcription + Cloud Cleanup (same plans; Intel Macs cloud-only)Audio to Deepgram: "When a user selects or is eligible for cloud transcription, Paraspeech may transmit selected audio and related request metadata to Deepgram for cloud transcription." With Cloud Cleanup on, "selected transcript text, prompt text, rewrite instructions, and related request metadata" go to Groq and Cerebras. Docs: Cloud Cleanup "uses cloud processing and requires internet access. It is available only when cloud processing is allowed."Not stated

The July 23, 2026 privacy policy names Deepgram, Groq, and Cerebras as processors but contains no statement on whether Paraspeech or those providers use audio or text for model training. Each provider's own terms govern; Paraspeech publishes no contractual no-training commitment. "Or is eligible for" covers Intel Macs, where every dictation is cloud by hardware rather than by choice.

Training11/25
Not published for audio or transcripts. Policy: provider-side handling "depends on the applicable provider and processing terms." Personal data kept "for as long as required by the purpose they have been collected for."
Retention11/25
No — audio leaves the Mac for Deepgram; Cloud Cleanup text leaves for Groq / Cerebras.
On-device3/25

EU GmbH controller, named processors, no incidents; no attestations, no BAA, policy silent on training.

Track record17/25
42/100WeakSep 10, 2026
Free trial (30 min) / Pro ($8.49/mo annual) / Lifetime ($260)Audio recorded on the device and relayed through Voicy's Heroku-hosted servers to Groq, which runs the open-source Whisper V3 model for transcription and post-processing. Name and email for billing; anonymized Mixpanel usage analytics (button clicks, recording counts, errors) that can be disabled; in some cases anonymized data about the website Voicy is used on.No (homepage only)

The homepage states: "We do not use your recordings to train an AI model or for any other purpose." Neither the Data Handling and Security Policy (v1.3, last updated 31/07/2025) nor the undated privacy policy contains the word "train" — the commitment is a marketing-page statement, not a policy clause. Groq-side: the privacy policy says "we have enabled their Zero Data Retention settings and we do not use Groq's batch or fine-tuning endpoints." No SOC 2, ISO 27001, or HIPAA BAA published.

Training16/25
Security policy: "No audio recordings are stored by Voicy or Groq - all audio data is permanently deleted immediately after processing" and "All transcribed content is immediately deleted after delivery to the user." Transcripts live only on the user's device. Account data "stored locally on the user's device, not in centralized databases." Hosting location conflicts between documents: the privacy policy says servers "are hosted in the EU and US. US servers are reserved for US users"; the security policy says "All processing occurs in USA-based infrastructure."
Retention21/25
No — cloud-only at every tier (Mac, Windows, Linux, Chrome extension, iOS, Android); no offline mode.
On-device3/25

No documented incidents. Security policy dated July 31, 2025 has not been updated to cover the 2026 iOS/Android apps or the EU hosting now mentioned in the privacy policy; no SOC 2 / ISO 27001 / HIPAA; entity Pishi LLC FZ (UAE) disclosed only in the security policy footer.

Track record15/25
55/100AdequateSep 10, 2026
Free / Pro ($9.99 App Store IAP) / AppSumo lifetime ($59 Tier 1)Per secondary reporting (vendor policy page unreachable on 2026-09-10 — Vercel security checkpoint, HTTP 429): microphone audio transmitted to Blip AI's cloud for GPT-powered transcription and formatting; anonymized usage statistics; account and billing data. The App Store privacy label states "The developer does not collect any data from this app," which conflicts with the policy-described usage statistics. The policy does not name its subprocessors, though a third-party LLM provider sits in the path.Not addressed (unverified)

Per our June 5, 2026 review of the policy, Blip AI's privacy policy does not make an explicit commitment that dictation is not used to train models, and it does not name the GPT provider that processes audio. The AppSumo listing states: "Your audio isn't stored, and your text stays completely private to you." We could not re-read the live policy on September 10, 2026; treat the training posture as not published. No SOC 2, ISO 27001, or named auditor; HIPAA compliance and a BAA "on request" are asserted per secondary reporting with no published third-party backing.

Training10/25
Per secondary reporting: audio deleted immediately after transcription ("typically within seconds"); transcribed text not stored on Blip AI servers, delivered only to the device; anonymized usage statistics retained up to two years. Unverified live on 2026-09-10.
Retention14/25
No — cloud-only (macOS 12+ Apple Silicon, Windows, Android); no offline mode.
On-device3/25

No documented incidents. Solo/bootstrapped developer (Ayush Bansal), product launched October 2025; App Store 2.0/5 from 2 ratings (v1.0.5, July 19); App Store label "Data Not Collected" conflicts with the policy; HIPAA claim without SOC 2 or named auditor; privacy page blocked to automated verification.

Track record12/25
39/100PoorSep 10, 2026
Local mode (default; free during early access)None transmitted during transcription: "By default, all speech-to-text processing happens entirely on your device using WhisperKit." Smart Typing cleanup "runs a local Llama 3.2 3B model entirely on your device via Apple's MLX framework." Transcript history, dictionary, snippets, settings, and API keys are stored in a local SQLite database and "never transmitted to our servers." Optional PostHog telemetry (random install ID, feature events, word counts) is described as "disabled by default."No (structural; policy silent)

The privacy policy (effective April 29, 2026) contains no training statement — the word "train" does not appear. In local mode nothing reaches a WisprType server, so there is no vendor-side data to train on. Caveat from our July 2026 hands-on test: v1.1.0 shipped with telemetry enabled on a fresh install despite the "disabled by default" wording; v1.1.0 is still the current build. No legal entity, terms of service, or compliance attestation is published.

Training22/25
Audio: "processed in real time and are not retained after transcription completes, unless you explicitly save them." Transcript text and target-app names persist in the local SQLite database until the user deletes them; nothing is synced to iCloud or external servers. Telemetry, when enabled, is retained by PostHog under PostHog's policy.
Retention23/25
Yes — WhisperKit speech-to-text and Llama 3.2 3B cleanup on Apple Silicon; no network calls during transcription in local mode (macOS 12+, Apple Silicon required).
On-device24/25

Closed-source; no legal entity, terms of service, or attestation; domain registered April 28, 2026; our July 2026 test found v1.1.0 telemetry enabled on install despite the policy's "disabled by default" statement; no App Store listing or third-party review surface.

Track record10/25
79/100StrongSep 10, 2026
BYOK cloud mode (OpenAI, Groq, or Deepgram — bring your own API key)Audio (WAV), model name, and optional prompt or keyword hints sent directly to the provider you configure; in cloud mode Smart Typing also sends raw transcript text and dictionary words to "the same cloud provider's chat LLM." "Cloud providers are never used unless you explicitly configure them by entering an API key and selecting cloud mode in Settings." API keys stay local.Depends on provider

WisprType has no server in the path; the training and retention question transfers to OpenAI, Groq, or Deepgram under your own API agreement. The policy links each provider's privacy page and says "Please review each provider's privacy policy to understand how they handle your audio data."

Training16/25
WisprType: audio not retained; transcripts stored locally. Provider-side retention follows the API terms of the provider you bring keys for.
Retention14/25
No — audio leaves the device for the configured provider.
On-device3/25

Same: no legal entity, no attestations, telemetry-default discrepancy found in v1.1.0.

Track record10/25
43/100WeakSep 10, 2026
Free (2,000 words/mo) / Pro ($7/mo or $69/yr; App Store IAP $7.99/mo, $79.99/yr)Microphone audio, dictated text, selected text, formatting instructions, vocabulary, snippets, and custom AI instructions, processed by DictaFlow and its providers; the August 14, 2026 policy names "OpenAI, Deepgram, and Groq for transcription and related AI processing," Apple/Google/Microsoft for sign-in, Google Cloud for hosting, Stripe for payments. Usage data includes word and minute totals, plan state, app version, a random installation identifier, and request identifiers. Local processing is offered with an optional cloud cleanup step; the vendor's facts page says the product is "never 'fully offline'." Operated by SmartBids.ai Corp., a Canadian corporation.No (homepage + App Store label)

The homepage states "your audio is never used to train models," and the App Store privacy label places Audio Data under "Data Not Linked to You." The privacy policy itself (last updated August 14, 2026) contains no training clause; it commits only that DictaFlow does "not sell personal information or use dictated audio or text for targeted advertising." The policy states the standard service is "not intended for medical dictation ... or any use involving protected health information (PHI)" and directs PHI users to DictaFlow Medical. No SOC 2 or ISO 27001 published.

Training15/25
"We do not permanently retain audio recordings on our servers after processing in the ordinary course" and "Audio is ordinarily discarded after processing." Transcription text "may be retained when needed to provide features you choose, synchronize supported devices, recover a requested result, investigate a support issue, or comply with law." No retention window in days is published. Local history and notes stay on the device until deleted; account deletion available via Settings → Delete account.
Retention15/25
Partial — local processing available on device; cloud cleanup (OpenAI, Deepgram, Groq) is optional per dictation. Windows, Mac, iPhone, iPad, Android.
On-device14/25

No documented incidents. Named Canadian corporation (SmartBids.ai Corp.) and named developer (Ryan Shrott); policy updated August 14, 2026 and it no longer names NVIDIA as a provider; App Store 4.4/5 from 12 ratings (v1.2.41, Aug 28). No SOC 2 / ISO 27001; consumer plan explicitly excluded from PHI use.

Track record17/25
61/100AdequateSep 10, 2026

Meeting Transcription

AI meeting assistants — bot-based notetakers and meeting transcription apps. This category had its privacy reckoning in 2025–2026: Otter, Fireflies, and now Granola have all been sued over recording meeting participants without consent (Granola's July 2026 wiretap class action targets its no-bot design — participants get no notice at all), and The Verge reported Granola's default-public note links and opt-out training. Audio + transcripts of meetings are uniquely high-stakes because participants often haven't consented individually, and 12 US states require all-party consent for recording.

ToolPlan tierData collectedTrains on your data?RetentionOn-deviceTrack recordTotalLast verifiedSource
Free / ProAudio recordings, transcripts, meeting metadata, speaker identification, account infoYes (opt-out)

Otter trains automatically on de-identified user data; recordings and transcripts are not manually reviewed by humans unless the customer gives explicit consent for support troubleshooting. Training data is encrypted. Neither the privacy & security page nor the privacy policy (effective June 16, 2026) describes an in-app training toggle: opt-out is by request via privacy@otter.ai, the trust-center form, or the "Do Not Sell or Share My Personal Information" link.

Training8/25
Conversations stored until manually deleted; trash holds for 30 days then auto-purges.
Retention8/25
No (cloud)
On-device3/25

In re Otter.AI Privacy Litigation, No. 5:25-cv-06911-EKL (N.D. Cal., Judge Eumi K. Lee) — four suits filed Aug–Sept 2025 (lead: Brewer v. Otter.ai), consolidated October 22, 2025, alleging unauthorized recording and use of meeting data to train AI without all-participant consent. On August 13, 2026 the court denied the motion to dismiss as to the federal Wiretap Act, CIPA § 631, both Illinois BIPA voiceprint claims, unjust enrichment, and the UCL, holding plaintiffs plausibly allege Otter independently collects, retains, and uses communications — including for model training — for its own commercial purposes; CFAA, CDAFA, Washington Privacy Act, and most intrusion claims were dismissed with leave to amend. The case proceeds toward discovery and class certification. Visible bot creates two-party consent issues. SOC 2 Type 2.

Track record6/25
25/100PoorSep 10, 2026
Free (800 min/mo) / ProAudio, video recordings, transcripts, speaker identification, meeting metadata. Voice characteristics used to separate speakers are processed by service providers and, per the policy, never received or stored on Fireflies' own servers; Illinois biometric data is destroyed within 3 years of last interaction.No (stated policy)

Fireflies' privacy policy (last updated March 6, 2026) states: "We do not use personal information for AI model training and we contractually prohibit our vendors from using this information for their own model training," and its Zero Data Retention section says meeting content is not used for training by third-party processors. Fireflies' help center repeats that it "does not use your meeting content or personal data to train AI models." Earlier versions of this row recorded a default-on training posture; the current policy text supersedes that.

Training20/25
Stored until manual deletion; account deletion processed within 30 days per the policy. Servers in the United States and other countries; transfers under DPF / SCCs.
Retention5/25
No (cloud)
On-device3/25

Cruz v. Fireflies.AI Corp., No. 3:25-cv-03399 (C.D. Ill., filed December 2025) — a BIPA class action alleging Fireflies captured all meeting participants' voiceprints without written consent — was terminated in March 2026 and effectively refiled in the Northern District of Illinois. Two related N.D. Ill. BIPA suits — Fricker v. Fireflies.AI Corp., No. 1:26-cv-02675, and Martinez v. Fireflies.AI Corp., No. 1:26-cv-03512 — were consolidated before Judge Steven C. Seeger with Fricker as the lead case and Werman Salas P.C. appointed interim class counsel; no ruling on the merits as of September 2026. The visible "Fireflies.ai Notetaker" bot creates two-party-consent issues. GDPR compliant.

Track record5/25
33/100PoorSep 10, 2026
Free / ProAudio is transcribed in real time on macOS/Windows and not stored; transcripts and user notes are stored on AWS. iOS uses temporarily cached audio.Yes (opt-out)

Granola uses de-identified user data to train internal AI models supporting its services by default; the privacy policy effective September 8, 2026 still says "We only use de-identified data to train AI models, which you can opt-out of within your Granola account settings." The opt-out exists only for Granola users: other meeting participants have no opt-out path and get no notice unless the user enables Granola's optional in-meeting chat notice or video watermark (added April 2026). Third-party LLM providers (OpenAI, Anthropic) are contractually prohibited from training on Granola data.

Training10/25
Audio not stored. Transcripts and notes stored in US-hosted AWS Virtual Private Cloud, encrypted at rest and in transit, backed up daily. Notes are accessible to anyone with the share link by default — link sharing is on, not off, until the user changes it.
Retention13/25
No. Despite real-time transcription on the user's device, audio is sent to cloud transcription providers (Deepgram, AssemblyAI) and summarization uses cloud LLMs (OpenAI, Anthropic).
On-device3/25

Chamberlain v. Granola, Inc., No. 3:26-cv-07926 (N.D. Cal., filed July 30, 2026) — putative class action against Granola, Inc. and Granola Labs Ltd. alleging Granola intercepts and records virtual-meeting participants without knowledge or consent and uses their communications to train its AI by default, under the federal Wiretap Act (ECPA), CIPA §§ 631–632, CDAFA, intrusion upon seclusion, the UCL, and unjust enrichment. Granola captures system audio without joining as a visible bot, so non-user participants get no notice; the complaint quotes Granola's marketing that "other people in the room won't know." The case is assigned to Judge Edward M. Chen; Granola, Inc. was served August 5, 2026, and on August 20, 2026 the parties stipulated to extend defendants' deadline to respond — no responsive pleading or ruling yet. Earlier, April 2 2026 — The Verge reported notes are accessible via share link by default and users are opted into internal AI training by default, contradicting the app's "private by default" marketing. No documented breaches.

Track record5/25
31/100PoorSep 10, 2026
AI Companion (included on paid Zoom plans; on by default, admin-controllable)Meeting audio, video, chat, screen share, attachments, transcripts, summaries, action items. Federated model approach uses Zoom-hosted models plus third-party subprocessors (OpenAI, Anthropic) and Perplexity for web search.No

Per Zoom's Privacy Statement and AI Companion Security & Privacy page, Zoom does not use customer audio, video, chat, screen sharing, attachments, or other communications-like content to train Zoom's or third-party AI models. Stated contractual prohibition, not an opt-out. The Privacy Statement was last updated July 27, 2026 (messaging-campaign data, EU/UK representative contacts, transfer mechanisms, DSAR-by-phone); the no-training language is unchanged.

Training20/25
Customer content processed to deliver features, retained for a period, shareable internally within the org account (host + admins), may flow to subprocessors. Zoom-hosted Models Only (ZMO) deployment limits processing to Zoom infra. BAA available for enterprise (covers specific configs only).
Retention12/25
No — cloud processing.
On-device3/25

August 2023: Zoom quietly added ToS language permitting training on customer content; after The Verge coverage and public outcry, walked it back and committed to no-training-without-consent, later hardened to an absolute prohibition. No known breaches of AI Companion content.

Track record12/25
47/100WeakSep 10, 2026
Free / Premium ($20/mo, $16/mo annual)Audio and video recordings, transcripts and speaker identification, meeting notes, calendar data (Google / Outlook), attendee names and emails, voiceprints in the iOS app (consent-based), device and usage data. All Fathom data is stored in the United States.Yes, de-identified (opt-out)

The privacy policy (last updated August 16, 2026) states: "Based on how you configure your account settings, we may use and create de-identified data generated from Meeting Content Information to improve our Service by training, improving, and customizing our in-house artificial intelligence models. You can opt out from this use of your data in your account settings." The help center confirms the default is on ("Fathom uses de-identified customer data to improve the accuracy of our proprietary AI models") with an in-app opt-out under Settings. Third-party AI sub-processors — Anthropic, OpenAI, Google — are "not contractually permitted to use our users' data to train their AI models."

Training9/25
No fixed window published: data kept "as long as necessary to fulfill the purposes." On account deletion, recording data and metadata are removed and backups purge after a further 7 days; deletion requests processed within 30 days.
Retention9/25
No (cloud). Recording is via a visible meeting bot on Zoom, Google Meet and Teams, plus a bot-free Chrome capture mode.
On-device3/25

No lawsuit, breach, or regulator action naming Fathom found as of September 10, 2026 (the 2026 AI-notetaker wiretap wave names Otter, Fireflies and Granola, not Fathom). Vendor states SOC 2 Type II, HIPAA and GDPR compliance; no ISO 27001. Trust center contents (SOC 2 report) unverified — page requires JavaScript. Visible bot creates two-party-consent obligations for the host.

Track record20/25
41/100WeakSep 10, 2026
Free / Pro (prices unverified — pricing page returned 404)Video and audio recordings, transcriptions, name, email, profile image, IP and location, device and usage data, billing data. Hosted primarily in the EEA (Google Cloud, Hetzner, Wasabi — Germany and Finland); AI processing region selectable US or EU. Named AI processors: Anthropic and Google Vertex AI; transcription: AssemblyAI and ElevenLabs.No (contractual, incl. third-party AI)

The privacy policy (last updated September 9, 2026) states Tldx Solutions GmbH does not use Customer Content "to train, fine-tune, or improve foundation models, large language models, or other generative AI models," and that third-party AI providers process Customer Content "solely to provide the requested functionality" and do not use it to train their general-purpose models. The security page repeats: "No customer data is used to train the AI." No toggle is needed because the exclusion is unconditional.

Training22/25
Free users: recordings deleted after 3 months. Paid users: retained until account deletion. Logs 10 days; data on non-client meeting participants up to 5 years or until a deletion request.
Retention15/25
No (cloud, EU-hosted). Visible meeting bot on Zoom, Google Meet and Teams.
On-device3/25

Security exposure, not a lawsuit: researcher bobdahacker reported on January 28, 2026 that a missing tenant-isolation rule on tl;dv's Firestore "meetings" collection let any authenticated account enumerate 181,874 meeting records across 84,312 users (creator email, conference ID, provider, recording status, timestamps), including roughly 1,000 joinable live calls and 1,000+ publicly shared meetings exposing 715 invitee emails. Follow-ups through July 22 went unresolved; public disclosure August 4, 2026. tl;dv told GIGAZINE (August 13, 2026) these were "two separate vulnerabilities," that no passwords, audio, transcripts, AI notes or billing data were reachable, the second vector was patched within 24 hours of discovery, and Firebase has since been removed from its infrastructure. No regulator notification is documented. No lawsuit found. SOC 2, ISO 27001, PCI DSS, GDPR stated; no HIPAA.

Track record9/25
49/100WeakSep 10, 2026
Free / Pro ($19.75/mo, $15/mo annual)Meeting audio and video, transcripts, meeting subject and description, and — when Google integrations are connected — Gmail messages (subject, body, recipients, senders, metadata), Google Calendar events and guest lists, and Google Chat messages. Servers in the United States.No by default (opt-in)

Read's privacy page states: "Contributing to model improvement is off by default. You decide - and can change your mind any time." The privacy policy (last updated June 29, 2026) describes a Customer Experience Program you "opt-in or opt-out" of in account settings, and states Read does not, and does not permit third-party AI tools to, use Google Workspace API data "to develop, improve, or train generalized/non-personalized AI and/or ML models." Third-party AI providers are not named in the policy.

Training21/25
Audio and video stored "in no case for longer than 2 years." Users can delete meeting reports they own; participants who type "opt out" have the meeting's data "immediately and permanently deleted." No shorter default window published.
Retention13/25
No (cloud). Bot-free options exist — a native Google Meet Media API integration (March 26, 2026, free for all customers) and desktop / mobile apps that record "locally or to the cloud without a bot" — but transcription and analysis remain cloud-side.
On-device3/25

No lawsuit, breach, or regulator action naming Read AI found as of September 10, 2026. Annual SOC 2 Type II audit, HIPAA technical safeguards with BAA on request, GDPR Article 28 processor role for eligible EU Workspace customers, AES-256 / TLS 1.2+. Read announces itself and requires host approval before joining. Support-center security articles and trust.read.ai could not be fetched (403 / JavaScript), so the SOC 2 report is unverified.

Track record20/25
57/100AdequateSep 10, 2026
Free (120 min/mo) / Pro ($13.61/mo, $8.17/mo annual) / Business ($27.78/$16.67)Audio, video, voice and text inputs for transcripts, summaries and translations; account, device, IP, location and usage data; integration identities from Apple, Google and Microsoft. Hosted on AWS (region not stated). Contracting entity is Notta PTE. LTD. (Singapore), or Notta 株式会社 for Japan users.Not published

Neither the privacy policy (effective September 2, 2025) nor the terms (effective September 2, 2026) nor the security page states whether Notta trains its own models on uploaded or recorded audio. The only training sentence is scoped to one data source: "Google Workspace APIs are not used to develop, improve, or train generalized/non-personalized AI and/or ML models." The policy lists "To analyze and improve our Services" as a purpose. Secondary blogs describe an opt-out setting or a support-request opt-out; that is unverified — no Notta page fetched shows a training toggle.

Training8/25
Not published in the privacy policy. Per a Notta help-center search snippet (the article itself returned 403), transcription data is stored indefinitely and is not automatically deleted unless the user, workspace or account deletes it; automatic file deletion is an Enterprise-only setting.
Retention7/25
No (cloud). Notta Desktop (public beta, macOS 13+ / Windows 10+) captures system and microphone audio locally without a bot, but transcription runs in Notta's cloud.
On-device3/25

No lawsuit, breach, or regulator action naming Notta found as of September 10, 2026. Security page states SOC 2 Type II via independent audit, ISO 27001 ("follows the guidelines"), GDPR, CCPA and HIPAA ("Follow HIPAA guidelines" — no BAA offered on any fetched page); AES-256 / TLS 1.2 on AWS. Deducted for a policy that does not address training on its own models, publishes no retention window and no data-center region, and for help-center pages that could not be fetched.

Track record15/25
33/100PoorSep 10, 2026
Personal account on Google AI Pro / AI UltraMeeting audio is transcribed by Gemini; a notes document (summary, action items) and optional transcript are saved to a Google Doc in the organizer's Google Drive "Google Meet" folder and emailed as a recap. All participants are notified on screen. Available to Google AI Pro and Ultra subscribers since June 29, 2026 (not AI Plus).Not published for Meet notes

No Google page fetched says whether "Take notes for me" transcripts on personal accounts are exempt from training. The Gemini Apps Privacy Hub (updated August 10, 2026) states Google uses Gemini Apps Activity — "chats and what you share with Gemini," "audio, Gemini Live videos and screenshares" — to "provide, develop, and improve its services (including training generative AI models)," and that "a subset of chats are reviewed by human reviewers." Whether Meet notes fall under that activity or only under the general Google Privacy Policy is unverified; the Workspace no-training commitment explicitly applies to work and school accounts, not personal ones.

Training10/25
The notes Doc and transcript persist in Drive until the organizer deletes them. Gemini Apps Activity, if applicable, auto-deletes after 18 months by default (adjustable to 3 or 36 months; 72 hours when activity is off).
Retention10/25
No (Google cloud).
On-device3/25

Thele v. Google LLC, No. 5:25-cv-09704 (N.D. Cal., Judge Noël Wise), filed November 11, 2025, alleges Google switched Gemini "smart features" on by default for Gmail, Chat and Meet on or about October 10, 2025, without consent, under CIPA, the Stored Communications Act, CDAFA, intrusion upon seclusion and the California Constitution. On July 7, 2026 the court granted Google's motion to dismiss the first amended complaint for lack of concrete harm, with leave to amend; plaintiffs filed a Second Amended Complaint on August 18, 2026 (adding Edward Goldstein) and Google stipulated to extend its response deadline on August 20. No ruling on the merits. Google states its smart features are optional and user-controlled. The suit concerns the smart-features default, not "Take notes for me" itself, which notifies all participants.

Track record12/25
35/100PoorSep 10, 2026

Clinical Documentation (Ambient AI Scribes)

Ambient scribes that listen to a patient visit and draft the clinical note. Every vendor here signs a Business Associate Agreement, so the differentiator is not HIPAA paperwork but what happens to the audio and the de-identified transcript afterwards: several vendors train their models on de-identified visit data by default, and the retention window for raw audio ranges from minutes to indefinite. Microsoft's Dragon Copilot is scored in the Dragon row of the Voice & Dictation table. Scores here reflect published policy, not a compliance assessment — verify the BAA scope before any PHI touches the tool.

ToolPlan tierData collectedTrains on your data?RetentionOn-deviceTrack recordTotalLast verifiedSource
Abridge for Clinicians app — no self-serve plan; requires a health-system deploymentVisit audio, transcripts, and generated clinical notes, all processed as Customer Data under the health system's BAA. The App Store privacy label lists audio data, email, phone number, photos/videos, user ID, and customer-support info as linked to the user; an optional voiceprint feature is retained "only as long as needed to provide the feature to you." Abridge sells to health systems only — there is no individual sign-up or published price.Yes — de-identified (contractual, no published opt-out)

Abridge states it trains on de-identified clinical data. Its privacy policy as archived (effective December 1, 2022) said: "We use de-identified data for research and development of new products or tools, to refine our algorithms and machine learning applications, and to improve the Services." The current policy (last updated August 27, 2026) removes that sentence and instead says Customer Data "is governed exclusively by the terms of our agreements with those customers, including our Business Associate Agreements (BAAs)" — so training terms are now contractual and not public. On June 11, 2026 Abridge and Nvidia announced a clinical-conversation model trained on de-identified doctor–patient conversations from Abridge's platform. Any opt-out is contractual at the health-system level; an individual clinician or patient has no published opt-out.

Training9/25
Not published by Abridge. Customer deployments publish their own windows: one pediatric group's FAQ states "Audio recordings and transcripts are automatically deleted after 30 days"; Kaiser Permanente told The Markup (June 16, 2026) recordings are "stored for no longer than 14 days." Finalized notes persist in the EHR.
Retention13/25
No — cloud (Google Cloud listed as a data vendor on the trust center).
On-device3/25

Abridge is not a defendant in any suit, but three class actions target its deployments: Saucedo v. Sharp HealthCare (San Diego Superior Court, filed Nov 26, 2025) alleges recording without all-party consent under CIPA and that false "patient consented" statements were inserted into charts; Washington et al. v. Sutter Health and MemorialCare (N.D. Cal., filed April 8, 2026) alleges CIPA, CMIA, UCL, and federal Wiretap Act violations. The Markup (June 16, 2026) reported Kaiser clinicians could not get answers on where recordings are stored or who can access them. Abridge did not respond to press requests in that reporting. SOC 2 Type 1 & 2, TX-RAMP. Vendor states HIPAA-aligned and signs BAAs.

Track record11/25
36/100PoorSep 10, 2026
Individual clinician — Starter $39/mo (40 notes), Core $79/mo, Premier $104/mo annual ($119 monthly); 7-day trialVisit audio ("Patient Recordings"), transcripts, generated notes, templates, and account data. Optional opt-in voice ID. PHI is encrypted at rest and in transit (TLS 1.2–1.3) and hosted on Microsoft Azure in the United States under Freed's own BAA with Microsoft.De-identified notes, per a Platform setting (default not published)

Freed states it trains on de-identified clinical notes. The security page says "AI is only trained on de-identified notes"; the Platform Terms of Use (last updated August 13, 2026) say Freed may use De-identified Data "to train the AI/ML models to improve the Platform ... with your consent provided via Platform settings," and will "link your De-identified Data with your customer ID and use it to customize and train our Platform based on your specific styles." The Terms also give Freed "the right in our sole discretion to use De-identified Data and to disclose such De-identified Data to third parties." Whether the consent setting is on or off at sign-up is not published — treat it as a setting to check on day one. PHI "is never used for AI training purposes."

Training13/25
Audio: "Patient recordings are saved only until the note has been completed, then automatically deleted"; the Terms add a settings option to "delete the Patient Recordings immediately once they are processed." Notes: user chooses 30-day auto-delete or retention for the term of the agreement.
Retention18/25
No — cloud (Microsoft Azure, US).
On-device3/25

SOC 2 Type 1 and Type 2; vendor states HIPAA- and HITECH-aligned and signs BAAs (the BAA is Schedule A of the Platform Terms). No documented breaches, lawsuits, or regulator actions found as of September 2026. Privacy policy (Aug 11, 2026) is short on specifics — it does not itself describe audio deletion or the training setting; those live in the Terms and marketing pages.

Track record20/25
54/100WeakSep 10, 2026
Free (unlimited transcription) / Clinician (14-day trial; price not shown on pricing page)Consultation audio (streamed for transcription, not stored), transcripts, generated notes and documents, Ask Heidi queries, account and payment data, device data. Health information is stored in the user's own jurisdiction (AU, UK, US, EU, CA); some non-sensitive functions may use offshore third-party services.No (stated) — de-identified / aggregate use reserved in policy

Heidi states it does not train on clinical data. The support article "How Heidi protects your data" (July 23, 2026) says: "Heidi does not use consultation data, session transcriptions, or generated notes to train its AI model in any way." The privacy FAQ says: "We don't use any of your sensitive health information for model training." Two caveats: the privacy policy (page shows October 2024) still permits Heidi to "de-identify and/or aggregate your personal information, including your health information" for platform development, and for Ask Heidi it says "Queries may be reviewed and used in de-identified form to improve the platform, but PHI will not be used for model training." No opt-out is needed for scribe data because no training is stated; there is no published opt-out for de-identified Ask Heidi query review.

Training19/25
Audio: "Heidi does not store audio recordings of patient consultations." Notes and transcripts: user-configurable "between 1 day and 'never delete'"; the FAQ describes the default as never delete. Privacy policy gives no fixed window ("as long as we need it or are lawfully required").
Retention15/25
No — cloud; "local, private servers in every region it operates in" (AWS and Azure environments named for enterprise).
On-device3/25

ISO 27001, ISO 42001, SOC 2 Type 2 per Heidi's pages; vendor states HIPAA-aligned and makes a BAA "available to every covered entity we work with" (no plan restriction stated). $65M Series B (2026). No documented breaches, lawsuits, or regulator actions found as of September 2026. Minor drag: the privacy-policy page still displays an October 2024 date while the site's help content has been revised through July 2026.

Track record20/25
57/100AdequateSep 10, 2026
Self-serve clinician plan (prices not published on nabla.com)Encounter audio streamed for transcription (not stored by default), transcripts, generated notes, and account data. Optional feedback submissions can include de-identified audio. Data is hosted on Google Cloud in the region chosen at organization creation (U.S. regions for U.S. clients, EU/Belgium for non-U.S.); Azure is used for speech-to-text in U.S. regions.Only on opt-in feedback (de-identified)

Nabla does not state that it trains on routine encounter data; its trust center says "By default, we don't store audio." The training path is the optional feedback feature: the help center says "All audio is anonymized immediately upon sharing and original audio replaces PHI with a 'beep' sound," "Only the de-identified file is stored," and that audio "is stored in our Feedback database indefinitely." The companion Feedback article (which returned 404 on September 10, 2026; wording per its search-index text, unverified) says feedback is used for training Nabla and troubleshooting. So training is opt-in per submission, on de-identified data, with indefinite retention of what you submit. Nabla's public privacy policy (May 16, 2023) covers website visitors only and says nothing about clinical data.

Training20/25
Audio: not stored by default. Clinical notes: "We retain clinical notes for a short period of time (14 days), which is configurable by client." Opt-in feedback audio: retained indefinitely in de-identified form.
Retention21/25
No — cloud (Google Cloud Platform; Azure speech-to-text in U.S. regions).
On-device3/25

SOC 2 Type II (Nov 2024–Oct 2025 period), ISO 27001 (Sept 2025), TX-RAMP Level 2 (Jan 2026); vendor states HIPAA- and GDPR-aligned. BAA availability on self-serve plans is not published (the HIPAA help article points to the Trust Center). The public privacy policy is dated May 2023 and does not cover the product. Press reporting (Verite News, Jan 22, 2026 — page 403'd, unverified) questioned an LCMC Health deployment's patient-consent practice; no suit names Nabla. No breaches or regulator actions found.

Track record18/25
62/100AdequateSep 10, 2026
No self-serve plan published — practice contract (pricing page not found)Encounter audio, transcripts, generated notes, dictation audio, and account data, processed as PHI under a BAA. Hosted on Google Cloud (Cloud SQL, Cloud Storage) in the United States; AES-256 at rest, TLS 1.2 in transit.Yes — de-identified (contractual, no published opt-out)

Suki states it trains on de-identified data and its Terms of Service grant the permission contractually. The developer security FAQ says: "Any data that is used for ML training and improving the product is de-identified," that audio is broken "into chunks and isolate[d] ... such that the original audio cannot be re-constructed," and that transcripts are de-identified "by removing all PII." The Terms of Service say Suki may access Customer Data "for the purpose of using the Customer Data for system tuning, grammar tuning, training of acoustic models and other models, tools and algorithms," and that Suki "owns all rights in and to any aggregated, non-identifiable data." No opt-out toggle or contractual opt-out is published; the privacy policy (May 16, 2025) does not mention model training.

Training9/25
Audio and transcripts: "permanently deleted after 30 days." Clinical notes: retained for the duration of the service contract. PHI returned or destroyed at termination per the BAA.
Retention17/25
No — cloud (Google Cloud, US).
On-device3/25

SOC 2 Type 2 per suki.ai; vendor states HIPAA-aligned and "We sign Business Associate Agreements (BAA) ... with our customers." No breaches, suits, or regulator actions found as of September 2026. Drag: /security and /pricing pages return 404 and the Trust Center is a JS portal, so the retention and training facts live only on the developer documentation site.

Track record19/25
48/100WeakSep 10, 2026
No self-serve plan — "Request a demo" onlyEncounter audio, transcripts, generated notes, and account/device data, processed as PHI under a BAA (Exhibit A of the Terms). Data is transferred to "servers of the Company and the authorized third parties ... located in the United States." Subprocessors are not named.Yes — contractual license (Terms of Use); de-identified data "for any legal purpose"

DeepScribe's Terms of Use (updated October 17, 2024) grant it the right to "use, store, host, perform, display and create derivative works from the Customer Data ... utilizing machine learning and artificial intelligence applications, for the purposes of ... operating, analyzing and improving DeepScribe's products and services," and to "retain, share, and use the De-Identified Data for any legal purpose including without limitation for purposes of operating, analyzing, improving or marketing the DeepScribe Services." That is a contractual license to train on customer data, broader than the de-identified-only wording used by Freed, Suki, or Abridge. The privacy policy (no date visible) says only that data is used to "improve the Services." No opt-out is published.

Training7/25
Not published. The privacy policy retains data "for as long as it remains necessary for the identified purpose or as required by law"; the BAA requires PHI to be returned or destroyed at termination "if feasible." No audio or transcript deletion window is stated anywhere fetched.
Retention8/25
No — cloud (US servers; provider not named).
On-device3/25

SOC 2 and HIPAA logos on the homepage; the security-practices page says DeepScribe holds "relevant certifications" without naming them, and the Trust Center is a JS portal that rendered no content. Vendor states HIPAA-aligned and incorporates a BAA into its Terms. No breaches, suits, or regulator actions found as of September 2026. Drag: undated privacy policy, unnamed subprocessors, no published retention.

Track record17/25
35/100PoorSep 10, 2026

Local LLM Runtimes

Tools for running large language models entirely on the user's own hardware. These tools score uniformly high on Training, Retention, and On-device because there is no vendor server in the inference path — the user supplies the model and the compute. They sit in the same architectural category as Voibe and earn the same scores for the same reason.

ToolPlan tierData collectedTrains on your data?RetentionOn-deviceTrack recordTotalLast verifiedSource
Open-source (MIT) — local install (optional Ollama Cloud models, Pro $20/mo)Local models: none on Ollama servers; models run locally behind an OpenAI-compatible local REST API on localhost:11434. Per the privacy policy (last updated March 2026), Ollama may collect limited device and usage metadata "such as app version and request counts" that does not include prompt or response content. Optional cloud-hosted models send prompts to Ollama-hosted servers in the US, Europe, and Singapore.No

Local mode has no Ollama-operated server in the inference path. For the optional cloud models the policy states: "We do not use your inputs or outputs to train any AI models."

Training25/25
Local: nothing leaves the user's machine beyond limited app-version / request-count metadata. Cloud models: "process this content transiently … not stored beyond the time required to fulfill the request."
Retention25/25
Yes (default). Models run entirely on local hardware; optional Ollama Cloud models are opt-in and cloud-hosted.
On-device25/25

Open-source; over 180,000 GitHub stars; transparent codebase; no documented incidents. Users who deliberately expose the local API beyond localhost create their own attack surface — that is a deployment choice, not an Ollama default.

Track record23/25
98/100ExcellentSep 10, 2026
Free desktop app — Mac / Windows / LinuxNone on LM Studio servers for local inference. Optional LM Link shares only device-list metadata (not chats) with LM Studio's backend for device discovery. Optional paid Cloud Services (cloud models, web search, the "Secure Cloud" behind the Bionic agent app) process data transiently. A third-party traffic audit (April 2026) observed analytics calls carrying app version, OS, GPU, and model names — no content — on an opt-out basis.No

LM Studio runs models locally with an optional OpenAI-compatible local server. The app privacy policy (effective June 2026) states that for local models "none of your messages, chat histories, and documents are ever transmitted," and that Cloud Services data is "processed transiently and not stored after the request completes" with all processors "under Zero Data Retention or substantially equivalent terms." LM Link uses end-to-end encrypted Tailscale mesh VPNs; chats remain local across linked devices.

Training25/25
Local inference: nothing uploaded. LM Link: device list metadata only. Optional Cloud Services: transient, not stored after the request, per the June 2026 policy.
Retention23/25
Yes by default; optional paid cloud models / web search are cloud-processed.
On-device25/25

Privately operated; no documented incidents; transparent about data flows on the LM Link page and the June 2026 app privacy policy. Closed-source application; opt-out (not opt-in) analytics observed by a third-party audit in April 2026.

Track record20/25
93/100ExcellentSep 10, 2026
Open-source (Apache 2.0) — local installNone on Jan servers. Optional cloud-provider integrations inherit the chosen provider's terms.No

Jan runs models locally; there is no Jan-operated model server. Optional cloud-provider integrations are explicitly opt-in.

Training25/25
Nothing leaves the device in default operation.
Retention25/25
Yes (default mode). Optional cloud integrations are opt-in.
On-device25/25

Open-source; auditable codebase; active community roadmap; no documented incidents.

Track record23/25
98/100ExcellentSep 10, 2026

Privacy Policy Quick Read: Does Each AI Tool Train on Your Data?

For each of the 57tools in the matrix above, here is what the vendor’s own privacy policy says about training, retention, and on-device support — quoted verbatim where the policy text supports a clean citation. Each entry links to the primary source we verified against on .

AI Assistants

Does ChatGPT train on my data?

Yes, by default — opt-out available.

ChatGPT's consumer plans (Free, Plus, Pro) train on user prompts, outputs, and uploaded files by default. To opt out, navigate to Settings → Data Controls and disable "Improve the model for everyone." Conversations are retained for 30 days after deletion. Temporary Chat is never used for training. ChatGPT Team, Enterprise, and API plans are explicitly excluded from training under OpenAI's enterprise terms — API users can optionally opt in via Playground feedback. Limited April–September 2025 data is preserved due to the NYT litigation order; OpenAI's standard 30-day retention practices resumed September 26, 2025.

Primary source: OpenAI privacy policy

Does Claude train on my data?

User choice required (since Aug 2025).

As of August 28, 2025, Anthropic shifted Claude's consumer plans (Free, Pro, Max) from "not used for training" to a user-choice model. New users must actively choose during signup whether to share data for training; existing users had until October 8, 2025. Users who opt in have their data retained for up to 5 years; users who decline keep the previous 30-day retention window. Flagged conversations are retained 2–7 years for trust & safety review. Claude for Work, the Claude API, Amazon Bedrock, and Google Vertex AI are all contractually excluded from training under Anthropic's Commercial Terms.

Primary source: Anthropic Aug 2025 update

Does Gemini train on my data?

Yes, by default — opt-out via the "Keep Activity" setting.

Free Gemini and Gemini Advanced (consumer) train on user conversations by default. Per Google's documentation, when Keep Activity (the renamed "Gemini Apps Activity" toggle) is on, "Google uses your activity to provide, develop, and improve its services (including training generative AI models)." To opt out, set Keep Activity to OFF, or use a Temporary Chat, which is never used for training — but even when off, future chats are saved for 72 hours so Gemini can respond and process feedback. Default retention is 18 months, adjustable to 3 months, 36 months, or never. Human-reviewed conversations are kept up to 3 years (disconnected from your Google Account). Vertex AI customer data is contractually excluded from training: "Google won't use your data to train or fine-tune any AI/ML models without your prior permission or instruction."

Primary source: Gemini Keep Activity controls

Does Perplexity train on my data?

Yes, by default — opt-out for logged-in users only.

Perplexity trains on user queries, prompts, and AI responses by default for Free, Pro, and Max plans. The "AI Data Retention" toggle in Account Settings → Preferences disables this. Logged-out users are trained on by default with no opt-out path — sign in to gain control. Threads are retained until manually deleted. On July 2, 2026 Perplexity replaced its privacy policy with a restructured Privacy Notice that dropped the written descriptions of the training opt-out, the 30-day account-deletion timeline (now "as long as reasonably necessary"), and the promise to notify users of material changes; the toggle and deletion still work in the product, but they are no longer commitments in the policy text. The same rewrite expanded stated collection to the Comet browser, voice input, and health and genetic data. The Sonar API offers Zero Data Retention with prompts and responses never stored. Third-party providers (OpenAI, Anthropic) are contractually prohibited from training on Perplexity's API data. Enterprise file uploads are deleted after 7 days.

Primary source: Perplexity data collection policy

Does DeepSeek train on my data?

Yes, by default — no in-app toggle; opt-out only by emailing privacy@deepseek.com.

DeepSeek's privacy policy (last updated February 10, 2026) states: "we directly collect, process and store your Personal Data in People's Republic of China" — covering prompts, outputs, uploaded files, IP addresses, device info, and account info (the keystroke-pattern language from the 2025 versions no longer appears). There is no in-app training toggle; the policy grants a right to opt out of training "or optimizing our technologies," exercisable by emailing privacy@deepseek.com. Retention is "for as long as necessary to provide our Services," with no fixed window. Italy's Garante imposed a 72-hour ban in January 2025 followed by 13 European jurisdiction probes; government device bans followed in Australia, Taiwan, South Korea, Czech Republic, the Netherlands, Germany, and multiple US federal agencies (Pentagon, NASA, US Navy). The bipartisan "No DeepSeek on Government Devices Act" is pending in the US Senate. A January 2025 database breach exposed over 1 million records. Open-source MIT-licensed weights can be self-hosted, which removes this concern entirely — but that is the self-hosted path, not the consumer chatbot.

Primary source: DeepSeek privacy policy (Feb 2026)

Does Apple Intelligence train on my data?

No — Apple's published policy excludes user interactions from foundation-model training.

Apple's published policy states Apple does not use users' private personal data or user interactions when training its foundation models. The on-device foundation model handles most tasks; for larger requests, Private Cloud Compute extends device security to Apple Silicon servers using a verifiable transparency model — signed binaries are publicly inspectable, there is no SSH or admin access, and researchers can audit via the Virtual Research Environment. PCC is stateless: data is processed only to fulfill the request and returned to the device; it is not stored or made accessible to Apple. Apple collects only request metadata (size, feature, duration), not content. The optional ChatGPT integration is a separate, opt-in boundary that routes through OpenAI's enterprise terms with IP obfuscation — Apple's guarantees do not extend to ChatGPT requests.

Primary source: Intelligence Engine privacy

Does Microsoft Copilot train on my data?

Consumer: yes by default — opt-out available. M365 Copilot business: no, contractually excluded.

First, the disambiguation: Microsoft Copilot is the consumer / Windows / Microsoft 365 assistant — it is distinct from GitHub Copilot (the developer product) which is tracked separately in the coding section. For Copilot Free, Windows Copilot, and copilot.com chat on Microsoft 365 Premium ($19.99/mo; Copilot Pro was retired on August 1, 2026), Microsoft trains by default on signed-in users' conversation activity. Microsoft Support states: "Except for certain categories of users or users who have opted out, Microsoft uses data from Bing, MSN, Copilot, and interactions with ads on Microsoft for AI training." Since the updated Copilot app of August 18, 2026 there are two toggles under Profile icon → Privacy: "Training on conversation activity" and "Training on voice conversations"; Microsoft says opting out excludes future conversation activity from model training, while conversations may still be used for other product improvements, advertising, digital safety, security, and compliance. Signed-out users, and Copilot inside Microsoft 365 apps on Personal / Family / Premium subscriptions, are not used for training. A law firm announced an investigation into the default-on training on August 26, 2026. Conversation activity is retained 18 months by default; user-deletable, but deletion is independent of the training opt-out. On Copilot+ PCs (NPU ≥40 TOPS), Recall and Click to Do run locally on-device with snapshots stored on-device; most Copilot chat still runs in the cloud. Microsoft 365 Copilot ($30/user/mo, enterprise) is contractually excluded from training under the Microsoft Products and Services Data Protection Addendum: "Your data isn't used to train foundation models... the prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation models." Retention is tenant-controlled via Microsoft Purview. On January 7 2026, Microsoft enabled Anthropic as a default subprocessor for M365 Copilot (used in Researcher, Copilot Studio, Office agents); EU and UK tenants have it disabled by default, processing occurs outside the EU Data Boundary.

Primary source: Microsoft Copilot privacy controls (Aug 2026)

Does Meta AI train on my data?

Yes, by default globally. EU / UK have an opt-out form; US / Australia / most non-GDPR regions do not.

Meta AI — the assistant available in WhatsApp, Instagram, Facebook, Messenger, Threads, and Ray-Ban / Oakley Meta — is opt-in by default for training on adult users' public Facebook and Instagram posts and on all Meta AI conversations. The regional split is load-bearing. EU / EEA users have opt-out via the "Right to Object" form in Privacy Center → "How Meta uses information for generative AI models and features" — Meta resumed EU training on May 27 2025 under GDPR Article 6(1)(f) "legitimate interest" after a one-year pause forced by Irish DPC + noyb pressure. UK opt-out is available after ICO concessions. US, Australia, and most non-GDPR jurisdictions have no general opt-out — Meta confirmed in Australian Senate hearings (Sept 2024) that it scrapes Australian public posts back to 2007 and will offer an opt-out only "if governments force it to." WhatsApp 1:1 and group messages remain end-to-end encrypted and are not used for training unless the user explicitly invokes @Meta AI in chat. On December 16 2025, Meta extended Meta AI conversations into ad personalization in addition to model training (US and most non-EU regions; EU, UK, South Korea carved out). Retention is indefinite unless the user manually deletes; Ray-Ban Meta voice recordings are retained up to 1 year by default. No on-device inference for the Meta AI assistant — WhatsApp "Private Processing" is server-side TEE confidential compute, not local.

Primary source: Meta Privacy Center — How Meta uses information for generative AI

Does Grok train on my data?

Consumer: yes by default — three separate opt-out paths. API: no, by default.

Consumer Grok — on X (Free, Premium $16/mo, Premium+ $40/mo), grok.com, and the Grok mobile app — is opt-in by default for training on both X posts and Grok prompts / outputs. The opt-out lives in three separate places: (1) Grok on X: Settings → Privacy and safety → Data sharing and personalization → Grok & xAI → uncheck "Allow your posts as well as your interactions, inputs, and results with Grok and xAI to be used for training and fine-tuning." (2) Grok mobile app: Settings → Data Controls → uncheck "Improve the model." (3) Grok web at grok.com: Settings → Data → uncheck "Improve the model." EU / EEA users were excluded from X-post training under the September 4 2024 DPC undertaking; the DPC opened a fresh statutory inquiry on April 11 2025. The January 15 2026 X Terms of Service update explicitly classifies Grok prompts and outputs as user "Content" available for AI training. Grok chats are retained indefinitely unless the user deletes them. The xAI API has a materially different posture: "xAI never trains on your API inputs or outputs without your explicit permission," matching the OpenAI and Anthropic API defaults. API retention is 30 days for abuse monitoring; Zero Data Retention is available for enterprise customers via sales@x.ai and on Oracle Cloud Infrastructure since June 17 2025. Track record: "MechaHitler" antisemitic-content failure (July 2025) drew EU Commission and Turkey enforcement; ~370,000 Grok shared chats were indexed on Google in August 2025 because the share feature generated unauthenticated URLs without `noindex`.

Primary source: xAI Privacy Policy

AI Coding Tools

Does Cursor train on my code?

Yes, by default for individuals — Privacy Mode opt-out.

For individual accounts Privacy Mode is off by default (Cursor no longer calls this "Share Data"). With it off, Cursor's Data Use page (last updated September 3, 2026) says it "may use and store codebase data, prompts, editor actions, code snippets, and other code data and actions to improve our AI features and train our models." Toggling Privacy Mode ON prevents training and discards plaintext after each request — cached files are encrypted with client-generated keys that exist on Cursor's servers only for the duration of each request. Enterprise workspaces default to Privacy Mode ON and admins can enforce it; Cursor states it holds zero-data-retention agreements with all model providers (naming SpaceXAI, OpenAI, Anthropic, and Meta, plus Google in the enterprise docs). Models that require provider-side retention, such as Claude Fable 5.1 and Fable 5, are blocked for Privacy Mode and Enterprise users until an admin approves them. On August 29, 2026 Cursor removed its disclosure of how indexing embeddings and metadata are retained, so that detail is now unpublished. Requests go through Cursor's backend even when you bring your own API key; Cursor can also be pointed at local Ollama or LM Studio models.

Primary source: Cursor data use

Does GitHub Copilot train on my code?

Yes, by default for consumer plans — opt-out as of April 24, 2026.

On April 24, 2026, GitHub began using Free, Pro, and Pro+ user interaction data — including code snippets — to train AI models by default. Existing opt-outs are honored. To disable training going forward, go to Settings → Privacy. User Engagement Data is retained for 2 years; Coding Agent session logs persist for the lifetime of the account. Private repository code at rest is NOT used for training, but in-flight interaction data IS. Business and Enterprise plans are explicitly prohibited from being used for training under GitHub's agreements: subscription Prompts and Suggestions are retained 28 days, and User Engagement Data 2 years.

Primary source: April 2026 policy change

Does Devin Desktop (formerly Windsurf / Codeium) train on my code?

Yes for individuals by default — paid plans can opt out, which also enables Zero Data Retention.

Windsurf was rebranded Devin Desktop by Cognition on June 2, 2026, and its legal pages now live on cognition.com and docs.devin.ai. Cognition's privacy policy (March 9, 2026) says User Content is used "to train, fine tune and improve the models that power our Services" unless you opt out. Per the Devin security docs: "If you're on a paid plan, you can opt out at any time on the Data Controls settings page. After you opt out, your data will not be used for training and Zero Data Retention will be enabled with our model providers." No opt-out is documented for the free tier. On Teams only an administrator can exercise the opt-out; Enterprise customers are not trained on without express prior written consent. With ZDR, Customer Data is "not saved to disk or otherwise persistently retained" and is deleted once the output is generated. Cognition offers a hybrid deployment and has placed its self-hosted deployment in maintenance mode.

Primary source: Devin security docs

Does Cline train on my code?

No — Cline operates no model server. Privacy depends on your chosen API provider.

Cline is an open-source VS Code extension that operates no model server of its own. User code is sent only to whichever API provider you configure (Anthropic, OpenAI, AWS Bedrock, Google Gemini, Cerebras, Groq, etc.) and is governed by that provider's terms. Cline's stated principle: "Code never leaves your machine" toward Cline servers. Anonymous telemetry (features used, task completion rates) is collected but can be disabled via the Cline Telemetry setting. Code, file contents, command arguments, and conversation content are explicitly NOT collected by telemetry. For fully on-device use, configure Cline with a local Ollama or LM Studio model.

Primary source: Cline telemetry docs

Does Claude Code train on my code?

Pro/Max account: yes by default — opt-out. API / Console / Enterprise: no, contractually excluded.

Claude Code is Anthropic's CLI for Claude — separate from Claude.ai consumer chat and from Claude for Work. The data-handling posture depends entirely on how the user is billed. A developer running Claude Code from a Claude.ai Pro or Max account inherits the August 28 2025 consumer-terms update, which made training default-on with user-choice opt-out. Anthropic stated explicitly: "These updates apply to users on our Claude Free, Pro, and Max plans... including when they use Claude Code from accounts associated with those plans." Opt-out lives at claude.ai/settings/data-privacy-controls. Retention with training enabled is 5 years; with training disabled, 30 days. A developer running Claude Code via an API key, Anthropic Console, Claude for Enterprise, AWS Bedrock, Google Vertex AI, or Microsoft Foundry is on Anthropic's Commercial Terms Section B: "Anthropic may not train models on Customer Content from Services." Reinforced in the Claude Code data-usage docs: "Anthropic does not train generative models using code or prompts sent to Claude Code under commercial terms." Bedrock and Vertex users cannot enroll in the Development Partner Program at all. Zero Data Retention is available specifically for Claude Code on Claude for Enterprise per the docs, enabled per-organization by the account team. Architecturally, Claude Code is partial on-device: the binary runs locally (file edits, shell execution, MCP servers, local session log), but inference is cloud-only via the Anthropic API.

Primary source: Claude Code data usage

Does Tabnine train on my code?

No — proprietary models trained only on permissively licensed open-source code.

Tabnine's proprietary completion models are trained only on permissively licensed open-source code (MIT, Apache 2.0). Customer code is never used to train Tabnine's models and is never shared with third parties. The code-privacy page states verbatim: "We never use your code to train any of our models. Code exposed to Tabnine is never stored or shared when using our proprietary models." Tabnine also implements zero code retention — requests are processed in-memory and immediately discarded after the server returns a suggestion. Operational metrics and logs (which contain no code and no PII) are retained roughly one week for support. The architecture scales further at the Enterprise tier, which supports SaaS, VPC, on-premises, and fully air-gapped deployment options where code never leaves the customer network. There is no free or individual plan; the entry tier is the Code Assistant Platform at $39 per user per month (Agentic Platform $59). Tricentis acquired Tabnine, announced July 30, 2026, and the August 28, 2026 privacy-policy edits were cosmetic. Compliance posture: SOC 2 Type II, ISO 27001, GDPR, with IP indemnification for enterprise customers. No documented privacy incidents. Tabnine sits alongside Cline as the rare "good actor" in a coding category that otherwise carries opt-out-by-default training (Cursor individual, Devin Desktop individual) or recent reversal (GitHub Copilot, since April 2026).

Primary source: Tabnine code privacy

Does OpenAI Codex train on my code?

Yes with a ChatGPT Free / Go / Plus / Pro sign-in unless you turn off two separate toggles; no with an API key, Business, Enterprise, or Edu.

Codex — the CLI, IDE extension, desktop app mode, and cloud agent — inherits the terms of whatever account you sign in with. OpenAI's help center (updated August 2026) names it directly: "When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models." The ChatGPT opt-out (Settings → Data Controls → "Improve the model for everyone") is documented as turning off training "for your ChatGPT conversations and Codex tasks." That is not the whole story: "Codex has separate controls for allowing training on full environments, which you can manage in the Codex Settings. Note that adjusting your settings in the ChatGPT interface or privacy portal will not affect these full-environment Codex settings." OpenAI does not publish the default of that environment control, so treat it as something to check by hand. Retention on consumer plans follows the ChatGPT privacy policy — deleted content is removed within 30 days unless kept for legal or abuse reasons — with no Codex-specific window published. Business, Enterprise, Edu, and API-key use fall under the enterprise privacy page (updated January 8, 2026): "By default, we do not use your business data for training our models." Enterprise admins control retention; API traffic is kept 30 days for abuse monitoring, and /v1/responses is eligible for Zero Data Retention. Architecturally, Codex is partial on-device: the CLI edits files and runs shell commands locally, but per OpenAI's own docs "Local execution does not mean offline or device-only model inference," and cloud tasks run in VM-backed sandboxes on OpenAI infrastructure. A GitHub-token command-injection flaw reported December 16, 2025 was patched February 5, 2026 with no exploitation reported.

Primary source: How your data is used to improve model performance

Does Gemini CLI train on my code?

Depends on the key: yes with a free Gemini API key (no toggle); no with a paid API key, Vertex AI, or a Gemini Code Assist Standard / Enterprise license.

Gemini CLI is Google's open-source (Apache 2.0) terminal agent, and its privacy posture is set entirely by how you authenticate. The picture changed in 2026. Google announced on May 19, 2026 that "Gemini CLI and Gemini Code Assist IDE extensions will stop serving requests for Google AI Pro and Ultra, as well as those using it free of charge using Gemini Code Assist for individuals" from June 18, 2026, steering those users to Antigravity CLI. The old "Gemini Code Assist for individuals" privacy notice, which offered an opt-out checkbox, now serves only a deprecation notice. What remains for individuals is a Gemini API key. On the unpaid tier the Gemini API Terms (effective March 23, 2026) state: "Google uses the content you submit to the Services and any generated responses to provide, improve, and develop Google products" and "human reviewers may read, annotate, and process your API input and output." There is no opt-out setting; the only off-switch is enabling paid billing, where "Google doesn't use your prompts ... or responses to improve our products" and logs are kept "for a limited period of time" for abuse detection — no number is published. Organization licenses fall under the Gemini for Google Cloud data-governance page (last updated September 3, 2026): "Gemini doesn't use your prompts or its responses as data to train its models." Separately, the CLI sends anonymous usage statistics by default (tool names, model, durations) but "We do not log the content of your prompts or the responses"; set privacy.usageStatisticsEnabled to false to stop it. Track record: a critical advisory (GHSA-wpqr-6v78-jr5g) on April 24, 2026 fixed workspace-trust and tool-allowlist bypasses in headless runs.

Primary source: Gemini CLI terms and privacy

Does Kiro train on my code?

Yes by default for Free and for paid plans signed in with GitHub, Google, or AWS Builder ID — opt-out in settings; no for Enterprise or IAM Identity Center sign-ins.

Kiro is AWS's agentic IDE and CLI. Its data-protection page (last updated August 4, 2026) is unusually explicit that the split is by sign-in method, not by price. "Kiro may use certain content from Kiro Free Tier and Kiro individual subscribers for service improvement," where individual subscribers are users "that have a paid Kiro subscription and access it through a social login provider (like GitHub or Google) or through AWS Builder ID." The content covered is "your questions to Kiro, other inputs you provide, and the responses and code that Kiro generates," and the stated purposes include "for model training." A $200-per-month Power subscriber on a GitHub login is therefore trained on by default; a $20 Pro subscriber on IAM Identity Center is not. To opt out in the IDE: Settings → User → Application → Telemetry and Content, then uncheck "Data Sharing and Prompt Logging: Content Collection for Service Improvement" and "Usage Analytics And Performance Metrics." The CLI exposes a telemetry.enabled setting; its default is not published. For Enterprise, "We do not use content from Kiro enterprise users for service improvement," and those users "are automatically opted out of telemetry and content collection." Retention: Free Tier inputs may be stored "for up to 60 days" for abuse detection, OpenAI GPT classifier-flagged traffic "for up to 30 days," and no window is published for service-improvement content. Inference runs in AWS on Bedrock-hosted Claude, OpenAI GPT, and open-weight models; there is no local-model option. Two caveats for buyers: the AWS Service Terms (updated September 1, 2026) carry no Kiro-specific no-training clause, a gap flagged in GitHub issue #2206 since August 2025, and 2026 brought three AWS security bulletins plus a Mindgard-disclosed prompt-injection exfiltration, all patched with no customer data loss reported.

Primary source: Kiro data protection

Voice & Dictation

Does Voibe train on my voice data?

No — on-device mode never transmits audio; cloud mode is zero-retention with no training.

Voibe has two user-selectable modes. In on-device mode (Apple Silicon Macs, M1 or later) audio is transcribed entirely on the Mac by Whisper models running on the Neural Engine, works with no internet, and nothing leaves the machine — so there is no training to opt out of. In zero-retention cloud mode (all Macs and Windows) audio is encrypted in transit and transcribed by Whisper Large Turbo, an open-source model running on zero-retention providers (Groq or Cloudflare); the audio is deleted the moment transcription completes and is never stored, sold, or used to train any model. Formatting is a second hop that sees only the text, never the audio, using GPT-OSS 120B, an open-weight model on Cerebras. No third-party AI lab is in the audio path and users never paste an API key. The speech-to-text API and MCP server launched in August 2026 run on the same pipeline, with transcripts readable for 24 hours and then deleted. Account holders provide an email (for authentication) and non-identifying usage analytics; crash reports exclude dictated content.

Primary source: Voibe AI & privacy

Does Wispr Flow train on my voice data?

Off by default since 2024 backlash — opt-in for training.

After 2024 community backlash, Wispr Flow shifted training to opt-in. Privacy Mode is OFF by default for Free users, meaning audio, transcripts, edits, and optional Context Awareness (screenshots of the active app's screen) are retained indefinitely. Data passed to third-party LLM providers (OpenAI, Meta) is retained for 30 days. Enterprise plans default to Privacy Mode ON with zero data retention by Wispr or any third party — audio is processed and immediately discarded after transcription. A Business Associate Agreement is available for Enterprise; once signed, Privacy Mode locks irreversibly. Transcription always happens in the cloud; even Privacy Mode is "zero-retention cloud," not local processing.

Primary source: Wispr Flow privacy policy

Deep dive: Is Wispr Flow Safe? — full investigation

Does Superwhisper train on my voice data?

No — verbatim from policy.

Superwhisper's privacy policy states explicitly: "Your data is not retained on Superwhisper servers" and "not used for training AI models or any other machine learning purposes." On-device modes (Fast, Nano, Standard Whisper, Parakeet — available on the Free plan and within Pro) process audio entirely locally; nothing is transmitted. Cloud modes (Ultra transcription, Super Mode LLMs — Pro tier) proxy audio through Superwhisper's infrastructure with no retention. One caveat: audio recordings are saved to local disk by default. Opt out in settings if local audio retention is a concern. Note: the privacy policy does not currently distinguish between on-device and cloud modes — verify cloud-mode specifics with the vendor before sensitive use.

Primary source: Superwhisper privacy

Deep dive: Is Superwhisper Safe? — full investigation

Does MacWhisper train on my voice data?

No — primarily on-device, with optional cloud + BYOK paths.

MacWhisper does not train its own models on user audio. The on-device transcription path uses local Whisper models that you can download for offline use; Apple Foundation Models also run on-device for AI features. MacWhisper's optional "Assistant" cloud transcription service and BYOK integrations (OpenAI Whisper API, ElevenLabs) inherit those providers' terms when used. The App Store version's privacy disclosure shows only "Usage Data" and "Product Interaction" as Data Not Linked to You. There is no separate enterprise tier; the data-handling architecture is identical for individuals and bulk-licensing customers.

Primary source: App Store listing

Does Aqua Voice train on my voice data?

Privacy policy does not explicitly address training — opt-out via Privacy Mode.

Aqua Voice's privacy policy does not explicitly state whether stored data is used for AI training. With Privacy Mode disabled, "we may securely store transcript data on our servers"; with Privacy Mode enabled, "transcript data is not collected" though session metadata (timestamps, device type, performance metrics) may still be. Aqua Voice is SOC 2 Type II certified by Advantage Partners. Teams and Enterprise plans support an org-wide Privacy Mode that applies the same protections across an entire organization. No HIPAA Business Associate Agreement is publicly advertised. Audio is cloud-processed; there is no on-device option.

Primary source: Aqua Voice privacy policy

Deep dive: Is Aqua Voice Safe? — full investigation

Does Typeless train on my voice data?

No, per the published privacy policy — but verify the architecture.

Typeless's privacy policy states: "Your data is never used to train these services and is configured for zero retention by the providers." Audio plus contextual information is "processed in real time on our cloud servers and immediately discarded once the result is returned to your device." Free and Pro tiers receive the same data-handling treatment. However, a November 2025 reverse-engineering analysis (covered in our Typeless privacy issues investigation) reported collection beyond what the published policy describes — including URL capture, window-title metadata via the macOS accessibility API, and broad permission requests. Verify the current subprocessor list at trust.typeless.com/subprocessors before relying on Typeless for sensitive content.

Primary source: Typeless privacy policy

Deep dive: Typeless Privacy Issues — full investigation

Does Apple Dictation train on my voice data?

Only if you opt in via "Improve Siri & Dictation."

Apple Dictation only uses your audio to improve its models if you have explicitly enabled "Improve Siri & Dictation" — the default at setup is to be asked. If opted in, audio and transcripts are retained under a rotating random ID for up to 6 months, then dissociated and kept for up to 2 years for improvement; a reviewed subset is retained beyond 2 years. If opted out, recordings are not retained for improvement. On Apple Silicon Macs running modern macOS or iOS, most languages process locally for general text fields (Notes, Mail, Messages). Server-side fallback applies to unsupported languages, search-box dictation, and some third-party Speech Recognition API uses. Apple does not sign a Business Associate Agreement for consumer Dictation, so it is not HIPAA-compliant.

Primary source: Ask Siri & Dictation policy

Does Willow Voice train on my voice data?

No by default — Private Mode is the default; opt-in required for training.

Willow Voice operates with Private Mode as the default for all users. The privacy policy states verbatim: "In private mode, Willow only collects basic technical and account-related data needed to run the app and nothing else. No voice, no dictated text." Training only occurs if the user explicitly opts into Opt-In Mode: "You can allow Willow to collect minimal usage data to help improve our text correction models." This is the inverse of Wispr Flow's pre-2024 posture, which had Privacy Mode off by default until community backlash forced the switch. With Opt-In Mode enabled, anonymized text and usage data are retained "only as long as needed to train and improve the app" — no specific retention window is published. Architecturally, Willow Voice is cloud-first; the privacy policy's reference to local storage applies to the UI display of past transcripts, not to local transcription. The current willowvoice.com pricing page does not advertise an offline transcription mode. The Enterprise tier is marketed as zero data retention with HIPAA and SOC 2 compliance and BAA availability, but the privacy policy itself only references GDPR and SOC 2 — request signed BAA terms in writing before processing PHI. Policy last updated April 30 2025.

Primary source: Willow Voice privacy policy

Does Monologue train on my voice data?

Unknown — the privacy policy does not address training either way.

Monologue's published data-privacy page makes a narrow set of claims: "No audio files or transcripts are saved on our servers," "Deep context screenshots are deleted immediately," "Zero LLM data retention," and "Custom modes and dictionaries stay on your device." The page does not explicitly state whether de-identified or aggregated data is used to improve models. Peer cloud dictation products — Wispr Flow, Typeless, Superwhisper, Willow Voice — all explicitly address the training question one way or the other. Monologue's silence is the load-bearing gap. The Notion-hosted privacy policy at modern-ton-234.notion.site is similarly sparse — no list of data categories collected, no subprocessor disclosures, no retention windows beyond the broad "not saved" claim, no SOC 2 / HIPAA / BAA mention. Offline transcription is listed as a feature on both the Free and Pro tiers of monologue.to, but the architecture (cloud-default with offline fallback vs on-device-default) is not documented. For regulated work, request explicit training and subprocessor disclosure in writing before use.

Primary source: Monologue data privacy

Does VoiceInk train on my voice data?

No — 100% on-device. Source code is auditable on GitHub.

VoiceInk is an open-source Mac dictation app under GPL v3.0. The README states verbatim: "100% offline processing ensures your data never leaves your device." The marketing site at tryvoiceink.com describes it as a "privacy-focused dictation app for macOS with local transcription." Transcription runs locally on Apple Silicon via whisper.cpp; there is no VoiceInk-operated server in the inference path, so there is no retention surface and no training path. No telemetry, BYOK cloud feature, or AI-command post-processing service is documented at this verification pass. The source code is auditable on GitHub at Beingpax/VoiceInk with 4,900+ stars, 675+ forks, and 119 releases (latest v1.76 on May 7 2026). Pricing is $39.99 one-time for the Pro license, or free if built from source under GPL v3. Because the codebase is open, the privacy claim is verifiable — the same reason Ollama, Jan, and LM Studio score highly in our Local LLM category.

Primary source: VoiceInk (tryvoiceink.com)

Deep dive: VoiceInk Review — full investigation

Does Dragon (Nuance / Microsoft) train on my voice data?

Depends entirely on which Dragon product. Desktop is offline; cloud tiers are not separately disclosed.

Dragon is now a multi-product line under Microsoft (acquired from Nuance in 2022) with materially different privacy postures across SKUs. Dragon Professional v16 is a locally installed Windows desktop application — audio is processed on the user's machine, voice profiles are stored locally, and nothing is transmitted to Nuance/Microsoft servers in normal desktop dictation operation. There is no training path because there is no upload. Dragon Anywhere (the iOS/Android mobile product) is cloud-based — audio is transmitted to Nuance/Microsoft for recognition. Microsoft's general Privacy Statement reserves the right to use customer data to develop and train AI models (with opt-out paths surfaced through Copilot privacy controls), and the Dragon-specific product pages do not separately disclose AI-training carve-outs for Dragon Anywhere at the consumer tier. Treat that cell as unverified until Microsoft publishes a Dragon-specific commitment. Dragon Medical One, Professional Anywhere, and Dragon Copilot are enterprise cloud products marketed as HIPAA-compliant when a BAA is in place. The umbrella Microsoft Products and Services Data Protection Addendum excludes Customer Data from foundation-model training, which is the contractual basis healthcare buyers rely on — but Dragon-specific commitments are still primarily contract-level rather than published product commitments. Track record context: the 2017 NotPetya attack disrupted Nuance's hosted services (including Dragon Medical) for weeks. Development on the desktop line has largely stalled since the Microsoft acquisition (v17 ≈ v16); the Mac version was discontinued in 2018 and Dragon Home in 2023. Non-medical Nuance product pages now redirect to Microsoft's health-solutions hub, creating a discovery gap that itself is a buyer risk when trying to verify privacy commitments.

Primary source: Dragon Medical One (Microsoft)

Deep dive: Dragon pricing — full breakdown

Does Spokenly train on my voice data?

No by Spokenly — but the answer depends on which of three modes you use.

Spokenly is a Mac and iOS dictation app from named solo developer Vadim Akhmerov (disclosed via the App Store, not the privacy policy) with three architectural modes that produce three different privacy postures. Local Only Mode runs OpenAI Whisper Large-v3 and NVIDIA Parakeet locally on Apple Silicon with no network calls during transcription — nothing is transmitted, so there is no training path. BYOK cloud mode routes audio to whichever provider you supply API keys for (OpenAI, Deepgram, Groq, Anthropic, or Google); Spokenly does not train on it, but the chosen provider's terms govern training and retention. Pro managed cloud ($9.99/mo) forwards audio to third-party transcription services — the privacy policy last updated July 7, 2026 names eleven: OpenAI, Deepgram, Soniox, Cartesia, ElevenLabs, Groq, Mistral AI, xAI, OpenRouter, fal.ai, and Cerebras. The policy's load-bearing claim is that Spokenly's backend "does not store these recordings," applied across all modes; on training it says only that Spokenly prefers "options that do not use customer data for model training where available," a hedged preference rather than a guarantee. Three structural caveats apply regardless of mode: no SOC 2 / HIPAA BAA / ISO 27001 / GDPR / CCPA attestations (disqualifying for regulated work), no disclosed corporate entity or jurisdiction (only admin@spokenly.app as contact), and an iOS keyboard caveat where the developer's own App Store reply recommends switching to online models for keyboard reliability — which defeats the on-device benefit. App Store rating is 4.4/5 from 56 ratings.

Primary source: Spokenly privacy policy (updated 2026-07-07)

Deep dive: Is Spokenly Safe? — full investigation

Does Handy train on my voice data?

No — there is no Handy server to send audio to, and the MIT source proves it.

Handy is a free, MIT-licensed push-to-talk dictation app for macOS, Windows, and Linux maintained by CJ Pais. Its privacy claim is one sentence on handy.computer and in the README: "Your voice stays on your computer. Get transcriptions without sending audio to the cloud." There is no privacy policy document at all — handy.computer/privacy returned HTTP 404 on September 10, 2026 — so the claim is backed by architecture and source rather than by a policy. Audio passes through bundled Silero voice-activity detection and a locally downloaded speech model (Whisper, Parakeet, and other families); no cloud transcription endpoint exists in the codebase, so there is no retention surface and no training path. The one default network call is a signed update check against GitHub releases, toggleable in Settings. Two items to re-check each release: the README roadmap still lists "Opt-in Analytics" ("Privacy-first approach with clear opt-in") as unshipped through v0.9.6 (August 24, 2026), and our July 2026 source audit found an optional bring-your-own-key LLM post-processing path that sends transcript text, never audio, to a provider you configure — it is off by default and not documented in the README. The repository has 31,300+ stars and 2,800+ forks with no CVEs or advisories. What Handy lacks is anything a procurement team can file: no legal entity, no DPA, no BAA, no SOC 2. For an individual who wants on-device dictation at $0, that is a fair trade; for regulated work it means you are relying on your own audit.

Primary source: Handy GitHub repo (MIT)

Deep dive: Is Handy Safe? — full investigation

Does OpenWhispr train on my voice data?

No, in writing across all three modes — but which mode you use decides who else sees your audio.

OpenWhispr is an MIT-licensed desktop client (7,900+ GitHub stars) from Gizmo Labs Inc., a Delaware corporation, that grew a hosted product in 2026. The privacy policy (last updated August 12, 2026) describes three paths. Local mode runs Whisper or NVIDIA Parakeet on-device and "no audio is transmitted anywhere." BYOK sends audio directly to the provider whose key you supply, under that provider's terms; the policy discloses that "API keys you enter for third-party providers are currently stored in plaintext in your local app data directory." OpenWhispr Cloud sends audio to OpenWhispr's API, where it is "processed in real time, and discarded" — the "0% Data Retention" claim — and forwards it to subprocessors named in DPA Annex 3: OpenAI OpCo, LLC, Parasail, Inc., and Groq, Inc. The no-training commitment is contractual, not just marketing: the Terms of Service (effective July 23, 2026) state "We will not use Customer Content — audio, transcriptions, notes, or AI agent prompts — to train any artificial intelligence or machine learning model," and DPA v1.2 (updated August 21, 2026) repeats it for voice audio, transcriptions, notes, prompts, and AI-agent conversations. Synced notes are deleted within 30 days of account deletion. The DPA lists a SOC 2 Type 2 report (May 1–July 31, 2026), ISO/IEC 27001:2022 valid through July 27, 2027, and HIPAA / GDPR attestations; the auditor is not named on the trust center. A BAA is available for Business ($13.33/user/mo billed annually) and above, and the DPA says covered entities must not process PHI through the cloud without one. Free is $0 with 2,000 cloud words/week; Pro is $6.67/user/mo billed $80/yr.

Primary source: OpenWhispr privacy policy (updated 2026-08-12)

Deep dive: Is OpenWhispr Safe? — full investigation

Does FluidVoice train on my voice data?

No — audio and text stay on the Mac; the only thing sent to the vendor by default is anonymous usage metadata.

FluidVoice is a free macOS dictation app whose app source is GPLv3 (since February 23, 2026; Apache 2.0 before that) at github.com/altic-dev/FluidVoice, with 11,400+ stars and releases through v1.6.9 on August 18, 2026. The README is the governing privacy text because no formal policy document exists: "Your voice, audio, and transcribed text never leave your machine unless you explicitly opt in to a cloud AI provider." Every speech model — Whisper, Parakeet, Nemotron, Cohere Transcribe, Apple Speech — is a local download, and the enhancement runtime Fluid Intelligence (the Fluid-1 model, about 3.5 GB) also runs offline. That runtime is the one closed component: "We're keeping Fluid Intelligence private for now so we can sustainably offer the core dictation experience for free." You can verify it sends nothing; you cannot audit what it changes. Two defaults matter. First, "Detailed anonymous analytics are enabled by default" and can be disabled at Settings → Share Detailed Anonymous Analytics; the payload is a random installation ID, activity date, app version, macOS platform label, and daily feature and model usage totals — the README states voice, audio, transcripts, prompts, and clipboard content are not collected, and no retention window for that metadata is published. Second, optional cloud enhancement can send transcript text to OpenAI, Groq, or a custom provider with your own key (stored in the macOS Keychain), under that provider's terms. Requirements: macOS 15 or later; Apple Silicon for everything except Whisper on Intel. There is no legal entity, DPA, BAA, or compliance attestation published, and no documented incidents.

Primary source: FluidVoice GitHub repo (GPLv3)

Deep dive: Is FluidVoice Safe? — full investigation

Does VoiceDash train on my voice data?

No, per its policy and DPA — but every dictation goes to OpenAI, and the retention story has an unverified piece.

VoiceDash is a cloud-only dictation app for Mac and Windows from Jumpstart Ventures FZE in Dubai Silicon Oasis (the entity is named in the DPA, not the privacy policy), founded February 2025. The privacy policy, which carries no effective date, states: "Your voice recordings are transmitted to OpenAI for the sole purpose of real-time transcription and response generation," "VoiceDash does not store your voice files or transcriptions on our servers" (the exception is content you save to Notes), and "Your voice data is strictly excluded from our machine learning training sets. Your input does not influence the development of VoiceDash AI models." The Data Processing Addendum effective May 26, 2026 adds that data is "processed transiently, structurally excluded from AI model training loops," commits to TLS 1.3 in transit and AES-256 at rest, and names nine subprocessors: OpenAI, Groq, Cloudflare, Hetzner, Sentry, Metabase, CustomerIO, RevenueCat, and Apple — Groq does not appear in the privacy policy. Downstream, OpenAI's API data-usage page confirms API inputs are not used for training by default and that abuse-monitoring logs are retained up to 30 days unless a Zero Data Retention arrangement is approved. The homepage advertises "Zero Data Retention processing through its OpenAI partnership," and the founder told AppSumo buyers a formal BAA had been signed with OpenAI, but neither claim appears in the policy or DPA, so treat both as unverified. No SOC 2, ISO 27001, or HIPAA attestation is published and the DPA contains no BAA language. Pricing: Free (1,000 words/mo), Pro $15/mo or $12/mo billed yearly, Teams $29/mo or $24/mo billed yearly for up to five members. AppSumo rating 4.42/5 from 234 reviews; no documented incidents.

Primary source: VoiceDash privacy policy

Deep dive: Is VoiceDash Safe? — full investigation

Does Paraspeech train on my voice data?

Not stated — the policy names its cloud processors but says nothing about training; local mode on Apple Silicon avoids the question entirely.

Paraspeech is a macOS 14+ dictation app (with an iOS Voice Keyboard) from Burlis Management GmbH, Philippstrasse 27, 52349 Dueren, Germany, priced at $14.99/mo, $99/yr, or a local-models-only Lifetime license. It has two data paths. In local mode on Apple Silicon, docs state "When the active speech-recognition backend is local, speech audio is transcribed on your Mac," and "Supported local modes can run offline after setup"; the privacy policy (last updated July 23, 2026) adds that "the last audio file is stored locally on the user's device until replaced by a new recording or deleted by the user." In cloud mode, "When a user selects or is eligible for cloud transcription, Paraspeech may transmit selected audio and related request metadata to Deepgram for cloud transcription," and the opt-in Cloud Cleanup rewrite sends "selected transcript text, prompt text, rewrite instructions, and related request metadata" to Groq and Cerebras. "Eligible for" covers Intel Macs, which "use cloud models with a subscription" and have no local option. The policy contains no statement on AI training in either direction, and cloud retention "depends on the applicable provider and processing terms" — no audio or transcript window is published. Analytics vendors are PostHog, Cloudflare Web Analytics, Ahrefs, Simple Analytics, and Dub.co; hosting is Hetzner, Supabase, and Apple; raw journey events are cleared after 90 days. No SOC 2, ISO 27001, HIPAA, or BAA is offered at any tier, including the custom Enterprise plan. The iOS app's App Store privacy label (v1.0.4) lists Audio Data under "Data Linked to You." No documented incidents.

Primary source: Paraspeech privacy policy (updated 2026-07-23)

Deep dive: Is Paraspeech Safe? — full investigation

Does Voicy train on my voice data?

No, per the homepage — but the commitment does not appear in either governing policy.

Voicy's homepage states: "We do not use your recordings to train an AI model or for any other purpose." That sentence does not appear in Voicy's Data Handling and Security Policy (version 1.3, last updated July 31, 2025) or in its undated privacy policy; neither document contains the word "train." What the policies do commit to is deletion: "No audio recordings are stored by Voicy or Groq - all audio data is permanently deleted immediately after processing" and "All transcribed content is immediately deleted after delivery to the user." Every dictation is relayed through Voicy's Heroku-hosted servers to Groq, which runs the open-source Whisper V3 model; the privacy policy adds that Voicy has "enabled their Zero Data Retention settings" at Groq and does not use Groq's batch or fine-tuning endpoints. Account data is stored on the user's device rather than in a central database, and Mixpanel analytics can be disabled. Two gaps remain as of September 10, 2026: the privacy policy now says servers are "hosted in the EU and US" while the security policy still says all processing is US-based, and the security policy predates the 2026 iOS and Android apps. Voicy is cloud-only at every tier, including Voicy for Teams ($6.79/user/mo, 3+ users), which carries no separate data terms. No SOC 2, ISO 27001, or HIPAA BAA is published; the operating entity is Pishi LLC FZ (UAE).

Primary source: Voicy Data Handling and Security Policy

Deep dive: Is Voicy Safe? — full investigation

Does Blip AI train on my voice data?

Not published — the policy is silent on training, and the policy page could not be re-verified on September 10, 2026.

Blip AI's privacy page at blipai.app/privacy returned a Vercel security checkpoint (HTTP 429) to every automated request on September 10, 2026, so the statements below rest on our June 5, 2026 investigation, which read the policy directly, plus the App Store and AppSumo listings fetched today. Per that investigation, the policy states that voice audio is deleted immediately after transcription, typically within seconds; that transcribed text is not stored on Blip AI's servers and is only delivered to your device; that anonymized usage statistics are retained for up to two years; and that Blip AI is HIPAA compliant with a Business Associate Agreement available on request. The policy does not make an explicit commitment that dictation is not used to train models, and it does not name its subprocessors, even though the product is GPT-powered and therefore routes audio or text through at least one third-party model provider. The AppSumo listing adds: "Your audio isn't stored, and your text stays completely private to you." The App Store listing (developer Ayush Bansal, version 1.0.5 released July 19, rated 2.0 of 5 from 2 ratings) carries the privacy label "The developer does not collect any data from this app," which conflicts with the usage statistics the policy describes. No SOC 2 Type II, ISO 27001, or named auditor backs the HIPAA claim. Blip AI is cloud-only on macOS (12+, Apple Silicon), Windows, and Android; AppSumo team tiers (7 or 15 members) carry no separate data terms. Treat the training posture as not published until the live policy can be read.

Primary source: App Store listing (privacy label)

Deep dive: Is Blip AI Safe? — full investigation

Does WisprType train on my voice data?

No in local mode by architecture — the policy never mentions training, and BYOK cloud mode inherits the provider's terms.

WisprType's privacy policy (effective April 29, 2026) does not contain the word "train." What it does state is structural: "By default, all speech-to-text processing happens entirely on your device using WhisperKit," and the Smart Typing cleanup step "runs a local Llama 3.2 3B model entirely on your device via Apple's MLX framework." Audio recordings "are processed in real time and are not retained after transcription completes, unless you explicitly save them." Transcript history, dictionary, snippets, settings, and any API keys live in a local SQLite database under ~/Library/Application Support/com.wisprtype.app/ and are "never transmitted to our servers." In local mode, WisprType has no server-side copy of your speech, so there is nothing for the vendor to train on. Cloud mode is opt-in: "Cloud providers are never used unless you explicitly configure them by entering an API key and selecting cloud mode in Settings," and the supported providers are OpenAI, Groq, and Deepgram; in that mode, your audio and, for Smart Typing, your raw transcript text go to the provider under your own API agreement, so training and retention follow that provider's terms, not WisprType's. Two caveats keep this personal-grade. First, the policy describes PostHog telemetry as "disabled by default," but our July 2026 hands-on test of v1.1.0, still the current build as of September 10, 2026, found it enabled on a fresh install; check Settings → Privacy. Second, no legal entity, terms of service, business tier, SOC 2, or BAA is published, so there is no party to contract with for regulated use. The app is free during early access and requires macOS 12+ on Apple Silicon.

Primary source: WisprType privacy policy (effective 2026-04-29)

Deep dive: Is WisprType Safe? — full investigation

Does DictaFlow train on my voice data?

No, per the homepage and App Store label — the privacy policy itself is silent on training and names OpenAI, Deepgram, and Groq as processors.

DictaFlow's homepage states "your audio is never used to train models," and its App Store privacy label places Audio Data under "Data Not Linked to You." The privacy policy, last updated August 14, 2026 and operated by SmartBids.ai Corp., a Canadian corporation, contains no training clause; its closest commitment is that DictaFlow does "not sell personal information or use dictated audio or text for targeted advertising." The same policy names "OpenAI, Deepgram, and Groq for transcription and related AI processing" (the July 2026 version we reviewed named NVIDIA; that reference is gone), Apple, Google, and Microsoft for sign-in, Google Cloud for hosting, and Stripe for payments. On retention: "We do not permanently retain audio recordings on our servers after processing in the ordinary course" and "Audio is ordinarily discarded after processing," while transcription text "may be retained when needed to provide features you choose, synchronize supported devices, recover a requested result, investigate a support issue, or comply with law." No retention window in days is published. DictaFlow offers local processing with optional cloud cleanup and describes itself as never fully offline. The policy states the standard service is "not intended for medical dictation, clinical documentation, patient care, diagnosis, treatment, or any use involving protected health information (PHI)." DictaFlow Medical Pro ($39/user/mo for 1–4 seats, $29/user/mo for 5+) is a separate build with a published subprocessor list (Deepgram, OpenAI, Groq on allowlisted routes; Railway/Firebase, Resend/Postmark, Stripe), a Medical privacy policy dated July 24, 2026, and a customer-executed BAA as a prerequisite. No SOC 2 or ISO 27001 is published for either product. Consumer pricing: Free with 2,000 words per month, Pro $7/mo or $69/yr.

Primary source: DictaFlow privacy policy (updated 2026-08-14)

Deep dive: Is DictaFlow Safe? — full investigation

Meeting Transcription

Does Otter.ai train on my voice data?

Yes, by default on de-identified data — opt-out in account settings.

Otter trains automatically on de-identified user data. Recordings and transcripts are not manually reviewed by humans unless the customer gives explicit consent for support troubleshooting; training data is encrypted. Opt-out is by request (privacy@otter.ai or the trust-center form) rather than an in-app toggle, per the June 16, 2026 privacy policy. Four suits filed in the Northern District of California in August–September 2025 (lead: Brewer v. Otter.ai) — alleging Otter "deceptively and surreptitiously" recorded private conversations and used meeting data to train AI models without all-participant consent, under the Electronic Communications Privacy Act, Computer Fraud and Abuse Act, and California Invasion of Privacy Act — were consolidated on October 22, 2025 as In re Otter.AI Privacy Litigation, No. 5:25-cv-06911-EKL, before Judge Eumi K. Lee. On August 13, 2026, Judge Lee granted the motion to dismiss in part and denied it in part: the Wiretap Act, CIPA § 631, both Illinois BIPA voiceprint claims, unjust enrichment, and UCL claims survive, on the reasoning that Otter plausibly "independently collects, retains, and uses communications for its own commercial purposes" — including training — and so may be a third-party eavesdropper rather than the host's tool; CFAA, CDAFA, Washington Privacy Act, and most intrusion claims were dismissed with leave to amend, and the case proceeds toward discovery and class certification. The visible "Otter.ai" bot creates consent issues in the 12 US states with two-party consent laws. Conversations are stored until manually deleted; trash holds for 30 days then auto-purges. Otter holds SOC 2 Type 2.

Primary source: Otter privacy & security

Does Fireflies.ai train on my voice data?

No, per the March 2026 privacy policy — the live litigation is about voiceprint consent, not training.

Fireflies' privacy policy (last updated March 6, 2026) states: "We do not use personal information for AI model training and we contractually prohibit our vendors from using this information for their own model training," and its Zero Data Retention section says meeting content is not used for training by third-party processors; earlier versions of this tracker recorded a default-on training posture, which the current policy text supersedes. Cruz v. Fireflies.AI Corp., No. 3:25-cv-03399 (C.D. Ill., filed December 2025), a class action under Illinois' Biometric Information Privacy Act (BIPA) alleging Fireflies captured all meeting participants' voiceprints without written consent, was terminated in March 2026 and refiled in the Northern District of Illinois; two related N.D. Ill. suits — Fricker v. Fireflies.AI Corp., No. 1:26-cv-02675, and Martinez v. Fireflies.AI Corp., No. 1:26-cv-03512 — were consolidated before Judge Steven C. Seeger with Fricker as the lead case and Werman Salas P.C. as interim class counsel; no ruling on the merits as of September 2026. The visible "Fireflies.ai Notetaker" bot creates consent complications in two-party consent jurisdictions. Audio, video recordings, transcripts, speaker identification, and meeting metadata are stored until manual deletion; voice characteristics used to separate speakers are processed by service providers and, per the policy, never stored on Fireflies' own servers. Account deletion is processed within 30 days. Fireflies is GDPR compliant.

Primary source: Fireflies privacy policy

Does Granola train on my meeting data?

Yes, by default for consumer — Enterprise plans have training off by default.

Granola uses de-identified user data to train internal AI models supporting its services by default. The opt-out is in account settings and is not surfaced prominently — and it exists only for Granola users; other meeting participants have no notice and no opt-out path. Enterprise users have training off by default. Third-party LLM providers (OpenAI, Anthropic) are contractually prohibited from training on Granola data. That training default is now the subject of litigation: Chamberlain v. Granola, Inc., No. 3:26-cv-07926 (N.D. Cal., filed July 30, 2026) is a putative class action alleging Granola intercepts and records virtual-meeting participants without their knowledge or consent and uses their communications to train its AI models by default, claiming violations of the federal Wiretap Act (ECPA), California Invasion of Privacy Act §§ 631–632, CDAFA, common-law intrusion upon seclusion, the UCL, and unjust enrichment. Because Granola captures system audio rather than joining meetings as a visible bot, participants who aren't Granola users get no indication recording is happening — the complaint quotes Granola's own marketing that "other people in the room won't know." Earlier, on April 2 2026, The Verge reported that Granola makes meeting notes accessible to anyone with the share link by default and opts users into internal AI training in default settings, contradicting the app's "private by default" marketing. Architecturally, audio is transcribed in real time on macOS/Windows and not stored, but audio is sent to cloud transcription providers (Deepgram, AssemblyAI) and summarization uses cloud LLMs (OpenAI, Anthropic). Transcripts and notes are stored in a US-hosted AWS Virtual Private Cloud, encrypted at rest and in transit, backed up daily. iOS uses temporarily cached audio.

Primary source: Granola privacy policy

Deep dive: The Granola lawsuit, explained — full analysis

Does Zoom AI Companion train on my meeting data?

No — contractual prohibition on training across customer content.

Zoom AI Companion does not train Zoom's or third-party AI models on customer audio, video, chat, screen sharing, attachments, or other communications-like content. The commitment is contractual rather than an opt-out — it is stated directly in both the Zoom Privacy Statement and the AI Companion Security & Privacy page. The architecture itself is the honest qualifier: AI Companion uses a federated approach that routes prompts through Zoom-hosted models plus third-party subprocessors (OpenAI, Anthropic, and Perplexity for web search). Customer content is processed to deliver features, retained for a period, and shareable inside the org account (hosts and admins can access transcripts and summaries from any meeting they organize or join). Enterprise admins can elect the Zoom-hosted Models Only (ZMO) deployment, which keeps processing inside Zoom infrastructure and removes external subprocessors from the path. BAA is available for specific HIPAA-regulated configurations. Track record context: in August 2023, Zoom quietly added ToS language that would have permitted training on customer content; after The Verge coverage and public outcry, Zoom walked it back and committed to no training without consent, later hardening that to an absolute prohibition. No known breaches of AI Companion content. The Feb 2 2026 Privacy Statement update expanded the definition of "Customer Content" to cover content created using Zoom products beyond meetings, webinars, and messages — the no-training commitment continued to apply, and a further July 27 2026 revision (messaging-campaign data, EU/UK representatives, transfer mechanisms) left it unchanged.

Primary source: Zoom AI whitepaper — data governance & privacy

Does Fathom train on my meeting recordings?

Yes, by default on de-identified meeting content for Fathom's own models — with an in-app opt-out; its third-party AI vendors may not.

Fathom's privacy policy (last updated August 16, 2026) states: "Based on how you configure your account settings, we may use and create de-identified data generated from Meeting Content Information to improve our Service by training, improving, and customizing our in-house artificial intelligence models. You can opt out from this use of your data in your account settings." The help-center article "Is Fathom secure?" confirms the default is on — "Fathom uses de-identified customer data to improve the accuracy of our proprietary AI models" — and that individuals opt out under Settings while organizations on Team Edition can opt out every user from Organization Settings. The same policy draws a line for outside vendors: "We do not authorize third parties (e.g., OpenAI, Anthropic, Google, etc) to use your personal information or Meeting Content Information to train their artificial intelligence models," and the help center names Anthropic, OpenAI and Google as the AI sub-processors. Fathom collects audio and video recordings, transcripts, speaker identification, calendar data, attendee names and emails, and voiceprints in its iOS app; all data is stored in the United States. There is no published retention window for active accounts — data is kept "as long as necessary" — but on account deletion recording data is removed and backups purge after a further 7 days. Pricing: Free, Premium $20/mo ($16 annual), Team $19/$15, Business $34/$25, Enterprise custom; a HIPAA BAA and custom retention appear only on Enterprise. Fathom states SOC 2 Type II, HIPAA and GDPR compliance and no ISO 27001. As of September 10, 2026 no lawsuit or regulator action names Fathom; the 2026 wiretap wave targets Otter, Fireflies and Granola.

Primary source: Fathom privacy policy (Aug 16 2026)

Does tl;dv train on my meeting recordings?

No — the policy excludes Customer Content from training by tl;dv and its AI providers; the 2026 concern is a six-month Firestore exposure, not training.

tl;dv's privacy policy (last updated September 9, 2026; Tldx Solutions GmbH, Aachen, Germany) states that Customer Content — meeting recordings, transcripts, notes and files — is not used "to train, fine-tune, or improve foundation models, large language models, or other generative AI models," and that third-party AI providers process it "solely to provide the requested functionality" without training their general-purpose models. The security page repeats: "No customer data is used to train the AI." Named processors are Anthropic and Google Vertex AI for AI features and AssemblyAI and ElevenLabs for transcription; hosting is on Google Cloud, Hetzner and Wasabi, primarily in Germany and Finland, with AI processing selectable in the US or EU. Retention: free users' recordings are deleted after 3 months; paid users' data is kept until account deletion; logs 10 days; data on non-client meeting participants up to 5 years or until a deletion request. Certifications stated: SOC 2, ISO 27001, PCI DSS, GDPR; no HIPAA. The track-record issue is security, not policy. Researcher bobdahacker reported on January 28, 2026 that a missing tenant-isolation rule on the Firestore "meetings" collection let any authenticated account enumerate 181,874 meeting records across 84,312 users — creator emails, conference IDs, recording status, timestamps — including roughly 1,000 joinable live calls; follow-ups through July 22 went unresolved and the write-up was published August 4, 2026. tl;dv told GIGAZINE on August 13, 2026 that these were "two separate vulnerabilities," that no audio, transcripts, AI notes, passwords or billing data were reachable, that the second vector was patched within 24 hours of discovery, and that Firebase has been removed from its infrastructure. No lawsuit or regulator action names tl;dv. The tldv.io pricing page returned 404 on September 10, 2026, so prices are not recorded here.

Primary source: tl;dv privacy policy (Sept 9 2026)

Does Read AI train on my meeting recordings?

Not unless you opt in — model improvement is off by default, and Workspace-API data is excluded from generalized training.

Read AI's privacy page states: "Contributing to model improvement is off by default. You decide - and can change your mind any time." The privacy policy (last updated June 29, 2026) implements this as a Customer Experience Program you "opt-in or opt-out" of in account settings, and adds: "We do not, and do not permit third party AI tools to, use user data collected via Google Workspace APIs to develop, improve, or train generalized/non-personalized AI and/or ML models." The policy does not name its third-party AI providers. What Read collects is broader than a typical notetaker: meeting audio and video, transcripts, and — when Google integrations are connected — Gmail message bodies and metadata, Google Calendar events and guest lists, and Google Chat messages. Audio and video are stored "in no case for longer than 2 years"; users can delete reports they own, and any participant who types "opt out" in chat has the meeting's data "immediately and permanently deleted." Read announces itself at the start of every meeting and requires host approval before joining. Since March 26, 2026 it also offers a botless Google Meet mode built on the Meet Media API, free for all customers, plus desktop and mobile apps that record "locally or to the cloud without a bot"; transcription and analysis still run in Read's US-hosted cloud. Pricing: Free, Pro $19.75/mo ($15 annual), Enterprise $29.75/$22.50, Enterprise+ $39.75/$29.75; HIPAA support, SAML/SCIM and a custom retention policy are Enterprise+ only. Read states an annual SOC 2 Type II audit, HIPAA technical safeguards with a BAA on request, and AES-256 / TLS 1.2+ encryption; its support-center security articles and trust portal could not be fetched, so the audit report is unverified. As of September 10, 2026 no lawsuit or regulator action names Read AI.

Primary source: Read AI privacy policy (June 29 2026)

Does Notta train on my recordings and transcripts?

Not published — Notta's policy excludes only Google Workspace API data from training and says nothing about its own models on your audio.

Notta's privacy policy (effective September 2, 2025; contracting entity Notta PTE. LTD. in Singapore, or Notta 株式会社 for Japan users) contains one sentence about model training, and it is scoped to a single data source: "Google Workspace APIs are not used to develop, improve, or train generalized/non-personalized AI and/or ML models." Nothing in the policy, the terms of service (effective September 2, 2026) or the security page states whether audio you upload or record, or the transcripts Notta generates, are used to train Notta's own models. The policy lists "To analyze and improve our Services" as a purpose, and the terms grant Notta a licence to "access, process, copy, export, and display Customer Data" for service purposes. Secondary blogs describe an opt-out setting or a support-request opt-out; no Notta page fetched on September 10, 2026 shows one, so that claim is unverified. Retention is also not published in the policy: a Notta help-center search snippet (the article itself returned 403) says transcription data is stored indefinitely unless the user, workspace or account deletes it, and that automatic file deletion (1–365 days) is an Enterprise-only setting. Notta collects audio, video, voice and text inputs, account, device, IP, location and usage data; it hosts on AWS with no region stated, and encrypts with TLS 1.2 in transit and AES-256 at rest. The security page states SOC 2 Type II through an independent audit, ISO 27001, GDPR and CCPA compliance, and that Notta "follow[s] HIPAA guidelines"; no BAA is offered on any fetched page. Notta Desktop (public beta, macOS 13+ / Windows 10+) records without a bot but transcribes in the cloud. Pricing: Free 120 min/mo, Pro $13.61/mo ($8.17 annual), Business $27.78/$16.67, Enterprise custom from 51 seats. No lawsuit, breach or regulator action names Notta.

Primary source: Notta privacy policy (effective Sept 2 2025)

Does Google train on my Meet "Take notes for me" transcripts?

Not on Workspace accounts, by contract; on personal Google AI Pro / Ultra accounts the exclusion is not published.

The answer splits by account type. For work and school accounts, the Google Workspace generative-AI privacy hub (updated August 14, 2026) states: "Workspace does not use customer data for training models without customer's prior permission or instruction" and "Your content is not human reviewed or otherwise used for Generative AI model training outside your domain without permission." Meet is listed among the covered apps, prompts are customer data under the Cloud Data Processing Addendum, Gemini in Meet is covered by the Google Business Associate Agreement, and admins can disable Gemini in Meet, require participant consent, and set retention ("90 days to indefinite, as determined by admins"). The notes document "is saved in the meeting organizer's Google Drive in the 'Google Meet' folder" and "follow[s] the Meet retention policy that your organization has configured." For personal accounts, Google opened "Take notes for me" to Google AI Pro and Ultra subscribers on June 29, 2026. No Google page fetched says whether those transcripts are exempt from training. The Gemini Apps Privacy Hub (updated August 10, 2026) says Gemini Apps Activity — including audio and Gemini Live video — is used to "provide, develop, and improve its services (including training generative AI models)," that "a subset of chats are reviewed by human reviewers," and that activity auto-deletes after 18 months by default; whether Meet notes are logged as Gemini Apps Activity is unverified. In both cases Meet notifies every participant that notes are being taken. Track record: Thele v. Google LLC, No. 5:25-cv-09704 (N.D. Cal., Judge Noël Wise), filed November 11, 2025, alleges Google switched Gemini "smart features" on by default across Gmail, Chat and Meet on or about October 10, 2025, in violation of CIPA and the Stored Communications Act. On July 7, 2026 the court dismissed the first amended complaint for lack of concrete harm with leave to amend; a Second Amended Complaint was filed August 18, 2026, and Google stipulated on August 20 to extend its response deadline. No merits ruling exists.

Primary source: Workspace generative-AI privacy hub (Aug 14 2026)

Clinical Documentation (Ambient AI Scribes)

Does Abridge train on my patient conversations?

Yes, on de-identified data — and since August 2026 the exact terms are in your health system's BAA, not a public policy.

Abridge is an ambient clinical scribe sold only to health systems; there is no individual plan. Its archived privacy policy (effective December 1, 2022) stated: "We use de-identified data for research and development of new products or tools, to refine our algorithms and machine learning applications, and to improve the Services." The policy last updated August 27, 2026 drops that sentence and says Customer Data "is governed exclusively by the terms of our agreements with those customers, including our Business Associate Agreements (BAAs)." On June 11, 2026, Abridge and Nvidia announced a clinical-conversation model trained on de-identified doctor–patient conversations from Abridge's platform of roughly 100 health systems. The training term is negotiable at the contract level — Kaiser Permanente told The Markup (June 16, 2026) it does not allow its data to be used for training — but no clinician- or patient-level opt-out is published. Abridge does not publish an audio-retention window; customer FAQs cite 30 days (CPCMG) and 14 days (Kaiser) for audio and transcripts, with the finalized note living in the EHR. Litigation targets Abridge deployments rather than Abridge: Saucedo v. Sharp HealthCare (San Diego Superior Court, filed November 26, 2025) alleges recording without all-party consent under California's CIPA and that the tool inserted "consented" statements into charts; Washington et al. v. Sutter Health and MemorialCare (N.D. Cal., filed April 8, 2026) alleges CIPA, CMIA, UCL, and federal Wiretap Act violations. Abridge holds SOC 2 Type 1 and 2 and TX-RAMP, states it is HIPAA-aligned, and signs BAAs. For comparison, Microsoft's Dragon Copilot puts its no-training commitment in the Microsoft Products and Services DPA, a published document.

Primary source: Abridge privacy policy (Aug 27, 2026)

Deep dive: Best AI medical scribe tools for doctors — full comparison

Does Freed train on my clinical notes?

Yes, on de-identified notes, gated by a Platform setting whose default Freed does not publish — check it on day one.

Freed is a self-serve AI scribe for individual clinicians (Starter $39/mo for 40 notes, Core $79/mo, Premier $104/mo on annual billing; 7-day trial; Groups on custom pricing). Its security page states: "AI is only trained on de-identified notes." The Platform Terms of Use (last updated August 13, 2026) are more specific: Freed may use De-identified Data "to train the AI/ML models to improve the Platform, including but not limited to improving the Platform accuracy, efficiency and quality of speech recognition and Output, with your consent provided via Platform settings," and will "link your De-identified Data with your customer ID and use it to customize and train our Platform based on your specific styles and requirements." The same Terms give Freed "the right in our sole discretion to use De-identified Data and to disclose such De-identified Data to third parties." Freed states "Protected health information is never used for AI training purposes." Whether the consent setting is on or off at sign-up is not published. On retention, Freed states "Patient recordings are saved only until the note has been completed, then automatically deleted," with a settings option to "delete the Patient Recordings immediately once they are processed"; notes follow a user-chosen 30-day auto-delete or term-of-agreement window. Data is hosted on Microsoft Azure in the United States. Freed holds SOC 2 Type 1 and 2, states it is HIPAA- and HITECH-aligned, and signs BAAs (the BAA is Schedule A of the Terms). No breaches, lawsuits, or regulator actions naming Freed were found as of September 2026.

Primary source: Freed Platform Terms of Use (Aug 13, 2026)

Deep dive: Best AI medical scribe tools for doctors — full comparison

Does Heidi Health train on my consultation data?

No, per Heidi's July 2026 statement — though the privacy policy still reserves de-identified and aggregate use.

Heidi Health is an Australian AI scribe with a free tier (unlimited transcription), a Clinician plan with a 14-day trial, and Teams and Enterprise contracts; the pricing page shows no dollar amounts. Heidi's support article "How Heidi protects your data" (July 23, 2026) states: "Heidi does not use consultation data, session transcriptions, or generated notes to train its AI model in any way." The privacy and security FAQ adds: "We don't use any of your sensitive health information for model training. We only use your data for the purpose it was collected." Two qualifications: the privacy policy (page dated October 2024) still lets Heidi "de-identify and/or aggregate your personal information, including your health information" for platform development, and for the Ask Heidi evidence feature it says "Queries may be reviewed and used in de-identified form to improve the platform, but PHI will not be used for model training." No opt-out is published for that de-identified query review. On retention, the FAQ states "Heidi does not store audio recordings of patient consultations" and that notes and transcripts are configurable "between 1 day and 'never delete'," with never delete as the default. Health information is stored in the user's own jurisdiction (Australia, UK, US, EU, Canada); US data is "locally hosted within the United States." Heidi lists ISO 27001, ISO 42001, and SOC 2 Type 2, states it is HIPAA-aligned, and says "We make a BAA available to every covered entity we work with." It raised a $65M Series B in 2026. No breaches, lawsuits, or regulator actions naming Heidi were found as of September 2026.

Primary source: How Heidi protects your data (Jul 23, 2026)

Deep dive: Best AI medical scribe tools for doctors — full comparison

Does Nabla train on my encounter audio?

Not by default — only on de-identified audio you choose to submit as feedback, which Nabla then keeps indefinitely.

Nabla is a French-founded ambient scribe hosted on Google Cloud, with U.S. client data in U.S. regions and non-U.S. data in EU (Belgium) data centers; Azure handles speech-to-text in U.S. regions. Its trust center states: "By default, we don't store audio." and "We retain clinical notes for a short period of time (14 days), which is configurable by client based on geographic region requirements" — the shortest published note window among the clinical scribes in this tracker. The only training path Nabla documents is optional feedback. Its help center says: "All audio is anonymized immediately upon sharing and original audio replaces PHI with a 'beep' sound," "Only the de-identified file is stored," and "Audio is stored in our Feedback database indefinitely." The companion Feedback article states that feedback is used for training Nabla and for troubleshooting; that article returned 404 on September 10, 2026, so the wording is per its search-index text and marked unverified. The public privacy policy (May 16, 2023) covers website visitors only and says nothing about clinical audio. Nabla lists SOC 2 Type II (November 2024–October 2025 period), ISO 27001 (September 2025), and TX-RAMP Level 2 (January 2026), states it is HIPAA- and GDPR-aligned, and lists BAA terms among its trust-center documents; BAA availability on self-serve plans is not published, and Nabla publishes no prices. University of Iowa Health Care's patient FAQ states "Nabla does not retain the audio recording." No breaches, lawsuits, or regulator actions naming Nabla were found as of September 2026; a January 2026 Verite News story questioned patient-consent practice at an LCMC Health deployment (page not retrievable, unverified).

Primary source: Nabla Trust Center

Deep dive: Best AI medical scribe tools for doctors — full comparison

Does Suki train on my dictation and visit data?

Yes, on de-identified data, under a Terms-of-Service license with no published opt-out.

Suki is an ambient scribe and dictation assistant sold to practices and health systems; no self-serve pricing page was found (the /pricing and /security URLs return 404). Suki's developer security documentation states: "Any data that is used for ML training and improving the product is de-identified." It describes the method: "For audio, we use a de-identification algorithm that breaks audio into chunks and isolates them such that the original audio cannot be re-constructed," and "We de-identify the transcript by removing all PII." The Terms of Service make the permission contractual: Suki may access Customer Data "for the purpose of using the Customer Data for system tuning, grammar tuning, training of acoustic models and other models, tools and algorithms," and "Suki owns all rights in and to any aggregated, non-identifiable data it develops or creates." The privacy policy (last updated May 16, 2025) does not mention model training, and no opt-out toggle or contractual opt-out is published — a customer that wants one must negotiate it. Retention is published: audio and transcripts are "permanently deleted after 30 days," and clinical notes are retained for the duration of the service contract. Data is hosted on Google Cloud in the United States with AES-256 at rest and TLS 1.2 in transit. Suki's homepage states it is "SOC2 Type 2 certified, and HIPAA compliant"; the developer docs add "We sign Business Associate Agreements (BAA) for patient data handling with our customers." No breaches, lawsuits, or regulator actions naming Suki were found as of September 2026.

Primary source: Suki developer docs — Security & Compliance

Deep dive: Best AI medical scribe tools for doctors — full comparison

Does DeepScribe train on my patient data?

Yes — its Terms of Use license customer data to machine learning for product improvement and de-identified data "for any legal purpose," with no published opt-out or retention window.

DeepScribe is an enterprise ambient scribe with no self-serve plan (the site offers "Request a demo" only). Its Terms of Use (updated October 17, 2024) grant DeepScribe the right to "use, store, host, perform, display and create derivative works from the Customer Data, including by combining Customer Data with data from third party sources and utilizing machine learning and artificial intelligence applications, for the purposes of (a) providing the DeepScribe Services ... (b) complying with applicable laws and regulations; and (c) operating, analyzing and improving DeepScribe's products and services." A separate clause lets DeepScribe "retain, share, and use the De-Identified Data for any legal purpose including without limitation for purposes of operating, analyzing, improving or marketing the DeepScribe Services." That is broader than the de-identified-only wording at Freed, Suki, or Abridge, and because de-identified data falls outside PHI, a standard BAA does not narrow it — a health system must do that in the master agreement. The privacy policy (no date visible on the page) says only that data is used to "improve the Services" and is retained "for as long as it remains necessary for the identified purpose or as required by law"; no audio or transcript deletion window is published anywhere. Data is processed on servers "located in the United States"; subprocessors are not named. The homepage shows SOC 2 and HIPAA badges, the security-practices page says DeepScribe holds "relevant certifications" without naming them, and the BAA is incorporated as Exhibit A of the Terms and "shall control" over the Terms on PHI. No breaches, lawsuits, or regulator actions naming DeepScribe were found as of September 2026.

Primary source: DeepScribe Terms of Use (Oct 17, 2024)

Deep dive: Best AI medical scribe tools for doctors — full comparison

Local LLM Runtimes

Does Ollama train on my data?

No — Ollama operates no model server.

Ollama is a local runtime: there is no Ollama-operated model server to send prompts or completions to. Models run entirely on local hardware. Ollama exposes an OpenAI-compatible local REST API on localhost:11434; nothing leaves the user's machine in default operation. There is no vendor server, so there is no retention surface and no training path. Ollama is open-source under the MIT license with over 100,000 GitHub stars; the codebase is transparent and there are no documented incidents. Caveat: users who deliberately expose the local API beyond localhost create their own attack surface — that is a deployment choice, not an Ollama default.

Primary source: Ollama GitHub repo

Does LM Studio train on my data?

No — inference runs locally; no chat data leaves the device.

LM Studio runs models locally on the user's hardware with an optional OpenAI-compatible local server. Inference happens entirely on-device; there are no LM Studio servers in the inference path. The optional LM Link feature uses end-to-end encrypted Tailscale mesh VPNs to connect a user's own devices for remote inference — chats remain local even when accessed from another linked device. LM Link shares only device-list metadata (not chats) with LM Studio's backend for device discovery. The application is closed-source but transparent about data flows on the LM Link product page; no documented incidents.

Primary source: LM Studio

Does Jan train on my data?

No — Jan operates no model server.

Jan is an open-source local LLM runtime under Apache 2.0. Models run locally; there is no Jan-operated model server, so there is nothing to train on and no retention surface. Optional cloud-provider integrations are explicitly opt-in — when enabled, the chosen provider's terms apply for that specific feature. In default operation, nothing leaves the device. The codebase is auditable on GitHub with an active community roadmap; no documented incidents.

Primary source: Jan

Frequently Asked Questions

What does "on-device" actually mean?
On-device means a tool can complete its core workflow without sending your input to a vendor's servers. For dictation, that means audio is captured, transcribed, and discarded entirely on your computer — nothing leaves the machine. Most AI assistants and coding tools are not on-device by default: they transmit your prompts and code to a cloud model, even if the vendor doesn't retain or train on it. "Partially on-device" means parts of the workflow are local but specific cases (unsupported languages, agentic operations, large models) fall back to the cloud. Apple Dictation and Cline (when paired with a local Ollama or LM Studio model) are examples of partial on-device.
Why does the same tool show different answers for Free vs Business?
Consumer and business tiers operate under separate contracts. Most major AI vendors train on consumer data (or did until very recently) and explicitly exclude business / API / enterprise data from training under their commercial agreements. The two tiers can use the same underlying model but with different data-handling guarantees. Conflating the two is the most common error in third-party comparison articles. The tier filter at the top of each table separates them so you can answer either question independently.
Can a tool "unlearn" my data after training?
Practically, no. Once a model has been trained on a piece of data, the parameters reflect that training and cannot be cleanly reverted on a per-record basis. Vendors offering deletion typically delete the conversation record but cannot remove its influence on the model that has already absorbed it. This is why the relevant question is "will it be used for training in the first place," not "can I delete it later." Pages like this one focus on the training question because the deletion question rarely changes the outcome for already-trained models.
How often is this updated?
Each row carries a Last verified date. We re-check every cell against its primary source on a roughly monthly cadence, plus immediately whenever a vendor announces a policy change. The Recent Changes timeline at the top of the page lists every dated change we have logged. If you find an outdated cell or a missing change, email hi@getvoibe.com — we'll update and credit the report.
How is each tool scored?
Every tier carries four independent scores on a 0–25 scale that sum to a 0–100 composite: Training (does the vendor train on your data by default), Retention (how briefly is data kept), On-device (can the workflow run without sending data to the vendor), and Track record (the vendor's documented incidents, breaches, and unfavorable policy changes). We separate the axes because a tool with a strong contract but a weak track record isn't strictly better or worse than the inverse — but we also publish the composite total because most readers ultimately want a single answer. Per-axis buckets: 22–25 architectural or contractual guarantee, 17–21 default-off or short retention, 11–16 user must take action or policy has caveats, 6–10 unfavorable default with opt-out, 1–5 no opt-out or indefinite retention, 0 the policy does not address that dimension. Composite buckets: 85+ Excellent, 70–84 Strong, 55–69 Adequate, 40–54 Weak, below 40 Poor.
What score should I look for?
The right floor depends on the work. For regulated industries (healthcare, legal, financial) or anything that triggers regulator notification on leak, look for 85+ in the business tier and verify the vendor signs a BAA or DPA. For sensitive but unregulated business work — proprietary code, internal docs, M&A drafts — 70+ is the floor; below that, you are relying on opt-out toggles that team members may not have flipped. For day-to-day drafting, research, and light coding on non-secret content, 55+ is fine. For personal low-stakes use (notes to self, brainstorming) any score works as long as you understand what the tool retains. The Use case fit section on this page lists the tools that meet each threshold.
How do I report an error or a missing tool?
Email hi@getvoibe.com with the tool name, the cell you think is wrong, and a primary-source link. We'll verify and update on the next pass. Tool requests are welcome, but to be added we need a vendor-published privacy or data-handling page that we can cite — marketing claims aren't enough.

This tracker is maintained by the team at Voibe. We built it because privacy is the central design constraint of our product, and we kept being asked these questions. Voibe is one of the tools listed — the methodology is the same for every row.

Related reading on this site: Is Wispr Flow safe? · Typeless privacy issues · Apple Dictation privacy · Voice data privacy · Cloud vs local dictation · HIPAA dictation · Zero data retention explained.