Wispr Built Its Own Voice Model. Whose Voice Trained Canto?
Wispr raised $280M and shipped Canto — a model tuned for noise, accents and Hinglish. Their own docs hint at whose voices taught it. Not the Fortune 500.
Where Did Wispr Get the Voice Data to Train Canto?

Wispr's new speech model is unusually good at exactly the things a clean training set cannot teach: wind, open-plan offices, heavy accents, and Hindi-English code-switching that has to come back in romanized script rather than Devanagari. Somebody's voice taught it that. The company's own documentation narrows down whose.
TL;DR: On August 17, 2026, Wispr announced a $280 million Series B at a $2 billion valuation led by Menlo Ventures, and previewed Canto, its first proprietary speech model. Wispr has not published a training-data disclosure for Canto, and I could not find one. What Wispr has published is a default: its security and compliance FAQ states that “Privacy Mode off (standard mode): audio and transcription data may be used to evaluate, train, and improve Wispr's models. This is the default for trial and standard accounts. Enterprise and HIPAA BAA customers run with Privacy Mode on by default.” That is an inverted consent structure — the Fortune 500 accounts are walled off from the corpus, and the free, trial and individual paying accounts are the corpus. More than 60 billion words have been dictated through Flow.
This piece is an analysis, not an accusation. I read the Series B post, the privacy policy, the Data Controls page and the security FAQ, and I lay out what each one actually says, what it does not say, and where the evidence stops and inference begins. The three most likely sources — default-on standard accounts, the correction signal Wispr calls “edits,” and a heavily subsidised Indian user base — are all consistent with what Wispr documents about itself. None of them are things Wispr has confirmed about Canto specifically.
Key Takeaway
Wispr has not disclosed what Canto was trained on. But its own security FAQ says model training is ON by default for trial and standard accounts and OFF by default for Enterprise and HIPAA customers — so whatever else went into Canto, the users least able to negotiate were the ones opted in.
Key Takeaways: Canto, the $280M Round, and the Data Question
Every figure below comes from Wispr's own pages or from named third-party reporting, retrieved August 25, 2026.
| Item | Detail |
|---|---|
| Announced | August 17, 2026 (Wispr blog, Tanay Kothari) |
| Round | $280 million Series B at a $2 billion valuation, led by Menlo Ventures |
| Total raised | $361 million to date |
| New model | Canto — preview of Wispr's first proprietary speech model |
| Headline claim | In the hardest conditions, word error rate falls from more than 30% to between 5% and 10%; roughly 30–35% fewer dictations need editing |
| Scale | More than 60 billion words written with Flow; used at almost all Fortune 500 companies and 10,000+ enterprises |
| Training-data disclosure | None published for Canto as of August 25, 2026 — the announcement does not name a single dataset, licence or vendor |
| Default for free/trial/standard | Model training ON — “the default for trial and standard accounts” (security & compliance FAQ) |
| Default for Enterprise/HIPAA | Model training OFF — “Privacy Mode on by default” |
| What gets used | “audio, transcript, edits” (Data Controls) |
| Where the toggle lives | Settings > Data and Privacy > “Improve the model for everyone” — renamed from “Privacy Mode” |
| India pricing | ₹320/month on annual billing versus $12/month in the US — about 71% less (TechCrunch) |
| India share | 14% of global downloads but 2% of in-app purchase revenue, October 2025 – April 2026 |
| Android free tier | “Free + unlimited during launch” (wisprflow.ai/android) — no word cap on the platform Wispr prioritised for India |
Two of those rows do the real work. Training is on by default for the accounts that pay little or nothing, and off by default for the accounts that pay the most. Everything else in this article is an attempt to work out what that structure produced.
What Wispr Announced on August 17, 2026

Wispr announced two things at once: money and a model. The Series B post by CEO Tanay Kothari confirms $280 million at a $2 billion valuation led by Menlo Ventures, with Notable Capital, NEA, Neo Ventures, 8VC and MVP Ventures returning, and Acrew, Activate, Forerunner, Goodwater, Peak XV, Together Fund and PLUS Capital joining. Total capital raised reaches $361 million. TechCrunch reported the round the same day.
The model is the more interesting half. Canto is Wispr's first proprietary speech model, shipped as a preview, and Kothari frames it against the industry's standard training practice in a passage worth quoting in full:
“Most speech models are trained and evaluated on clean recordings made in a quiet room with a good microphone and an accent the model has heard many times before. Almost nobody lives in those conditions. You're in a car, on a street, or sitting a foot away from someone else's conversation in an open office. We built this model for where people actually use Flow.”
That last sentence is the whole question in seven words. “Where people actually use Flow” is not a dataset you can licence. It is a description of Flow's own users, in their own cars and offices and streets. The claimed result: in the hardest conditions — background noise, wind, heavy accents or music — error rates fall from more than 30% of words to between 5% and 10%, and across everyday use Wispr expects 30–35% fewer dictations to need editing. Those are Wispr's own internal numbers, not a third-party benchmark, and no evaluation set is named.
For the product itself rather than the model — features, plans, third-party ratings and the reliability record — see our Wispr Flow review.
Canto also targets code-switching, and the Hinglish example in the announcement is strikingly specific:
“Hindi speakers usually write in Devanagari, but when they're speaking Hinglish they expect it back in romanized script, so hearing every word correctly still isn't enough to give someone something they can send. The same is true of the vocabulary particular to your own life, which is why the model draws on your dictionary and the names of the people around you.”
Knowing that Hinglish speakers want romanized output is not a fact you extract from an audio corpus. It is a fact you learn by watching people reject the Devanagari you gave them. Hold that thought — it comes back below.
Canto's Own Description Is the Best Clue to Its Training Set
When a lab will not say what trained a model, the model's capabilities are the next best evidence. A speech model can only be good at conditions its training data contained, so Canto's feature list doubles as a partial description of its corpus.
Two of the four headline capabilities have a plausible non-user explanation:
- Background noise, wind and music. Noise augmentation is a standard, decades-old technique: take clean speech, mix in recorded noise, train on the result. This capability genuinely can be manufactured, and any competent speech team would do exactly that.
- Heavy accents. Partly addressable through public and licensed multi-accent corpora. Mozilla Common Voice and commercial data vendors both sell accented speech. It would not be cheap or complete, but it exists.
The other two do not have one:
- Hinglish returned in romanized script. Getting the words right is a corpus problem. Knowing which script the user wants back is a preference problem, and preferences live in user behaviour, not in audio files. You learn this by shipping Devanagari to Hindi-English speakers and watching them fix it.
- Your dictionary and the names of the people around you. Wispr states plainly that the model “draws on your dictionary and the names of the people around you.” No licensed dataset contains your colleagues' names. That capability is constituted by user data — there is no other way to build it.
So the honest reading is a split one: Canto's noise robustness could be synthetic; its personalisation and its code-switching behaviour almost certainly are not. The interesting question is not whether Wispr used user data — the dictionary feature means it did, by construction. The question is which users, and under what consent.
Source 1: The Default That Turns Standard Accounts Into a Corpus
The clearest answer Wispr gives is in its security and compliance FAQ, and it is worth reading twice:
“Privacy Mode off (standard mode): audio and transcription data may be used to evaluate, train, and improve Wispr's models. This is the default for trial and standard accounts. Enterprise and HIPAA BAA customers run with Privacy Mode on by default. If you handle confidential or privileged information, enable Privacy Mode before dictating.”
The same document confirms the enterprise side separately: “Enterprise defaults. Data sharing defaults to off (Privacy Mode effectively on).”
Laid out as a table, the structure is hard to miss:
| Account type | Typical price | Model training by default |
|---|---|---|
| Basic (free) | $0 — 2,000 words/week on Mac and Windows | On |
| 14-day trial | $0, no card required | On |
| Pro (standard) | $15/month, or $144/year | On |
| Enterprise | Quoted; contracted | Off |
| HIPAA BAA customers | Contracted | Off |
Read down the right-hand column. The protection tracks the contract, not the sensitivity of the speech. A solo therapist on Pro at $144/year dictating session notes is opted in by default; a Fortune 500 marketing team dictating press releases is opted out by default. Wispr's own privacy page says “You decide whether data can be used to train or improve our models” — true in the sense that a toggle exists, less true in the sense that most people never find it, and the page does not tell you which way it is pointing.
Scale turns the default into a corpus. Wispr says more than 60 billion words have been written with Flow. Even a modest fraction of that, from accounts that never opened Settings, is one of the larger real-world speech datasets in existence. For the wider pattern of what Flow collects beyond audio, see our breakdown of what Wispr Flow's founder revealed about user tracking and the full incident history in Is Wispr Flow safe?
Warning
Privacy Mode is not the same as zero retention. Turning it on stops training but leaves audio, transcripts and dictation history on Wispr's servers. Zero Data Retention requires Privacy Mode on AND Cloud Sync off — two separate toggles.
The Toggle That Changed Its Name (and Its Polarity)
The setting that controls all of this was renamed. Wispr's Data Controls page, last updated August 18, 2026 — the day after the Canto announcement — states: “This setting was previously called 'Privacy Mode.' If you had Privacy Mode enabled, your data is not shared for model training. This behavior remains unchanged.”
The behaviour is unchanged. The framing is inverted. Compare what the two labels ask of you:
| Old label | New label | |
|---|---|---|
| Name | Privacy Mode | Improve the model for everyone |
| What ON means | Your data is protected | Your data is shared |
| What OFF means | Your data is shared | Your data is protected |
| What the name appeals to | Your self-interest | Your generosity |
A toggle called “Privacy Mode” that is off invites the question “should I turn privacy on?” A toggle called “Improve the model for everyone” that is on invites the question “why would I turn that off?” Same switch, same wiring, opposite social pressure. Nothing about this is unlawful or even unusual — it is ordinary consent-interface design, of the kind regulators have started calling dark patterns when it runs one way and best practice when it runs the other.
What the Data Controls page does not say, anywhere, is which position the switch ships in. That fact lives only in a separate help-centre FAQ aimed at security reviewers. If you want to know whether you are currently in the training set, the marketing privacy page will not tell you, the Data Controls page will not tell you, and the privacy policy — last updated August 19, 2026, two days after the announcement — describes training only in opt-in terms: “If you opt to share your content with us for model training…”
I want to be precise about what I can and cannot show here. I can show that these three pages describe the same setting at different levels of candour, and that two of them were updated in the 48 hours after Canto was announced. I cannot show what changed in those updates, because Wispr does not publish diffs, and I have no archived copies of the previous versions.
Wispr's own India landing page carries a pull-quote from Kothari: “Any company that wants to sit this close to people's work should give you upfront control of your privacy.” I agree with the sentence. “Upfront” is the word doing the work, and a default that is documented only in a compliance FAQ is not upfront.
Source 2: “Edits” Is the Most Valuable Word on the Data Controls Page
Wispr's Data Controls page specifies exactly what the training toggle governs: “your data (i.e. audio, transcript, edits) may be used to evaluate, train, or improve AI models, by Wispr.”
Audio and transcript are the obvious pair. Edits are the valuable one, and I think it is the most underrated word in the whole disclosure.
Supervised speech training needs labelled data: an audio clip paired with the correct text. Producing that normally means paying human annotators to transcribe recordings, which is one of the largest costs in building a speech model. An edit collapses that cost to zero. When Flow transcribes “bishop” and you retype “Biswaroop,” you have just produced a perfectly labelled training example — the audio, the model's wrong answer, and the human-verified right answer — for free, in the course of doing your own work.
Wispr collects this at a second point too. Its Data Controls page describes an Auto-add to Dictionary feature in Settings > Personalization: “Flow monitors the text box where it pastes text to detect any edits you made to transcribed words. If you change the spelling of a word, it is automatically added to your dictionary.”
Now put that next to the metric Wispr says it manages by. From the Series B post: “Internally, we measure ourselves against what we call zero edit rate, which is the share of everything you say that comes back right the first time and needs nothing from you at all.”
The metric and the training signal are the same stream. Wispr scores itself on the dictations you do not edit and learns from the ones you do. That is a genuinely elegant flywheel, and it is also the mechanism by which the Hinglish insight in the announcement is most plausibly discovered: you ship Devanagari, Hindi-English speakers retype it in Latin script, and the correction stream tells you what the user wanted without anyone having to run a survey.
To be explicit: Wispr has not said this is how Canto learned romanized Hinglish. I am saying it is the mechanism their own documentation describes, applied to the example their own announcement chose.
Source 3: India, ₹320, and the Problem You Cannot Synthesise
India is where the commercial logic and the training logic point at the same users. Per TechCrunch's May 2026 report, India is Wispr Flow's second-largest market after the United States by both users and revenue, and the numbers underneath that are lopsided: India accounted for 14% of global downloads between October 2025 and April 2026 but only 2% of in-app purchase revenue. India supplies about seven times more share of usage than of revenue.
The pricing is deliberately subsidised. Wispr Pro in India is ₹320 per month on annual billing — roughly $3.50 — against $12 per month in the United States, a discount of about 71%. Kothari told TechCrunch the eventual target is ₹10–20 per month, or 10 to 20 US cents. On Android, the platform Wispr prioritised for India, the free tier is advertised as “Free + unlimited during launch” — no weekly word cap at all, against 2,000 words per week on Mac and Windows and 1,000 on iPhone.
| Market / platform | Price or cap | Versus US Pro annual ($12/mo) |
|---|---|---|
| US Pro, annual | $12/month ($144/year) | Baseline |
| India Pro, annual | ₹320/month (~$3.50) | About 71% less |
| India, stated target | ₹10–20/month (~$0.10–$0.20) | About 98–99% less |
| Android free tier | Unlimited words during launch | $0 |
| Mac / Windows free tier | 2,000 words per week | $0 |
Now line that up against what Canto is specifically good at. Canto handles code-switching “inside a single sentence” and returns Hinglish in romanized script. Kothari told TechCrunch that Indian users increasingly dictate in personal apps like WhatsApp, where they “frequently switch between Hindi and English while speaking.” Neil Shah, VP at Counterpoint Research, told the same reporter that “India is the ultimate stress test for voice AI,” citing linguistic, accent and contextual friction.
So: a market that is 14% of downloads and 2% of revenue, priced at a 71% discount heading toward a 98% discount, on a free tier with no word cap, whose defining speech pattern is the exact one the new model was built to handle. Wispr is not hiding the strategic interest — the India page and the Android page are both public. The part that is not public is whether the audio flowing back from that subsidy is what taught Canto, and whether users paying ₹320 understood that the standard-account default put them in the training set.
My read, stated as a read: a company does not price at 71% off, ship an uncapped free tier on the platform that market uses, hire a local go-to-market team, and then build a model whose flagship feature is that market's dialect, without the users and the model being connected. That is an inference from public facts, not a disclosure, and Wispr is entitled to rebut it with one.
The Honest Counter-Case: What Could Explain Canto Without User Audio
An argument is only worth reading if it survives its own best counter-argument, so here is the strongest version of the other side. Several of these are things a competent speech team would do regardless, and I have no evidence Wispr did not do them.
- Noise augmentation is standard and cheap. Mixing recorded wind, traffic, cafe babble and music into clean speech is a textbook technique. The “30% to 5–10%” improvement in noisy conditions is exactly what good augmentation plus a modern architecture delivers. This claim needs no user audio at all.
- Licensed and public corpora are real options. Mozilla Common Voice, LibriSpeech and commercial data vendors supply accented, multilingual speech under clear licences. A company with $361 million raised can simply buy data, and buying it is far less legally fraught than repurposing customer audio.
- Paid collection scales fine now. Vendors will record thousands of hours of Hinglish to spec. It is expensive, but $280 million buys a lot of specification.
- Open-weight starting points exist. “Proprietary” does not mean trained from scratch. Fine-tuning an existing multilingual foundation model on a comparatively small in-domain set is the normal path, and it shifts most of the data question upstream to whoever trained the base model.
- Personalisation may be inference-time, not training-time. This is the strongest counter to my dictionary point. “The model draws on your dictionary and the names of the people around you” can describe biasing decoding at inference against your local word list — a technique that touches your data on your request, per session, without that data ever entering a training run. If that is what Wispr means, my fourth row is wrong.
- Wispr did tighten its practices after being criticised. Following the 2025 backlash, the company changed settings and policy language and its CTO publicly acknowledged the handling had been poor. Enterprise and HIPAA defaults being training-off is a real protection that many competitors do not offer at all.
Where does that leave the case? Weaker on noise, materially weaker on personalisation, and essentially untouched on the consent question. Even if every byte of Canto's pre-training came from licensed corpora, the documented default still routes audio, transcripts and edits from free, trial and standard accounts into model improvement, and still exempts the enterprise accounts. That structure exists whether or not it is the thing that made Canto good.
Info
Nothing here alleges that Wispr broke a law or its own policy. The criticism is narrower and, I think, harder to dismiss: the disclosure that matters most is the least prominently published, and the users least protected by it are the ones with the least bargaining power.
What Wispr Has Not Said About Canto
Absences are evidence too, as long as you are precise about them. As of August 25, 2026, across the Series B announcement, the privacy policy, the Data Controls page and the security and compliance FAQ, I could not find answers to any of the following:
- Was customer audio used to train Canto specifically? The policies describe what may happen to data in general. Neither names Canto.
- Which datasets, licences or vendors contributed? No dataset is named anywhere in the announcement.
- How much of the corpus is customer-derived? No proportion is given.
- Were the free tier and paid tiers treated differently as training sources? The tier distinction Wispr documents is enterprise versus everyone else, not free versus paid.
- Was India-sourced audio used to build the Hinglish capability? Not addressed.
- What happens to a model already trained on your voice if you later opt out? Opting out stops future use. Model weights are not retroactively unlearned, and no policy claims they are.
- What changed in the August 18 and August 19 policy updates? No changelog is published.
That last point deserves its own sentence. Trained weights are not revocable. Every other privacy control on this list is prospective — you can stop the next recording from being used. You cannot remove your voice from a model that has already learned from it. This is the same structural problem at the centre of the Granola wiretap lawsuit, where the complaint alleges training-by-default is irreversible once done.
I emailed no one for this piece and it is not investigative reporting; it is a close reading of public documents. If Wispr publishes a training-data disclosure for Canto, I will update this page and say so.
The Consent Asymmetry, Stated Plainly
Strip out the speculation and one fact remains, fully documented by Wispr itself: the protection from model training is allocated by contract size, not by the sensitivity of what you are saying.
An enterprise buyer has a procurement team, a security questionnaire and a lawyer. They get Privacy Mode on by default. A freelancer on the free tier, a student on the trial, a Pro subscriber in Bengaluru paying ₹320 a month — none of them have a procurement team, and all of them get Privacy Mode off by default. The speech is often more sensitive at the small end, not less: the solo therapist, the immigration lawyer with one paralegal, the doctor in a two-person practice.
This is not unique to Wispr. It is close to an industry norm, and that is precisely why it deserves naming. The pattern is that consent defaults follow purchasing power. The users whose data is most useful for reaching new markets — accented, code-switched, recorded in noisy real-world conditions — are also the users with the least ability to negotiate the terms under which it is taken. A $280 million round raised partly on the strength of a model that handles those conditions is the clearest illustration of that trade I have seen this year.
Call it the subsidy-for-signal trade: the discount is real, the product is genuinely good, and the price of the discount is paid in training data by people who were never shown the invoice. Whether that is a fair deal is a judgement call, and reasonable people land differently on it. What is not a judgement call is whether the deal was clearly disclosed. For the broader argument about why this is a property of cloud architecture rather than of any one vendor, see cloud vs local dictation and why offline dictation matters.
Key Takeaway
Wispr's documented defaults allocate training protection by contract size, not by sensitivity of speech. Enterprise and HIPAA accounts are opted out; free, trial and standard accounts — including subsidised markets like India — are opted in.
Willow Shipped Its Own Model Too — and Said the Opposite Out Loud
Wispr is not the only dictation company that built its own speech model this year, which makes the comparison useful: it separates what is forced by the economics from what is a choice.
In July 2026, Willow Voice launched two models of its own — Frontier Pro and Frontier Mini — and then gave Mini away as free, unlimited dictation, describing it on its own pricing page as the “weaker speech-to-text model.” The pitch was explicitly a shot at the category's subscription norm: “Why are you paying for dictation? We're releasing free, unlimited AI dictation.” We covered the offer, the funnel behind it and the retention nuance in our review of whether Willow's free unlimited dictation is really free.
Same structural moment as Wispr: proprietary model, consumer tier priced at or near zero, real money expected from enterprise. The obvious question — is the free tier buying training data? — got two very different answers.
Willow answered it directly. Its launch announcement stated: “No, you are not becoming the product. You are not becoming the data.” Wispr, announcing Canto, said nothing about training data at all. Whatever you make of Willow's claim — and our own review flags that zero data retention appears as an Enterprise feature on Willow's pricing page, which does not obviously line up with a blanket free-tier promise — the point stands that a competitor facing the same incentive chose to address the question in public. Wispr's silence is a decision, not an industry necessity.
Where the two companies converge is the tier structure, and this is the part that supports the pattern rather than the exception. Per Willow's own plan matrix, privacy mode is listed on every Willow tier including Free, while zero data retention, SOC 2, SSO and admin privacy controls are Enterprise-only. Wispr's split runs the same direction with a sharper edge: the toggle exists on every tier, but the default points at training for everyone who is not on a contract.
| Wispr Flow | Willow Voice | |
|---|---|---|
| Own model shipped | Canto, August 2026 | Frontier Pro and Frontier Mini, July 2026 |
| Consumer tier | Free: 2,000 words/week (unlimited on Android during launch) | Free: unlimited on Frontier Mini, the weaker model |
| Training default, non-enterprise | On, per its security FAQ | Privacy mode listed on every tier, including Free |
| Zero data retention | Privacy Mode on plus Cloud Sync off | Enterprise plan only |
| Said anything about training data at model launch? | No | Yes — “you are not becoming the data” |
Two vendors, one category, one shared habit: the strongest guarantees sit behind a signed contract. That is the pattern worth naming, and it is bigger than either company. The difference between them is whether anyone told you.
If you are choosing between the two products rather than reading about their models, our Willow Voice vs Wispr Flow comparison puts the pricing, platforms and privacy defaults side by side.
How to Check Whether Your Voice Is in the Next Training Run
If you use Wispr Flow, here is the five-minute audit. Do all four steps — the first one alone is the mistake most people make.
- Open Settings > Data and Privacy in the Wispr Flow desktop app.
- Find “Improve the model for everyone.” This is the renamed Privacy Mode. If it is on, your audio, transcripts and edits can be used for training. Turn it off to opt out. On a trial or standard account, assume it is on until you have looked.
- Turn off Dictation Cloud Storage (Cloud Sync) as well. Opting out of training does not delete anything — Privacy Mode on plus Cloud Sync off is what Wispr calls Zero Data Retention. One toggle without the other leaves your transcripts and dictation history on their servers.
- Check Context Awareness in the same panel. Accessibility-text context is on by default and reads text from your active window; Screen OCR is off by default and captures a screenshot of the display containing your cursor. Decide on each one deliberately.
If you want the general version of this pattern — why zero retention began life as an enterprise contract term rather than a consumer feature, and the six clauses that quietly undo it — see our guide to zero data retention.
Two limits worth knowing. First, Wispr states that “Wispr may collect usage statistics such as the number of words you have dictated, regardless of your data controls” — some telemetry is not covered by any toggle. Second, “Transcription always occurs on the cloud”: even with everything turned off, your audio still leaves your machine to be transcribed. Privacy Mode governs what happens to it after it arrives, not whether it travels.
If you are on an Enterprise plan, your administrator sets this and you cannot override it — in that direction you are probably already protected. If you are on a free, trial or Pro account, nobody set it for you.
Tip
Opting out is prospective only. It stops future dictations from being used; it does not remove your voice from a model that has already trained on it. The earlier you check, the more it is worth.
If You'd Rather Not Be in Anyone's Training Set
The architectural answer to “whose voice trained this model” is to use a dictation app where the audio never leaves your machine in the first place. A toggle is a policy promise that can be renamed, re-defaulted or re-scoped; on-device processing is a property of the software that no policy update can reverse.
| Approach | Where audio is processed | Can a policy change put you in a training set? | Price |
|---|---|---|---|
| Wispr Flow | Always cloud | Yes — governed by a toggle and its default | Free tier; $15/mo or $144/yr Pro |
| Voibe, on-device mode | Your Mac (Apple Silicon) | No — audio never leaves the machine | $149 one-time |
| Voibe, private cloud mode | Zero-retention private cloud | No — nothing is retained to train on | $149 one-time |
Voibe — the app we build — runs on Mac and Windows. The Windows app, launched in July 2026, uses Voibe's private zero-retention cloud; the fully on-device mode is Mac-only (Apple Silicon). It is $149 one-time against Wispr Flow Pro's $144 per year, so it costs roughly one year of Wispr and then stops costing anything: over three years that is $149 versus $432, a saving of $283, or 65%. There is no training toggle in Voibe because in on-device mode there is nothing on our side to train on.
Voibe is not the only option, and for some readers it is not the right one — if you need iOS or Android, Wispr Flow remains the more complete cross-platform product and I would say so plainly. Our roundup of the best offline dictation apps covers the full on-device field including open-source choices, privacy-focused Wispr Flow alternatives is the direct swap list, and our Wispr Flow pricing breakdown has the full three-year cost maths if you are weighing the subscription on its own terms.
The Bottom Line
Wispr raised $280 million and shipped a model that is, by its own account, good at the messy real-world conditions clean corpora do not contain. It has not said what taught it that. Its own documentation says model training is on by default for trial and standard accounts and off by default for Enterprise and HIPAA customers, that the data covered includes “audio, transcript, edits,” and that the toggle governing it was renamed from “Privacy Mode” to “Improve the model for everyone” the day after Canto was announced.
From those facts, the most likely sources are the ones the user base makes cheapest to reach: free and trial accounts running on the default, the correction stream from every user who ever retyped a misheard word, and a heavily subsidised Indian market supplying exactly the code-switched, accented speech Canto is advertised as handling. Noise robustness could well be synthetic. Romanized Hinglish and your colleagues' names could not be.
I would use Flow. It is a good product and the people building it are not villains. I would also open Settings > Data and Privacy before I said anything into it I would not want in a training corpus — and I would notice that the company never quite told me I needed to.
Frequently Asked Questions
What is Canto, Wispr's new voice model?
How much did Wispr raise, and at what valuation?
Has Wispr said what data Canto was trained on?
Does Wispr Flow train on my voice by default?
Is model training opt-in or opt-out on Wispr Flow?
What exactly does Wispr use — just audio, or the text too?
Where is the setting that stops Wispr Flow training on my data?
Is turning off model training enough to keep my dictation private?
If I opt out now, is my voice removed from models already trained on it?
Why does Wispr Flow cost so much less in India?
Does the free tier get used for training more than paid tiers?
How do I use dictation without being in any training set?
Have other dictation companies said whether they train on user voice data?
Ready to type 5x faster?
Voibe is the fastest, most private dictation app for Mac and Windows. Try it today.
- On-device or private cloud
- Free to try
- No subscription
- Mac + Windows
- 90+ languages
Prefer to go Pro? Save 20% on any plan with code VOIBE20 View pricing →
Related Articles
Is Wispr Flow Safe? Seven Incidents and Two Toggles You Need to Know
Screenshots, a keystroke tap, a fake-audit scandal, and LinkedIn posts built from user dictations. Every Wispr Flow incident, and the two settings that help.
What Wispr Flow's Founder Revealed About User Tracking (2026)
On a Think School podcast, Wispr Flow's CEO described an analytics engine that ties your dictation — word counts, which apps, name, employer — to your identity. What it means.
Is Willow Voice Safe? Private Mode, HIPAA & Enterprise Verdict (2026)
Is Willow Voice safe? Private Mode default-on for individuals, opt-in training, HIPAA marketed but absent from policy text, SOC 2 referenced. Full safety review.

