Limited time: Save up to 33% on every planView pricing
Voibe Logovoibe Resources
wispr flowcantowispr cantowispr fundingseries bvoice modelspeech recognitionai training datavoice data privacyprivacy modehinglishindiacloud dictationon-deviceprivacywillow voicewillow frontier mini

Wispr Built Its Own Voice Model. Whose Voice Trained Canto?

Wispr raised $280M and shipped Canto — a model tuned for noise, accents and Hinglish. Their own docs hint at whose voices taught it. Not the Fortune 500.

Where Did Wispr Get the Voice Data to Train Canto?

Animated diagram of Wispr Flow's two default settings. Two account cards appear side by side. On the left, a card labelled Free, Trial and Standard accounts, with a Privacy Mode toggle that settles in the OFF position and a red chip reading training ON by default; a stream of audio, transcript and edits flows out of it into a bin labelled Canto training corpus. On the right, a card labelled Enterprise and HIPAA BAA, with a toggle that settles ON and a green chip reading training OFF by default; its stream is blocked by a green barrier and never reaches the bin. A counter fills to 60 billion words written with Flow. The closing stamp reads: the accounts that pay the most are the ones excluded from the corpus.
Two accounts, two opposite defaults — and only one of them ends up in the training set.

Wispr's new speech model is unusually good at exactly the things a clean training set cannot teach: wind, open-plan offices, heavy accents, and Hindi-English code-switching that has to come back in romanized script rather than Devanagari. Somebody's voice taught it that. The company's own documentation narrows down whose.

TL;DR: On August 17, 2026, Wispr announced a $280 million Series B at a $2 billion valuation led by Menlo Ventures, and previewed Canto, its first proprietary speech model. Wispr has not published a training-data disclosure for Canto, and I could not find one. What Wispr has published is a default: its security and compliance FAQ states that “Privacy Mode off (standard mode): audio and transcription data may be used to evaluate, train, and improve Wispr's models. This is the default for trial and standard accounts. Enterprise and HIPAA BAA customers run with Privacy Mode on by default.” That is an inverted consent structure — the Fortune 500 accounts are walled off from the corpus, and the free, trial and individual paying accounts are the corpus. More than 60 billion words have been dictated through Flow.

This piece is an analysis, not an accusation. I read the Series B post, the privacy policy, the Data Controls page and the security FAQ, and I lay out what each one actually says, what it does not say, and where the evidence stops and inference begins. The three most likely sources — default-on standard accounts, the correction signal Wispr calls “edits,” and a heavily subsidised Indian user base — are all consistent with what Wispr documents about itself. None of them are things Wispr has confirmed about Canto specifically.

Key Takeaway

Wispr has not disclosed what Canto was trained on. But its own security FAQ says model training is ON by default for trial and standard accounts and OFF by default for Enterprise and HIPAA customers — so whatever else went into Canto, the users least able to negotiate were the ones opted in.

Key Takeaways: Canto, the $280M Round, and the Data Question

Every figure below comes from Wispr's own pages or from named third-party reporting, retrieved August 25, 2026.

ItemDetail
AnnouncedAugust 17, 2026 (Wispr blog, Tanay Kothari)
Round$280 million Series B at a $2 billion valuation, led by Menlo Ventures
Total raised$361 million to date
New modelCanto — preview of Wispr's first proprietary speech model
Headline claimIn the hardest conditions, word error rate falls from more than 30% to between 5% and 10%; roughly 30–35% fewer dictations need editing
ScaleMore than 60 billion words written with Flow; used at almost all Fortune 500 companies and 10,000+ enterprises
Training-data disclosureNone published for Canto as of August 25, 2026 — the announcement does not name a single dataset, licence or vendor
Default for free/trial/standardModel training ON — “the default for trial and standard accounts” (security & compliance FAQ)
Default for Enterprise/HIPAAModel training OFF — “Privacy Mode on by default”
What gets used“audio, transcript, edits” (Data Controls)
Where the toggle livesSettings > Data and Privacy > “Improve the model for everyone” — renamed from “Privacy Mode”
India pricing₹320/month on annual billing versus $12/month in the US — about 71% less (TechCrunch)
India share14% of global downloads but 2% of in-app purchase revenue, October 2025 – April 2026
Android free tier“Free + unlimited during launch” (wisprflow.ai/android) — no word cap on the platform Wispr prioritised for India

Two of those rows do the real work. Training is on by default for the accounts that pay little or nothing, and off by default for the accounts that pay the most. Everything else in this article is an attempt to work out what that structure produced.

What Wispr Announced on August 17, 2026

Wispr Flow, the cloud AI dictation app that raised $280 million in August 2026 and previewed its own speech model, Canto
Wispr Flow: 60 billion words dictated, $361 million raised, and no published account of what taught its new model.

Wispr announced two things at once: money and a model. The Series B post by CEO Tanay Kothari confirms $280 million at a $2 billion valuation led by Menlo Ventures, with Notable Capital, NEA, Neo Ventures, 8VC and MVP Ventures returning, and Acrew, Activate, Forerunner, Goodwater, Peak XV, Together Fund and PLUS Capital joining. Total capital raised reaches $361 million. TechCrunch reported the round the same day.

The model is the more interesting half. Canto is Wispr's first proprietary speech model, shipped as a preview, and Kothari frames it against the industry's standard training practice in a passage worth quoting in full:

“Most speech models are trained and evaluated on clean recordings made in a quiet room with a good microphone and an accent the model has heard many times before. Almost nobody lives in those conditions. You're in a car, on a street, or sitting a foot away from someone else's conversation in an open office. We built this model for where people actually use Flow.

That last sentence is the whole question in seven words. “Where people actually use Flow” is not a dataset you can licence. It is a description of Flow's own users, in their own cars and offices and streets. The claimed result: in the hardest conditions — background noise, wind, heavy accents or music — error rates fall from more than 30% of words to between 5% and 10%, and across everyday use Wispr expects 30–35% fewer dictations to need editing. Those are Wispr's own internal numbers, not a third-party benchmark, and no evaluation set is named.

For the product itself rather than the model — features, plans, third-party ratings and the reliability record — see our Wispr Flow review.

Canto also targets code-switching, and the Hinglish example in the announcement is strikingly specific:

“Hindi speakers usually write in Devanagari, but when they're speaking Hinglish they expect it back in romanized script, so hearing every word correctly still isn't enough to give someone something they can send. The same is true of the vocabulary particular to your own life, which is why the model draws on your dictionary and the names of the people around you.”

Knowing that Hinglish speakers want romanized output is not a fact you extract from an audio corpus. It is a fact you learn by watching people reject the Devanagari you gave them. Hold that thought — it comes back below.

Canto's Own Description Is the Best Clue to Its Training Set

When a lab will not say what trained a model, the model's capabilities are the next best evidence. A speech model can only be good at conditions its training data contained, so Canto's feature list doubles as a partial description of its corpus.

Two of the four headline capabilities have a plausible non-user explanation:

  • Background noise, wind and music. Noise augmentation is a standard, decades-old technique: take clean speech, mix in recorded noise, train on the result. This capability genuinely can be manufactured, and any competent speech team would do exactly that.
  • Heavy accents. Partly addressable through public and licensed multi-accent corpora. Mozilla Common Voice and commercial data vendors both sell accented speech. It would not be cheap or complete, but it exists.

The other two do not have one:

  • Hinglish returned in romanized script. Getting the words right is a corpus problem. Knowing which script the user wants back is a preference problem, and preferences live in user behaviour, not in audio files. You learn this by shipping Devanagari to Hindi-English speakers and watching them fix it.
  • Your dictionary and the names of the people around you. Wispr states plainly that the model “draws on your dictionary and the names of the people around you.” No licensed dataset contains your colleagues' names. That capability is constituted by user data — there is no other way to build it.

So the honest reading is a split one: Canto's noise robustness could be synthetic; its personalisation and its code-switching behaviour almost certainly are not. The interesting question is not whether Wispr used user data — the dictionary feature means it did, by construction. The question is which users, and under what consent.

Source 1: The Default That Turns Standard Accounts Into a Corpus

The clearest answer Wispr gives is in its security and compliance FAQ, and it is worth reading twice:

“Privacy Mode off (standard mode): audio and transcription data may be used to evaluate, train, and improve Wispr's models. This is the default for trial and standard accounts. Enterprise and HIPAA BAA customers run with Privacy Mode on by default. If you handle confidential or privileged information, enable Privacy Mode before dictating.”

The same document confirms the enterprise side separately: “Enterprise defaults. Data sharing defaults to off (Privacy Mode effectively on).”

Laid out as a table, the structure is hard to miss:

Account typeTypical priceModel training by default
Basic (free)$0 — 2,000 words/week on Mac and WindowsOn
14-day trial$0, no card requiredOn
Pro (standard)$15/month, or $144/yearOn
EnterpriseQuoted; contractedOff
HIPAA BAA customersContractedOff

Read down the right-hand column. The protection tracks the contract, not the sensitivity of the speech. A solo therapist on Pro at $144/year dictating session notes is opted in by default; a Fortune 500 marketing team dictating press releases is opted out by default. Wispr's own privacy page says “You decide whether data can be used to train or improve our models” — true in the sense that a toggle exists, less true in the sense that most people never find it, and the page does not tell you which way it is pointing.

Scale turns the default into a corpus. Wispr says more than 60 billion words have been written with Flow. Even a modest fraction of that, from accounts that never opened Settings, is one of the larger real-world speech datasets in existence. For the wider pattern of what Flow collects beyond audio, see our breakdown of what Wispr Flow's founder revealed about user tracking and the full incident history in Is Wispr Flow safe?

Warning

Privacy Mode is not the same as zero retention. Turning it on stops training but leaves audio, transcripts and dictation history on Wispr's servers. Zero Data Retention requires Privacy Mode on AND Cloud Sync off — two separate toggles.

The Toggle That Changed Its Name (and Its Polarity)

The setting that controls all of this was renamed. Wispr's Data Controls page, last updated August 18, 2026 — the day after the Canto announcement — states: “This setting was previously called 'Privacy Mode.' If you had Privacy Mode enabled, your data is not shared for model training. This behavior remains unchanged.

The behaviour is unchanged. The framing is inverted. Compare what the two labels ask of you:

Old labelNew label
NamePrivacy ModeImprove the model for everyone
What ON meansYour data is protectedYour data is shared
What OFF meansYour data is sharedYour data is protected
What the name appeals toYour self-interestYour generosity

A toggle called “Privacy Mode” that is off invites the question “should I turn privacy on?” A toggle called “Improve the model for everyone” that is on invites the question “why would I turn that off?” Same switch, same wiring, opposite social pressure. Nothing about this is unlawful or even unusual — it is ordinary consent-interface design, of the kind regulators have started calling dark patterns when it runs one way and best practice when it runs the other.

What the Data Controls page does not say, anywhere, is which position the switch ships in. That fact lives only in a separate help-centre FAQ aimed at security reviewers. If you want to know whether you are currently in the training set, the marketing privacy page will not tell you, the Data Controls page will not tell you, and the privacy policy — last updated August 19, 2026, two days after the announcement — describes training only in opt-in terms: “If you opt to share your content with us for model training…

I want to be precise about what I can and cannot show here. I can show that these three pages describe the same setting at different levels of candour, and that two of them were updated in the 48 hours after Canto was announced. I cannot show what changed in those updates, because Wispr does not publish diffs, and I have no archived copies of the previous versions.

Wispr's own India landing page carries a pull-quote from Kothari: “Any company that wants to sit this close to people's work should give you upfront control of your privacy.” I agree with the sentence. “Upfront” is the word doing the work, and a default that is documented only in a compliance FAQ is not upfront.

Source 2: “Edits” Is the Most Valuable Word on the Data Controls Page

Wispr's Data Controls page specifies exactly what the training toggle governs: “your data (i.e. audio, transcript, edits) may be used to evaluate, train, or improve AI models, by Wispr.

Audio and transcript are the obvious pair. Edits are the valuable one, and I think it is the most underrated word in the whole disclosure.

Supervised speech training needs labelled data: an audio clip paired with the correct text. Producing that normally means paying human annotators to transcribe recordings, which is one of the largest costs in building a speech model. An edit collapses that cost to zero. When Flow transcribes “bishop” and you retype “Biswaroop,” you have just produced a perfectly labelled training example — the audio, the model's wrong answer, and the human-verified right answer — for free, in the course of doing your own work.

Wispr collects this at a second point too. Its Data Controls page describes an Auto-add to Dictionary feature in Settings > Personalization: “Flow monitors the text box where it pastes text to detect any edits you made to transcribed words. If you change the spelling of a word, it is automatically added to your dictionary.

Now put that next to the metric Wispr says it manages by. From the Series B post: “Internally, we measure ourselves against what we call zero edit rate, which is the share of everything you say that comes back right the first time and needs nothing from you at all.

The metric and the training signal are the same stream. Wispr scores itself on the dictations you do not edit and learns from the ones you do. That is a genuinely elegant flywheel, and it is also the mechanism by which the Hinglish insight in the announcement is most plausibly discovered: you ship Devanagari, Hindi-English speakers retype it in Latin script, and the correction stream tells you what the user wanted without anyone having to run a survey.

To be explicit: Wispr has not said this is how Canto learned romanized Hinglish. I am saying it is the mechanism their own documentation describes, applied to the example their own announcement chose.

Source 3: India, ₹320, and the Problem You Cannot Synthesise

India is where the commercial logic and the training logic point at the same users. Per TechCrunch's May 2026 report, India is Wispr Flow's second-largest market after the United States by both users and revenue, and the numbers underneath that are lopsided: India accounted for 14% of global downloads between October 2025 and April 2026 but only 2% of in-app purchase revenue. India supplies about seven times more share of usage than of revenue.

The pricing is deliberately subsidised. Wispr Pro in India is ₹320 per month on annual billing — roughly $3.50 — against $12 per month in the United States, a discount of about 71%. Kothari told TechCrunch the eventual target is ₹10–20 per month, or 10 to 20 US cents. On Android, the platform Wispr prioritised for India, the free tier is advertised as “Free + unlimited during launch” — no weekly word cap at all, against 2,000 words per week on Mac and Windows and 1,000 on iPhone.

Market / platformPrice or capVersus US Pro annual ($12/mo)
US Pro, annual$12/month ($144/year)Baseline
India Pro, annual₹320/month (~$3.50)About 71% less
India, stated target₹10–20/month (~$0.10–$0.20)About 98–99% less
Android free tierUnlimited words during launch$0
Mac / Windows free tier2,000 words per week$0

Now line that up against what Canto is specifically good at. Canto handles code-switching “inside a single sentence” and returns Hinglish in romanized script. Kothari told TechCrunch that Indian users increasingly dictate in personal apps like WhatsApp, where they “frequently switch between Hindi and English while speaking.” Neil Shah, VP at Counterpoint Research, told the same reporter that “India is the ultimate stress test for voice AI,” citing linguistic, accent and contextual friction.

So: a market that is 14% of downloads and 2% of revenue, priced at a 71% discount heading toward a 98% discount, on a free tier with no word cap, whose defining speech pattern is the exact one the new model was built to handle. Wispr is not hiding the strategic interest — the India page and the Android page are both public. The part that is not public is whether the audio flowing back from that subsidy is what taught Canto, and whether users paying ₹320 understood that the standard-account default put them in the training set.

My read, stated as a read: a company does not price at 71% off, ship an uncapped free tier on the platform that market uses, hire a local go-to-market team, and then build a model whose flagship feature is that market's dialect, without the users and the model being connected. That is an inference from public facts, not a disclosure, and Wispr is entitled to rebut it with one.

The Honest Counter-Case: What Could Explain Canto Without User Audio

An argument is only worth reading if it survives its own best counter-argument, so here is the strongest version of the other side. Several of these are things a competent speech team would do regardless, and I have no evidence Wispr did not do them.

  1. Noise augmentation is standard and cheap. Mixing recorded wind, traffic, cafe babble and music into clean speech is a textbook technique. The “30% to 5–10%” improvement in noisy conditions is exactly what good augmentation plus a modern architecture delivers. This claim needs no user audio at all.
  2. Licensed and public corpora are real options. Mozilla Common Voice, LibriSpeech and commercial data vendors supply accented, multilingual speech under clear licences. A company with $361 million raised can simply buy data, and buying it is far less legally fraught than repurposing customer audio.
  3. Paid collection scales fine now. Vendors will record thousands of hours of Hinglish to spec. It is expensive, but $280 million buys a lot of specification.
  4. Open-weight starting points exist. “Proprietary” does not mean trained from scratch. Fine-tuning an existing multilingual foundation model on a comparatively small in-domain set is the normal path, and it shifts most of the data question upstream to whoever trained the base model.
  5. Personalisation may be inference-time, not training-time. This is the strongest counter to my dictionary point. “The model draws on your dictionary and the names of the people around you” can describe biasing decoding at inference against your local word list — a technique that touches your data on your request, per session, without that data ever entering a training run. If that is what Wispr means, my fourth row is wrong.
  6. Wispr did tighten its practices after being criticised. Following the 2025 backlash, the company changed settings and policy language and its CTO publicly acknowledged the handling had been poor. Enterprise and HIPAA defaults being training-off is a real protection that many competitors do not offer at all.

Where does that leave the case? Weaker on noise, materially weaker on personalisation, and essentially untouched on the consent question. Even if every byte of Canto's pre-training came from licensed corpora, the documented default still routes audio, transcripts and edits from free, trial and standard accounts into model improvement, and still exempts the enterprise accounts. That structure exists whether or not it is the thing that made Canto good.

Info

Nothing here alleges that Wispr broke a law or its own policy. The criticism is narrower and, I think, harder to dismiss: the disclosure that matters most is the least prominently published, and the users least protected by it are the ones with the least bargaining power.

What Wispr Has Not Said About Canto

Absences are evidence too, as long as you are precise about them. As of August 25, 2026, across the Series B announcement, the privacy policy, the Data Controls page and the security and compliance FAQ, I could not find answers to any of the following:

  • Was customer audio used to train Canto specifically? The policies describe what may happen to data in general. Neither names Canto.
  • Which datasets, licences or vendors contributed? No dataset is named anywhere in the announcement.
  • How much of the corpus is customer-derived? No proportion is given.
  • Were the free tier and paid tiers treated differently as training sources? The tier distinction Wispr documents is enterprise versus everyone else, not free versus paid.
  • Was India-sourced audio used to build the Hinglish capability? Not addressed.
  • What happens to a model already trained on your voice if you later opt out? Opting out stops future use. Model weights are not retroactively unlearned, and no policy claims they are.
  • What changed in the August 18 and August 19 policy updates? No changelog is published.

That last point deserves its own sentence. Trained weights are not revocable. Every other privacy control on this list is prospective — you can stop the next recording from being used. You cannot remove your voice from a model that has already learned from it. This is the same structural problem at the centre of the Granola wiretap lawsuit, where the complaint alleges training-by-default is irreversible once done.

I emailed no one for this piece and it is not investigative reporting; it is a close reading of public documents. If Wispr publishes a training-data disclosure for Canto, I will update this page and say so.

Strip out the speculation and one fact remains, fully documented by Wispr itself: the protection from model training is allocated by contract size, not by the sensitivity of what you are saying.

An enterprise buyer has a procurement team, a security questionnaire and a lawyer. They get Privacy Mode on by default. A freelancer on the free tier, a student on the trial, a Pro subscriber in Bengaluru paying ₹320 a month — none of them have a procurement team, and all of them get Privacy Mode off by default. The speech is often more sensitive at the small end, not less: the solo therapist, the immigration lawyer with one paralegal, the doctor in a two-person practice.

This is not unique to Wispr. It is close to an industry norm, and that is precisely why it deserves naming. The pattern is that consent defaults follow purchasing power. The users whose data is most useful for reaching new markets — accented, code-switched, recorded in noisy real-world conditions — are also the users with the least ability to negotiate the terms under which it is taken. A $280 million round raised partly on the strength of a model that handles those conditions is the clearest illustration of that trade I have seen this year.

Call it the subsidy-for-signal trade: the discount is real, the product is genuinely good, and the price of the discount is paid in training data by people who were never shown the invoice. Whether that is a fair deal is a judgement call, and reasonable people land differently on it. What is not a judgement call is whether the deal was clearly disclosed. For the broader argument about why this is a property of cloud architecture rather than of any one vendor, see cloud vs local dictation and why offline dictation matters.

Key Takeaway

Wispr's documented defaults allocate training protection by contract size, not by sensitivity of speech. Enterprise and HIPAA accounts are opted out; free, trial and standard accounts — including subsidised markets like India — are opted in.

Willow Shipped Its Own Model Too — and Said the Opposite Out Loud

Wispr is not the only dictation company that built its own speech model this year, which makes the comparison useful: it separates what is forced by the economics from what is a choice.

In July 2026, Willow Voice launched two models of its own — Frontier Pro and Frontier Mini — and then gave Mini away as free, unlimited dictation, describing it on its own pricing page as the “weaker speech-to-text model.” The pitch was explicitly a shot at the category's subscription norm: “Why are you paying for dictation? We're releasing free, unlimited AI dictation.” We covered the offer, the funnel behind it and the retention nuance in our review of whether Willow's free unlimited dictation is really free.

Same structural moment as Wispr: proprietary model, consumer tier priced at or near zero, real money expected from enterprise. The obvious question — is the free tier buying training data? — got two very different answers.

Willow answered it directly. Its launch announcement stated: “No, you are not becoming the product. You are not becoming the data.” Wispr, announcing Canto, said nothing about training data at all. Whatever you make of Willow's claim — and our own review flags that zero data retention appears as an Enterprise feature on Willow's pricing page, which does not obviously line up with a blanket free-tier promise — the point stands that a competitor facing the same incentive chose to address the question in public. Wispr's silence is a decision, not an industry necessity.

Where the two companies converge is the tier structure, and this is the part that supports the pattern rather than the exception. Per Willow's own plan matrix, privacy mode is listed on every Willow tier including Free, while zero data retention, SOC 2, SSO and admin privacy controls are Enterprise-only. Wispr's split runs the same direction with a sharper edge: the toggle exists on every tier, but the default points at training for everyone who is not on a contract.

Wispr FlowWillow Voice
Own model shippedCanto, August 2026Frontier Pro and Frontier Mini, July 2026
Consumer tierFree: 2,000 words/week (unlimited on Android during launch)Free: unlimited on Frontier Mini, the weaker model
Training default, non-enterpriseOn, per its security FAQPrivacy mode listed on every tier, including Free
Zero data retentionPrivacy Mode on plus Cloud Sync offEnterprise plan only
Said anything about training data at model launch?NoYes — “you are not becoming the data”

Two vendors, one category, one shared habit: the strongest guarantees sit behind a signed contract. That is the pattern worth naming, and it is bigger than either company. The difference between them is whether anyone told you.

If you are choosing between the two products rather than reading about their models, our Willow Voice vs Wispr Flow comparison puts the pricing, platforms and privacy defaults side by side.

How to Check Whether Your Voice Is in the Next Training Run

If you use Wispr Flow, here is the five-minute audit. Do all four steps — the first one alone is the mistake most people make.

  1. Open Settings > Data and Privacy in the Wispr Flow desktop app.
  2. Find “Improve the model for everyone.” This is the renamed Privacy Mode. If it is on, your audio, transcripts and edits can be used for training. Turn it off to opt out. On a trial or standard account, assume it is on until you have looked.
  3. Turn off Dictation Cloud Storage (Cloud Sync) as well. Opting out of training does not delete anything — Privacy Mode on plus Cloud Sync off is what Wispr calls Zero Data Retention. One toggle without the other leaves your transcripts and dictation history on their servers.
  4. Check Context Awareness in the same panel. Accessibility-text context is on by default and reads text from your active window; Screen OCR is off by default and captures a screenshot of the display containing your cursor. Decide on each one deliberately.

If you want the general version of this pattern — why zero retention began life as an enterprise contract term rather than a consumer feature, and the six clauses that quietly undo it — see our guide to zero data retention.

Two limits worth knowing. First, Wispr states that “Wispr may collect usage statistics such as the number of words you have dictated, regardless of your data controls” — some telemetry is not covered by any toggle. Second, “Transcription always occurs on the cloud”: even with everything turned off, your audio still leaves your machine to be transcribed. Privacy Mode governs what happens to it after it arrives, not whether it travels.

If you are on an Enterprise plan, your administrator sets this and you cannot override it — in that direction you are probably already protected. If you are on a free, trial or Pro account, nobody set it for you.

Tip

Opting out is prospective only. It stops future dictations from being used; it does not remove your voice from a model that has already trained on it. The earlier you check, the more it is worth.

If You'd Rather Not Be in Anyone's Training Set

The architectural answer to “whose voice trained this model” is to use a dictation app where the audio never leaves your machine in the first place. A toggle is a policy promise that can be renamed, re-defaulted or re-scoped; on-device processing is a property of the software that no policy update can reverse.

ApproachWhere audio is processedCan a policy change put you in a training set?Price
Wispr FlowAlways cloudYes — governed by a toggle and its defaultFree tier; $15/mo or $144/yr Pro
Voibe, on-device modeYour Mac (Apple Silicon)No — audio never leaves the machine$149 one-time
Voibe, private cloud modeZero-retention private cloudNo — nothing is retained to train on$149 one-time

Voibe — the app we build — runs on Mac and Windows. The Windows app, launched in July 2026, uses Voibe's private zero-retention cloud; the fully on-device mode is Mac-only (Apple Silicon). It is $149 one-time against Wispr Flow Pro's $144 per year, so it costs roughly one year of Wispr and then stops costing anything: over three years that is $149 versus $432, a saving of $283, or 65%. There is no training toggle in Voibe because in on-device mode there is nothing on our side to train on.

Voibe is not the only option, and for some readers it is not the right one — if you need iOS or Android, Wispr Flow remains the more complete cross-platform product and I would say so plainly. Our roundup of the best offline dictation apps covers the full on-device field including open-source choices, privacy-focused Wispr Flow alternatives is the direct swap list, and our Wispr Flow pricing breakdown has the full three-year cost maths if you are weighing the subscription on its own terms.

The Bottom Line

Wispr raised $280 million and shipped a model that is, by its own account, good at the messy real-world conditions clean corpora do not contain. It has not said what taught it that. Its own documentation says model training is on by default for trial and standard accounts and off by default for Enterprise and HIPAA customers, that the data covered includes “audio, transcript, edits,” and that the toggle governing it was renamed from “Privacy Mode” to “Improve the model for everyone” the day after Canto was announced.

From those facts, the most likely sources are the ones the user base makes cheapest to reach: free and trial accounts running on the default, the correction stream from every user who ever retyped a misheard word, and a heavily subsidised Indian market supplying exactly the code-switched, accented speech Canto is advertised as handling. Noise robustness could well be synthetic. Romanized Hinglish and your colleagues' names could not be.

I would use Flow. It is a good product and the people building it are not villains. I would also open Settings > Data and Privacy before I said anything into it I would not want in a training corpus — and I would notice that the company never quite told me I needed to.

Frequently Asked Questions

What is Canto, Wispr's new voice model?

Canto is Wispr's first proprietary speech recognition model, previewed on August 17, 2026 alongside the company's $280 million Series B. Wispr says it was built for real-world conditions rather than clean recordings, and claims that in the hardest conditions — background noise, wind, heavy accents or music — word error rates fall from more than 30% to between 5% and 10%, with roughly 30–35% fewer dictations needing edits across everyday use. These are Wispr's own internal figures; no third-party benchmark has been published.

How much did Wispr raise, and at what valuation?

Wispr raised $280 million in a Series B at a $2 billion valuation, announced August 17, 2026 and led by Menlo Ventures. Existing investors Notable Capital, NEA, Neo Ventures, 8VC and MVP Ventures returned; Acrew, Activate, Forerunner, Goodwater, Peak XV, Together Fund and PLUS Capital joined. Total capital raised is $361 million.

Has Wispr said what data Canto was trained on?

No. As of August 25, 2026, Wispr has not published a training-data disclosure for Canto. The Series B announcement does not name a single dataset, licence or data vendor, and neither the privacy policy nor the Data Controls page mentions Canto by name. What Wispr does document is that audio, transcripts and edits from accounts without Privacy Mode enabled may be used to evaluate, train and improve its models.

Does Wispr Flow train on my voice by default?

If you are on a free, trial or standard (Pro) account, yes, unless you have changed the setting. Wispr's security and compliance FAQ states: “Privacy Mode off (standard mode): audio and transcription data may be used to evaluate, train, and improve Wispr's models. This is the default for trial and standard accounts. Enterprise and HIPAA BAA customers run with Privacy Mode on by default.”

Is model training opt-in or opt-out on Wispr Flow?

It depends on your account type, which is why sources conflict. Wispr's privacy policy describes training in opt-in language (“If you opt to share your content with us for model training…”), and that matches the Enterprise and HIPAA BAA experience, where data sharing defaults to off. For trial and standard accounts, the security FAQ is explicit that training is the default — which is opt-out in practice. Check the toggle rather than relying on either description.

What exactly does Wispr use — just audio, or the text too?

Per Wispr's Data Controls page, the data covered is “audio, transcript, edits.” Edits are the corrections you type when Flow gets a word wrong, which makes them labelled training examples. Separately, the Auto-add to Dictionary feature monitors the text box Flow pastes into to detect spelling changes you make. Wispr also states it “may collect usage statistics such as the number of words you have dictated, regardless of your data controls.”

Where is the setting that stops Wispr Flow training on my data?

Settings > Data and Privacy, under “Improve the model for everyone.” Turn it off to opt out. This setting was previously called “Privacy Mode” — Wispr's Data Controls page confirms the rename and states the behaviour is unchanged, but note the polarity flipped: Privacy Mode ON meant protected, while “Improve the model for everyone” ON means shared.

Is turning off model training enough to keep my dictation private?

No. Opting out of training stops your data being used to improve models, but it does not stop storage. Wispr's Zero Data Retention state requires two settings: Privacy Mode on and Cloud Sync (Dictation Cloud Storage) off. With only the first, your audio, transcripts and dictation history remain on Wispr's servers. Transcription also always occurs in the cloud, so your audio leaves your machine regardless.

If I opt out now, is my voice removed from models already trained on it?

No. Opting out is prospective: it stops future dictations from being used. Trained model weights are not retroactively unlearned, and no Wispr policy claims otherwise. This is why the timing of the default matters more than the existence of the toggle.

Why does Wispr Flow cost so much less in India?

Wispr prices Pro in India at ₹320 per month on annual billing — roughly $3.50 — against $12 per month in the US, a discount of about 71%, with a stated eventual target of ₹10–20 per month. Per TechCrunch, India is Wispr's second-largest market and accounted for 14% of global downloads but only 2% of in-app purchase revenue between October 2025 and April 2026. Wispr presents this as market expansion. It also happens to source exactly the accented, code-switched speech Canto is built to handle.

Does the free tier get used for training more than paid tiers?

Wispr does not document a free-versus-paid distinction. The distinction it does document is enterprise versus everyone else: trial and standard accounts default to training on, while Enterprise and HIPAA BAA accounts default to training off. A paying Pro subscriber is in the same default bucket as a free user. On Android, the free tier is advertised as unlimited during launch, meaning no word cap on the platform Wispr prioritised for India.

How do I use dictation without being in any training set?

Use an app that transcribes on your own device, so there is no cloud copy to train on. Voibe runs on Mac and Windows — the Windows app uses a private zero-retention cloud, and the fully on-device mode is Mac-only (Apple Silicon) — for $149 one-time against Wispr Flow Pro's $144 per year, a saving of $283 over three years. MacWhisper, VoiceInk and Handy are other on-device options. The general rule: a toggle is a promise that can be re-defaulted, while on-device processing is a property of the software.

Have other dictation companies said whether they train on user voice data?

Willow Voice did. Launching its own Frontier Pro and Frontier Mini models in July 2026 alongside free unlimited dictation, Willow stated publicly: “No, you are not becoming the product. You are not becoming the data.” Wispr, announcing Canto in August 2026, made no statement about training data. Both companies reserve zero data retention for enterprise plans — on Willow's plan matrix it is Enterprise-only, and on Wispr it requires Privacy Mode on plus Cloud Sync off.

Ready to type 5x faster?

Voibe is the fastest, most private dictation app for Mac and Windows. Try it today.

  • On-device or private cloud
  • Free to try
  • No subscription
  • Mac + Windows
  • 90+ languages

Prefer to go Pro? Save 20% on any plan with code VOIBE20 View pricing →