DevNews

Wispr Canto: word errors and clean dictations differ

On this page
  1. What was announced, and what was expected
  2. The same word-error total can produce different experiences
  3. Test the real microphone and vocabulary

Wispr announced its $280 million Series B on August 17 alongside a preview of Canto. The useful distinction is between words recognized correctly and complete dictations that require no editing.

Original fictional corpus: each square is a ten-word dictation, 100 per group. Ten substitutions spread across ten dictations leave 90 clean; ten in one dictation leave 99 clean. Both have 10 errors per 1,000 words (1%).
Original fictional corpus: each square is a ten-word dictation, 100 per group. Ten substitutions spread across ten dictations leave 90 clean; ten in one dictation leave 99 clean. Both have 10 errors per 1,000 words (1%). Chart : PeopleAreGeek. Data source.
View full-size image

What was announced, and what was expected

The official announcement describes a preview, not an unrestricted release to every user. It reports a $2 billion valuation and says noisy-condition word error rates fall from over 30% to 5-10%. It expects 30-35% fewer dictations to need editing. That last claim concerns the number of affected dictations, not necessarily 30-35% fewer individual edit operations.

These are Wispr's figures and expectations. They do not specify the outcome for your microphone, accent, language mix or technical vocabulary. An announcement of model improvements should not be presented as an independent test of every Flow platform.

The same word-error total can produce different experiences

Consider an original fictional corpus of 100 dictations, each ten words long: 1,000 reference words in total. Suppose the only errors are ten substitutions. Both of the following patterns have a 1% word error rate:

Distribution of the ten errorsError-free dictations
One error in each of ten dictations90 out of 100
All ten errors in one dictation99 out of 100

The cover depicts the two distributions. It is an arithmetic example, not a Wispr result. Average word errors alone cannot recover the distribution of clean utterances without additional assumptions. Real zero-edit rate also depends on punctuation, formatting and what the user intended to send, not just a transcript's word matching.

One correction does not necessarily erase the entire time benefit of dictation either. Measure total completion time, including review and editing, instead of assuming that any intervention makes typing faster.

Test the real microphone and vocabulary

Use a fixed set of ordinary messages, names, technical terms and mixed-language sentences, with comparable recording conditions. Record which outputs need correction, the nature of the errors and the time to an acceptable final result. Review significant mistakes separately: one wrong number can matter more than several harmless punctuation changes.

The product changelog adds useful context after the announcement. August 21 enables choosing virtual microphones such as noise-suppression inputs; September 4 adds dictionary sharing at organization or department level for relevant business accounts. Record the selected input and dictionary configuration in comparisons, since those can change the result independently of the model. The changelog entries do not, by themselves, establish a universal Canto rollout date.

Correct Canto preview versus general shipment and expected reduction in dictations needing edits; compare word error and clean utterances without independence assumptions; add mic/dictionary updates.