Wispr raises $280M to push AI voice technology beyond dictation
The funding round was led by Menlo Ventures, an early investor known for backing technology companies such as Uber, Roku, and Anthropic. Existing investors Notable Capital, NEA, Neo Ventures, 8VC, and MVP Ventures also participated.
US-based artificial intelligence startup Wispr has raised $280 million in a Series B funding round at a $2 billion valuation, as the company behind the voice-based writing platform Flow seeks to improve speech recognition and expand its broader human and AI interaction technology.
The funding round was led by Menlo Ventures, an early investor known for backing technology companies such as Uber, Roku, and Anthropic. Existing investors Notable Capital, NEA, Neo Ventures, 8VC, and MVP Ventures also participated, alongside new backers Acrew, Activate, Forerunner, Goodwater, Peak XV, Together Fund and PLUS Capital.
The latest round brings Wispr’s total funding to $361 million.
Founded to make voice a practical alternative to typing, Wispr develops AI tools that convert spoken language into written text. Its flagship product, Flow, is designed to help users write emails, messages and documents by speaking naturally. The company recently introduced Notetaker, a meeting transcription tool that automatically captures conversations and generates notes.
The newly-raised capital will be directed largely towards research and development, particularly in improving speech recognition accuracy and expanding the deployment of Wispr’s technology across the places where people already communicate and work.
Chief executive and co-founder Tanay Kothari said accuracy would determine whether voice technology becomes a genuine replacement for typing rather than remaining a novelty. He noted that voice systems fail when users must repeatedly stop to correct mistakes, because interruptions break concentration and reduce trust in the technology.
Alongside the funding announcement, Wispr unveiled a preview of Canto, its first proprietary speech model. Speech models are AI systems trained to recognise spoken language and convert it into text. According to the company, many existing systems are evaluated using clean recordings produced in quiet environments, despite the fact that most users interact with voice tools in cars, on busy streets or in open offices. Wispr said Canto was developed specifically for those real-world conditions.
The company claims that, in particularly noisy environments, the model can reduce word error rates from more than 30% to between 5 and 10%. In everyday use, it expects the number of voice-generated texts requiring manual editing to fall by 30 to 35%.
One of Canto’s distinguishing features is its ability to handle multilingual speech. This is especially relevant in countries such as India, where many speakers routinely switch between languages within a single conversation, a behaviour known as code-switching. Wispr used Hinglish, a blend of Hindi and English, to illustrate how traditional speech systems often struggle to recognise mixed languages and convert them into the correct written form. The company said its model also incorporates personal dictionaries and frequently used names to improve accuracy.
The emphasis on multilingual AI reflects a wider industry trend. India’s AI ecosystem has increasingly prioritised language technology through initiatives under the IndiaAI Mission, which aims to strengthen domestic AI infrastructure, datasets and language models. Government support for regional language computing has also encouraged investment in speech and conversational AI technologies. Recent funding activity in India includes fresh investments in Sarvam and Gnani, while enterprise automation platforms continue to attract capital as businesses seek more efficient customer service and workflow tools.
Globally, competition in AI voice technology has intensified. Companies such as ElevenLabs, Deepgram, and Speechmatics are developing increasingly sophisticated systems for speech synthesis, transcription and conversational AI. Advances in large language models have accelerated the shift from simple voice commands towards AI assistants capable of understanding context and maintaining more natural interactions.
Kothari said the conversation around voice technology has changed significantly over the past year. Instead of questioning whether voice tools work, users now focus on how much time they save and what new capabilities they want.
Wispr said users have already written more than 60 billion words through Flow, while the platform is used at almost all Fortune 500 companies and by more than 10,000 enterprises.
Edited by Megha Reddy


