Skip to content

This Is What Running a $2B AI Startup Looks Like

Tanay Kothari shares the journey of WhisperFlow, a $2B AI startup revolutionizing voice interfaces and competing with tech giants like Google.

Ask about this video. Answers come from its transcript only — with the timestamp, so you can check them.

Generated from the transcript and can be wrong — check the timestamp.

Key Takeaways

  • Voice will become the primary interface for human-computer interaction, replacing typing.
  • Integrating voice technology seamlessly across applications is critical for user adoption.
  • Building AI startups requires balancing innovation with competitive threats from tech giants.
  • User feedback is essential to evolve products, such as adding meeting recording and summarization features.
  • A collaborative and open company culture fosters innovation and high performance.

What the video covers

  • Tanay Kothari founded WhisperFlow five years ago, initially as a hardware company focused on voice interfaces.
  • WhisperFlow replaces traditional keyboards by enabling voice-to-text input across all applications and devices.
  • The company culture is warm, friendly, and intense, emphasizing collaboration across teams without echo chambers.
  • WhisperFlow developed a wearable brain-computer interface that detects silent speech using electromyography sensors.
  • After three years of hardware development, they built Flow, a standalone voice operating system product.
  • The product integrates with tools like Gmail and Slack to provide live transcription, note-taking, and meeting summaries.
  • WhisperFlow faces competition from major players like Google and Apple, highlighting the shift from typing to voice interaction.
  • The company emphasizes speed, accuracy, and natural formatting in its voice recognition technology.
  • Tanay discusses the challenges and opportunities in building AI software businesses in 2026.
  • The video includes behind-the-scenes insights into product development, office culture, and future plans.

Answers

Questions about this video

What is WhisperFlow and what problem does it solve?

WhisperFlow is an AI startup that replaces traditional typing with voice-to-text technology, enabling users to interact with any application through voice, improving speed and accuracy.

How does WhisperFlow’s wearable brain-computer interface work?

The device uses electromyography sensors to detect tiny electrical impulses from muscles involved in speech, allowing users to silently speak to themselves and have their thoughts transcribed.

How does WhisperFlow compete with major tech companies like Google?

WhisperFlow focuses on delivering a seamless, fast, and accurate voice interface product that integrates across all applications and devices, positioning itself as a leader in the emerging voice-first computing space.

Full Transcript — Download SRT & Markdown

00:00
Speaker A
Google is coming for us. My name is Tanay. Five years ago, I founded WhisperFlow, which initially started off as a hardware company, but is today replacing millions of people's keyboards with their voice. But this wasn't my first voice app. When I was 11, me and a
00:15
Speaker A
friend of mine built one of the world's first voice assistants with 2 and 1/2 million users, but Google wasn't a big fan and they sent us a cease and desist.
00:23
Speaker A
We had to shut that thing down. 16 years later, we're back. I'm running a $2 billion company and Google has just joined the same race as WhisperFlow. So yeah, things are starting to get a bit more interesting.
00:41
Speaker A
Hey guys, welcome home. This is the Whisper office. We just moved to this space about a month ago. We used to be a hardware company and so we've had this office but the floor below for the last few years. Then we just moved upstairs.
00:53
Speaker A
We have everything in the pantry, so eat, drink, whatever you want and then just uh relatively disorganized area for people to hang out with. Usually during lunchtime, we all eat lunch here together. And then in the evening, if
01:07
Speaker A
you come here about like after 5:00, people are just going to be lying around doing their thing. I don't know, it just makes the office feel a little more like home. The culture at Whisper is very warm, friendly, uh but also intense. If
01:19
Speaker A
you're here, you're part of a very high-power team that has each other's back, but also we push each other to do our best work. Throughout the office all day, you'll just hear people laughing.
01:29
Speaker A
We don't take ourselves too seriously. It's just fun to build this company, to build a product. That's generally the the vibe of the office. So this is the whole open office area where everybody's working. The one thing that I really
01:40
Speaker A
like that we said from day zero was we generally have the whole teams kind of spread out. Sales people and marketing folks and engineers and designers all just intermingle together cuz you don't want to create just echo chambers within
01:51
Speaker A
the company. The biggest room of all, still being set up, is going to be our giant boardroom. This room we actually had to break down this wall to extend space more because we ran out of space for the all hands in the middle of the
02:04
Speaker A
office being constructed. And so now we have a lot more space. We're going to extend this out. And so this just becomes a fun area where the team spends time together at least once a week.
02:13
Speaker A
Whisper Flow, right now it's a really simple product. You speak [music] and it writes for you perfectly in every single application without needing to do single other thing. And say you're in Gmail and you want to write an email. All you have
02:25
Speaker A
to do is you're going to press this button, speak, let go, and text will show up perfectly formatted. So, Hey Sage, there's three things Whisper does incredibly well. It's insanely fast, it's insanely accurate, and it gets names right. And that got eight
02:40
Speaker A
perfectly. [music] It gets his name right and also formats things like a nice bulleted list. It has a little colon here. It's all the small things that really matter for people. It works in any single application. So if
02:52
Speaker A
I'm then going and messaging my team, "Hey guys, fantastic work." Fire emoji, it'll do that just as well. This works across all your applications, across all your devices. And that is what operating systems always like to get into. That is
03:06
Speaker A
what browsers always like to get into. [music] Whisper is attracting these like large competitors like Google, like Apple because I think that they realize that this is the future, you [music] know?
03:18
Speaker A
Like typing is going to be seen as like a very old like mechanism of interacting with machines, I think very soon.
03:28
Speaker A
[music] When we started the company, we believed voice would be the primary interface over time. And we wanted to solve the problem of, "Hey, how are we going to use voice when we're around other people?" And so we started the company
03:40
Speaker A
building a wearable brain computer interface. All right, cool. So these are some of the research prototypes of the devices that we built.
03:48
Speaker A
It's a [music] headset that you can wear. It can understand when you're speaking silently to yourself. The way it worked was using these sensors called electromyography sensors. So, these pick up really tiny electrical impulses that your brain sends to the muscles that
04:01
Speaker A
produce speech, and we can pick up those impulses and figure out what you're thinking about saying, not just what you're actually saying out loud. Early gen one type of prototype of how might people use it, how do we get it to work
04:12
Speaker A
when it's big and bulky, and um I mean, the first version of it was a bunch of taped to my face. Kind of got down from there into this like really cool prototype. The goal over time was to
04:22
Speaker A
miniaturize it so it would be like an earbud that you can wear over here, and you could use it literally anywhere.
04:27
Speaker A
So, we were working on this this hardware device, and we were building it for 3 years to really be able to speak anywhere. So, it actually allowed for for silent speech. You could just think about something and it would write for
04:39
Speaker A
you. Took us 3 years, we built it, then we tried to connect it with ChatGPT, Siri, Alexa, and turns out they all just sucked. We decided we needed to build a product that goes from your mental RAM into something that was structured and
04:52
Speaker A
ready to send, and that we built as Flow, which is the operating system for the hardware device and now is a standalone product by itself. And so, the thing that everybody asks for is, [music] "Hey Tanay, why can't I use it to record
05:03
Speaker A
my meetings? Cuz I'm meeting with somebody and I just want to like record all the notes that I'm saying, or I'm having a Zoom meeting and I just wanted it to record [music] everything. The accuracy is so good, it works
05:13
Speaker A
everywhere, it's so non-invasive. I would just rather use this than any other product." At some point, we just realized we're being stupid. Like, people really want was for us to do this, and so we listened, and now you
05:23
Speaker A
have a note-taker [music] where I can just click this button, and it knows that, "Hey, we have this this meeting right now where you and I are chatting." It says that, "Okay, Alex wants to keep it casual, so like I'm not I'm not being
05:35
Speaker A
super formal here." This is really meta, actually, cuz you get to actually see like the process behind the process.
05:40
Speaker A
Whisper also tells me like, "Hey, this [music] is what the meeting's about, this is what you need to know to be really prepared for it." It is capturing every single word I'm saying live and it's perfectly formatted. I can even
05:51
Speaker A
[music] talk to it. So, if I say like, "What did I just say?" It is going to just tell me that and so with all the context here, with all the context across like my Slack and Gmail, it's
06:02
Speaker A
able to pull all that in. This just makes it so that within a meeting, I don't have to focus on taking all the notes cuz after the meeting, it generates the the summary as well.
06:19
Speaker A
Yeah, I do bounce. I'm going to Menlo Ventures for a thing. This is now starting to feel like the behind the diary kind of Yeah.
06:27
Speaker A
[laughter] No, literally. [music] Uh I'm Tanay. [music] I'm the CEO and co-founder of WhisperFlow.
06:41
Speaker A
My audience always loves it whenever I use Whisper. Like that's like the tool I use every day in whatever I'm doing. I'm prompting, I'm asking anything, I'm typing my emails, it's all Whisper. You spoiled me at this point. Like my
06:54
Speaker A
fingers just don't want to do anything. Yeah. [laughter] People who are watching this, they want to learn about AI, they want to start a business in AI. We have a investor who represents Menlo and has invested into a
07:07
Speaker A
lot of AI companies. You who's built a massive AI company and we have the builder, Davia, who's building an AI company as well. Where are we today in the world of AI in 2026?
07:18
Speaker A
The way businesses were built 10 years ago is fundamentally very different from now. Specifically, software businesses.
07:24
Speaker A
Before, most software businesses found a moat in how difficult [music] it was to build software. It was expensive, you had to hire people, you had to go and build things [music] and, you know, maybe that's not true for some of the
07:37
Speaker A
most exceptional companies from that era, you know, quality of Google's and Facebook's. The fundamental value of software came [music] from how difficult it was to do and how capital intensive it was because very few people had the skill to do that. Fundamentally, when
07:49
Speaker A
you look grapple with what's happening with AI and intelligence today, [music] the cost of building software has come down rapidly. What are the different bottlenecks? What are the challenges?
07:58
Speaker A
[music] The ability to distribute and acquire users is immensely challenging. The ability to do software at the very pinnacle [music] of where the training data doesn't exist is is a mode. It's still quite challenging.
08:12
Speaker A
And And number three, the ability to [music] convince anybody to even join your company is extremely extremely hard today [music] in a world where, you know, being a founder is is no longer a niche high-risk activity. [music] It's
08:27
Speaker A
something everyone wants to do. Tanay is truly like one of the smartest people I know. Tanay is a IOI silver medalist.
08:35
Speaker A
People for it's the International Olympiad for Informatics. He's one of the smartest people I would say in his grade if not beyond that in all of India. So, just to contextualize that for people like there's this misconception that's, you know, people
08:47
Speaker A
sort of come to America cuz it's the easy way out and maybe in some cases [music] that's true.
08:51
Speaker A
Um that's definitely not the case at least for for Tanay. So, just to contextualize for the audience.
08:57
Speaker A
[laughter] [music] All of these are files and you can see the users flying around and every time they do a little zap, it's a change.
09:10
Speaker A
Like a file change, like a deletion addition. And it's actually really sick. If you can see, we're in August of 2025.
09:16
Speaker A
This is actually when I joined. In the last couple months, this thing is going crazy. There's laser zaps everywhere. It's like it's a flurry of activity. It's really cool to see because that's what it feels like in the last couple of months as
09:29
Speaker A
we've gotten more and more smart people working here. Yep. Now it's starting to blow up.
09:36
Speaker A
It goes crazy. See, this makes me want to code more. I don't see my name up there anymore.
09:41
Speaker A
I would make every single PR in every single different folder just so you see my name bounce around.
09:46
Speaker A
We can I'm only going to put a When I have [laughter] Look, we can do a focus mode that follows an individual user. And so you instead of looking at the whole thing, you're following them around.
09:56
Speaker A
Mhm. So we'll do that. We'll spotlight you. You have to put a token in though.
10:01
Speaker A
Uh Mhm. You have to put a token credit card. Yeah, credit card. It goes to me.
10:05
Speaker A
[laughter] You can use a ramp card if you want. How would you guys describe WhisperFlow?
10:12
Speaker A
Um Yeah, like that. Why I came here was because of the feeling that I felt when I first started using it. My ideas come out like [music] as I'm speaking like this. And to be able to speak like this and get all
10:25
Speaker A
those ideas out brought my [music] ability to like problem-solve to a different level because I'm just like creating more information that I would.
10:33
Speaker A
I'm like forming more detailed accounts of my thoughts that I can then bring to life because the information can all be organized and executed on and yeah, it like totally [music] changed my my workflow. And I would say that's what it
10:46
Speaker A
is. It's just like oh, WhisperFlow is oh, I can speak normally to my computer like I'm speaking right now. What do you think about that?
10:54
Speaker A
So why is Google a threat to WhisperFlow? I don't really try to think of it that way so much as there are lots of big incumbents who also see that voice is going to be a big part of the future.
11:05
Speaker A
Who also want to build the interface that becomes the primary consumer technology that people use all the time to interface their devices. Google is one of those companies that would care about that kind of a problem.
11:16
Speaker A
Whisper right now has such a significant moat because of how much [music] effort we've invested into like those real-world examples. A lot of like ASR like voice models are benchmarked off of like one person speaking in a silent
11:29
Speaker A
room. Um but Whisper is [music] built for like actual real-world use-cases and optimizing around that is like not something you can do overnight.
11:40
Speaker A
We are building in a space where if we build it right, we have the opportunity to be a company that's potentially of that size and scale. That also means that companies of that size and scale are a big threat to us. I like to run
11:53
Speaker A
towards fear, not away from it, if I have fear. And now, over to Note Taker. So, we did a presentation about 4 weeks ago where we were talking about, "Hey, why are we building this new product? Why are we
12:06
Speaker A
prioritizing it? How are we starting to think about it?" And a lot has started to take shape. So, what I wanted to talk about today was what the next month looks like for Note Taker. Showcase a few different things on like what we
12:19
Speaker A
have built so far, what to expect, and spend some time actually diving into some design decisions. There's some people here who've been deeply involved with this, so you probably know all the all the nooks and crannies, but for a
12:30
Speaker A
lot more people who've been a little less involved, I want this to be the place where we can go and answer all the questions. And so, I'm going to go through the stuff I talked about fast, and then spend a majority of it on any
12:42
Speaker A
and all Q&A that comes up. So, so no dumb questions. There's some dumb questions, but mostly no dumb questions.
12:49
Speaker A
So, in terms of the timeline, this is a lot more finalized now. Marketing launch, this is big, splashy GA all over Twitter, LinkedIn, social media, everything else. What we want to do before that is spend at least 2 weeks
13:01
Speaker A
doing a GA rollout. Uh what we're finally going to be ready for today is rolling it out to beta.
13:08
Speaker A
How are you How are you feeling about launching in what, 2 weeks? Dude, I am uh excited and stressed because there's like some some core parts of it that we still haven't built yet, which is again a very high-risk surface, but this is
13:21
Speaker A
coming along so beautifully. Like now when I use the product, I used to call it jank even like 2 days ago. And now like it just feels it just feels smooth.
13:29
Speaker A
It feels so delightful. So, I am I'm super psyched, and so is the team, and it's going to be a very stressful like next 4 weeks as we like prepare for everything with launch.
13:38
Speaker A
But, I think we're on track to make it happen. Every single day over the last week has had one major thing that [music] I'm fighting up against.
13:59
Speaker A
And so, uh I had to have to actually like open up my calendar to see what even I was doing this week, cuz wow, this has been a this has been a long one. So, we're hiring for our chief
14:08
Speaker A
revenue officer, who uh for people who don't know is basically the person who runs the sales and like enterprise sales team. And just figures out all the motions and like architects the whole whole system up. So, I interviewed 40
14:19
Speaker A
people, and I found this one guy who was brilliant. But, he was also being chased very heavily by three other companies.
14:29
Speaker A
Thursday evening, and we finally got him. It's like a huge kind of weight lifted off my shoulders. There's actually nine people that report to me right now, who I would hand over to him to build and scale our our enterprise
14:41
Speaker A
sales team. The next thing that I'm really thinking about is everything [snorts] with the launch of the note taker.
14:49
Speaker A
Initially, when we started the software, it was like, "Okay, it's going to be a meeting note taker." Right? You press a button, it records a meeting, it shows you a summary, end of story. Really simple. And so, as we get started having
14:59
Speaker A
more with people, like the list kept getting bigger. From just one surface, it became five different surfaces [music] that just did their own whole sets of things uh that were just completely new.
15:10
Speaker A
All those need designs, a good amount of innovation on the ML side and the UI side, cuz we're building and shipping things that nobody has built before.
15:17
Speaker A
We're we're starting to get a lot of feedback from people, which is great. It's so much better than getting no feedback. And whenever you fix something and you just like make this change, like it's it's like you're seeing this this
15:27
Speaker A
like little baby of yours grow. There's this this like deep energy between everybody on the team, honestly. The question is, in the next 22 [music] days, can we build the best note taker on the planet? You know, one of my
15:42
Speaker A
favorite moments from the last week has been So, there were there were three sales calls I was a part of. We're we're talking to potential customers. And I just mentioned in passing that we were building a note taker. And I I kid you
15:55
Speaker A
not, in these three meetings, the the main person leading the call messaged their procurement team right then. It's like, "Hey, stop our other meeting recorder pilot." And I was like, "No, no, no. Like, this is not ready yet.
16:06
Speaker A
This is going to take another 4 weeks to be ready for you." [music] And they're like, "No.
16:10
Speaker A
We would rather not have a meeting recorder for 4 weeks if we can get the one that you're building." I I had a strong hypothesis about brand. That if you build a great product and deliver a great experience to people, they want
16:23
Speaker A
you to do more for them. With with Siri, you don't want Siri to do a lot more for you because it has failed you and it has broken your trust. But with Whisper, because it has like promised one thing
16:34
Speaker A
and then done it so [music] well, people know that the next time Whisper promises something, they're going to deliver really well on it. And this started to show this to me in action. And so, now we have a long list of customers who are
16:47
Speaker A
ready and excited to try this product when we have it. It's not just that we have to build a great note taker product. Like, no, we have to meet those expectations that our customers have from us because that is the one thing
17:01
Speaker A
that is very hard to recover from. And the last thing that's been one of the hardest things about the note taker is actually sequencing everything.
17:09
Speaker A
Because they're moving so fast, right? We come up with a new feature, we ship it in the product. But we now we need to make sure uh the marketing team that's making the new website and the assets knows about it because guess what the
17:20
Speaker A
website needs to be live in a week. So, you're just speed running all of these things together at basically the limits at which everything starts to break.
17:29
Speaker A
Now, what I have is I have the whole system architecture, I have the whole product spec, and I have the whole marketing roadmap in my head all together. What I'm thinking about every single day is which of these have
17:42
Speaker A
updates that need to be made and communicated to the right people. Which of these still has some scope for improvement and optimization and do that day after day every other day. Now, that we passed this challenge going up
17:57
Speaker A
against some of the biggest and most renowned startups in the valley for this one candidate we were hiring, time for the next challenge.
18:08
Speaker A
So, do you see WhisperFlow transitioning back into hardware ever? It's always a possibility.
Topics:WhisperFlowvoice recognitionAI startupTanay Kotharibrain-computer interfacevoice-to-textFlow OSAI product developmentmachine learningstartup culture

Get More with the SozAI App

Transcribe recordings, audio files, and YouTube videos — with AI summaries, speaker detection, and unlimited transcriptions.

Or transcribe another YouTube video here →