Does anyone know just how much of the total functionality of Mycroft is actually running on the Raspberry Pi? I asked this question four years ago on Reddit (I'll paste the response below) and now I wonder if things have changed, particularly with regard to speech to text.
There are several ‘layers’ to a voice assistant;
Wake Word - that detects when you are speaking to the device. This is local to the device and we use PocketSphinx.
Speech to text - that detects what you say to determine Intents - we currently use a cloud service for this
Intent matching - this is done locally using our own open source software - Adapt and Padatious
Skills - Intents then match to Skills. Some Skills require internet connectivity.
Text to Speech - We use our own software called Mimic for this, it’s local to the device.
> In order to provide an additional layer of privacy for our users, we proxy all STT requests through Mycroft's servers. This prevents Google's service from profiling Mycroft users or connecting voice recordings to their identities. Only the voice recording is sent to Google, no other identifying information is included in the request. Therefore Google's STT service does not know if an individual person is making thousands of requests, or if thousands of people are making a small number of requests each.
Well, unless Google does voiceprint analysis. But Google wouldn't do that, would they? /s
Beyond that, if I'm reading right, local STT will still require a separate STT server. It won't run on the Mark II itself, right?
The blogpost linked in this submission says the following:
> Mimic 3: Mycroft’s newer, better, privacy-focused neural text-to-speech (TTS) engine. In human terms, that means it can run completely offline and sounds great. To top it all off, it’s open source.
If "skills" are what I think they are (something like external commands, for example "Play X on Spotify"), then my understanding would be that everything but those runs offline and local-only.
But if things like `speech to text` requires internet connection and sends the data to some cloud service, then the entire value proposition of this product falls apart.
I hope that's really not the case, as that would be outright lying and false advertisement.
The raspberry pi 3 that is used in older products doesn't have enough power to be all offline. Maybe you could setup a server at home (but they won't help you!), but you cannot do it on the hardware they have. The next gen mycroft 2 (should ship this fall - first announced many years ago) will have a pi 4 which might have enough power to run offline, this isn't clear yet.
> might have enough power to run offline, this isn't clear yet
It's very unclear and misleading to put "it can run completely offline" if you're not 100% sure it can actually run "completely offline", hardware be damned.
If they're marketing it as "fully offline", they ought to be doing the speech to text bit locally now. Worked on part of a platform which could use Rasa for this a couple of years ago, but running on something a bit more powerful than a Raspberry Pi!
There are several ‘layers’ to a voice assistant;
Wake Word - that detects when you are speaking to the device. This is local to the device and we use PocketSphinx.
Speech to text - that detects what you say to determine Intents - we currently use a cloud service for this
Intent matching - this is done locally using our own open source software - Adapt and Padatious
Skills - Intents then match to Skills. Some Skills require internet connectivity.
Text to Speech - We use our own software called Mimic for this, it’s local to the device.