Yes.The ability to do speech-to-text and text to-speech with a relatively high level of speed and accuracy is impressive. And it has at least a few potentially positive real world applications.
Similar tools have existed for years but the quality and user friendliness of those has generally been pretty lacking.
is it that impressive? I’m guessing it’s just a layer of speech-to-text then token parsing like usual.
Yes.The ability to do speech-to-text and text to-speech with a relatively high level of speed and accuracy is impressive. And it has at least a few potentially positive real world applications.
Similar tools have existed for years but the quality and user friendliness of those has generally been pretty lacking.