I hate AI and I sometimes have issues with hearing. I honestly wish there were smart glasses and smart captions with disability-friendly, but anti gen AI features.
This is very definitely one of those problems where you really need to define, mostly for yourself, exactly what you mean by “without AI.”
What you’re describing is automation, and it is, by necessity, artificially intelligent automation. That’s because human beings don’t make the same sounds the same way every single time, and those sounds aren’t going to be made against the same background noise every single time. So there is no way that you can ever deterministically transcribe spoken word to text. It requires guesswork, and when you’re talking about computers and guesswork, you’re necessarily talking about technology that falls within the general purview of “artificial intelligence.” This has been true for as long as computer assisted transcription has been a thing.
So what exactly is it that you’re trying to avoid here? Is it specifically your data going to an AI company that concerns you (ie, would a local model be an acceptable solution)? Are you perhaps trying to avoid giving money to certain AI companies? These are the sort of issues that might be addressable. But just a broad “no AI” prescription will basically rule out any useful answers. I hope that makes sense.
I mean… I don’t consider ‘old school’ ML (TTS, STT, OCR, photo tagging) as ‘AI’
Yeah, I mean generative AI.
You mean you want speech to text? Yeah you have to figure out what exactly you object to. The better STT stuff does use language models, but I wouldn’t call it “generative” in the chatbot sense.
You mean have a human do it for you?
I just mean without generative AI



