If you've got RTX 5090, maybe try ninfer (https://github.com/Neroued/ninfer). Folks over on /r/localllama have been reporting wild prefill/token gen speeds with ninfer (NVFP4; 256k ctx).
Looks like slope benchmarks and results, and as usually, people are mixing MTP numbers with non MTP numbers.
Or just 100 token input benchmarks.
Or just failed ones as actual measures.
I was kinda wondering too, and did a (very shallow) dive into the JavaScript on that page. I'm almost positive they are using Deepgram(dot com)'s speech-to-text service.
I ran whisper.cpp on that audio file on my laptop, and it does a reasonably well job too.
Standard audiometry, usually, only produces a test that goes up to 8 kHz (functional speech recognition). An extended high frequency audiometry may, or may not be more revealing as it can go higher. It would have no bearing on your medical treatment though.