Transcribe an interview audio to text with AI
Verbatim, with the minute of every sentence for quoting, and without the interview leaving your PC. For feature stories and theses.
Updated: October 2026 · 6 min read
To transcribe a recorded interview to text with AI without uploading it anywhere: open AppGrabbit › Download › Transcribe with AI, drop the audio (m4a, mp3, wav, ogg) or the video, pick the Accurate model and click Transcribe. Out comes a verbatim .txt, just as it was said, and an .srt with the minute of every sentence for quoting. Everything runs on your PC: nobody else hears the interview. What it does not do: separate speakers. 1 credit · 5 free a month
Why an interview should not go to a website
An interview almost always holds something that is not public: a source, a patient, a person who trusted you. “Free” websites upload it to their servers and keep it; some use it for training. And they charge by the minute past the first 30. Here the audio stays on your computer and an engine running right there transcribes it; you can even unplug the internet.
What you get, and what you don't
You get the verbatim transcript: what was said, with fillers, pauses and repetitions, which is what qualitative research asks for and what a feature story needs to quote properly. You get the .srt with timing, so every sentence has its minute and you can go back to the audio to confirm. You do not get who is speaking: the text is running, with no per-person labels. For a two-person interview it reads fine; for a five-person panel, it is not the tool.
Tips that change the result
- Accurate model — about 5 words wrong in 100 against 8 for Fast; on names and technical terms it shows.
- Pin the language — if interviewer and interviewee mix languages, pin the main one before starting.
- Audio close to the mouth — a phone recorder on the table, a meter away, transcribes far better than a video from the back of the room.
- Review with the .srt — proper names fail most: the minute takes you to the audio in one click.
AppGrabbit against the other ways
| AppGrabbit | Transcription websites | Typing it yourself | Human service | |
|---|---|---|---|---|
| The interview never leaves your PC | ✓ Never | ✗ Goes to their servers | ✓ Never | ✗ Goes to a third party |
| Verbatim, as said | ✓ Yes | ✓ Yes | ✓ Yes | ~ As ordered |
| Separates speakers | ✗ No | ~ Some, paid | ✓ Yes | ✓ Yes |
| One hour of interview | 3 to 12 minutes | Minutes, with an account | 4 to 6 hours | 1 to 3 days |
| Cost | 1 credit per transcript (5 free a month) · Pro $8.99/mo · 7-day pass $3.99 | Free minutes, then per minute | Your time | USD 1 to 2 per minute |
Step by step
- 1Install AppGrabbit for Windows — Download the free installer (everything included). Windows will say it protected your PC because the installer is not code-signed: click More info, then Run anyway.
- 2Drop the interview audio or video — The phone recorder (m4a), the voice note, the recorder's mp3, or the mp4 if it was on video. If the interview is on YouTube, paste the link in the same tab.
- 3Pick Accurate when there are names and terms — For an interview the Accurate model is worth it: fewer mistakes on proper names and technical vocabulary. Pin the language if there are two.
- 4No prompt, or “Key ideas” — For a thesis or a feature story “No prompt” usually fits: the .txt comes out clean for quoting. “Key ideas” works for a first pass.
- 5Open the .txt and the .srt — They land in the Transcripts folder inside your downloads folder, and in the Library. The .txt pastes into Word or into your AI; the .srt carries the minute of every sentence.
Frequently asked questions
No. The text is running, without “Interviewer:” and “Interviewee:”. In a two-person interview you tell them apart by content and by the .srt: the minute tells you where to look. If you need per-person labels, it does not add them today.
The .srt carries the exact minute of every sentence: “(Interview, 14:32)”. And since the audio never left your PC, you can listen again in the app or any player to confirm the quote.
The engine transcribes what it hears, fillers and repetitions included: that is the verbatim one qualitative research asks for. The clean version you make on top of it, in Word.
No problem: there is no cap. With an NVIDIA card it takes about 6 minutes; CPU only, about 25 with the Fast model.
The engine detects the spoken language. If languages mix, pin one (Spanish, English, Portuguese and more) before transcribing. It comes out in the language of the audio, not translated.
1 credit per transcript. Free comes with 5 credits a month and, at 0, you can claim 1 more every 48 hours. The 7-day pass (USD 3.99, no renewal) removes credits for a week; Pro removes them for good at USD 8.99 a month. No minute cap and no file-length cap.
Nothing. The audio is transcribed by Whisper (whisper.cpp) running on your computer, on your NVIDIA card if you have one or on the CPU if not. You can unplug the internet and it keeps working. Only the first time does it download the engine and the model.
Related guides
Transcribe your interview, on your PC
Windows 10/11 · 1 credit per transcript, 5 free a month · Verbatim, with minutes, nothing uploaded