Transcribing interviews for qualitative research, end to end
By Andrey ChmerevI build Kekoso, a local transcription app, which is relevant to one section and irrelevant to the rest. The import requirements for NVivo, ATLAS.ti and MAXQDA are quoted from each vendor's own manual and linked, read on 4 September 2026. Nothing here is ethics advice — your committee's rules are the ones that count.

Transcribing in qualitative research is not the same job as transcribing a meeting. The output is not a document to skim — it is the data you will code, quote and defend, and the decisions you make before you start determine whether it can carry that weight.
Four of them matter, in this order: which verbatim, where the audio is allowed to go, how the file has to look for your analysis package, and when anonymisation happens.
To transcribe research interviews properly, the first question is not which tool but which transcript your method demands. Transcribing interviews for research is a methodological choice before it is a technical one.
Which verbatim your method requires
Three conventions are used in qualitative research, and the choice between them is methodological rather than aesthetic.
Verbatim keeps everything: fillers, repetitions, false starts, stammers, audible pauses, overlapping speech. It is what conversation analysis, discourse analysis and discursive psychology work on, because in those traditions how something was said is data.
Clean verbatim removes fillers and false starts while keeping every substantive word and the speaker’s own phrasing. Most thematic and framework analysis runs on this.
Intelligent verbatim tidies grammar into readable prose. It is a reporting format, and using it as your analysis corpus quietly deletes the hesitations and self-corrections you might later want to interpret.
Decide before transcription starts, write the choice into your method section, and apply it consistently. A corpus half-cleaned is worse than either convention applied uniformly, because you cannot tell later whether a missing “um” was absent or removed.
One practical note: machine transcription produces something close to clean verbatim by default. Models drop most fillers on their own. If you need full verbatim, you will be adding them back by ear, and that is slower than correcting a clean transcript — worth knowing before you plan the timeline.
Where the audio is allowed to go
Interview transcription for dissertation work usually has this written into the ethics approval, which makes it the constraint that picks the tool.
This is where transcribing qualitative interviews differs most from every other use of transcription, and where the answer is often “not there”.
An interview recording is identifiable personal data. Uploading it to an online transcription service transfers that data to a third party. In the EU and UK that usually makes the service a processor under GDPR. A written agreement then has to be in place before the first upload, along with a lawful basis and whatever your institution’s data protection office adds on top.
Many ethics committees restrict third-party processing outright for sensitive topics. And some participant information sheets promise that recordings will not be shared — a promise the upload breaks, even if nobody notices.
Three things follow, and none of them are optional:
Ask before you record, not after. The approval you have describes a specific handling plan. If it does not mention a transcription service, adding one is a change to the plan.
“Deleted after 30 days” is a retention promise, not an absence of transfer. The data still went, and the commitment is only as durable as the company.
Transcribing research interviews on your own machine sidesteps the question rather than answering it. If the audio never leaves the laptop, there is no third party, no processor agreement, and nothing to declare beyond the storage arrangements you already have. That is the reason local transcription has an audience in research disproportionate to its share elsewhere — including mine, which is the interested part of this article.
If you use a service anyway, that can be entirely legitimate — with the agreement signed and the committee informed.
If your method quotes with timings, know that they are approximate: two engines on the same file disagree by 130 ms at the median.
What your analysis package will accept
This is the part that wastes the most time, because people discover it after transcribing everything. All three major packages take transcripts with timestamps, but not in the same formats — read from each vendor’s manual on 4 September 2026:
| Subtitle files (SRT/VTT) | Word / RTF | Plain text | |
|---|---|---|---|
| ATLAS.ti | Yes, directly | Yes | Yes |
| MAXQDA | Yes, directly | Yes | Yes |
| NVivo | No | Yes, as a table | Yes, tab or comma delimited |
ATLAS.ti states plainly that “The automated transcripts are provided in form of VTT or SRT files”, and imports them against the linked media file.
MAXQDA has a dedicated path — Import, Transcripts, From VTT files — and converts the timestamps into its own internal ones, removing them from the visible text for readability. It reads SRT the same way.
NVivo is the one that needs preparation. Its manual specifies plain text in comma- or tab-separated format, rich text, or Word. Timestamps go in as hh:mm:ss or hhmmss, timespans are joined by a hyphen or forward slash. For a Word document it wants a table with a timestamp or timespan column and, optionally, a speaker column. For a text file, tab-delimited lines with the same number of tabs on every line.
The practical consequence is worth stating simply: if you use ATLAS.ti or MAXQDA, export subtitles and you are done. Any transcription tool that writes SRT or VTT — including the automatic captions built into most transcription apps — produces a file those two import with timings intact. If you use NVivo, plan a conversion step into a two- or three-column table.
Speakers, and the thing software will not do for you
Attribution is the other place where research transcripts differ from meeting notes. A quote without a reliable speaker label is not usable, and in qualitative research the quote is the point.
Machine transcription of a single mixed recording does not do this. Splitting one audio track by speaker is diarization, a separate problem from recognition, and most local tools do not attempt it. Two ways around it:
Record so the problem does not arise. Two microphones, two files, or a call recorded with each side on its own track. Then attribution is a fact of the recording rather than a guess about it.
Attribute during the correction pass. You are listening to the whole recording anyway to check the text. Marking speakers as you go costs almost nothing extra; doing it afterwards means listening twice.
The same pass is where you mark what the notation of your method requires — pauses, overlaps, [inaudible 14:32], non-verbal events. Do it with the audio in your ears, not from the text alone.
Anonymisation happens before analysis, not before publication
A common and expensive mistake is treating pseudonymisation as a step before submitting a paper. By then the identifiable names are in your codes, your memos, your quotes file and three backups.
Do it as the transcript becomes final: replace names, employers, place names and any detail that identifies a small population, consistently, with a documented convention. Keep the key in a separate location from the data, governed by your data management plan. Everything downstream then works on the pseudonymised corpus.
Watch for the identifiers that are not names — a job title in a small organisation, a rare condition, a specific date, the one person who does that role in that town.
The whole workflow, in order
- Decide the verbatim convention and write it into the method.
- Check where the audio may go before recording, and record accordingly.
- Record so speakers can be attributed — separate tracks if the setting allows.
- Transcribe, by hand or by machine.
- Correct with the audio playing, attributing speakers and adding method notation in the same pass.
- Pseudonymise, keeping the key separate.
- Export in your package’s format — subtitles for ATLAS.ti and MAXQDA, a timestamped table for NVivo.
- Code.
The order matters more than the tools. Steps 2 and 3 are cheap before the interview and impossible afterwards, and step 6 gets harder every day you postpone it.
On doing it by hand
There is a real argument for manual transcription that has nothing to do with cost: in several traditions, transcribing is where you come to know the data, and outsourcing it to software or to a service removes a step of the analysis rather than a chore before it.
If that is your position, hold it — the four to six hours per hour of audio is buying you something. If it is not, the honest version of the machine route is that it does not remove the listening. It moves the effort from typing to correcting, and the correction pass still means hearing every minute of the recording. What you save is the typing, which is real, and what you keep is the immersion, which is the part that mattered.
Import requirements above were read from the NVivo, ATLAS.ti and MAXQDA manuals on 4 September 2026 and are linked at each claim. Ethics and data protection rules vary by institution and country: this describes the shape of the question, and your committee gives the answer.
Questions people ask
What is the difference between verbatim and clean verbatim transcription?
Verbatim keeps everything — every 'um', repetition, false start, stammer and audible pause. Clean verbatim removes fillers and false starts while keeping the words and the meaning. Intelligent verbatim goes further and tidies grammar into readable sentences. Which one you use is a methodological decision, not a style preference: conversation analysis and discursive psychology need full verbatim, while thematic analysis usually works fine on clean verbatim.
Can I use an online transcription service for research interviews?
Only if your ethics approval and your data agreement allow it, and that is a question about your specific project. Uploading a recording sends identifiable personal data to a third party, which in EU and UK research usually means the service is a data processor and needs an agreement in place before the first file goes up. Many committees restrict it outright for sensitive topics. Check before you record, not after.
Which formats can I import into NVivo, ATLAS.ti and MAXQDA?
ATLAS.ti and MAXQDA both take VTT and SRT subtitle files directly, so a transcript exported as subtitles goes straight in with its timestamps. NVivo does not: its manual lists Word, rich text and tab or comma delimited text, with timestamps in hh:mm:ss and, for Word, a table with a timestamp column and optionally a speaker column.
How long does it take to transcribe an hour of interview?
By hand, the widely cited range is four to six hours per hour of audio for a competent typist working from a clear recording, and longer for full verbatim or poor audio. Machine transcription changes the shape of the work rather than removing it: the first pass takes minutes, and what remains is correction, speaker attribution and anonymisation.
Do I need to anonymise the transcript or the recording?
Both, at different points. The recording is the identifiable artefact and is usually governed by your data management plan, including where it is stored and when it is destroyed. The transcript is what you code and quote from, so pseudonymisation happens before analysis rather than before publication — replacing names, places and employers consistently, and keeping the key separate from the data.
Should I transcribe interviews myself or use software?
Transcribing by hand is a form of immersion in the data, and some methodological traditions treat it as part of the analysis rather than a chore preceding it. If that applies to your method, the time is not overhead. If it does not, machine transcription followed by a careful correction pass gets you the same document with the listening pass preserved, because correcting still means listening to the whole recording.