qualtranscribe logo

Transcription

Translation

qualtranscribe logo

6 mins

How to Write a Transcription Data Management Plan for IRB Submission

Most IRB protocols get sent back for revision on the data security and management sections, and within those sections, transcription is the specific area where protocols most frequently fall short. The reasons are consistent: researchers describe their interview methodology in careful detail, specify their analysis software, and then write a single sentence about transcription that doesn't address who's doing it, what compliance standards the vendor meets, how files will be transferred, where they'll be stored, or when they'll be deleted.

Illustration of a tilted, IRB-approved Data Management Plan document beside a checklist of required sections — data storage, vendor access controls, de-identification, retention timeline, data sharing agreements, and incident response — for writing a transcription data management plan for IRB submission

TL;DR

30 sec read

Here’s what you need to know

The transcription section is the most commonly revised part of an IRB data management plan, because it's where researchers are most likely to say something vague, miss a required specification, or fail to address how a third-party vendor handles participant data. This guide covers every element IRBs look for, with sample language you can adapt, verified encryption and retention standards, and a complete template for the transcription section of your DMP.

Best for researchers, compliance teams, and operations leaders evaluating transcription vendors.

Read the full guide ↓

IRB reviewers read hundreds of protocols. They know what's missing within seconds.

This guide fixes that.

What IRBs Are Actually Evaluating

Before getting into specific language, it helps to understand what the board is trying to assess. Federal regulations (45 CFR 46 for human subjects research, 21 CFR 56 for FDA-regulated research) require IRBs to evaluate whether researchers have adequate provisions to protect participant privacy and maintain data confidentiality. Audio and video recordings are the highest-risk data category in most qualitative research because they are directly identifiable: a voice recording or video file cannot be de-identified the way a survey response can.

Penn State's IRB Guideline XI makes this explicit: when a transcription service will be used, the protocol must state which service and describe the confidentiality controls in place. The University of Michigan's IRB data security guidelines require that storage platforms be IT-vetted, with access restricted to named IRB-approved research personnel. Columbia TC IRB directs researchers to use institutional platforms specifically because access can be controlled and audited.

Vague language like "data will be securely stored" does not satisfy these requirements. The six elements below are what IRBs are looking for, and each needs to be addressed specifically.

1. Data Classification and Sensitivity

Start by classifying what your recordings actually contain. Most IRBs require researchers to categorize data by sensitivity level before specifying handling procedures. The classification determines everything else in the DMP, from encryption requirements to retention timelines.


Data Type

Examples

Risk Level

IRB Scrutiny

Non-identifiable audio

Generic opinions on public topics, no names or locations

Minimal

Low

Identifiable but non-sensitive

Name, employer, general location mentioned

Moderate

Standard

Sensitive identifiable

Health information, immigration status, trauma, sexuality

High

Elevated

Protected Health Information

Any patient or clinical data

PHI

HIPAA required

Also define the technical scope of your data collection: file formats (WAV, MP3, MP4), approximate recording lengths, total estimated volume in hours, and number of individual recordings. IRBs use this to assess the scale of the confidentiality risk.

What the IRB looks for: A classification that matches the actual content of your recordings. Classifying sensitive health interviews as "minimal risk" is one of the most common reasons protocols are returned. When in doubt, classify up.

Sample language: "Interview audio will be recorded in MP3 format, averaging 60 minutes per session across 20 participants, generating approximately 20 hours of total audio. Recordings contain identifiable information including participant names, employer details, and health-related disclosures and are therefore classified as sensitive identifiable data."

2. De-Identification and Anonymization

Specify exactly how participant identifiers will be removed from transcripts, and when. IRBs want to see a structured approach, not a general statement about anonymization.


Identifier Type

Examples

Replacement Standard

Direct identifiers

Full name, email, phone, SSN, medical record number

Replace with participant code (P01, P02)

Geographic identifiers

Street address, ZIP code, specific city

Replace with region or generalize (e.g., "Southeast US")

Organizational identifiers

Named employer, specific hospital unit, named school

Replace with descriptor (e.g., "a regional hospital")

Indirect identifiers

Job title + location + age that combined could identify

Generalize or omit

Contextual background detail

Unintended references to colleagues, family members

Flag and redact during transcription

Contextual redaction deserves specific mention in your DMP. Participants frequently disclose identifying information about third parties during interviews, mentioning a colleague's name, a specific incident, or a location, without realizing it. Define a clear rule: flag all such references during transcription and replace with a generalized descriptor before the transcript is shared beyond the immediate research team.

What the IRB looks for: Specific replacement conventions (participant codes, not just "names will be changed"), and a clear rule for handling indirect and contextual identifiers. "Names will be replaced with pseudonyms" is insufficient. HIPAA Safe Harbor de-identification requires removal of all 18 identifier categories, not just names.

Sample language: "All direct identifiers will be replaced with sequential participant codes (P01, P02, etc.) during transcription. Geographic identifiers will be generalized to region level. Organizational identifiers will be replaced with descriptive categories. Any contextual references to identifiable third parties will be flagged and redacted before transcript distribution. De-identification will be completed before any transcript leaves the primary research environment."

3. Storage, Encryption and Access Security

This section requires specific technical standards, not general assurances. IRBs at multiple institutions have confirmed they require named encryption protocols and explicit access restrictions.

In-transit encryption: All audio files must be transferred using encrypted channels. The standard is TLS 1.2 or higher for web-based file transfer. Emailing audio recordings as attachments or sharing via unencrypted cloud links does not meet this standard.

At-rest encryption: Data stored on devices or servers must be encrypted. The University of Maine and multiple other institutions specify AES-128 or AES-256 as the required standard. On Windows, BitLocker meets this requirement. On Mac, FileVault meets it. Institutional server environments should meet or exceed this standard.

Access controls: Access must be restricted to named, IRB-approved research personnel. University of Michigan's IRB guidance requires that access be limited to the PI and trained research staff only. If the institution uses a shared drive, it must be access-controlled, not publicly or department-wide accessible.

Platform vetting: Storage platforms must be IT-vetted by the institution. Columbia TC IRB specifically permits institutional Google Drive with access restrictions because it has been evaluated for compliance. Personal Google accounts, Dropbox, and consumer cloud storage do not meet this requirement at most institutions.

What the IRB looks for: Specific encryption standards by name (AES-256, TLS 1.2), named storage locations (not "a secure cloud"), and an explicit list of who has access. "Only the research team" is not specific enough. Name the roles.

Sample language: "Audio recordings will be transferred to the transcription service via encrypted portal using TLS 1.2 encryption in transit. Recordings and transcripts will be stored on [Institution]-approved encrypted storage (AES-256) accessible only to the Principal Investigator and [named research assistant role]. No recordings will be stored on personal devices, portable drives, or consumer cloud services. Access will be restricted to IRB-approved personnel only."

4. Vendor Compliance and Processing Standards

If you are using a third-party transcription service, the IRB needs to know who they are, what compliance standards they meet, and what safeguards are in place. This is the section most researchers underspecify.


Compliance Requirement

What It Means

Required Documentation

HIPAA compliance

Required for any health-related data

Signed Business Associate Agreement (BAA)

GDPR compliance

Required for EU participant data

Data Processing Agreement (DPA)

PIPEDA compliance

Required for Canadian participants

Documented compliance confirmation

APPI compliance

Required for Japanese participants

Documented compliance confirmation

NDA with transcriptionists

All personnel with file access are bound by confidentiality

Signed NDA per transcriptionist

Zero AI training

Recordings not used to train models

Written policy confirmation

File deletion policy

Files deleted after project completion

Defined timeline in writing

Data residency

Files stored in specified region

Confirmation of storage location

Two requirements deserve specific attention because IRBs are now asking about them directly.

AI model training: Many general-purpose AI transcription tools use uploaded audio to train their models. This is disclosed in their terms of service, often in fine print. For research involving human subjects who consented to a specific study use of their data, this creates a direct conflict with your consent agreement. Your DMP must explicitly state whether any AI transcription is used and confirm that the platform does not train on uploaded data.

Zero retention after delivery: Some platforms retain audio files after transcription is complete. Your DMP should specify that recordings will be permanently deleted from the vendor's systems within a defined window of project completion.

Qualtranscribe meets every requirement in the table above: HIPAA compliance with BAA available, GDPR and PIPEDA data processing agreements, APPI coverage for Japanese research, NDA signed with every transcriptionist, zero AI training on any plan, and files deleted or anonymized within 30 days of project completion by default. Data residency is specified per project: US data in us-east-1 Northern Virginia, EU data in eu-central-2 Frankfurt, Japan data in ap-northeast-1 Tokyo.

What the IRB looks for: The vendor's name, specific compliance certifications, and confirmation that a BAA or DPA is in place before any file transfer. A statement that the service "follows best practices" does not satisfy this requirement.

Sample language: "Transcription will be conducted by Qualtranscribe (qualtranscribe.com), a HIPAA-compliant transcription service operating under a signed Business Associate Agreement for this study. All transcriptionists sign NDAs prior to accessing research files. Recordings are transferred via encrypted portal and are not used for AI model training under any plan. Files will be deleted from Qualtranscribe's systems within 30 days of project completion. US participant data is stored in us-east-1 (Northern Virginia)."

5. Data Retention and Destruction

Two separate timelines need to be specified: one for raw audio recordings and one for de-identified transcripts.

Audio recordings are the highest-risk data and should be deleted as soon as transcripts have been verified for accuracy. Multiple IRB guidance documents confirm this standard practice: "delete audio recordings once transcripts have been verified against them." The DMP should specify a maximum retention period for audio files and name the person responsible for confirming deletion.

Transcripts have a longer required retention period. The IRB standard across most institutions is three years after project completion, as specified in regulations. NIH-funded research typically requires seven years. If a funder or institution requires a longer period, that requirement overrides the baseline. Check your specific funder's data retention policy before writing this section.

What the IRB looks for: Two separate timelines for audio and transcripts, a named person responsible for deletion, and confirmation that the timeline aligns with funder requirements. Saying "data will be retained per institutional guidelines" without specifying the actual period is insufficient.

Sample language: "Raw audio recordings will be permanently deleted from all research systems within 30 days of transcript verification. De-identified transcripts will be retained for [3/7] years following project completion, consistent with [IRB/NIH/funder name] requirements, then destroyed via secure deletion. The Principal Investigator is responsible for confirming deletion and documenting the date of destruction."

6. Informed Consent Alignment

Every data handling commitment in your DMP must match the exact language provided to participants in their informed consent form. This is one of the most frequently cited revision reasons: the DMP promises something the consent form doesn't mention, or the consent form makes a commitment the DMP doesn't support.

Check each of these specifically:

  • The consent form should state that interviews will be transcribed and whether a third-party service will be used

  • If AI transcription will be used, the consent form should disclose this

  • The retention period stated in the consent form must match the DMP

  • The de-identification approach described to participants must match what the DMP specifies

  • Any promise to provide participants with a copy of their transcript must appear in both documents

What the IRB looks for: Direct language correspondence between the DMP and the consent form. Reviewers check both documents together. Discrepancies between them are a primary reason protocols get sent back.

Sample language for consent form: "Your interview will be audio-recorded. Recordings will be transcribed by a professional transcription service operating under a confidentiality agreement. Your name and any identifying information will be replaced with a participant code before transcripts are shared with any member of the research team beyond the primary researcher. Recordings will be deleted once transcripts have been verified. De-identified transcripts will be retained for [X] years following the study."

7. Addressing AI Transcription in Your DMP

IRBs are now asking specifically whether AI transcription tools are used in the research workflow. If they are, address it directly rather than hoping the reviewer doesn't ask.

What to confirm before writing AI transcription into your DMP:

  • Does the platform hold a BAA (if health-related) or DPA (if EU participants)?

  • Does the platform explicitly commit to not training AI models on uploaded recordings?

  • Where are files processed and stored?

  • How quickly are files deleted after delivery?

  • Can you produce documentation of any of the above if asked?

If you cannot confirm all of these, the platform should not appear in your IRB protocol. Using a general-purpose AI transcription tool that trains on uploaded data, for research where participants consented only to a specific study use of their recordings, creates an ethical and legal exposure that no IRB will approve.

Sample language: "Where AI-assisted transcription is used for preliminary analysis, [platform name] is used. This platform is [HIPAA/GDPR] compliant, does not use uploaded recordings for AI model training under any plan, and deletes files within [X days] of delivery. A [Business Associate Agreement/Data Processing Agreement] is in place before any files are transferred. Human transcription is used for all final transcripts that will be coded or quoted in research outputs."

Complete Sample DMP Transcription Section

Important: The following template is a starting point. You must verify it against your institution's specific IRB requirements, your funder's data management policy, and the actual compliance documentation available from your transcription vendor before submitting. Do not copy it verbatim without reviewing it with your IRB office.

Data Collection and Classification

Interviews will be audio-recorded in [MP3/WAV] format, averaging [X] minutes per session across [N] participants, generating approximately [X] hours of total audio. Recordings contain [identifiable/sensitive identifiable/PHI] information and are classified accordingly.

De-Identification

All direct identifiers (names, contact details, organizational affiliations) will be replaced with sequential participant codes (P01, P02, etc.) during transcription. Geographic identifiers will be generalized. Indirect identifiers will be reviewed and redacted or generalized before transcript distribution. De-identification will be completed before any transcript leaves the primary research environment.

Transcription

Transcription will be conducted by [Vendor name], a [HIPAA/GDPR]-compliant service operating under a signed [BAA/DPA] for this study. All transcriptionists sign NDAs prior to accessing research files. Files are transferred via encrypted portal (TLS 1.2+) and are not used for AI model training. [If AI transcription used: AI-assisted transcription via [platform] is used for preliminary analysis only. Human transcription is used for all final coded transcripts.]

Storage and Access

Recordings and transcripts will be stored on [institution-approved encrypted platform] using AES-256 encryption at rest. Access is restricted to the Principal Investigator and [named roles]. No recordings will be stored on personal devices, portable drives, or non-institutional cloud services.

Retention and Destruction

Raw audio recordings will be permanently deleted within 30 days of transcript verification. De-identified transcripts will be retained for [3/7] years following project completion consistent with [funder/IRB] requirements, then destroyed via secure deletion. The Principal Investigator will document the date of destruction.

Consent Alignment

All data handling procedures described above are consistent with the informed consent form provided to participants, which discloses transcription, third-party processing, de-identification procedures, and retention timelines.

Five Reasons IRB Protocols Get Sent Back on Transcription

1. Vague storage language. "Secure cloud storage" without naming the platform, the encryption standard, or the access controls. Fix: name the platform, specify AES-256 at rest, list who has access by role.

2. No vendor compliance documentation mentioned. Saying you'll use a transcription service without confirming HIPAA compliance, BAA, or NDA arrangements. Fix: name the vendor and list the specific compliance documents in place.

3. Mismatch between DMP and consent form. The DMP mentions third-party transcription but the consent form doesn't disclose it, or the retention period differs between documents. Fix: read both documents side by side before submitting.

4. No deletion timeline for audio recordings. Audio files are the most identifiable data in qualitative research and IRBs expect a specific deletion schedule. Fix: commit to deleting audio within a defined period after transcript verification, with a named person responsible for confirming it.

5. AI transcription undisclosed. Using an AI tool in the workflow without mentioning it in the protocol. Fix: disclose any AI transcription use, confirm the platform's compliance status, and specify its role in the workflow.

Need transcription that satisfies every requirement in this DMP? Qualtranscribe provides HIPAA, GDPR, PIPEDA, and APPI compliance as standard, with signed NDAs, BAAs on request, zero AI training, and documented file deletion timelines. Get started here.

FAQ

Does my IRB require me to name my transcription vendor in the protocol? Many do. Penn State's IRB Guideline XI explicitly requires naming the transcription service and describing its confidentiality controls. Check your institution's specific requirements, but naming the vendor and providing compliance details is best practice regardless.

What encryption standard does my IRB require? The University of Maine and multiple other institutions specify AES-128 or AES-256 for data at rest as a minimum standard. TLS 1.2 or higher is the standard for data in transit. Check your institution's IT security policy for the specific requirement.

How long do I need to retain transcripts? The minimum across most institutions is three years after project completion. NIH-funded research typically requires seven years. Your funder's data management requirements override the institutional baseline. Check both before specifying a retention period in your DMP.

Can I use AI transcription for IRB-approved research? Yes, if the platform meets your compliance requirements and you disclose its use in the protocol. The key requirements: the platform must not train on your data, it must meet applicable compliance standards (HIPAA, GDPR), and your consent form must disclose AI processing. See our full guide on AI transcription for IRB-approved research.

What is the difference between de-identification and anonymization in my DMP? De-identification removes specific identifiers but may not eliminate re-identification risk entirely. Anonymization permanently severs the link between data and the individual. IRBs typically accept de-identification for most qualitative research, but some require anonymization for particularly sensitive data. See our detailed guide on de-identification, anonymization, and pseudonymization.

What should my consent form say about transcription? It should disclose that interviews will be transcribed, whether a third party will conduct the transcription, that identifying information will be replaced with codes, how long recordings will be retained, and when they will be deleted. Every commitment in the consent form must match a corresponding provision in your DMP.

Related Reading

Turn your recordings into analysis-ready transcripts.

Human Transcription

Clean verbatim and full verbatim transcripts, delivered by specialist transcriptionists

AI Transcription

Instant Draft powered by AI, with Smart Insights for analysis-ready output

Translation Services

Accurate translation across 99+ languages for multilingual research workflows

Keep reading

Related articles

Illustration of a growth chart showing +42% qual research market growth from 2020–2026, remote and AI-tools icons, and a warning card highlighting the transcription gap — studies scaling faster than the teams turning recordings into data.

Qualitative Research in 2026: Growth, Remote Studies, AI, and the Transcription Gap

Qualitative research has come a long way from the backroom focus group. In 2026, it is a remote-first, AI-assisted, globally distributed discipline operating under more compliance pressure than ever before. The methodologies have evolved. The tools have changed. The participant pools have expanded to markets that didn't exist in most research programs five years ago. What hasn't kept pace is the step that turns all of that collected audio and video into something analyzable. Transcription, and the accuracy, compliance, and formatting it requires for professional research use, remains the operational bottleneck most research teams have not fully solved.

Read article

Illustration of a Zoom meeting call grid with host and participant tiles, recording indicator, and Zoom logo, connected to a checklist for running focus groups on Zoom — tech checks, breakout rooms, note-taker, and local backup recording

How to Conduct a Focus Group on Zoom: A Complete Guide

Zoom has become the default platform for virtual focus groups, and for good reason. Participants already know how to use it. It handles groups of six to twelve comfortably. Breakout rooms let you run sub-group exercises. And the recording quality, on good internet with decent microphones, is good enough for professional transcription. But running a focus group on Zoom is not the same as running a standard meeting on Zoom. The setup decisions that don't matter for a team standup matter a lot when you need accurate speaker attribution, clean audio for six participants talking over each other, and a recording that will be professionally transcribed for qualitative analysis. This guide covers every step

Read article

Illustration of KOL interviews and patient insights converging into a ranked list of top pharma market research firms, covering therapeutic-area specialization, global KOL access, HIPAA-compliant patient panels, regulatory-aware moderators, and terminology-accurate transcripts.

Top Pharma Market Research Companies: From KOL Interviews to Patient Insights

Pharmaceutical market research moves faster than most research categories and carries higher stakes at every stage. A misread KOL interview before a launch can redirect a commercial strategy. A payer research program that doesn't reflect actual formulary decision-making logic produces positioning that doesn't survive first contact with market access. A patient insights study that wasn't conducted under proper compliance protocols can create problems that outlast the study itself.

Read article

qualtranscribe logo