Aristotelis Zervos
Aristotelis Zervos, Editorial Director at 2B Advice, combines legal and journalistic expertise in Data protection, IT-Compliance and AI regulation.
The increasing use of transcription features in digital communication environments raises classic data protection issues in practice. New technologies such as Microsoft Teams Intelligent Speakers, Voice Match, and speaker identification and speaker recognition methods add an extra layer of complexity. What to keep in mind when using these features.
From Transcription to Identification
The purpose of transcribing meetings is to Documentation, quality assurance, or preservation of evidence.
But modern systems can do even more: They not only convert speech into text, but also attempt to identify individual speakers.
This is where technologies like Voice Match come into play. This technology compares a person’s voice with stored voice profiles to uniquely identify who is speaking. When used in combination with „Microsoft Teams Intelligent Speakers,” this feature is particularly useful in hybrid meetings where multiple people are participating together in the same room. The devices analyze acoustic characteristics and assign the transcription to individual participants.
These methods can be classified as „speaker identification” (attributing a statement to a specific person) and „speaker recognition” (recognizing a voice based on its characteristics).
Personal Reference and the Biometric Dimension in Transcription
Even the technical basis of such systems—that is, the temporary storage and analysis of audio data—constitutes a Processing personal data, since speech can be directly attributed to an identifiable individual. It is therefore directly subject to the requirements of the GDPR as well as the relevant criminal law provisions.
In the following transcript, it is often not possible to unambiguously identify a natural person. However, it is only through the addition of automated speaker identification that the level of intrusion increases significantly.
If voice characteristics are processed specifically for the purpose of uniquely identifying a person, this constitutes, pursuant to Art. 4 No. 14 GDPR biometric data. According to the case law of the European Court of Justice and Recital 51, the prerequisite is GDPRthat the Processing is carried out specifically for identification purposes—which is typically the case with speaker recognition. In this case, the stricter requirements of Article 9 apply GDPR, in particular the general prohibition on processing, with Reservation of permission.
In any case, if statements are clearly personalized, the risk to those affected increases significantly—for example, with regard to performance or behavioral monitoring in an employment context.
Legitimate Interest in Transcription
Particularly in the case of internal meetings, training sessions, or structured interviews, simple transcription may be based on Art. 6(1)(f) GDPR (legitimate interest).
Provided that the Responsible persons If a controller relies on a legitimate interest as the legal basis, a careful and transparent balancing of interests is required.
- First, it must be determined whether the transcription is actually necessary to achieve the intended purpose. Its use is permissible only in cases where no less intrusive, equally effective means is available. If the preparation of a manual transcript is sufficient, then there is no Necessity. However, in cases involving extensive material that requires a verbatim transcript, automatic transcription may be necessary.
- A balancing of interests must then take place. Pursuant to Art. 6(1)(f) GDPR is the Processing is permissible only if the legitimate interest of the controller outweighs the interests of the data subjects and their rights are not disproportionately infringed. In particular, the following aspects Confidentiality, to take into account potential drawbacks of recording as well as the risk of performance or behavioral monitoring.
- After all, a transparent Documentation the weighing of interests that was carried out is necessary in order to provide a clear and comprehensible rationale for the decision in the case in question.
In the context of employment, Art. 6(1)(f) GDPR However, its applicability is limited. To the extent that the transcription serves to document the conduct of the employment relationship, § 26 BDSG, as the more specific provision, must be applied first. For purposes beyond this, such as quality assurance or Documentation, the general balancing test under Article 6(1)(f) applies GDPR, although the structural dependency relationship typically complicates matters.
Is transcription permitted? Consent as the legal basis
Especially with technologies such as speaker recognition, there is often a Consent required because the Processing is particularly labor-intensive.
The Consent must be voluntary, informed, and unambiguous. To ensure that the Consent is valid, the affected A person must take an active step to indicate their consent. This can be done, for example, by clicking a button or in some other way. Simply remaining silent or accepting preset options or automatically activated features in conferencing tools is not sufficient for this purpose.
However, particularly in the context of an employment relationship, there are significant doubts as to whether the Consent. Additional regulations are needed here. Under Section 26(4) of the Federal Data Protection Act (BDSG), company agreements are expressly recognized as an independent legal basis for data processing in the employment context and offer, compared to the Consent the advantage that they are collectively negotiated and can partially offset the structural power imbalance between employers and employees.
Others Legal basis Although these are conceivable, they play a minor role in practice:
- Fulfillment of the contract is usually out of the question, as transcriptions are rarely absolutely necessary.
- As a rule, there are no legal obligations.
For systems with speaker identification, it is also necessary to check whether a Data Protection Impact Assessment according to Art. 35 GDPR is required. Since biometric data is processed, this is generally the case according to the positive lists of the German supervisory authorities.
Reading tip: Consent to the processing of personal data
Criminal Law Risks: Section 201 of the German Criminal Code (StGB)
In addition to data protection law, criminal law must also be taken into account. Under Section 201 of the German Criminal Code (StGB), the unauthorized recording of private conversations is a criminal offense. Since transcription systems generally require an audio recording, there is a significant risk involved if one does not have the appropriate authorization.
An effective Consent may serve as a justification under both data protection law and criminal law. However, it is important to distinguish between the two: While in criminal law, consent that precludes criminal liability may, under certain circumstances, be implied, the GDPR a clear and documented Consent, in the case of biometric data pursuant to Article 9(2)(a) GDPR In addition, an explicit one.
For systems that use speaker recognition, this means that simply participating in a meeting is not sufficient to ensure legal identification.
Technical and Organizational Measures
The use of such technologies requires comprehensive Technical and organizational measures. These include, in particular:
- Privacy by Design and by Default, for example, by disabling default features for recording and speaker identification.
- Access Restrictionsn transcripts and audio files.
- Encryption during storage and transmission.
- Training for employees.
Another key element is a Fire Suppression Plan. In particular, audio recordings should be deleted after the transcripts have been created, unless there is another purpose for retaining them.
Additional Regulatory Requirements for the Use of AI
Speaker recognition is often based on AI systems. Depending on the context of use, such systems may be subject to additional regulatory requirements, particularly under the AI Regulation.
Biometric identification systems for natural persons are classified as high-risk AI systems under Annex III, No. 1 of the AI Regulation, unless they involve real-time remote biometric identification, which is largely prohibited under Article 5 of the AI Regulation. In the typical context of a meeting, this means that speaker recognition systems fall into the high-risk category, with the corresponding obligations regarding conformity assessment, technical Documentation, Transparency toward those affected and human oversight.
This is to be distinguished from the use of emotion-recognition systems in the workplace: Such use is generally prohibited under Article 5(1)(f) of the AI Regulation. Exceptions apply only for strictly limited purposes, such as protecting the health and safety of employees. In the context of meetings, the prohibition is therefore likely to apply as a general rule.
The regulatory requirements of the AI Regulation thus come on top of data protection and criminal law provisions, significantly increasing the compliance burden on companies.
High Standards for the Legally Compliant Use of Transcription
The legally compliant use of transcription and speaker identification requires careful review on several levels: the legal basis under data protection law, criminal law safeguards under Section 201 of the German Criminal Code (StGB), and—when using AI-based systems—the requirements of the AI Regulation. Companies should always assess whether less invasive alternatives can equally fulfill the intended purpose. Where speaker identification is actually necessary, explicit Consent of those affected, accompanied by a Company agreement, the most reliable way to ensure legal certainty.
Legal requirements are complex, but putting them into practice doesn’t have to be. We help you integrate transcription solutions into your processes in a legally compliant manner: from selecting the appropriate legal basis to drafting workplace agreements and implementing the technical and organizational aspects.
Please contact us to discuss your specific application.
Aristotelis Zervos is Editorial Director at 2B Advice, a lawyer and journalist with profound expertise in data protection, GDPR, IT-Compliance and AI governance. He regularly publishes in-depth articles on AI regulation, GDPR-Compliance and risk management. You can learn more about him on his Author profile page.
Questions and Answers
Is the automatic transcription of meetings permitted under data protection laws?
An automatic transcription is being processed personal data and therefore requires an appropriate legal basis. For internal meetings, training sessions, or structured interviews, a legitimate interest under Article 6(1)(f) may apply under certain conditions GDPR be considered. In particular, it is necessary to examine the Necessity, a careful balancing of interests and a transparent Documentation.
When do voice recordings become biometric data?
Voice recordings may constitute biometric data within the meaning of the GDPR occur when voice characteristics are specifically processed to uniquely identify a person. This is particularly relevant in the context of automated speaker recognition systems. In this case, the stricter requirements for special categories of personal data under Article 9 also apply GDPR.
Is participation in a meeting sufficient to constitute consent to transcription or speaker identification?
No. Simply participating in a meeting is not sufficient to ensure compliance with data protection laws regarding automated transcription or speaker identification. A Consent must be clear and unambiguous. In the case of Processing The use of biometric data for identification purposes generally requires explicit Consent required.
What should be considered regarding transcription and speaker identification in an employment relationship?
In an employment relationship, particular care must be taken to determine the legal basis for the Processing takes place. Due to the relationship of dependency between employer and employee, the voluntary nature of a Consent be problematic. Workplace agreements can therefore serve as an important foundation for the regulated use of such technologies.
Is a data protection impact assessment required for speaker identification systems?
For systems that process biometric data to identify individuals, it must be determined whether a Data Protection Impact Assessment according to Art. 35 GDPR is required. Due to the increased risk of Processing The processing of biometric data may be necessary for such verification, particularly in automated speaker identification systems.
What measures should companies take when using transcription and speech recognition systems?
Among other things, companies should implement privacy-friendly default settings, clear access restrictions, Encryption, training sessions, and a Fire Suppression Plan provide for. Audio recordings should be deleted after the transcript has been created, provided they are no longer needed for their original purpose. In addition, depending on the technology used, criminal law provisions and the requirements of the AI Regulation must also be taken into account.




