Speech and audio processing, spoken language understanding

The research group focuses on speech and audio signal processing, leveraging advanced machine learning techniques to develop and analyze speech and speaker recognition systems, multilinguality, low-resourced scenarios, speech generation, and voice cloning technologies. Its expertise further encompasses detection of synthetic and manipulated audio (deepfakes), paralinguistic analysis of speech, sign language processing, and pathological speech processing, with applications in human–machine interaction,security, and healthcare.
Led by Dr. Petr Motlicek and Dr. Mathew Magimai-Doss, the group tackles a range of research problems, with each leader focusing on specific areas of expertise, including speech analytics, voice intelligence, analysis of spoken conversations, paraliguistics in speech, analyses of pathological speech, sign language processing, etc.
Artificial Intelligence (AI) has become a powerful and pervasive technology in recent years, influencing numerous aspects of our daily lives. We encounter AI through recommendation algorithms in online stores, voice-activated smartphone assistants, or the widespread use of technologies like ChatGPT. However, its rapid growth and integration into society raise complex questions and concerns among the public. Public opinions on AI vary widely; while some people are enthusiastic about its potential to revolutionize industries and enable breakthroughs in fields such as medicine, others fear that AI could lead to undesirable outcomes, such as a loss of human control or privacy. These views are often shaped by media narratives and the competing interests of different stakeholders, which play a significant role in influencing both public opinion and policy decisions.
Our team of scientists and communication experts aims to enhance the understanding of AI technologies among the Swiss people, with a particular focus on teenagers and female students, to create a positive societal impact. Building on the foundations of our previous project, NewsOnAI, we will expand beyond traditional media such as newspapers and employ diverse methods, including artistic performances and interactive exhibitions. We will design these activities to be highly interactive, encouraging active participation and dialogue. Activities will include themed theater plays that explore AI’s impact on everyday life, exhibitions where participants can interact with AI tools, and workshops specifically designed for teenagers and female students to discuss AI’s future role in society. Feedback collected will include real-time audience reactions, structured questionnaires, and focus group discussions, which will be analyzed to continuously refine and adapt our engagement strategies.
Our primary audience includes Swiss citizens interested in cultural activities, particularly teenagers who are keen to follow new trends. Additionally, we are committed to addressing gender aspects by designing content and activities that specifically appeal to female students. We aim to inspire and empower young women to take on more prominent roles in shaping the digital world, acknowledging that they have historically been underrepresented in these fields. As societal attention shifts toward greater inclusion, our project will contribute to fostering a more balanced and equitable digital future.
While many individuals in our target groups may lack in-depth technical knowledge of AI, they often encounter new AI products, companies, and social issues through various media channels, including newspapers and science fiction movies. As a result, they may be aware of recent developments but also susceptible to misunderstandings and controversies related to technologies such as ChatGPT, Elon Musk's brain-chip startup, and other emerging AI applications. It is crucial to recognize that media portrayal significantly influences public opinion on AI, both positively and negatively. Media creators, even if they are not experts in AI, often produce content that captures public attention, which high-profile figures, including entrepreneurs, CEOs, and politicians, may leverage to advance their agendas. This can sometimes lead to skewed public perceptions, whether intentionally or unintentionally. Given this landscape, it is essential for AI scientists to collaborate with media creators, providing evidence-based insights to ensure accurate and balanced information is shared with the public. Our project fosters such collaboration, ensuring that both the potential and limitations of AI are clearly communicated. By sharing our findings through diverse media outlets, we aim to reach a broad audience, extending beyond Switzerland. Furthermore, our proactive engagement efforts will foster dynamic, two-way communication between scientists and the public, using interactive methods in exhibitions and theater plays to engage teenagers and female students specifically. Analyzing the feedback from these initiatives will provide invaluable insights into public perspectives on emerging technologies. This understanding will guide scientists in pursuing research directions that effectively address societal concerns, demonstrating the tangible benefits of our project for both scientific advancement and societal well-being. We anticipate that our efforts will have a multiplying social impact over time, promoting informed public discourse and a deeper understanding of AI technologies.
ELOQUENCE is focused on the research and development of innovative technologies for collaborative voice/chat bots. Voice assistant-powered dialogue engines have previously been deployed in a number of commercial and governmental technological pipelines, with a diverse level of complexity. In our concept, such a complexity can be understood as a problem of analysing unstructured dialogues. ELOQUENCE’s key objective is to better comprehend those unstructured dialogues and translate them into explainable, safe, knowledge-grounded, trustworthy and bias-controlled language models. We envision to develop a technology capable of learning by its own, by adapting from a very data-limited corpora to efficiently support most of the EU languages; from a sustainable computational framework to efficient and green-power architectures and, in essence, that may serve as a guidance for all European citizens whilst being respectful and showing the best of our European values, specifically supporting safety-critical applications by involving humans-in-the-loop.
Overall, ELOQUENCE’s project considers building on top and to improve of prior achievements in the domain of conversational agents, e.g. recently launched and public-domain Large Language Models (LLMs), such as chatGPT (e.g., more recent versions), or LaMDa most of them developed in non-EU countries. While including key industrial enterprises from Europe (i.e., Omilia, Telefonica, Synelixis), ELOQUENCE will validate the developed technology through (i) safety-critical scenarios with human-in-the-loop for security-critical applications (i.e., emergency services in call centres) and (ii) smart home assistants via information retrieval and fact-checking against an online knowledge base for lesser risky autonomous systems (i.e., home-assistants). ELOQUENCE will target the R&D of these novel conversational AI technologies in multilingual and multimodal environments and demonstrated in several pilots.
The ASR system will predict controller commands to reduce the
recognition engine’s search space resulting in reduced command
recognition error rates and checks the recognition output for
plausibility.
The recent emergence of Artificial Intelligence (AI) methods and audio signal processing techniques open new perspectives on the use of voice to detect or monitor diseases. The vocal biomarker research field is booming but is facing, like other AI-driven fields, a reproducibility and generalisability crisis. To move this promising field of research forward, there is an urgent need to develop a common research framework in Europe, with standardization principles and definitions of good practices and guidelines to help it reach its full potential. A multidisciplinary approach is necessary to tackle this complex task impacting all stakeholders in healthcare virtually.
eVoiceNet aims to establish Europe as a leader in vocal biomarkers, create a collaborative and multidisciplinary network, and facilitate knowledge sharing, standardisation, and the development of privacy-aware solutions, by creating an international network of clinicians, AI experts, speech/voice processing specialists, voice pathologists, privacy/data protection experts, go-to-market specialists, venture capitalists and other providers of financial resources, patients organisations, policymakers as well as industrial partners and other stakeholders (end users, regulators, and other decision-makers), to overcome the major challenges and boost the integration of voice technologies into clinical practice.
Artificial Intelligence (AI) has become a powerful and pervasive technology in recent years, influencing numerous aspects of our daily lives. We encounter AI through recommendation algorithms in online stores, voice-activated smartphone assistants, or the widespread use of technologies like ChatGPT. However, its rapid growth and integration into society raise complex questions and concerns among the public. Public opinions on AI vary widely; while some people are enthusiastic about its potential to revolutionize industries and enable breakthroughs in fields such as medicine, others fear that AI could lead to undesirable outcomes, such as a loss of human control or privacy. These views are often shaped by media narratives and the competing interests of different stakeholders, which play a significant role in influencing both public opinion and policy decisions.
Our team of scientists and communication experts aims to enhance the understanding of AI technologies among the Swiss people, with a particular focus on teenagers and female students, to create a positive societal impact. Building on the foundations of our previous project, NewsOnAI, we will expand beyond traditional media such as newspapers and employ diverse methods, including artistic performances and interactive exhibitions. We will design these activities to be highly interactive, encouraging active participation and dialogue. Activities will include themed theater plays that explore AI’s impact on everyday life, exhibitions where participants can interact with AI tools, and workshops specifically designed for teenagers and female students to discuss AI’s future role in society. Feedback collected will include real-time audience reactions, structured questionnaires, and focus group discussions, which will be analyzed to continuously refine and adapt our engagement strategies.
Our primary audience includes Swiss citizens interested in cultural activities, particularly teenagers who are keen to follow new trends. Additionally, we are committed to addressing gender aspects by designing content and activities that specifically appeal to female students. We aim to inspire and empower young women to take on more prominent roles in shaping the digital world, acknowledging that they have historically been underrepresented in these fields. As societal attention shifts toward greater inclusion, our project will contribute to fostering a more balanced and equitable digital future.
While many individuals in our target groups may lack in-depth technical knowledge of AI, they often encounter new AI products, companies, and social issues through various media channels, including newspapers and science fiction movies. As a result, they may be aware of recent developments but also susceptible to misunderstandings and controversies related to technologies such as ChatGPT, Elon Musk's brain-chip startup, and other emerging AI applications. It is crucial to recognize that media portrayal significantly influences public opinion on AI, both positively and negatively. Media creators, even if they are not experts in AI, often produce content that captures public attention, which high-profile figures, including entrepreneurs, CEOs, and politicians, may leverage to advance their agendas. This can sometimes lead to skewed public perceptions, whether intentionally or unintentionally. Given this landscape, it is essential for AI scientists to collaborate with media creators, providing evidence-based insights to ensure accurate and balanced information is shared with the public. Our project fosters such collaboration, ensuring that both the potential and limitations of AI are clearly communicated. By sharing our findings through diverse media outlets, we aim to reach a broad audience, extending beyond Switzerland. Furthermore, our proactive engagement efforts will foster dynamic, two-way communication between scientists and the public, using interactive methods in exhibitions and theater plays to engage teenagers and female students specifically. Analyzing the feedback from these initiatives will provide invaluable insights into public perspectives on emerging technologies. This understanding will guide scientists in pursuing research directions that effectively address societal concerns, demonstrating the tangible benefits of our project for both scientific advancement and societal well-being. We anticipate that our efforts will have a multiplying social impact over time, promoting informed public discourse and a deeper understanding of AI technologies.
ELOQUENCE is focused on the research and development of innovative technologies for collaborative voice/chat bots. Voice assistant-powered dialogue engines have previously been deployed in a number of commercial and governmental technological pipelines, with a diverse level of complexity. In our concept, such a complexity can be understood as a problem of analysing unstructured dialogues. ELOQUENCE’s key objective is to better comprehend those unstructured dialogues and translate them into explainable, safe, knowledge-grounded, trustworthy and bias-controlled language models. We envision to develop a technology capable of learning by its own, by adapting from a very data-limited corpora to efficiently support most of the EU languages; from a sustainable computational framework to efficient and green-power architectures and, in essence, that may serve as a guidance for all European citizens whilst being respectful and showing the best of our European values, specifically supporting safety-critical applications by involving humans-in-the-loop.
Overall, ELOQUENCE’s project considers building on top and to improve of prior achievements in the domain of conversational agents, e.g. recently launched and public-domain Large Language Models (LLMs), such as chatGPT (e.g., more recent versions), or LaMDa most of them developed in non-EU countries. While including key industrial enterprises from Europe (i.e., Omilia, Telefonica, Synelixis), ELOQUENCE will validate the developed technology through (i) safety-critical scenarios with human-in-the-loop for security-critical applications (i.e., emergency services in call centres) and (ii) smart home assistants via information retrieval and fact-checking against an online knowledge base for lesser risky autonomous systems (i.e., home-assistants). ELOQUENCE will target the R&D of these novel conversational AI technologies in multilingual and multimodal environments and demonstrated in several pilots.
The ASR system will predict controller commands to reduce the
recognition engine’s search space resulting in reduced command
recognition error rates and checks the recognition output for
plausibility.
The recent emergence of Artificial Intelligence (AI) methods and audio signal processing techniques open new perspectives on the use of voice to detect or monitor diseases. The vocal biomarker research field is booming but is facing, like other AI-driven fields, a reproducibility and generalisability crisis. To move this promising field of research forward, there is an urgent need to develop a common research framework in Europe, with standardization principles and definitions of good practices and guidelines to help it reach its full potential. A multidisciplinary approach is necessary to tackle this complex task impacting all stakeholders in healthcare virtually.
eVoiceNet aims to establish Europe as a leader in vocal biomarkers, create a collaborative and multidisciplinary network, and facilitate knowledge sharing, standardisation, and the development of privacy-aware solutions, by creating an international network of clinicians, AI experts, speech/voice processing specialists, voice pathologists, privacy/data protection experts, go-to-market specialists, venture capitalists and other providers of financial resources, patients organisations, policymakers as well as industrial partners and other stakeholders (end users, regulators, and other decision-makers), to overcome the major challenges and boost the integration of voice technologies into clinical practice.