Human-centered AI

Including social computing, activity understanding, multi-modal processing, computer vision, assistive technology, human-robot interaction, robot learning & interaction, and demographic fairness.

Ongoing projects

ALIGNAI

Large Language Models (LLMs) are trained on broad data, using self-supervision at scale, to complete a wide range of tasks. Wider use of LLMs has risen in recent months due to applications such as ChatGPT. Although LLMs bring many opportunities to improve our everyday lives, the impacts on humans and society have not yet been prioritized or fully understood. Given the rapid development of these tools, the risk of negative implications is significant if LLMs are not developed and deployed in a way that is aligned with human values and responds to individual needs and preferences. To mitigate any negative consequences, academia, in close collaboration with industry, needs to train the next generation of researchers to understand the complexities of the socio-technical implications surrounding the use of LLMs.
The alignAI Doctoral Network will train 17 doctoral candidates (DCs) to work in the international and highly interdisciplinary field of LLM research and development. The core of the project focuses on the alignment of LLMs with human values, identifying relevant values and methods for alignment implementation. Two principles provide a foundation for the approach. First, explainability is a key enabler for all aspects of trustworthiness, accelerating development, promoting usability, and facilitating human oversight and auditing of LLMs. Second, fairness is a key aspect of trustworthiness, facilitating access to AI applications and ensuring equal impact of AI-driven decision-making. The practical relevance of the project is ensured by three use cases in education, positive mental health, and news consumption. This approach allows us to develop specific guidelines and test prototypes and tools to promote value alignment. We follow a unique methodological approach, with DCs from social sciences and humanities “twinned” with DCs from technical disciplines for each use case (9 DCs in total), while the other 8 DCs carry out horizontal research across the use cases.

CAR-LORO

Au cours des vingt dernières années, les robots ont progressivement quitté les environnements industriels classiques pour entrer dans les espaces de vie et de travail des humains. De nouveaux robots ont ainsi été conçus spécifiquement pour interagir avec des personnes, ce qui implique des capacités avancées de manipulation, de déplacement autonome, de programmation intuitive, et d’interaction sociale afin d’être réellement utiles et acceptés par leurs utilisateurs. Ces évolutions ont donné naissance à la robotique d’assistance, un domaine de recherche qui étudie comment les robots peuvent soutenir les humains dans des contextes variés, allant de l’aide aux personnes âgées ou en situation de handicap jusqu’au soutien des travailleurs en entreprise. Toutefois, cette recherche pose des défis importants, notamment l’implication de patients et d’utilisateurs réels dans la conception des technologies, ainsi que la création d’environnements de développement et d’évaluation qui reproduisent fidèlement les conditions réelles de vie et d’activité. Pour renforcer le réalisme et l’impact de ses futurs projets, l’Idiap souhaite développer un nouveau Centre de Robotique d’Assistance, destiné à centraliser les recherches en robotique de l’Idiap, à recréer des conditions d’usage proches du réel, à accueillir des participants locaux et à mieux communiquer les avancées scientifiques auprès du public.

CHASPEEPRO

Oral verbal communication represents the main communication channel among humans. In most communication contexts, speakers must speak clearly and accurately in order to be intelligible. Intelligible speech can be disrupted in a variety of conditions of motor speech disorders (MSD). MSD in adults refers to a broad set of altered speech dimensions (articulation, speech rate, voice, prosody) in the course of several neurological diseases, which can dramatically impact patients’ communication. MSDs are due to disruption in the processes transforming a linguistic message intoarticulated speech, i.e., (a) the retrieval/encoding, contextualization and coordination of speech goals into a speech plan, (b) the preparation of motor programs with detailed neuromuscular specifications, and (c) the execution of these programs. Impairments atthese different stages have been associated with different MSDs, with apraxia of speech(AoS) associated with impairments at the first stage, i.e., the planning stage, and dysarthria associated with impairments at the programming or at the execution stage. Nonetheless, defining planning and programming stages, as well as distinguishing impairments at these two levels in terms of speech features and clinical differential diagnosis, is far from being clear-cut. This proposal builds on the successful outcomes of the Sinergia MoSpeeDi project (2017-2021, https://www.unige.ch/fapse/mospeedi/) led by the same multidisciplinary consortium. Thanks to the complementary expertise in speech and language pathology, psycholinguistics, neurology, phonetics, and speech engineering, we have collected an impressive database of MSD speech, have developed procedures sensitive enough to assess and classify mild and moderate MSD, and have obtained converging experimental evidences for the characterization of processes occurring at the planning and motor programming stages. This knowledge gained from carefully designed experiments and laboratory settings should now be expanded to speech production elicited in a more natural clinical setting. The distinction between speech planning and programming processes should also be further tackled to overcome the difficulty in defining and operationalizing processes at these two stages. With the overarching goal of understanding and modelling speech planning and programming and their related disorders, we will pursue our synergic approach based on the integration of methods and on the convergence of evidence obtained with experimentally induced speech behaviours, electrophysiological brain signals, and acoustic analyses of typical and impaired speech. Based on the results and expertise developed in the ongoing project to pinpoint speech planning and programming and to classify speakers and speech samples, in this project we propose to (a) develop assessment and classification methods applicable to realistic clinical constraints and needs, (b) build on the convergence of phonetic knowledge-based approaches and knowledge-free approaches, (c) enrich our set of acoustic descriptors in order to capture alterations at different scales of speech organization, and (d) complement acoustic-based characterization of speech planning, programming, and MSD classification with EEG signals. The outcomes of the project will rely on substantial data of disordered speech collected from over 180 French speaking participants with different types of MSDs including AoS and subtypes of dysarthria following stroke or neurodegenerative diseases. Results will be used to challenge current models of speech production which need to integrate data from MSD and will contribute to the development of speech assessment systems adapted to atypical speech and to the needs of clinical practice.

DEMO-AI

Context: Access to factual information is essential for democratic decision-making, public trust, and civic engagement, yet artificial intelligence (AI) enables large-scale creation and dissemination of manipulated content, fabricated narratives, and content amplification that can distort public perception, erode confidence in democratic institutions, and polarize political discourse. These risks threaten to reshape political debates, influence electoral outcomes, and undermine public trust in media sources in Switzerland. Democratic values can be upheld by developing AI tools and governance frameworks to counter disinformation and monitor media framing.


Goals: DEMO-AI is an interdisciplinary research project, driving advances in computing to enhance the resilience of democracy, integrating expertise from law, journalism and communication studies, media and information literacy to ensure that AI-supported solutions align with democratic values and regulations. Four project goals include: AI tools for analyzing news media framing; AI tools for detecting manipulation of audio-visual media; legal research on regulatory frameworks for AI and disinformation in Switzerland; and engaging both the public and professionals in evaluating and testing media tools.


Expected Impact: DEMO-AI will produce tools to analyze issue framing and related narratives in Swiss media, facilitate the detection of audio-visual disinformation, and understand legal challenges. These tools will be designed, tested, and refined in collaboration with the general public and professionals, placing their specific needs at the center, thus ensuring real-world applicability. Through societal impact activities, the project extends beyond technology, addressing key challenges across AI, democracy, and policy.

Past projects

2000LAKES

Alpine lakes (those located above the 2000 m tree line) are excellent sentinels of climate change as their chemistry and biology respond rapidly to environmental forcing. The Swiss alps are host to over 1500 alpine lakes, many of which have been newly mapped and thus never been studied2. Microorganisms play major ecological roles in these ecosystems, including primary production, cycling of elements, and attenuation of contaminants, but it is uncertain how physical climatic changes may affect microbial communities and their activities in alpine lakes. This project aims to: (i) record and monitor the unexplored microbial diversity in Swiss alpine lakes, and (ii) engage citizens in science and spread awareness about environmental conservation through participation in our field campaigns. In summary, 2000LAKES is a project of alpine citizen science aiming to understand the ecological impacts of climate change in alpine lakes and to promote the conservation of alpine microbial ecosystems joining forces between scientists and citizens.

3D2CUT

The main objective of this project is to verify the feasibility of using innovative artificial intelligence algorithms to analyze vine based on vineyard images. The aim is to automatically extract the essential components and use them to recommend appropriate pruning.

ADEL

The goal of the here submitted proposal is to finance a first year of research as a concrete first step towards the creation of the Center for Leadership and New Technologies (Unil, Idiap/EPFL, IMD). The Center that we aim to create long term will include AI and virtual reality among other technologies in relation to leadership. The Center will develop tools for assessing and developing leadership, conduct research with respect to new technologies related to leadership, as well as showcase our developments and empirical results for the corporate world (e.g., writing white papers, organizing symposia and conferences). Ideally, firms would turn to the Center for advice, training, and thought leadership on the topic of new technologies and leadership. IMD will be crucial in creating the link with companies and will be able to use the new technologies for their teaching and training. A first concrete project for which we ask for seed funding from the Trans4 consortium concerns the development of a collection of software modules that will be able to automatically detect leadership skills from videotaped speeches using voice and body language information. The algorithms developed will be able to automatically detect perceived leadership based on voice and video samples. We will train an algorithm to infer leadership (e.g., trustworthiness, competence as a strategic leader, competence as a transformational leader etc.) automatically based on vocal cues and body language automatically detected by the machine. We will train the algorithms with ground truth data that we will collect from a panel of evaluators (e.g., MTurk workers) on either selfpresentation videos (e.g., video CVs on YouTube) or on public speaking videos (e.g., TED Talks). Given that the quality of the algorithm depends on the quality of the training data (i.e., ground truth), we will put extra care and effort in producing this training data. The so developed software modules can then be used for leadership skill assessment and for leadership skill training and development. It can be seen as a stand‐alone outcome but at the same time it can be incorporated to the Charismometer algorithm that John and Philip have already developed and it can be added to work Daniel and Marianne have been doing in the past (on automatic extraction of nonverbal behavior from video). Basing the new development on existing work ensures that we do not start from scratch and that we can achieve the goal within one year of funding. The seed money project is thus at the same time a continuation of existing work and an important extension of it.

ADIVA

This project investigates social interaction in personnel selection interviews enhanced by digital technology. We will create a database of applicants participating in video interviews (applicants receive a list of interview questions from a recruiter online and then record themselves answering those questions), which are a newly emerging interview format. We will develop automated procedures for extracting relevant behavioral features from streams of applicants’ verbal and nonverbal behavior in these interviews. This information will be (1) linked to external criteria (e.g., hireability ratings by expert recruiters), (2) used to train machine learning algorithms, and (3) fedback to the applicants. We will assess applicants’ perceptions of this feedback, whether and how they use it to improve their performance in a second video interview a day later, and how they perceive data privacy issues related to the use of their data. The project addresses three issues mentioned in the call. First, how is digitalization transforming social ties? The selection interview is the gateway to employment and thus the potential beginning of one of the fundamental social ties in modernity: the work relationship. We explore a new format by which selection interviews are conducted in an online, asynchronous manner. Second, how is digitalization transforming the economy? The selection interview is an important personnel selection procedure, which itself is an important component of strategic talent management. The digitalization of talent management is rapidly expanding in practice, but is currently poorly understood in research. Third, how is digitalization transforming our subjective experience? Video interviews are a novel experience for many applicants. Machine learning techniques can be used to extract the applicants’ behaviors recorded on the videos and to some degree infer their personality and social skills. This information can then be fed back to the applicants, potentially changing their subjective experience of the video interview. However, questions like how such feedback is best provided and how the applicant apprehends and uses it are largely unexplored. The study will yield four main sets of outputs. First, the primary data from the study will lead to publications in scientific journals or conference proceedings in human-computer interaction and organizational psychology or human resources. Second, the data will be used to adapt an existing data collection platform and to improve the quality of algorithms to infer verbal and nonverbal behavior from videos. Third, data about user experiences will inform the development of evidence-based coaching programs for improving applicants’ performance. Fourth, the rich set of data and experience generated will constitute fruitful avenues for further research by the applicant team.

Don't miss a Step - Join us
Whether you want to join our team, become part of our community, support us through a donation, or explore a partnership, you’ll find all the ways to connect with us right here.