Skip to content
Leading Medicine Guide logo

Expert Interviews

Artificial Intelligence in Otolaryngology

Alexandra_Pfitzmann.jpg

Alexandra Pfitzmann · March 2, 2026

Artificial intelligence (AI) has become one of the most dynamic fields of innovation in medicine. It helps doctors identify patterns in vast amounts of data, refine diagnoses, tailor treatment plans more precisely, and make procedures safer. From radiology to oncology to surgery, AI is changing the way diseases are understood and treated—often quietly in the background, but with tangible benefits for patients. It is particularly exciting to see how deeply these technologies have now penetrated the field of otolaryngology, where they are taking diagnostics, navigation, and surgical precision to a new level.

The editorial team of the Leading Medicine Guide spoke with ENT specialist Professor Dr. med. Marco Domenico Caversaccio about this topic.

Professor Marco Domenico Caversaccio, M.D.Copyright "Gianni Pauciello, Inselspital ENT Clinic"

AI technology is primarily used for research, text processing, and similar tasks. AI has long since become a business—no longer an experiment—and the major providers are now earning substantial profits from it. At the same time, the boundaries between artificial intelligence, reality, and virtual content are becoming increasingly blurred.

“Despite all its possibilities, working with AI remains challenging. We must not rely on it blindly, because the human touch is missing. Direct conversation yields different insights, nuances, and follow-up questions—something machines cannot replace. Moreover, a lot of nonsense circulates on the internet, which makes a critical approach all the more important. It’s interesting to see how earlier predictions have been put into perspective. About twenty years ago, for example, it was claimed that radiology would be largely automated by AI. In fact, the number of radiologists in the U.S. has risen by four percent. There, AI has tended to create new business opportunities rather than replace jobs. Technical developments remain fascinating nonetheless, especially in the medical field. In otolaryngology as well, the question arises as to how AI can help in the future—for example, in the earlier or more precise detection of tumors in the head and neck region through pattern recognition in imaging or pathology. Initial approaches already exist, and researchers are continuing to work intensively on this,” notes Prof. Dr. Caversaccio at the beginning of our conversation.

AI can detect tumors in the head and neck region earlier and more precisely because, at several stages of the diagnostic process, it possesses capabilities that far exceed what the human eye or traditional evaluation methods can achieve. In imaging, AI models analyze CT, MRI, or PET scans not just superficially, but mathematically, point by point. 

“This data holds enormous potential for digital biopsies and pattern recognition, which continue to evolve year after year. The goal is to link radiological information with pathological data to detect tumors earlier and more accurately. If superficial or small tumors could be reliably identified through imaging, invasive procedures such as panendoscopy might become less necessary in the future. This is not yet a reality, but various research groups—such as those in Würzburg and Heidelberg—are working intensively to integrate radiomics, pathomics, and big data analytics. The idea behind this is clear: the more data a system processes, the better it can recognize patterns. Similar to large data platforms, AI models learn from each new examination. Nevertheless, practical implementation is still in its infancy. Many concepts are technically impressive, but the transition to broad clinical application is complex. Tumors differ greatly from one another, which makes it difficult to develop universal models. In some areas, such as the planning and fitting of cochlear implants, progress has already been made—there, diagnostic data and intraoperative strategies can already be more closely integrated today, explains Prof. Dr. Caversaccio.


Radiomics, pathomics, and big data analytics represent a new approach in modern medicine: they transform image and tissue data into highly precise, actionable information. Radiomics extracts thousands of features from CT, MRI, or PET images and reveals patterns that allow conclusions to be drawn about tumor biology, aggressiveness, or response to therapy. Pathomics operates on the same principle, but based on digitized tissue samples in which AI detects the finest cellular structures and changes. Big data analytics link this information with genetic data, laboratory values, and clinical courses. This creates a comprehensive data profile that makes diagnoses more precise, enables more personalized treatment planning, and can help avoid unnecessary procedures in the long term.


AI is fundamentally changing preoperative planning in ENT surgery because it can extract significantly more information from complex image data than would be possible through human interpretation alone.

Photo: Robot-Assisted Electrode Implantation with OTODRIVE and OTOARM
Copyright "Gianni Pauciello, Inselspital ENT Clinic"

For procedures at the skull base or in the middle ear—where millimeters determine safety and the preservation of function—AI analyzes radiological images with such precision that even the finest anatomical variations, high-risk structures, or potential narrowings become visible. This results in virtual models that precisely map the individual course of nerves, blood vessels, or air-filled spaces, making planning significantly more reliable. Especially in the middle ear, where the ossicles, the facial nerve, and the inner ear are extremely close together, this greatly increases surgical safety.

Prof. Dr. Caversaccio explains: “Various digital methods are used in surgical planning today, which are primarily based on precise segmentation. Structures such as the vestibular nerve, the facial nerve, or the auditory nerve can thus be visualized in detail—partly semi-automatically, partly fully automatically, depending on the system. This information can either be mentally incorporated into the surgical strategy or fed directly into a navigation system. During the procedure, it is then possible to determine with millimeter precision how close one is to sensitive structures. This requires highly precise navigation systems, because as soon as one moves away from the bone and into mobile areas such as the subarachnoid space, accuracy drops to about two millimeters—a factor that must always be taken into account. Such navigation and segmentation methods are routinely used in everyday clinical practice, often in close collaboration with neurosurgery. In addition, there are other tools such as rapid-prototyping models that visualize complex skull base situations in three dimensions and assist in selecting the optimal approach. Even though these models are not considered AI, they significantly facilitate spatial orientation. At the same time, digital concepts such as simulations, virtual training systems, and so-called digital twins continue to evolve. They make it possible to plan surgical approaches in advance, identify risks, or control robotic systems with greater precision. In the future, augmented reality technologies could display additional information directly in the surgeon’s field of view—similar to modern cockpits. The increasing digitization of the operating room environment is leading to ever-greater amounts of available data that must be meaningfully integrated: imaging, navigation, physiological measurements, or AR overlays. Despite all these developments, one thing remains clear: the responsibility always lies with the surgeon. Digital systems can support, visualize, and warn—but they do not replace surgical experience, anatomical understanding, and situational decision-making in the operating room.”

For many surgeons, the increasing use of technology in the operating room can indeed be overwhelming. The older generation, in particular, is strongly shaped by the principle of “learning by doing”—through anatomical specimens, hands-on experience, and an intuitive spatial understanding. Digital systems, simulations, and AI-powered tools are fundamentally changing this world. Nevertheless, one thing remains clear: Good surgery continues to rely on anatomical knowledge, spatial awareness, and fine motor precision. Some people possess a highly developed 3D visual perception and a high degree of tactile sensitivity—skills that even the most advanced technology cannot replace.

Navigation systems support many procedures today, but they remain passive tools. They display positions, much like a car navigation system, but do not take over active control. The next step would be systems with sensor-actuators that detect tissue, measure forces, and provide feedback—but this technology is still in its infancy. At the same time, augmented reality solutions, intraoperative imaging, and digital cockpit interfaces are being developed, bringing more and more information into the operating room. The challenge lies in integrating this data effectively without overwhelming the surgeon. At the same time, reliance on technology carries risks. Just as with driving, where many people can barely find their way without a navigation system, a surgeon’s own expertise can atrophy if they rely too heavily on devices. That is why basic surgical training remains just as important as ever: knowledge of anatomy, hands-on practice with models or specimens, and the development of manual skills continue to form the foundation. While the younger generation is, of course, growing up in a digital world, this does not replace the need to train one’s own brain and be able to make independent decisions. AI systems can recognize patterns, compare data, and provide support—but they rarely create anything new. Emotional intelligence, intuition, and creative problem-solving remain human strengths. This is also evident in fields such as tinnitus research, where machine-learning models—despite years of work—can barely extract reliable patterns from EEG data because individual perceptions vary so widely. At the same time, there are fields where digital support is already enabling major advances today: for example, in cochlear implants, where diagnostics, intraoperative measurements, and postoperative fitting strategies can be coordinated with ever-greater precision. This is where true personalized medicine is taking shape. Overall, digital systems, AI, and robotics serve as valuable add-ons—as support, not as a replacement. The responsibility remains with the surgeon, who decides, interprets, and acts,” says Prof. Dr. Caversaccio, adding:

“Through modern measurement techniques such as impedance analysis (a method for measuring the electrical resistance of tissue or electrodes) and neural response measurements (methods used to determine whether the auditory nerve responds to electrical stimuli), it is already possible during cochlear implant surgery to estimate what level of hearing may be achievable later on. The cochlea has a tonotopic structure, which facilitates the mapping of frequencies—from low 125 hertz up to 8 kilohertz. This structure makes it possible to gain early indications of future function based on intraoperative data. Whether AI leads to fewer complications overall depends heavily on the specific field of application. In some areas, studies already exist that show AI-supported pattern recognition can help refine diagnoses or avoid unnecessary procedures. In the case of chronic rhinosinusitis, for example, radiological patterns could in the future provide clues to specific causes such as cystic fibrosis, thereby shortening the diagnostic process. AI could then help determine whether a biopsy is necessary, which therapy appears appropriate, or whether surgery can be avoided. Such systems are capable of deriving recommendations from large amounts of data—for example, which medications would be appropriate or which treatment strategy promises the greatest benefit. This creates an additional decision-making framework that can support everyday clinical practice without replacing medical responsibility.”

AI-supported navigation systems offer significant added value during endoscopic sinus surgeries because they not only display static image data but also actively interpret it and compare it in real time with the patient’s individual anatomy. 

“Practical training in the use of modern surgical technology traditionally begins with anatomical specimens and models. For procedures such as those on the paranasal sinuses, special training systems are available, such as the FACON model or anatomical specimens as prescribed by the FMH (Foederatio Medicorum Helveticorum, the Swiss Medical Association). Only once these fundamentals have been mastered can surgeons gradually begin working in the operating room. Navigation systems have been standard there for over twenty years, but they require technical understanding: Accuracy must be checked regularly, and surgeons must be able to assess whether the system is functioning reliably. As experience grows, additional digital tools are introduced—such as 3D visualizations, the marking of tumors or critical structures like the optic nerve, or features like “Navigator Control,” which protect specific areas. This level of proficiency is typically not reached until after three to four years. In cases of complex tumors, particularly at the base of the skull, multiple specialties collaborate, such as neurosurgery and oral and maxillofacial surgery. Digital planning also plays a major role here: In mandibular reconstructions, the required bone structure—often the fibula—is digitally measured and planned in advance and later precisely adjusted intraoperatively. Such planning and navigation aids facilitate orientation and increase precision, but are not used for every patient. Some procedures, such as digital planning for fibula grafts, are covered by health insurance and are therefore well established. Other technologies are used only in complex cases. Overall, however, it is clear that imaging, simulations, and digital planning tools are becoming increasingly important and serve as a valuable complement to surgical work, explains Prof. Dr. Caversaccio.


AI can generate precise 3D models from CT and MRI data, mark anatomical variations, and automatically highlight high-risk structures. This results in significantly more intuitive and safer navigation in the operating room. During the procedure, the system detects even the slightest deviations from the planned instrument path and provides early warnings of critical areas, thereby reducing complications. At the same time, the AI learns from large amounts of data to predict typical danger zones and dynamically adapts the navigation model to changes such as bleeding or mucosal swelling—an advantage that traditional static image datasets do not offer.


AI can identify intraoperative risks in ENT surgery much earlier and more precisely because it continuously analyzes a wide range of data sources during the procedure and derives indications of potential hazards that would often only become apparent later to the human eye.

Photo: Electrocochleography - Electrical Measurement of Cochlear Activity (Inner Ear)
Copyright "Gianni Pauciello, ENT Clinic, Inselspital"

During surgeries at the skull base or in the immediate vicinity of sensitive structures such as the facial nerve, the cochlear nerve, or the internal carotid artery, AI compares the current instrument position in real time with the 3D model created preoperatively and detects even minimal deviations from the planned path. As soon as an instrument approaches a critical structure, the system can issue a warning before the structure even appears in the endoscopic image. At the same time, AI is capable of automatically identifying anatomical landmarks intraoperatively—even when bleeding, mucosal swelling, or tissue displacement make orientation difficult.

“Whether AI can actually prevent complications remains to be seen. In some areas, digital methods are already helping to better assess risks—for example, through automatic or semi-automatic segmentations that make nerves, vessels, or other sensitive structures visible. These visualizations support planning and improve orientation in the operating room, but they do not replace surgical experience. Similar to plastic surgery, where preoperative photos are used to visualize potential outcomes, AI-supported models can also provide predictions. But as with a rhinoplasty, the reality remains complex: soft tissues, facial expressions, and individual healing processes cannot be fully predicted. A plan may be 70 or 80 percent accurate—absolute precision is virtually unattainable. AI can also support perioperative management, for example, in selecting appropriate medications or assessing risks. In the future, such systems could provide recommendations based on comparable cases and large datasets. However, many of these approaches are still in the “proof of concept” stage. The key hurdle is translating them into approved medical devices. Regulatory requirements are stringent: systems must demonstrate very high sensitivity and specificity before they can be approved. It becomes particularly complex when AI is combined with robotics—here, strict Class III MDR regulations apply. For such applications, the large datasets required for approval are often lacking. As a result, many AI prototypes remain in the research phase, while only a few make their way into routine clinical practice. Low-threshold digital tools, simulations, or training apps, on the other hand, are widely used because they do not face high regulatory hurdles. However, true AI-supported systems that influence surgical decisions or guide surgical procedures remain rare. The technology is advancing rapidly, but the path to practical application is challenging—especially in highly sensitive areas such as skull base surgery, where maximum precision is essential,” notes Prof. Dr. Caversaccio.


To ensure that AI-supported systems in ENT medicine do not lead to overreliance, their role in everyday clinical practice must be clearly defined from the outset: support, yes; decision-making authority, no. A key point is transparency in both technical implementation and content.


AI is playing an increasingly central role in cochlear implantations because it improves both the technical precision of the procedure and the individual hearing prognosis.

Photo: Robot-Assisted Cochlear Implantation HEARO (MED-EL)
Copyright “Gianni Pauciello, Inselspital ENT Clinic”

In the field of hearing aids, AI already plays a significant role today. Modern systems can selectively filter out background noise, emphasize speech signals, and create individual sound profiles using neural networks. The devices adapt to pitch, environments, and personal hearing preferences—from amplifying specific frequencies to optimizing the perception of music or speech. Development is advancing rapidly, and the industry is working intensively to make hearing aids and speech processors increasingly personalized. Since this technology is non-invasive, it reaches the market relatively quickly and offers many people a significant improvement in quality of life. At the same time, challenges remain, such as in cases of unilateral hearing loss or in tinnitus research. Tinnitus is extremely individual, which makes it difficult to develop reliable AI-based solutions. For unilateral hearing loss, crossover systems or cochlear implants are frequently used today, provided the necessary conditions are met. Yet here, too, it is evident that while many ideas exist and numerous studies are underway, implementation remains complex. One particularly active field of research concerns the voice. The vocal folds and phonetics provide valuable clues to conditions such as diabetes, neurodegenerative diseases, or early tumor changes. Stroboscopy, video analysis, and AI-based pattern recognition are opening up new diagnostic possibilities that are currently being intensively investigated. The combination of imaging, acoustic data, and AI analysis opens up exciting prospects. At the same time, other technologies are developing, such as the fusion of CT and MRI data, genomics, and pathomics. Robotics also remains an important topic—though the decisive step has not yet been achieved: an AI that independently suggests the optimal surgical approach, assesses risks, and reliably prevents complications. In medicine, this is far more difficult than in industrial applications because every procedure on a human being involves individual variables,” Prof. Dr. Caversaccio points out, concluding our conversation by emphasizing:

“Robotic systems therefore continue to serve as assistants, not as autonomous actors. They provide support, limit risks, and increase precision, but they do not perform entire surgeries. Responsibility remains with humans. At the same time, increasing digitalization is changing everyday work: Information systems, standard processes, and software solutions are increasingly dictating how procedures function. This simplifies many tasks, but it also means that people must adapt to the logic of the systems. In a world that is becoming increasingly standardized, there is a growing risk of becoming a bit more ‘robotic’ ourselves. Nevertheless, one thing remains clear: In medicine, people are at the center—both as patients and as decision-makers in the operating room. Technology can support, structure, and enhance precision, but it does not replace the surgeon’s experience, intuition, and responsibility.”

Thank you very much, Professor Dr. Caversaccio, for these informative insights into AI in the field of ENT!

Share this article

Alexandra_Pfitzmann.jpg

About the medical author

Alexandra Pfitzmann

Editor

Alexandra Pfitzmann – medical author: expert knowledge, professional articles and medical insights in the Leading Medicine Guide.

More about the medical author

Expert Interviews

Read next

Portrait of Prof. Dr. med. Marco Domenico Caversaccio

Prof. Dr. med. Marco Domenico Caversaccio

Bern