ITT Italia
“Innovate to include”: this need gave rise to the new project developed by Cloudia Research for ITT Italia Srl, a leading company in the mobility technology sector and part of the international ITT group, present in over 35 countries with more than 11,700 employees.
ITT’s mission is summed up in its motto, “We solve it”: transforming unique ideas into innovative and sustainable solutions, working alongside customers to solve their daily challenges.
Designs and develops customized technological solutions for the transportation and industrial markets, with a strong focus on research and development.
Intelligent transcription recognising the speaker
It is precisely this innovation culture that drove ITT to come up with a concrete solution to improve accessibility and inclusion within its business processes.
The idea arose from a real need for inclusion: an employee with hearing impairment who, during company meetings, had difficulty identifying who was speaking in the automatic transcripts.
This specificity led to the creation of a web app based on speech-to-text technology, designed to automatically recognize the speaker based on context, without any voice sampling.
A solution that combines artificial intelligence and inclusion, improving communication and making the experience more accessible to everyone.
What it is and why it’s a game changer
The project’s goal is clear: making speech transcription a smooth, accurate, and fully automated process.
When a user says, “Good morning, I’m Marco”, the web app doesn’t just transcribe the sentence: it understands that this is an introduction and correctly associates the name with the speaker.
It is not based on voice samples or audio archives: recognition occurs entirely through understanding the linguistic context.
How it works
The system is based on GPT-4.1 mini, an advanced language model capable of understanding natural language and the semantic relationships between words.
Unlike traditional speech-to-text systems, this app does not use continuous machine learning: it does not evolve over time, but operates consistently and accurately thanks to its ability to interpret context in real time.
The interface is simple, modern, and intuitive. During a conversation, shows the live transcript and list of participating guests, identifying them as they speak.
Once completed, the transcript can be exported in DOCX format, ready to be shared or used in reports and official documentation.
A perfect model for companies like ITT that need reliable, secure tools that can be easily integrated into their daily workflows.
Smart handling of namesakes
One of the most advanced aspects of the platform is its ability to manage homonyms.
When multiple interlocutors share the same name, artificial intelligence does not get confused: by analyzing context, sequence of interventions, and dialogue structure, it correctly distinguishes each participant.
In this way, the transcription remains consistent and accurate, even in complex situations or with multiple similar voices.
This result is possible thanks to the system’s linguistic and contextual approach, recognising not the voice, but the meaning.
Accessibility and inclusion
This web app also represents a step forward in terms of accessibility.
Thanks to contextual recognition of speakers, the system also allows those with hearing difficulties, such as deaf or hard-of-hearing people, to accurately follow who is speaking during a conversation.
Each intervention is identified and associated with the correct speaker, allowing the reader of the transcript to immediately understand the flow of the dialogue and clearly distinguish between the different voices.
In this way, the platform not only simplifies communication, but also makes information more accessible and inclusive, ensuring a truly universal experience.
An approach that demonstrates how artificial intelligence, when carefully designed, can break down barriers and encourage participation in any context.
Why it’s a revolution
True innovation lies in the ability to recognize speakers without any vocal basis.
AI relies entirely on understanding language and context, a perspective that overturns the paradigm of traditional speech recognition systems.
The benefits are clear:
-
No voice data collection – privacy is fully protected
-
Contextual recognition – AI understands who is speaking based on the sentences and dynamics of the conversation
-
Immediate efficiency – no training or preliminary configuration required
-
Professional output – the transcript is ready to be downloaded and shared







