AIzac's proprietary STT and LLM were trained on more than 80,000 hours of English speech data.
General data was collected from video and podcasts (~70%), audiobooks (~25%), and additional
English audio datasets (~5%, including government-supported initiatives such as AI Hub), along
with recorded seminars.
The dental dataset comprises original clinical recordings spanning nine specialty departments.
Additionally, portions of the dental dataset were developed using real-world clinical audio
provided by U.S. DSO partners (referenced as C*, A*, W*, and H*), anonymized and used under
appropriate data-use agreements.
These contributed to significantly improved domain accuracy, procedure-specific terminology
recognition, and real-time charting performance.