Kakao Enhances AI Voice Generation Technology with Kanana-o

By BAEK SEO HYUN Posted : August 4, 2026, 14:40 Updated : August 4, 2026, 14:40

Kakao has advanced the voice generation technology of its proprietary omni artificial intelligence (AI) model, Kanana-o. The improvements allow the system to control not just natural reading but also tone, emotion, speed, and volume based on user natural language commands.


On August 4, Kakao revealed the voice generation capabilities of Kanana-o through its tech blog. The new model generates speech that reflects the desired speaking style input by users in natural language.


For instance, it can understand commands like "read very quickly," "read in a low voice," "read in a sad voice," and "read in a Gyeongsang dialect," adjusting speed, volume, pitch, emotion, and intonation accordingly.


It can also perform role-based commands such as "like a sports commentator," "like an announcer," or "like reading a storybook." The system can handle complex commands that include multiple conditions, such as "read quickly in a low, sad voice." While primarily trained in Korean, it can also execute the same speech commands in English.


Performance has also improved. Kanana-o scored 94.50 on the Korean benchmark for evaluating speech command execution, InstructTTSEval, surpassing the GPT-4 Mini TTS, which scored 91.10.


The speed and efficiency of voice generation have been enhanced as well. Kakao applied its proprietary voice tokenizer, LM-SPT, which compresses speech into fewer tokens, reducing the amount of data the AI needs to process and increasing voice generation speed.


LM-SPT has demonstrated excellent performance in evaluations comparing the speech understanding and generation capabilities of various language models in both Korean and English.





* This article has been translated by AI.

Copyright ⓒ Aju Press All rights reserved.