Chinese artificial intelligence (AI) startup DeepSeek has officially launched the V4 Pro, significantly enhancing its agent capabilities.
On August 13, DeepSeek updated the V4 Pro to its API (Application Programming Interface), as reported by the China Business Journal. In benchmarks for agent functionality, the V4 Pro demonstrated performance comparable to or even surpassing that of Anthropic's latest AI model, Claude Phable 5, in some evaluations.
The most notable aspect of its performance is in agent tasks. The V4 Pro scored 87.9 points in the 'terminal bench,' which involves executing actual commands and modifying files in a terminal environment. This score is just 0.1 points shy of Claude Phable 5's 88.0, indicating virtually equivalent performance.
DeepSeek's V4 Pro also excelled in cybersecurity agent assessments, such as CyberGym, and in high-difficulty task automation evaluations like AutomationBench, with some tests indicating it outperformed Phable 5.
The V4 Pro is designed as a mixed-expert (MoE) model with a total of 1.6 trillion parameters, of which 49 billion are active parameters used in calculations. It also supports a massive context of 1 million tokens, allowing for the input of lengthy documents or large codebases for extended tasks.
Its coding capabilities are impressive as well. According to DeepSeek's technical documentation, the V4 Pro achieved a score of 93.5 in LiveCodeBench and 3,206 in Codeforces. Notably, in Codeforces, it surpassed some of the latest models from OpenAI and Google.
Moreover, the V4 Pro offers overwhelming price competitiveness. The API pricing for the official version is 3 yuan (approximately $0.63) per 1 million tokens for input and 6 yuan ($1.26) per 1 million tokens for output. In contrast, Phable 5 charges $10 (about 14,000 won) for input and $50 (70,000 won) for output per 1 million tokens. This means DeepSeek's V4 Pro is priced at one-fiftieth of Phable 5's cost.
The China Business Journal notes that the V4 Pro demonstrates that Chinese AI is moving beyond merely creating low-cost models. As the competition shifts from chatbots that answer questions to agents capable of performing actual tasks such as coding, security, and automation, DeepSeek has proven the potential of Chinese AI agents.
* This article has been translated by AI.
Copyright ⓒ Aju Press All rights reserved.
