StartLux Small AI Rivals Trillion-Parameter Systems in Key Test
- tech360.tv

- 4 minutes ago
- 2 min read
StartLux, a Shanghai-based artificial intelligence organisation, achieved second place in the China Academy of Information and Communications Technology MCP test. Its 27 billion parameter local model surpassed DeepSeek-V4-Flash, a considerably larger model, performing at a level often associated with trillion parameter systems. This outcome suggests a notable capability for smaller AI architectures.

A testing report from the China Academy of Information and Communications Technology, or CAICT, issued a new assessment, as reported by Pandaily. This report highlighted StartLux-V1.0-27B-Preview. The Shanghai-based AI firm's model secured second position in the trusted-AI benchmark lineup's MCP special test. It performed ahead of DeepSeek-V4-Flash, despite possessing only 27 billion parameters.
The MCP test assesses six specific functions: location navigation, web search, browser automation, financial analysis, code repository management, and 3D design. A general assessment also featured, focusing on multi tool coordination, complex task execution, and interaction within operational settings. And, the models under examination comprised DeepSeek-V4-Pro, with 1.6 trillion parameters; DeepSeek-V4-Flash-0731, at 284 billion; Step-3.7-Flash, at 198 billion; StartLux-27B-260715, at 27 billion; Qwen-3.6-27B, also at 27 billion; and AgentCPM-Explore, with 4 billion parameters.
StartLux-V1.0-27B-Preview achieved a score of 39.25, securing second position. This score placed it above DeepSeek-V4-Flash, which has 284 billion parameters, and Step-3.7-Flash, with 198 billion parameters. Compared to Qwen-3.6-27B, a model of identical parameter size, StartLux gained an advantage of 5.34 points. The model ranked first in location navigation. It also either tied for first or achieved first position in browser automation and financial analysis, demonstrating performance equivalent to the trillion parameter DeepSeek-V4-Pro in certain instances.
The StartLux model employs Qwen3.6-27B as its base, augmented by specific post training enhancements. StartLux implemented an approach it terms "AI trains AI", referred to internally as Auto Research. This method involves autonomous execution of training experiments and subsequent strategy refinement based on received feedback. And the company states this is the first deployment of such a method for a local agent model within China.
This performance indicates a wider change in the industry. As advanced model capabilities meet practical requirements, the competitive pursuit of larger parameter counts has become less distinct. User priorities now focus on practical problem solving, data security, and cost management, rather than solely on benchmark results. Major technology organisations globally also show movement in this direction. Google's Gemma 4, Meta's open Muse Glimmer, and Nvidia's Nemotron 3.5 Lightning are examples of models designed for local deployment.
StartLux-V1.0-27B-Preview operates on standard consumer personal computers. The company intends to introduce its initial range of local intelligence solutions before the end of this year. So, the CAICT findings suggest to businesses that smaller, locally deployable models can effectively compete with much larger cloud based systems for essential tasks within actual operational environments. This development may influence future procurement choices for organisations.
StartLux-V1.0-27B-Preview achieved second place in the CAICT MCP test.
The model, with 27 billion parameters, outperformed larger models from DeepSeek and Step.
It demonstrated performance matching trillion parameter models in specific tasks.
StartLux utilised an "AI trains AI" method for the model's development.
The results signify a shift towards compact, locally deployable AI solutions.
Source: Pandaily


