IFLYTEK Releases Edge AI Models Supporting Million-Token Context
- tech360.tv

- 8 minutes ago
- 3 min read
iFLYTEK's wholly owned subsidiary recently released and open sourced Spark X2.5-4B and Spark X2.5-1.7B. These are presented as the first edge models to natively support up to one million tokens of context, according to the Chinese AI firm. This development aims to significantly expand the memory window for on device large models.

Earlier in Sept., the subsidiary made these two edge models available. Both models employ a hybrid attention architecture. They are calibrated for agent functions, code generation, mathematical operations, and instruction following. iFLYTEK asserts their performance surpasses comparable open source models of a similar size. This initiative addresses a recognised shortcoming in local AI systems, where previous edge models often required lengthy documents to be segmented. Such segmentation frequently led to a loss of earlier contextual information, increasing the potential for inaccurate responses.
So, with a one million token window, these new models process extensive documents and technical material effectively. The models were trained using high quality data. They can ingest an entire after sales manual in a single operation. During a company demonstration, after a full manual was loaded, the model combined details regarding return policy, fault handling, and exception rules from various chapters. It then determined if a device purchased ten days prior with third party consumables qualified for replacement. This capability eliminates the need for piecemeal querying, maintaining full context.
Beyond information retrieval, the models perform specific actions. In typical office environments, Spark X2.5-4B is capable of generating a script to compile sales data, extract key indicators, and produce a bilingual departmental report of approximately 3,000 characters. It also verifies the report's structure and layout. This allows a process to move from raw data to a completed report without exiting the local machine. For coding tasks, the four billion parameter model performs comparably to cloud based models that are two to three times larger in terms of algorithm implementation and completion. It additionally provides low latency, offline functionality, and local management of code assets. Devs can integrate it with open source harnesses such as DeepSeek Harness, OpenCode, Codex, and Pi.
And, smart home and robotics applications are also suitable for these on device models. In tests conducted on a Domux smart home set, the 1.7B model achieved 90.3 percent end to end execution accuracy for control commands. The average response time was 0.85 seconds. These models are deployable on robot bodies or other edge devices for tasks involving manipulation, tracking, and navigation. This reduces dependency on external cloud connections, offering greater autonomy for local operations.
Both models underwent training on a wholly domestic compute platform. This training utilised an estimated 20 trillion tokens. The models operate across various hardware platforms, including Nvidia, Huawei, and Hygon, among others. They also support vLLM, SGLang, and llama.cpp. Model weights and corresponding code are currently accessible on Hugging Face and GitHub. The model APIs are live on the iFLYTEK Spark MaaS platform, where they are offered free of charge for a limited duration.
iFLYTEK's subsidiary launched Spark X2.5-4B and X2.5-1.7B edge AI models.
These models are stated to be the first to natively support one million tokens of context on devices.
They can process entire lengthy documents, addressing a common limitation of previous local AI.
Applications include office tasks, coding, smart home systems, and robotics.
Models were trained on a domestic platform and are available via Hugging Face, GitHub, and the Spark MaaS platform.
Source: Pandaily


