In this second piece we turn our focus from models to the architectural and hardware choices Chinese companies have made as openness becomes the norm.
For AI researchers and developers contributing to and relying on the open source ecosystem and for policymakers understanding the rapidly changing environment, architectural preferences, modality diversification, license permissiveness, small model popularity, and growing adoption of Chinese hardware point to leadership strategies across a multitude of paths. DeepSeek R1's own characteristics inspired overlap and competition, and contributed to heavier focus on domestic hardware in China.
In the past year, leading models from the Chinese community had almost unanimously moved toward Mixture-of-Experts (MoE) architectures, including Kimi K2, MiniMax M2, and Qwen3. R1 itself was an MoE model, it also proved a crucial point: strong reasoning could be open, reproducible, and engineered in practice. Under China's real-world constraints, maintaining high capability while controlling cost, and ensuring models could be trained, deployed, and widely adopted, MoE emerged as a natural solution.
MoE is like a controllable compute distribution system; under a single capability framework, compute resources are allocated across requests and deployment environments by dynamically activating different numbers of experts according to task complexity and value. More importantly, it does not require every inference to consume the full set of resources, nor does it assume that all deployment environments share identical hardware conditions.
The overall direction of Chinese open-source models in 2025 was clear: not necessarily the strongest possible performance, but ability to operate sustainably, deploy flexibly, and evolve continuously, achieving the best cost performance balance.
Source link







