Skip to content
AI情报2026年8月13日AI情报
文章

WeChat AI Team Details WeLM Models Scaling to 617B Parameters

Tencent’s WeChat AI team has detailed a new scaling approach for its WeLM model family. The team trained WeLM-HD4-80B and WeLM-HD4-617B models using a method called Hidden Decoding, which expands each token into multiple internal computation streams without increasing the main Transformer backbone. The 80B model activates 3 billion parameters, while the 6...

Frontier 编辑部来源: TechNode
01

来源简报

WeChat AI Team Details WeLM Models Scaling to 617B Parameters: Tencent’s WeChat AI team has detailed a new scaling approach for its WeLM model family. The team trained WeLM-HD4-80B and WeLM-HD4-617B models using a method called Hidden Decoding, which expands each token into multiple internal computation streams without increasing the main Transformer backbone. The 80B model activates 3 billion parameters, while the 6...

02