有体验过了的么,觉得如何?除了官方现在有哪些 MaaS 已经支持了的可以分享下
Parameter |
Specification |
|---|
Architecture |
Mixture-of-Experts (MoE) |
Total Parameters |
2.8T |
Activated Parameters |
104B |
Number of Layers |
93 |
Number of Dense Layers |
1 |
Attention-Layer Composition |
69 KDA + 24 Gated MLA |
Attention Hidden Dimension |
7,168 |
Number of Attention Heads |
96 |
Latent MoE Dimension |
3,584 |
MoE Hidden Dimension (per Expert) |
3,072 |
Number of Experts |
896 |
Selected Experts per Token |
16 |
Number of Shared Experts |
2 |
Vocabulary Size |
160K |
Context Length |
1,048,576 |
Attention Mechanism |
KDA & Gated MLA |
Activation Function |
SiTU-GLU |
Vision Encoder |
MoonViT-V2 |
Parameters of Vision Encoder |
401M |
Quantization |
MXFP4 weights / MXFP8 activations (quantization-aware training) |
Modality |
Text, Image |