fp8-training

Tag

Cards List
#fp8-training

@_akhaliq: A.X K2 just dropped on Hugging Face Large-Scale Sparse MoE (688B / 33B Active) https://huggingface.co/skt/A.X-K2

X AI KOLs Following · 4d ago Cached

SKT released A.X K2, a 688B-parameter sparse MoE language model with 33B active parameters, natively trained in FP8 and featuring Think/Non-Think reasoning modes, on Hugging Face.

0 favorites 0 likes
← Back to home

Submit Feedback