Tag
An individual discusses building a server with 768GB VRAM for running frontier AI models but is concerned that new open-source models like GLM6 are becoming too large, prompting consideration of downsizing to smaller flash models.
The article announces that additional sizes of the Qwen 3.8 model family are coming, giving developers more options for deployment.
The author argues that releasing smaller Qwen models (27B, 35B, 122B, 397B) would better serve the local AI community than focusing on trillion-parameter behemoths, which are impractical for most users.
Analysis of the trend in AI model sizes, noting a gap in the 100-120B parameter range with recent releases focusing on smaller (25-35B) or larger (200B+) models.