@MaximeRivest: this was essentially me looking for ways to have a jev at home before they made it.

X AI KOLs Following News

Summary

The author shares their past efforts to achieve efficient single-expert loading for high-quality token-level classification at scale, echoing Maxime Rivest's thoughts on the utility of domain-specific expert models.

this was essentially me looking for ways to have a jev at home before they made it.
Original Article
View Cached Full Text

Cached at: 09/17/26, 08:22 PM

this was essentially me looking for ways to have a jev at home before they made it.

Maxime Rivest 🧙‍♂️🦙🐧 (@MaximeRivest): I hope I can load only one expert and do 1 token high-quality classification at scale. If the same expert can be used in 1 domain. It will still be very very useful for data applications.

Similar Articles

@nrehiew_: For the visual learners

X AI KOLs Timeline

A tweet describes a large mixture-of-experts model with 975B total parameters (41B active) trained on 45T tokens of multimodal data, featuring 6 routed experts and 2 shared experts, with comparisons to DeepSeek-V3.