Tag
A user optimized a 4x3060ti GPU rig for AI inference using tensor parallelism with Exl3 and vllm, achieving up to 120 tokens per second with large context windows.
A user announced they trained a custom 78M parameter AI architecture from scratch using 1B tokens on an 8x3090 rig over 12 hours, claiming the model can now speak.