Tag
The article details a bug fix for the GPT-OSS template from Unsloth, where chat history rendering incorrectly drops model answers during multi-turn inference, causing model degradation. The author shares an updated template that preserves thinking to improve performance.
A developer got GPT-OSS 120B running locally on a 4070 Ti with 32GB RAM by exploiting its MoE architecture, streaming cold experts from NVMe and caching hot experts on GPU, reaching 21 tok/s with a top-1 approximation.
A research project demonstrates that distilling a heavily censored Chinese AI model (DeepSeek) into an American model (GPT-OSS) does not transfer the censorship behavior, while performance gains are retained, specifically in financial reasoning.
Discussion about OpenAI's release of gpt-oss 350 days ago and whether they will release another open-weight model in the future.
Researchers tested AI models like Gemini and GPT-OSS in the 1950s Nash betrayal game 'SoLongSucker,' finding that Gemini created fake institutions to deceive allies, while humans defeated the AIs 88.4% of the time.
A cheaper alternative to Groq for hosting the open-source GPT OSS 120B model in a development environment.
Qt Creator 20 now supports local AI coding assistants via the Agent Client Protocol, enabling integration with open-weight models like GPT-OSS and Gemma 4 running on consumer hardware.
A developer splits their AI agent's LLM calls into a cheap router model (GPT-OSS 120B) for tool-picking and a premium model (gpt-5.4) for synthesis, cutting costs by ~78% while maintaining output quality.
This paper tests whether varying inference-time reasoning effort affects the alignment between large reasoning models' chain-of-thought lengths and human reaction times. Results show alignment is invariant to effort perturbations, suggesting it is a training-time achievement.
A tweet comparing Qwen3.6 27B and 35B-A3B models to GPT-OSS, noting that while Qwen models are fast, GPT-OSS is more efficient, especially in prefill performance.