Tag
A user reflects on GPT-4 as the first LLM that felt useful and practical, four years after its training finished, expressing excitement for what comes next.
Greg Brockman notes that GPT-4 finished training four years ago today, reflecting on the milestone.
This arXiv paper investigates implicit bias in LLM-based chat AI toward people with intellectual disabilities by generating and analyzing 25,000 stories across five LLMs. Findings reveal negative biases such as infantilization, paternalism, and dependency, highlighting the need for bias mitigation in AI development.
Sam Altman expresses surprise that OpenAI's models have become proficient at design tasks.
A new study in PLOS One finds that AI chatbots impersonating 112 UK public figures produced responses rated as more authentic, coherent, and relevant than the actual humans, warning of potential societal harm from political deception.
This article details the reverse engineering of Microsoft's Windows Copilot to create a free API compatible with OpenAI's interface, allowing access to GPT-4 without an API key or billing.
The article exposes a plagiarized version of John Koenig's 'The Dictionary of Obscure Sorrows' that copies the entire book text and replaces original illustrations with AI-generated images, also using GPT-4 to let users create new 'sorrows'. The incident raises concerns about AI-facilitated plagiarism and copyright infringement.
Dan Shipper argues that AI models like GPT-4 can replace human intuition in fields like psychology where scientific explanations are lacking, advocating for using AI to drive progress even without full understanding.
Introducing Caltext, an open-source calorie tracking tool built for iMessage, leveraging GPT-4.1 vision and a modern tech stack including Bun, Turborepo, Hono, and Vercel.
Stanford researchers led by James Zou found that AI models from OpenAI, Anthropic, and Google cite the wrong sources about 30% of the time, even when answers are mostly correct. The study highlights a critical mismatch between text generation and accurate citation, posing risks for fields like medicine and law.
A philosophical discussion questioning whether AI models truly 'understand' or if we are projecting human-like cognition onto pattern-matching systems, referencing Searle's Chinese Room, 'stochastic parrots', and GPT-4's performance.
A personal reflection on the rapid evolution of AI over the past three years, from early ChatGPT and GPT-4 quotas to BabyAGI, DALL·E, and voice cloning.
OpenAI releases GABRIEL, an open-source toolkit that uses GPT to convert unstructured qualitative data (text, images) into quantitative measurements for social scientists and economists. The tool enables researchers to analyze large-scale qualitative datasets more efficiently by automating repetitive labeling tasks while preserving the richness of human data.
Higgsfield is a generative media platform that uses GPT-4.1, GPT-5, and Sora 2 to turn simple product links or ideas into cinematic short-form social videos, generating roughly 4 million videos per day with a 'cinematic logic layer' that translates user intent into structured video plans.
OpenAI announces a strengthened safety ecosystem through external third-party testing and evaluations of frontier AI models, including independent assessments, methodology reviews, and subject-matter expert probing. The company commits to transparency by publicly sharing third-party assessment results and supporting independent evaluations since GPT-4's launch.
Blue J demonstrates how to scale AI expertise in complex regulated domains by combining GPT-4.1 with retrieval-augmented generation over curated tax documents, achieving <0.14% error rates and 70% weekly user engagement through rigorous feedback loops and domain-specific optimization.
Intercom shares three lessons from rapidly adopting AI to transform their customer service platform: testing models early and deeply, building AI-first from the ground up rather than bolting it on, and using rigorous evaluation processes to quickly adopt new models like GPT-4.1.
Morgan Stanley has successfully deployed AI solutions powered by GPT-4 across its wealth management division, with over 98% of advisor teams using the internal AI Assistant chatbot. The deployment was enabled by a robust evaluation framework that tests AI performance on real-world use cases like document summarization and multilingual translation before production rollout.
Arco Educação, Brazil's largest educational operating system, is partnering with OpenAI to launch the Teacher Assistant, an AI tool powered by GPT-4 that helps teachers create personalized lesson plans for students with diverse learning needs, significantly reducing time spent on administrative tasks and lesson planning.
Ada uses GPT-4 and a multi-agent system powered by OpenAI's API to improve customer service quality, doubling resolution rates from 30% to 60-80% while maintaining high containment rates, establishing a new industry standard beyond traditional metrics.