Tag
Jakob Porchman from Black Forest Labs presented at AI Engineer Paris on customizing the Flux video generation model for robot control and gaming, highlighting methods like prompt upsampling and content moderation.
The paper introduces a two-model architecture called Summarize-Judge-Refine (SJR) for multimodal content moderation, which decouples content understanding and policy learning via natural language summaries, enabling significant performance gains and few-shot policy adaptation.
Meta banned ads for a Virginia Woolf play in Spain because the content related to feminism, highlighting issues with automatic content moderation systems on social media platforms.
Jev Wrapped is an open-source tool that analyzes public Telegram channels by reading up to 1,500 posts to detect ads and clickbait, providing a monthly breakdown and links to top-scored posts.
HuggingFace has started censoring content, with disabled messages on certain posts, leading to discussions on Reddit about the context and specificity of these actions.
A US immigration policy targeting noncitizen technology researchers monitoring online hate speech is facing legal challenges, with court rulings sparking appeals and debates over free speech and deportation threats.
Show HN post introduces unslop.news, a tool that filters AI-related content from Hacker News to present a curated list of non-AI news items.
Hacker News is reducing the priority of AI-driven content on its platform to address concerns about such material.
A user vents that a subreddit for AI agents is dominated by bots posting and replying, with humans mostly being AI automation course buyers seeking validation.
Instagram is updating its labeling system for AI-generated profiles and will limit the reach of accounts that do not properly disclose AI usage, aiming to increase transparency in response to user concerns.
xAI faces a lawsuit alleging that it trained its Grok models on child sexual abuse material, including both real and AI-generated images, highlighting critical issues in AI ethics and data governance.
This paper presents the most comprehensive benchmarking study of LLM safety to date, evaluating 53 models across 11 datasets in various safety scenarios and providing practical guidance for model selection to mitigate different harms.
The article raises concerns about AI models potentially training on data generated by AI itself, which could lead to issues like model collapse, citing examples such as Deezer removing millions of AI-generated songs and the high proportion of AI-created blog articles.
Gergely Orosz discusses LinkedIn's shift from encouraging AI-generated content to attempting to remove it, citing an example of an AI influencer that rapidly gained followers and views.
This article argues that AI detectors and watermarking are problematic because they dismiss human effort behind work and degrade AI-generated text quality.
Meta allowed advertisements for an AI nudify app targeting female politicians, raising concerns about ethical AI use and platform policy failures against nonconsensual intimate imagery.
A user explains their YouTube ban in a linked external content.
The author shares their experience using Pangram, an AI content detection tool, to identify and block AI-written replies on social media in order to maintain human interaction.
Facebook is reportedly paying controversial creators to generate rage-bait content, raising concerns about algorithmic incentives and platform moderation.
This paper proposes BRACE, a method that encodes an Ordered Reasoning Chain to detect ever-shifting harmful chat dialogue, achieving high harm-type F1 scores with both encoder and decoder backbones.