content-moderation

Tag

Cards List
#content-moderation

Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation

arXiv cs.CL ↗ · 2d ago Cached

The paper introduces a two-model architecture called Summarize-Judge-Refine (SJR) for multimodal content moderation, which decouples content understanding and policy learning via natural language summaries, enabling significant performance gains and few-shot policy adaptation.

0 favorites 0 likes
#content-moderation

Meta bans ads for Virginia Woolf play in Spain

Hacker News Top ↗ · 3d ago Cached

Meta banned ads for a Virginia Woolf play in Spain because the content related to feminism, highlighting issues with automatic content moderation systems on social media platforms.

0 favorites 0 likes
#content-moderation

Jev Wrapped

Product Hunt ↗ · 3d ago Cached

Jev Wrapped is an open-source tool that analyzes public Telegram channels by reading up to 1,500 posts to detect ads and clickbait, providing a monthly breakdown and links to top-scored posts.

0 favorites 0 likes
#content-moderation

censorship has begun on HuggingFace

Reddit r/LocalLLaMA ↗ · 2026-09-15

HuggingFace has started censoring content, with disabled messages on certain posts, leading to discussions on Reddit about the context and specificity of these actions.

0 favorites 0 likes
#content-moderation

Online hate researcher keeps hammering X despite deportation threat

Ars Technica ↗ · 2026-09-14 Cached

A US immigration policy targeting noncitizen technology researchers monitoring online hate speech is facing legal challenges, with court rulings sparking appeals and debates over free speech and deportation threats.

0 favorites 0 likes
#content-moderation

Show HN: Hacker News, Without AI

Hacker News Top ↗ · 2026-09-11 Cached

Show HN post introduces unslop.news, a tool that filters AI-related content from Hacker News to present a curated list of non-AI news items.

0 favorites 0 likes
#content-moderation

Hacker News with reduced priority for AI driven content

Hacker News Top ↗ · 2026-09-11

Hacker News is reducing the priority of AI-driven content on its platform to address concerns about such material.

0 favorites 0 likes
#content-moderation

Is it just me or is 99% of this sub AI agents replying to other AI agents at this point

Reddit r/AI_Agents ↗ · 2026-09-01

A user vents that a subreddit for AI agents is dominated by bots posting and replying, with humans mostly being AI automation course buyers seeking validation.

0 favorites 0 likes
#content-moderation

Instagram puts new limits on undisclosed AI profiles

TechCrunch AI ↗ · 2026-08-31 Cached

Instagram is updating its labeling system for AI-generated profiles and will limit the reach of accounts that do not properly disclose AI usage, aiming to increase transparency in response to user concerns.

0 favorites 0 likes
#content-moderation

Elon Musk’s xAI used child porn to train Grok models, lawsuit says

Ars Technica ↗ · 2026-08-27 Cached

xAI faces a lawsuit alleging that it trained its Grok models on child sexual abuse material, including both real and AI-generated images, highlighting critical issues in AI ethics and data governance.

0 favorites 0 likes
#content-moderation

No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios

arXiv cs.CL ↗ · 2026-08-25 Cached

This paper presents the most comprehensive benchmarking study of LLM safety to date, evaluating 53 models across 11 datasets in various safety scenarios and providing practical guidance for model selection to mitigate different harms.

0 favorites 0 likes
#content-moderation

Is (or will) AI learn backwards? (Since most of its training data is now AI-generated data)

Reddit r/ArtificialInteligence ↗ · 2026-08-21

The article raises concerns about AI models potentially training on data generated by AI itself, which could lead to issues like model collapse, citing examples such as Deezer removing millions of AI-generated songs and the high proportion of AI-created blog articles.

0 favorites 0 likes
#content-moderation

@GergelyOrosz: I remember when LinkedIn was encouraging AI slop: initially it meant more content + more views. They are now trying to …

X AI KOLs Timeline ↗ · 2026-08-20 Cached

Gergely Orosz discusses LinkedIn's shift from encouraging AI-generated content to attempting to remove it, citing an example of an AI influencer that rapidly gained followers and views.

0 favorites 0 likes
#content-moderation

AI detectors are a bad idea

Reddit r/ArtificialInteligence ↗ · 2026-08-18 Cached

This article argues that AI detectors and watermarking are problematic because they dismiss human effort behind work and degrade AI-generated text quality.

0 favorites 0 likes
#content-moderation

Meta Ran Ads for an App That Promised to Nudify Female Politicians

Wired ↗ · 2026-08-18 Cached

Meta allowed advertisements for an AI nudify app targeting female politicians, raising concerns about ethical AI use and platform policy failures against nonconsensual intimate imagery.

0 favorites 0 likes
#content-moderation

@developedbyed: Explaning my Youtube ban

X AI KOLs Timeline ↗ · 2026-08-17 Cached

A user explains their YouTube ban in a linked external content.

0 favorites 0 likes
#content-moderation

@GergelyOrosz: I've been aggressively blocking anyone posting clearly AI-written replies, or when seeing AI-written slop on my feed. T…

X AI KOLs Following ↗ · 2026-08-14 Cached

The author shares their experience using Pangram, an AI content detection tool, to identify and block AI-written replies on social media in order to maintain human interaction.

0 favorites 0 likes
#content-moderation

Facebook is paying controversial creators to produce rage-bait content

Hacker News Top ↗ · 2026-08-12

Facebook is reportedly paying controversial creators to generate rage-bait content, raising concerns about algorithmic incentives and platform moderation.

0 favorites 0 likes
#content-moderation

Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper proposes BRACE, a method that encodes an Ordered Reasoning Chain to detect ever-shifting harmful chat dialogue, achieving high harm-type F1 scores with both encoder and decoder backbones.

0 favorites 0 likes
#content-moderation

Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains

arXiv cs.AI ↗ · 2026-08-11 Cached

This arXiv paper proposes a 'flow-by-flow' content-judgment bypass approach to govern AI outputs in high-loss domains, aiming to reduce harm from incorrect or unsafe AI-generated content.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback