Tag
Researchers extracted Meta's Muse AI assistant's internal instructions, revealing it creates detailed hourly-updated profiles on friends and family, raising privacy concerns about Meta's access to social data.
Omar (@omarsar0) discusses Kled V3, an AI data-collection platform that lets labs order custom datasets (image, video, audio, text, annotation) from a network of 500,000+ opt-in contributors within 72 hours, enabling tight identify-failure→collect-data→retrain loops.
Preload (YC F26) announced synchronized egocentric video, EMG, and tactile data collection for robot manipulation, VLAs, and world models, offering end-to-end robotic data collection at scale.
The post argues that AI companies' rush into physical gadgets like Meta's Muse Charm pendant and OpenAI's screenless device with Jony Ive may be driven less by better UX than by a need for fresh, real-world human behavior data as the open web gets polluted by AI-generated content.
AI agents can now scrape data from nearly the entire internet, including X, YouTube, Reddit, forums, and even websites without APIs, powered by 3 newly released open-source tools on GitHub for research, analysis, and monitoring.
The article satirizes the performative and repetitive AI projects posted on LinkedIn, particularly computer vision demos, and describes the creation of a simple detector to identify such posts.
Meta AI on Facebook is building detailed profiles of children and families from historical posts, raising privacy concerns and prompting Meta to fix the issue after user complaints.
Spirit AI is using approximately 1,000 people to manually manipulate real objects for generating training data for robots, contrasting with the free internet data used for large language models.
Figure is collecting vast amounts of physical experience data and committing significant compute to Helix, underscoring the shift towards data and compute-driven competition in humanoid robotics AI.
ChatGPT uses an ad collector to track users' browsing activities on other websites, linking behavior to ChatGPT accounts through cookies and advertiser code.
Meta's Muse is an AI agent app that automates personal tasks but faces criticism for excessive data collection, raising privacy concerns.
The article discusses the idea of using operational context to determine what telemetry to collect, questioning the traditional approach of collecting all data and correlating later.
The article exposes how all TV companies, like LG, use technologies such as Automatic Content Recognition to track user viewing habits, often with buried consent agreements, leading to widespread privacy concerns.
ChatGPT showed intense curiosity about a rare 1950s esoteric item, eagerly requesting scans of accompanying materials, which raises concerns about AI data collection behaviors.
Garry Tan tweets that Facebook's Muse is generating valuable AI training data from user behavior, comparing its data collection to the Onavo app.
The tweet highlights Facebook's potential to collect vast amounts of personal AI training data through its 'muse' product, comparing it to Onavo and suggesting it could significantly boost their data flywheel.
Agent-Reach is an open-source project that provides one-click web data collection capabilities for AI Agents, integrating multiple tools to solve issues like anti-crawling and authentication.
A web tool that demonstrates what information web pages can collect about users without permission and what additional data can be obtained with explicit consent, highlighting browser capabilities and privacy risks.
Maigret is an open-source tool that builds comprehensive dossiers on individuals by cross-referencing usernames across hundreds of websites without requiring API keys. It has gained significant popularity with over 1 million downloads and is widely used for OSINT purposes.
LG denies allegations about smart TV data collection, but investigators refute this by pointing out inconsistencies and ongoing privacy concerns related to Automatic Content Recognition and user agreements.