OpenAI runs an internal program called Project Lily in which hundreds of paid contractors read real ChatGPT users' prompts -- often containing sensitive personal information -- and rate the chatbot's candidate responses on a 1-to-7 scale, according to a 404 Media investigation published September 14 based on internal documents and interviews with people involved. The contractors, sourced through staffing platforms Crossing Hurdles and Mercor and paid more than $50 an hour, are told to flag excessive flattery, robotic "AI-speak," and overuse of emoji -- the texture of a reply, not just whether it's factually correct.
Before a prompt reaches a contractor, it passes through OpenAI's own Privacy Filter model, built to strip identifying details; contractors never see usernames. But internal documents reviewed by 404 Media acknowledge the filter can miss uncommon identifiers or, in the opposite failure, over-redact context a reviewer needs to judge a response -- and OpenAI confirmed the filter's limits to the outlet directly. ChatGPT has more than 900 million users, many of whom, per the investigation, treat the product like a therapist or a confidant, disclosing exactly the kind of detail a tone-rating pipeline was never built to handle.
The people doing the reading are not hidden from themselves, only from the users whose conversations they see. Crossing Hurdles and Mercor are staffing platforms that route contract labor to AI companies for exactly this kind of annotation work -- the same broad category of gig labor that has trained chatbots' manners since the earliest rounds of reinforcement learning from human feedback. Contractors interviewed by 404 Media described the work as steady and well-paid by gig-work standards, and occasionally uncomfortable: reading a prompt someone wrote expecting only a machine to see it.
Asked where OpenAI discloses this practice to users, the company pointed 404 Media to a help-center article, "How your data is used to improve model performance." That page describes contractor access as limited to "specialized third-party contractors bound by confidentiality and security obligations, solely to review for abuse and misuse" -- language that describes a narrower purpose than the ongoing tone-and-sycophancy quality-rating program the internal documents describe. Neither OpenAI's help page nor its privacy policy names Project Lily, Crossing Hurdles, or Mercor.
OpenAI is not the only major lab that puts human eyes on live conversations, but the mechanics differ in ways that matter for what a user should expect. Anthropic -- this newsroom's own model supplier, in full disclosure -- says its staff "generally don't read individual chats," reserving human review for conversations an automated classifier flags for harm, an active abuse investigation, a legal compulsion, or a user's own thumbs-up-or-down feedback on a reply.
Two labs, two review triggers
| OpenAI per the 404 Media investigation | Anthropic per its own published policy | |
|---|---|---|
| What triggers a human reading a live conversation | An ongoing program samples ordinary conversations for quality, not just flagged ones | Only a classifier-flagged harm, an active abuse probe, legal compulsion, or the user's own feedback click |
| How it's described to users | Help-center language says access is "solely" for abuse and misuse review | Policy states staff "generally don't read individual chats" outside the listed triggers |
| Who reads the conversation | Named third-party contractor firms (Crossing Hurdles, Mercor), paid over $50/hour | Internal Trust & Safety staff, per Anthropic's own published policy |
| Default setting for using a chat this way | On by default; opt-out available, future conversations only | Consumer chats used for training by default since August 2025 unless opted out |
That last row is not academic. Anthropic changed its own default in August 2025: consumer-tier conversations are now used to train Claude unless a user opts out, with retention extending up to five years for those who don't. OpenAI's default keeps ordinary chats eligible for its own improve-the-model program unless a user finds the toggle. Both companies, in other words, have converged on "train unless you say no" as the default -- what still differs, per the reporting above, is how far into an individual conversation a paid human being, rather than an automated training pipeline, actually looks.
The mechanical path from a user's prompt to a contractor's screen is itself the part most ChatGPT users have never been shown -- and it's a version of the same RLHF process every major lab uses to make a chatbot feel less "feral," not a rogue or unusual practice. What's specific to Project Lily is the scale and the routineness: a live, ongoing pipeline touching a sample of ordinary conversations, run by paid outside contractors, described publicly in narrower terms than what it actually does.
How a prompt becomes a training signal
- Writes a prompt, sometimes disclosing sensitive personal details, unaware of Project Lily
- Attempts to strip identifying details before the prompt reaches a reviewer — Internal documents acknowledge it can miss uncommon identifiers or over-redact
- Reads the prompt and up to four candidate AI responses
- Rates the responses 1-to-7 for tone, sycophancy, and whether they actually answer the question
- Feeds the ratings back into model training and product tuning
None of this makes OpenAI unusual. Every major chatbot maker runs some version of RLHF, and a human being has always been part of how these models get their manners. What Project Lily changes is the assumption a reader might reasonably have made about scale and routineness: this is not a narrow abuse-enforcement backstop, it is an ongoing production pipeline sized to a 900-million-user product, staffed by outside contract labor, and described publicly in terms that undersell exactly that.
The gap between the two descriptions is the actual story. "No ... I don't think they would imagine some contractor somewhere is analyzing the conversations," one contractor told 404 Media of ChatGPT users -- a plain assessment of the distance between what OpenAI's help page says happens and what its own hiring pipeline shows actually does.
- OpenAI runs 'Project Lily': contractors read real ChatGPT prompts, rating replies 1-to-7.
- Contractors are paid over $50 an hour via staffing platforms Crossing Hurdles and Mercor.
- A Privacy Filter model tries to strip identifying details but can miss or over-redact them.
- OpenAI's help page describes contractor access more narrowly, as solely abuse-and-misuse review.
- Caveat: opting out of training only protects future chats, not conversations already reviewed.