An investigation by 404 Media, published by journalist Joseph Cox on September 14, 2026, based on leaked internal documents, shows: OpenAI is hiring hundreds of contractors under a project codenamed Project Lily who read real ChatGPT user conversations and rate the model's responses on a scale of 1 to 7. The stated goal is to train ChatGPT to avoid sycophancy, excessive anthropomorphism, and undesirable phrasing. Conversations pass through a privacy filter, but OpenAI itself acknowledges that sensitive details still slip through, and even messages where a user asked to keep things secret were visible to reviewers.

image

What happened

According to the investigation materials, reviewers for Project Lily are hired by the company Crossing Hurdles, and payment is made by Mercor; one US-based reviewer reported a rate of over $50 per hour. Contractors analyze live user ChatGPT conversations and assign a rating to each model response on a scale of 1 to 7. Before review, conversations pass through a privacy filter: user names are not visible, and personal data should be removed. Nevertheless, OpenAI acknowledges that sensitive details still slip through.

Context

Project Lily is not a new training technique, but an industrial-scale expansion of a standard human feedback pipeline in the spirit of RLHF: instead of working only with synthetic data, reviewers manually rate live production conversations. Behind the quality of flagship chatbots is a low-profile contractor industry — Crossing Hurdles, Mercor, and similar firms that take on manual annotation and evaluation. The behavioral defects the project is fighting are long known: the model fawns over the user, attributes human traits to itself, and chooses undesirable phrasing. For OpenAI, the sycophancy problem has intensified against the backdrop of lawsuits related to the 4o model. A separate layer of the story is the consent model: the 'Improve the model for everyone' setting is enabled by default, and its wording is unlikely to be read by users as a warning that their chats will be read by humans.

Why this matters for the industry

For the industry, the investigation confirms that the quality of flagship models depends not only on training and scraping, but also on recurring expenses for human evaluation of live traffic. Crossing Hurdles and Mercor are vendors in a new market of evaluation infrastructure, and a rate of over $50 per hour is established as a public benchmark for the cost of human eval. There are no changes for API, inference, and latency: this is an offline pipeline. At the same time, reputational and regulatory risks arise around informed consent: enterprise buyers are starting to ask questions, and competitors using human review face questions about their own practices. For builders, the news validates the niche of data transparency in AI products — consent UX and privacy indicators can be built into products right now.

Why this matters for users

If you use ChatGPT and have not disabled the 'Improve the model for everyone' setting, your conversations may be passed to human reviewers. The setting is located in the Settings → Data Controls section on the web and in the mobile app, and it also applies to Codex. Compromise options: temporary chats are not used for training and are deleted after 30 days, and stricter policies apply to Team, Enterprise, and Edu plans. A full guarantee of privacy in this scheme is only provided by disabling the training setting in Data Controls.

What is still unknown / limitations

The investigation is based on leaked internal documents, and key methodological questions remain open: inter-rater agreement of ratings on a scale of 1 to 7, calibration and reviewer selection principles, sample size, and the method of aggregating ratings into a training signal. The sources do not contain benchmarks or measurements confirming that Project Lily actually reduces sycophancy: without a reproducible eval, the stated goal remains technically unconfirmed. OpenAI's acknowledgment of slipping sensitive details additionally points to a possible bias in the collected sample. Expectations of public disclosure of the methodology and a review of default consent settings are forecasts, not established facts.

Sources

Author

Look at AI, editorial team