OpenAI Is Hiring Hundreds of Contractors to Secretly Read ChatGPT Conversations

OpenAI is paying hundreds of contractors to read real conversations that users have with ChatGPT and rate the chatbot’s replies as part of an internal effort to improve its models, according to a report by 404 Media based on leaked internal documents.

The program, known inside the company as Project Lily, involves contractors who examine actual user prompts and full exchanges rather than synthetic examples. They summarize what the user appears to want and then score multiple generated responses on a scale of one to seven, from unacceptable to nearly impossible to improve.

404 Media obtained reviewer instructions, Slack discussions, scoring materials and examples of the prompts themselves. The work focuses on reducing behaviors OpenAI wants to curb, including excessive agreement with users, language that makes the model sound human, unnecessary emojis and other “AI-speak.”

Contractors are recruited through a firm called Crossing Hurdles and paid through Mercor, an AI training company. One reviewer based in North America told 404 Media they earned more than $50 an hour for the work.

Usernames are stripped and OpenAI runs an automated privacy filter intended to remove personal details before the material reaches human reviewers. The company has acknowledged that the filter can miss uncommon identifiers or fail when context is limited. Reviewers sometimes also see a summary of the user’s stored memories, which can include details about past topics or approximate location.

Some of the conversations examined by 404 Media contained sensitive personal information. In certain cases users explicitly asked ChatGPT to keep the discussion private. A person who worked with the prompts told the outlet they did not believe most users realized humans were analyzing the exchanges.

The reviews under Project Lily are separate from the company’s public safety processes that examine chats flagged for potential harm. ChatGPT has more than 900 million weekly active users, according to figures cited in the reporting.

Users on Free, Plus and Pro plans can prevent future conversations from being used to improve models by turning off the “Improve the model for everyone” setting in ChatGPT’s Data Controls. The option is enabled by default on those plans and applies only to new chats. Temporary Chat mode also excludes conversations from training. On Enterprise, Business and Edu plans the setting is off by default.

404 Media asked OpenAI where it informs users that humans may read chats to improve responses. The company did not answer directly at the time. After publication it pointed the outlet to a help page that addresses human review of content for model improvement.

Similar human-review programs exist at other AI companies, including Anthropic and Google, according to the 404 Media report. OpenAI’s privacy policy allows use of personal data to improve models, and the company states that deleted chats are purged within 30 days except for anonymized material retained with consent for training.

46 web pages

Latest Posts

[democracy id="16"] [wp-shopify type="products" limit="5"]