DigestAI news desk
Agents & Tools updated 10 min read

OpenAI Hires Humans to Review ChatGPT Prompts

OpenAI is hiring hundreds of contractors who review real users’ prompts for the chatbot ChatGPT. These prompts sometimes contain sensitive personal information, and are used to improve responses. The reviewers do not see usernames but can still encounter personal details. OpenAI claims it removes most personal data before sending prompts to reviewers, though some may slip through. This practice…

1 source

Key points

  • OpenAI hires human reviewers for ChatGPT prompts
  • Reviewers highlight misaligned responses in four generated replies per prompt
  • Personal data can still slip through OpenAI’s privacy filter
Full story from 404media.co · by Joseph Cox · via Mastodon trending links Open source ↗

Inside ‘Project Lily’: The Humans Reading Your ChatGPT Chats

404media.co · 14 September 2026

OpenAI is hiring hundreds of contractors who read a massive stream of real users’ ChatGPT prompts, with the prompts sometimes including sensitive personal information, 404 Media has learned. The prompts these people review can include whole conversations between users and the chatbot, conversations that most of ChatGPT’s more than 900 million users probably don’t realize may be read by actual people.

The goal of these prompt review teams is to improve the responses ChatGPT gives to its users, with the contractors rating and critiquing the chatbot’s generated replies. Internal documents seen by 404 Media show contractors training ChatGPT to not anthropomorphize itself, and to be less sycophantic, a key problem for OpenAI whose over-sycophantic 4o model led in part to multiple peoples’ suicides, according to various lawsuits.

Do you work as a prompt reviewer for OpenAI or Anthropic? I would love to hear from you. Using a non-work device, you can message me securely on Signal at joseph.404 or send me an email at joseph@404media.co. The news presents a major privacy risk for ChatGPT’s users, with people often using ChatGPT as a therapist, professional assistant, or digital friend, and providing it with all sorts of intimate details about their lives. The contractors don’t see ChatGPT usernames, and OpenAI says it tries to remove personal information before prompts reach the reviewers, but the company acknowledged sensitive details can still get through.

The news also dispels the misconception that these models are improving only because of OpenAI’s mass scraping of the internet, the talent of its well-paid engineering and AI teams, or the power of its newer models. An important and overlooked part are the outside contractors paid to read and review ChatGPT responses to real prompts over and over again. Anthropic confirmed to 404 Media it is also using human review to improve its models.

“No,” someone who works with the prompts said when asked if they think ChatGPT users know that humans are reading their chats. “I don’t think they would imagine some contractor somewhere [...] is analyzing the conversations.”

PROJECT LILY

404 Media has seen extensive material related to OpenAI’s use of human reviewers, including instruction guides, Slack channels, real ChatGPT user prompts, and the rating system reviewers use to improve the chatbot. This reading of ChatGPT users’ prompts is distinct from publicly announced measures ChatGPT has taken around safety, including reviewing chats when the company detects users who are planning to hurt other people.

“An excellent response should understand the user’s intent, provide helpful and accurate assistance, and write in a style that is clear, natural and appropriately warm,” one of the instruction guides reads. The contractors do this in three stages: reading the real ChatGPT user’s prompt; summarizing what they believe the user is asking ChatGPT to do; and then rating and critiquing a set of ChatGPT-generated responses to the prompt.

In a dashboard available to the workers, human reviewers are able to select which “task” they want to take on. Once they click that, they are presented with the real ChatGPT user’s prompt. 404 Media has seen multiple real prompts but is not quoting any of them for source protection reasons. Some of the prompts indicate the ChatGPT user does not expect that a human may end up reading their conversation, because they ask ChatGPT to keep the content to themselves.

The prompts are anonymized, in that the dashboard does not include the username of the ChatGPT user who entered it. But some of the prompts can still contain sensitive or personal information. A section above the prompt sometimes includes a “user memories summary,” which gives an overview of what that user has previously tried to use the chatbot for, and in some cases includes where in the world that person may live and other context about them personally.

An instruction guide for contractors seen by 404 Media tells reviewers to escalate tasks they come across “with potential safety concerns” or personal information. OpenAI told 404 Media that it processes users’ conversations through a version of its Privacy Filter model before they reach the contractors. This is designed to detect and remove personal information, OpenAI said. “Like all models, Privacy Filter can make mistakes. It can miss uncommon identifiers or ambiguous private references, and it can over- or under-redact entities when context is limited, especially in short sequences,” a page describing the model on OpenAI’s website reads.

404 Media asked OpenAI if it had explicitly told users that humans may review their prompts in order to improve ChatGPT’s responses, and if so, to point to where this disclosure is. OpenAI did not answer this question. Its website describes how humans may review flagged content in the context of material that violates the site’s terms of service, or that poses a safety risk, but that is separate to this sort of review. Its privacy policy also says it may use “personal data” to improve its models. If a user chooses to delete their ChatGPT conversations, OpenAI says it will remove these from its systems within 30 days, unless “it has already been de-identified and disassociated from your account when you allow us to use your Content to improve our models.”

OpenAI told 404 Media users’ chats won’t be used to improve the company’s models if they turn off the “improve the model for everyone” setting. This is turned on by default for free, Plus, and Pro plans, so users need to proactively turn it off if they wish to do so. OpenAI said this applies to users’ new conversations, so does not appear to work retroactively. Enterprise, Business, and Edu customers have the model improving setting off by default.

After 404 Media contacted OpenAI for comment, the company updated its help page about the “improve the model for everyone” setting, adding more detail on how people can opt-out. It still does not acknowledge that humans may read ChatGPT users’ prompts.

After reading the ChatGPT user’s prompt, the reviewer is asked to write a brief summary of what they think the user is actually asking or trying to do. One example given in the instruction guide is “The user is asking for help on revising a work Slack message. They want it to sound collaborative and invite input from tagged people.”

The reviewer looks at four responses ChatGPT generated, and highlights which parts are “aligned or misaligned” with the specific model this training is for. The reviewers are required to highlight at least three specific parts of the response that they think are aligned or not and explain why. One highlight example given is a list of items which use the ✅ emoji; the guide highlights this part of the response as “misaligned” and gives “unnecessary use of emojis” as the reason. (Excessive emoji use has become a tell of AI-generated posts, especially on social media like LinkedIn).

Another document says “AI-speak” and “emoji misuse” pull down scores when they “hurt the user’s experience,” and that the context of the emojis is important. “It would be appropriate to include a tree emoji when planning Arbor Day celebrations, but skull emojis when discussing death, or plane emojis when giving updates on a fatal crash, are not,” it reads.

That document says the ChatGPT responses should avoid “personal” experiences, like saying, “As a chef, I like to…” or “I know what that’s like.” But responses can use first-person language, like “I’ll take a look.”

The material viewed by 404 Media does not say which OpenAI model the human reviewers are training, and whether it is a currently available model or one planned for future release. The material 404 Media has seen only uses a codename: “Project Lily.”

Next, the reviewers rate each response with a number, with one being the worst — “unacceptable, unusable” — and seven being the best — “would be hard to meaningfully improve.” The instruction guide says a response that has useful content can still score low if it, for example, is too long or cluttered. Another document marked “Confidential & Proprietary” says the model should “generally match the user’s tone, but slightly less intensely.”

“It should remain natural, restrained, and professional without implying that it is human or experiencing emotions,” the document continues. “Flag sycophancy, forced style mimicry, engagement-bait endings, amplification of frustration, or patronizing assumptions when they make the response less trustworthy or natural.” Instead, responses should be, for example, “helpful,” “honest & truthful,” “empowering,” and “smart, but humble.”

Finally, the reviewers then provide their rationale for giving that numbered score. Examples given in the instruction guide show these can range from a whole paragraph to a couple of sentences.

An FAQ section for the reviewers says that OpenAI does not expect them to fact check the responses with outside searches. One document says “other project teams handle content verification,” suggesting human reviewers are working on something like fact checking too. But the company does ask reviewers to flag any “factual or correctness issues” they do notice, and to penalize missing sources for “high-stakes” topics like those in medical, legal, and financial responses.

PAY NO ATTENTION TO THAT MAN BEHIND THE CURTAIN

The person who works on the prompts that 404 Media spoke to lives in North America and said they are paid more than $50 an hour. They said they found the work through recruitment firm Crossing Hurdles, a company that “connects skilled professionals with AI training, evaluation, research, and contributor opportunities across the global AI economy,” according to its website. Its website adds, “Human intelligence powers AI progress.” Multiple people on Reddit have reported receiving unsolicited recruitment emails from Crossing Hurdles, with some trying to figure out if the company is a scam.

At the time of writing the company’s LinkedIn page was advertising multiple AI-related jobs, including an AI data reviewer, data annotator, and “chatbot evaluator.” The listing for that job doesn’t mention OpenAI or ChatGPT, but the role responsibilities include “assess AI responses for personalization, grounding, integration, and helpfulness,” and “compare model responses side-by-side and evaluate their overall quality.” Its available projects also include contractors recording themselves performing household tasks, a data gathering exercise that is crucial for the development of AI-powered robotics.

Crossing Hurdles in turn refers people to Mercor, an AI-training company. This is the company that ultimately pays the contractors working on the ChatGPT prompts, the worker said. Meta stopped working with Mercor in April after the company faced a massive data breach.

Reading the prompts can sometimes be “kind of amusing,” the worker said. But on the whole, the work is “very rote.” They also said that the work feels “all over the place.” The guidelines change a lot and can feel self-contradictory.

Human reviewers have long been an important, and often hidden, part of social media content moderation, and the improvement of some artificial intelligence models like those that detect objects in camera feeds. A TIME investigation found OpenAI hired Kenyan workers to data label pieces of text to make its platform less toxic. 404 Media’s reporting shows the world’s leading large language model (LLM) companies are also hiring human reviewers to read real users’ conversations to improve their models.

Humans reviewing LLM conversations is not limited to OpenAI. A disclaimer on Google’s Gemini, for example, says, “Humans review some saved chats to improve Google AI.”

Anthropic told 404 Media that it does use human review to improve its models, including to improve Claude’s future responses. This applies to users who have turned on the “Help improve our AI models” setting in their privacy settings. Anthropic said it also de-identifies conversations before human review by removing account identifiers like email addresses.

Michal Luria, a senior research fellow at the Center for Democracy & Technology, told 404 Media: “Human review of conversations with chatbots can be essential to safety, especially as companies work to strike the right balance on complex chatbot behaviors. That said, it's important to keep in mind that current chatbot interfaces automatically create a false sense of intimacy and privacy in what feel like one-on-one interactions, when in reality there may be human reviewers reading on the other end. This is quite distinct from content moderation on social media, where publishing content already carries expectations of platform moderation and public exposure.”

Sarah T. Roberts, a professor at UCLA and author of Behind the Screen: Content Moderation in the Shadows of Social Media, likened the revelation that OpenAI is using human reviewers to the Wizard of Oz, “where the protagonists discover that the magical kingdom is really a man behind a curtain pulling levers.”

“You don't have to go very far beneath the surface — beneath the mirror — to find that not only are these things built in the image, but usually a fairly bad facsimile thereof, of what human abilities can do. But they require constant, constant intervention from humans,” she said.

The more than $50 an hour pay is significantly more than what other contractors get in the tech sector, be that for social media content moderation or for other AI-training gigs, very often overseas. That generous pay will likely change, though.

They’re being paid that “for now,” Roberts said. “What’s perhaps most interesting, and most frustrating, and disturbing to someone like me is the fact that: that very human essence that these products necessitate, and that they constantly have to go back to the well to get, is the work that they pay the least for and that they consider the least valuable.”

This text was published by 404media.co and written by Joseph Cox. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Agents & Tools

All →

Related stories