UK FACT CHECK POLITICS

UK FACT CHECK POLITICS

Independent reporting, transparently verified by objective AI fact-checking
Menu
Get Involved
Account
cjr.org 11 July 2026 at 08:30

Ask Anika: Should AI write my headline?

View original article →
90
Trust Score

Highly Reliable — Accurate Opinion/Analysis Column

Confidence: High

Standard
Emotional Tone Low
How emotionally charged the language is (low is neutral)
Reading Level Advanced
Suitable for age 18+ readers (grade 13)
Article Length Long
1,062 words
Caps & Emphasis Moderate
2.3% of words are capitalised (high can indicate sensationalism)

Executive Summary

This is a CJR (Columbia Journalism Review) 'Ask Anika' advice column by Anika Collier Navaroli on whether journalists should feed unpublished drafts into generative AI tools to write headlines. It is explicitly an opinion/analysis piece with a clear normative recommendation (proceed with caution; avoid public LLMs for unpublished drafts), but its factual scaffolding is unusually strong and verifiable. Every checkable factual claim — the Washington Post C4 investigation finding that half the top-ten training sites were news outlets (with washingtonpost.com ranked eleventh), Anthropic's practice of buying, scanning and destroying millions of books, A.G. Sulzberger's 'brazen theft... unprecedented scale' quote, the September/October 2025 Stanford HAI study on all six frontier developers using chat data by default and some retaining it indefinitely, the CSU–OpenAI $16.9M deal with a no-training-on-CSU-data term, and the AP's 2023 OpenAI licensing deal plus its guidance not to enter confidential/sensitive information — was corroborated by primary or reputable secondary sources. Low emotional register (emotion_score 0.12) and modest capitalisation ratio are consistent with measured analysis. The high self-reference count (22) reflects the first-person advice-column format and author self-disclosure rather than bias. Points are deducted only for interpretive framing (e.g., calling data use a 'feedback loop to hell') and the inherent opinion nature of the recommendations, not for factual error.

Factual Verification

Verified Claims

  • A Washington Post investigation of Google's C4 dataset found that half of the top 10 sites overall were news outlets: nytimes.com No.4, latimes.com No.6, theguardian.com No.7, forbes.com No.8, huffpost.com No.9; washingtonpost.com ranked No.11 — corroborated by the Washington Post interactive (2023). [https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learning/]
  • Anthropic spent tens of millions of dollars to buy, scan and destroy (slice the spines off) millions of physical books to train its models — corroborated by Futurism and Tech Brew reporting on unsealed lawsuit documents (Project Panama). [https://futurism.com/future-society/anthropic-destroying-books ; https://www.techbrew.com/stories/2026/01/28/anthropic-ai-books-lawsuit]
  • A.G. Sulzberger, publisher of The New York Times, described AI's use of intellectual property as 'a brazen theft of intellectual property that has occurred at an unprecedented scale' (WAN-IFRA World News Media Congress) — corroborated by The New Republic and Media Copilot. [https://newrepublic.com/article/212095/sulzberger-big-tech-thief-liar]
  • A Stanford HAI study (led by Jennifer King; published circa Sept–Oct 2025) found all six leading US frontier developers (Amazon, Anthropic, Google, Meta, Microsoft, OpenAI) use users' chat data by default to train models, and that some retain this data indefinitely — corroborated by Stanford Report and HAI. [https://news.stanford.edu/stories/2025/10/ai-chatbot-privacy-concerns-risks-research]
  • The California State University system entered a deal with OpenAI for which the Chancellor's Office paid $16.9 million, and the contract prevents OpenAI from training its models on data provided through CSU EDU accounts — corroborated by CSU Bakersfield's ChatGPT Edu FAQ. [https://www.csub.edu/ai/chatgpt/faq.shtml]
  • The Associated Press entered a licensing agreement with OpenAI in 2023 and urges its staff not to put confidential or sensitive information into AI tools — corroborated by Nieman Lab and The Hill. [https://www.niemanlab.org/2023/08/not-a-replacement-of-journalists-in-any-way-ap-clarifies-standards-around-generative-ai/ ; https://thehill.com/policy/technology/4095868-associated-press-licensing-news-archive-to-openai/]

Unverified Claims

  • That the AI industry is 'reportedly worth some three trillion dollars' — plausible as an aggregate market-cap/valuation figure and hedged with 'reportedly', but the specific figure was not independently confirmed in this research pass; treat as an approximate, source-dependent estimate.
  • The author's characterisation of a 'feedback loop to hell' from synthetic data and her 2024 warnings — these are the author's own prior published framing/opinion, not independently verifiable factual claims; model-degradation from synthetic data is a documented research concern but the specific rhetorical claim is interpretive.
  • That Lex discloses it 'may collect your conversations in the AI Chat including both your inputs and the AI assistants outputs' for 'improving, upgrading, or enhancing' the platform, and that Lumo stores chats encrypted and only locally — quoted from the products' own policies; the underlying quotes were not separately re-verified in this pass and depend on current product terms which can change.
  • The existence/status of a proposed 'Newsroom Tooling Alliance' — presented as a proposal by researchers; not independently verified here.

Disputed / False Claims

  • None identified. No claim in the article was found to be false against available primary or reputable secondary sources.

Bias & Presentation

Detected Biases:

  • Advocacy/normative framing: the column argues a position (journalists should avoid feeding unpublished drafts into public LLMs), consistent with an advice/opinion column rather than neutral reporting.
  • Author perspective/self-disclosure: heavy first-person framing (self_reference_count 22) and the author's professional background in trust & safety and journalism ethics shapes emphasis toward risk and IP protection.
  • Selection/emphasis bias: examples chosen (Anthropic book destruction, 'theft' quote) emphasise harms of AI companies over any benefits, though counterpoints (enterprise protections, local LLMs, CSU no-train clause) are acknowledged.
  • Rhetorical/loaded language in limited instances ('feedback loop to hell', 'hoovered up', 'brazen theft' — the last being a directly attributed quote).

Language Patterns

Emotional manipulation: 0.15

Quality Assurance

Limitations: Embedded hyperlinks were stripped from the supplied verbatim text, so exact linked targets could not be read; corroboration relied on independent location of the described sources. Product terms (Lex, Lumo) and valuation/cost figures are time-sensitive and were not exhaustively re-verified.

Confidence

Level: High

All six high-priority, verifiable factual claims were independently corroborated by primary sources (Washington Post interactive, Stanford HAI/Stanford Report, CSU FAQ, Nieman Lab) or multiple reputable secondary sources (Futurism, Tech Brew, The New Republic, EdSource, The Hill). No claim required a 'False' verdict. Residual uncertainty is confined to time-sensitive figures (industry valuation, deal costs) and product-specific policy quotes, which were appropriately marked Unverified. The piece is transparently labelled opinion/analysis, reducing the risk of misclassifying advocacy as factual error.

Search Journal

Query: Washington Post inside the secret list websites that make AI chatbots sound smart

Query: Washington Post AI training data top websites news outlets NYT ranked news sites

Query: Anthropic destroyed books scan copyright lawsuit millions

Query: Sulzberger AI theft intellectual property unprecedented scale

Query: Stanford HAI chatbot user data training default study 2025

Query: Stanford HAI study retain chat data indefinitely six developers train by default

Query: California State University OpenAI 16.9 million OpenAI may not train model on CSU data

Query: Associated Press OpenAI licensing agreement 2023 confidential sensitive information AI tools

Query: AP standards generative AI journalists not put confidential sensitive information into AI tools

Article Content

_Sign up for the daily [**CJR newsletter**](

Hey, y’all. These days, journalists hear often about new technology, and what they can (and shouldn’t) do when it comes to tough ethical and safety issues around AI, social media, and more. But there’s a lot of conflicting advice out there. I’m here to help. With this new column, I’ll address common questions about the ethics, legal considerations, and best practices of adopting new technologies. Trained as a journalist and media lawyer, I’ve spent over a decade developing policies for emerging technologies inside civil society organizations and at companies such as Twitter and Twitch. I’m also a professor at Columbia Journalism School and lead the Craig Newmark Center for Journalism Ethics and Security. Send me your questions at [askanika@cjr.org](mailto:askanika@cjr.org).

**Q: I’m a busy, tired, and overworked journalist. Should I use AI to help me write headlines?**

I’ve always thought that if there was one solid use case for AI-powered chatbots like ChatGPT and Claude, it could be headline writing. The latest generation of large language models has ingested practically every headline ever published, making it especially suited to the task of suggesting which I should use for my little story.

Yet like many journalists navigating the nexus of craft and technology, I’ve been wary of feeding any of my unpublished work into generative AI models. Why? Because each time a journalist submits unpublished work to an LLM, whether to ask for headline help or power an investigation, they face significant questions about how their intellectual property will be used, and they open themselves up to novel legal and security risks.

That’s because those of us not working inside of an AI company actually have no idea what happens when we copy the text of our draft stories, paste them into a chat box, and press enter. And we have reasons to be concerned.

The AI industry, reportedly worth some three trillion dollars, was built upon the work of journalists whose words were hoovered up without compensation or consent. Half of the top ten websites used in chatbot training data were news outlets, [according]( to an investigation by the _Washington Post_ (whose own website ranked eleventh). Anthropic, for its part, tried to “destructively scan all the books in the world,” which involved slicing the spines off millions of volumes, for which it paid tens of millions of dollars. Thus, the creation of large language models was, as A.G. Sulzberger, the publisher of the _New York Times_, recently [described]( it, “a brazen theft of intellectual property that has occurred at an unprecedented scale.”

Then, when AI companies effectively ran out of words to use to keep improving their word-prediction machines, one solution was “synthetic data,” or, as [I wrote in 2024]( “information generated by AI itself, rather than humans, to continue to train their systems.” I [warned]( of a “feedback loop to hell” that this scenario could create as AI doubled down on its own biases and hallucinated falsities.

AI companies have managed to find a steady source of new, better-quality, human-written words to continue to train their new models on: the text entered into a chatbot’s prompt box. In a September 2025 [study]( of publicly available chatbots from Amazon, Anthropic, Google, Meta, Microsoft, and OpenAI, Stanford University researchers discovered “that all six developers appear to employ their users’ chat data to train and improve their models by default, and that some retain this data indefinitely.” Even Lex, an AI chatbot that is marketed as having been developed for writers and that [boasts]( an editing feature, [discloses]( that it “may collect your conversations in the AI Chat including both your inputs and the AI assistants outputs” (_sic_) for the purpose of “improving, upgrading, or enhancing” the platform. With AI companies’ prior scraping and current concerning use of data, I recommend journalists follow the lead of _The Markup_’s [newsroom policies]( and set a prohibition against feeding unpublished drafts into publicly available generative AI bots.

Enterprise-level (paid) models may be more secure, and basic agreements claim that they are. We can learn from the California State University system’s $16.9 million deal with OpenAI that these deals may provide more protections. According to recent [reporting]( from the _Times_ magazine, “the terms of the deal stipulate that OpenAI may not train its model on data from the C.S.U.” However, once content is entered into a commercial LLM, its ultimate use remains known only to the company itself. Journalists filing unpublished drafts into enterprise LLM models should follow the guidance of news outlets like the Associated Press, which entered a [licensing agreement]( with OpenAI in 2023 and still “[urges]( its journalists “to not put confidential or sensitive information into AI tools.”

The option with the least risk for a journalist’s unpublished drafts is a local large language model, a type of LLM that can run on a computer or server owned by a newsroom or journalist. The owners of local LLMs then have the power to determine how data like new text inputs are used or deleted. However, local LLMs are not common. So it would be interesting to see the [fruition]( of a real “Newsroom Tooling Alliance” (which was proposed last year by researchers) as a safer alternative for unpublished drafts. Another promising tool is [Lumo]( with chats that are encrypted and stored only locally.

But until these solutions are more widespread, the journalists whose work built the modern mythos of AI are being tendered limited options. Journalists can scrutinize the privacy policies of publicly available models, negotiate or opt out of allowing companies to use chat inputs for new training, and—at best—proceed with caution. But in the end, when a journalist clicks enter in an LLM chat box, whatever happens next is anyone’s guess.

_This piece was produced with support from the Craig Newmark Center for Journalism Ethics and Security._

_Has America ever needed a media defender more than now? Help us by [joining CJR today](

**Anika Collier Navaroli is an award-winning writer, lawyer, and researcher focused on journalism, social media, artificial intelligence, trust and safety, and technology policy. She is the Craig Newmark Assistant Professor of Professional Practice and director of the Craig Newmark Center for Journalism Ethics and Security at Columbia Journalism School. Over the past decade, her work has ranged from law firms and think tanks to advocacy organizations and senior policy positions at Twitter and Twitch.**

Share this fact check

← Check another article or image