AI citation accuracy on the CJR benchmark
Citation error rates reported by the Tow Center and Columbia Journalism Review, alongside PropensityAI's evaluated score of 22%.Breakdown: 152 correct · 24 partially incorrect · 20 completely incorrect · 4 not provided
Citation error rate
All evaluated systems
Model comparison
CJR citation error rates
Systems are ordered from lowest to highest error rate. PropensityAI's 22% score represents the lowest error rate among all tested systems.
| Rank | System | Citation error rate | Incorrect responses | Status |
|---|---|---|---|---|
| #1 | PAI PropensityAI Lowest error rate among all systems tested — evaluation complete 152 correct · 24 partially incorrect · 20 completely incorrect · 4 not provided | 22% | 44 of 200 | Evaluated |
| #2 | PPLX Perplexity Lowest reported error rate among the eight systems in the study | 37% | 74 of 200 | Published by CJR |
| #3 | GPT ChatGPT Search Incorrectly identified 134 articles | 67% | 134 of 200 | Published by CJR |
| #4 | GROK Grok 3 (beta) Highest reported error rate among the three systems shown here | 94% | 188 of 200 | Published by CJR |
Competitor results are reproduced from the Tow Center for Digital Journalism and Columbia Journalism Review article “AI Search Has a Citation Problem”. The source reports error rates, so lower percentages indicate better performance. Counts shown for Perplexity and Grok 3 are the corresponding percentages applied to the 200-prompt study size; ChatGPT's 134 incorrect identifications are explicitly reported by CJR. PropensityAI's 22% error rate (44 of 200) comes from the same reproducible evaluation harness, with the detailed grading: 152 correct, 24 partially incorrect, 20 completely incorrect, and 4 not provided.
Methodology
Identifying news from excerpts
The Tow Center supplied each AI search system with excerpts from news articles and asked it to identify the source article, including its headline, publisher, publication date and URL. The study covered 200 prompts drawn from 20 publishers.
Same prompt volume
Each evaluated system received the full citation-identification set.
Publisher breadth
The corpus spans a range of news organizations and access policies.
Citation task
Systems had to retrieve the correct story and provide accurate source details.
Headline metric
An incorrect identification counts against the system; lower is better.
FAQ
Citation benchmark questions
How to interpret the CJR results and PropensityAI's 22% score.
What does the CJR citation benchmark measure?+
The Tow Center for Digital Journalism tested whether AI search systems could identify a news article from an excerpt and provide accurate citation details. Each system received 200 prompts based on articles from 20 publishers.
Is a higher or lower score better?+
Lower is better. The published percentages are error rates: the share of prompts answered incorrectly. They should not be read as citation accuracy percentages.
How does PropensityAI's 22% error rate compare?+
PropensityAI's 22% error rate is the lowest among all systems tested in this benchmark, outperforming Perplexity (37%), ChatGPT Search (67%), and Grok 3 (94%). The breakdown: 152 correct, 24 partially incorrect (counted as errors), 20 completely incorrect, and 4 not provided. The error rate is calculated as (24+20)/200 = 22%.
Are these results from the original CJR study?+
Competitor results are from the original Tow Center and Columbia Journalism Review study published in March 2025. PropensityAI's score of 22% is from the same reproducible 200-prompt evaluation, completed on 14.08.2026.
Try source-backed search
Ask PropensityAI a current question and inspect the cited sources.