PropensityAI on the FreshQA benchmark
A transparent look at PropensityAI's performance on fresh, changing and false‑premise knowledge alongside published results from leading language models including PerplexityAI and Google.
Strict accuracy
PropensityAI vs published models
Model comparison
FreshQA results
Systems are ordered by relaxed accuracy. PropensityAI leads with scores from the official grading.
| Position | System | Relaxed accuracy | Strict accuracy | Sampled on |
|---|---|---|---|---|
| #1 | PAI PropensityAI All 500 questions graded | 95.6% | 94.0% | 2026‑08‑06 |
| #2 | PPLX PerplexityAI | 66.2% | 52.2% | 2023‑04‑26 |
| #3 | GOOG Google | 47.4% | 39.6% | 2023‑04‑26 |
| #4 | GPT GPT-4 | 46.4% | 28.6% | 2026‑05‑27 |
| #5 | GPT ChatGPT | 41.4% | 32.0% | 2026‑05‑27 |
| #6 | GPT GPT-3.5 | 32.4% | 26.0% | 2026‑05‑27 |
| #7 | Codex OpenAI Codex | 25.6% | 25.0% | 2026‑05‑27 |
| #8 | UN Flan-PaLM 540B | 23.6% | 23.4% | 2026‑05‑27 |
| #9 | UN PaLM 540B + chain-of-thought | 22.8% | 15.4% | 2026‑05‑27 |
| #10 | UN PaLM 540B + few-shot | 20.2% | 20.0% | 2026‑05‑27 |
| #11 | UN PaLM Chilla 62B | 15.0% | 12.2% | 2026‑05‑27 |
| #12 | UN PaLM 62B + few-shot | 14.2% | 12.8% | 2026‑05‑27 |
All scores are sampled from the FreshQA benchmark. Competitor scores are reproduced from the BenchmarkList repository; PropensityAI scores are from the official FreshQA evaluation and grading. PerplexityAI and Google scores are from published results (sampled April 26, 2023).
Methodology
What the evaluation covers
The evaluation ran PropensityAI’s retrieval-and-answer pipeline against every question in the FreshQA TEST split. Each response and its retrieved sources were stored and then graded under the official relaxed and strict rubrics.
False-premise
Tests whether the system corrects an invalid premise instead of accepting it.
Never-changing
Covers stable facts that should not depend on current web information.
Slow-changing
Measures knowledge that changes over months or years.
Fast-changing
Targets facts where current retrieval is especially important.
FAQ
FreshQA benchmark questions
A quick guide to interpreting this evaluation and its two accuracy measures.
What is the FreshQA benchmark?+
FreshQA evaluates whether language models and AI search systems can answer questions whose facts may change over time, while also detecting false premises and handling stable knowledge.
What is relaxed accuracy?+
Relaxed accuracy accepts answers that convey the required factual content even when wording or supporting detail differs from a strict reference format.
What is strict accuracy?+
Strict accuracy applies a tighter correctness standard. It is useful for measuring whether an answer fully satisfies the benchmark reference without material errors or omissions.
What are PropensityAI's scores on FreshQA?+
PropensityAI achieved 95.6% relaxed accuracy and 94.0% strict accuracy on the full 500‑question TEST set. These scores were obtained after completing all runs and applying the official FreshQA grading rubric (sampled on 2026‑08‑06).
How do PerplexityAI and Google compare on FreshQA?+
PerplexityAI scored 66.2% relaxed and 52.2% strict accuracy, while Google scored 47.4% relaxed and 39.6% strict accuracy on the FreshQA benchmark (both sampled on April 26, 2023).
Try the search system behind the evaluation
Ask a current, source-backed question with PropensityAI.