AI search evaluation · Full TEST set

PropensityAI on the FreshQA benchmark

A transparent look at PropensityAI's performance on fresh, changing and false‑premise knowledge alongside published results from leading language models including PerplexityAI and Google.

Strict accuracy

PropensityAI vs published models

Higher is better
PropensityAI94.0%
PerplexityAI52.2%
Google39.6%
ChatGPT32.0%
GPT-428.6%
GPT-3.526.0%
OpenAI Codex25.0%
Competitor scores sampled on 2026‑05‑27 (GPT‑4, ChatGPT, etc.) and 2023‑04‑26 (PerplexityAI, Google); PropensityAI sampled on 2026‑08‑06 via official FreshQA grading.
500
TEST questions
4
knowledge categories
98.8%
responses completed
95.6% / 94.0%
relaxed / strict accuracy

Model comparison

FreshQA results

Systems are ordered by relaxed accuracy. PropensityAI leads with scores from the official grading.

FreshQA relaxed and strict accuracy comparison
PositionSystemRelaxed accuracyStrict accuracySampled on
#1
PAI
PropensityAI
All 500 questions graded
95.6%94.0%2026‑08‑06
#2
PPLX
PerplexityAI
66.2%52.2%2023‑04‑26
#3
GOOG
Google
47.4%39.6%2023‑04‑26
#4
GPT
GPT-4
46.4%28.6%2026‑05‑27
#5
GPT
ChatGPT
41.4%32.0%2026‑05‑27
#6
GPT
GPT-3.5
32.4%26.0%2026‑05‑27
#7
Codex
OpenAI Codex
25.6%25.0%2026‑05‑27
#8
UN Flan-PaLM 540B
23.6%23.4%2026‑05‑27
#9
UN PaLM 540B + chain-of-thought
22.8%15.4%2026‑05‑27
#10
UN PaLM 540B + few-shot
20.2%20.0%2026‑05‑27
#11
UN PaLM Chilla 62B
15.0%12.2%2026‑05‑27
#12
UN PaLM 62B + few-shot
14.2%12.8%2026‑05‑27

All scores are sampled from the FreshQA benchmark. Competitor scores are reproduced from the BenchmarkList repository; PropensityAI scores are from the official FreshQA evaluation and grading. PerplexityAI and Google scores are from published results (sampled April 26, 2023).

Methodology

What the evaluation covers

The evaluation ran PropensityAI’s retrieval-and-answer pipeline against every question in the FreshQA TEST split. Each response and its retrieved sources were stored and then graded under the official relaxed and strict rubrics.

False-premise

Tests whether the system corrects an invalid premise instead of accepting it.

124 questions

Never-changing

Covers stable facts that should not depend on current web information.

125 questions

Slow-changing

Measures knowledge that changes over months or years.

121 questions

Fast-changing

Targets facts where current retrieval is especially important.

130 questions

FAQ

FreshQA benchmark questions

A quick guide to interpreting this evaluation and its two accuracy measures.

What is the FreshQA benchmark?+

FreshQA evaluates whether language models and AI search systems can answer questions whose facts may change over time, while also detecting false premises and handling stable knowledge.

What is relaxed accuracy?+

Relaxed accuracy accepts answers that convey the required factual content even when wording or supporting detail differs from a strict reference format.

What is strict accuracy?+

Strict accuracy applies a tighter correctness standard. It is useful for measuring whether an answer fully satisfies the benchmark reference without material errors or omissions.

What are PropensityAI's scores on FreshQA?+

PropensityAI achieved 95.6% relaxed accuracy and 94.0% strict accuracy on the full 500‑question TEST set. These scores were obtained after completing all runs and applying the official FreshQA grading rubric (sampled on 2026‑08‑06).

How do PerplexityAI and Google compare on FreshQA?+

PerplexityAI scored 66.2% relaxed and 52.2% strict accuracy, while Google scored 47.4% relaxed and 39.6% strict accuracy on the FreshQA benchmark (both sampled on April 26, 2023).

Try the search system behind the evaluation

Ask a current, source-backed question with PropensityAI.

Open PropensityAI