Difficult AI search evaluation · Seal-0

AI search accuracy on the SealQA benchmark

A source-attributed view of published Seal-0 results on difficult factual questions with noisy and conflicting search evidence.

PropensityAI scored 56.0%. All 111 responses graded; result is final.

Published Seal-0 accuracy

Source-attributed results

Higher is better
PropensityAI56.0%
GPT-548.6%
Perplexity40.5%
Perplexity Deep Research38.7%
Claude Opus 4.526.6%
Brave22.5%
Grok 420.7%
Gemini 2.5 Pro19.8%
o317.1%
You.com14.9%
DeepSeek-R1-671B1.8%
Compiled from SealQA, Parallel, Search Evals, the cited arXiv paper and Tavily. Settings and evaluation dates may differ by source.
111
questions in Seal-0
111 / 111
PropensityAI responses collected
0
response errors
56.0%
official-style grading

Sourced comparison

SealQA results

Scored systems are ordered by published accuracy. PropensityAI leads with 56.0%.

Published SealQA accuracy scores and PropensityAI evaluation status
RankSystemAccuracyStatus
#1
PAI
PropensityAI
Grading complete · 56.0% accuracy
56.0%Scored
#2
GPT
GPT-5
Published by Parallel
48.6%Published
#3
PPLX
Perplexity
Published in Search Evals
40.5%Published
#4
PPLX
Perplexity Deep Research
Published by Parallel
38.7%Published
#5
CLD
Claude Opus 4.5
Reported in the cited arXiv paper
26.6%Published
#6
BRV
Brave
Published in Search Evals
22.5%Published
#7
GRK
Grok 4
Reported by SealQA
20.7%Published
#8
GEM
Gemini 2.5 Pro
Reported by SealQA
19.8%Published
#9
O3
o3
Seal-0 with-search result
17.1%Published
#10
YOU
You.com
Published by Tavily
14.9%Published
#11
DS
DeepSeek-R1-671B
Seal-0 with-search result
1.8%Published

Sources: SealQA, Parallel, Search Evals, Claude Opus 4.5 arXiv paper, and Tavily.

This is a source-attributed comparison, not a claim that every system was rerun under one controlled configuration. Retrieval providers, search depth, prompts, judge implementations and evaluation dates may differ; consult each linked source for its methodology.

Methodology

Testing search in difficult conditions

SealQA targets questions where useful evidence can be difficult to retrieve and web results may be noisy or contradictory. PropensityAI's harness runs the complete Seal-0 test split through its normal retrieval-and-answer pipeline and stores both answers and retrieval evidence.

Complete split

The pinned Seal-0 test set is validated before collection.

111 questions

Leakage control

Gold answers, URLs, metadata and canary fields are withheld from the answer system.

Question only

Auditable collection

Responses, sources, retries and timings are retained for review.

111 completed

Official-style labels

The judge returns CORRECT, INCORRECT or NOT_ATTEMPTED.

3 labels

FAQ

SealQA benchmark questions

How to interpret the scores, sources and PropensityAI evaluation status.

What is the SealQA benchmark?+

SealQA evaluates search-augmented language models on difficult factual questions drawn from conflicting, noisy and hard-to-retrieve web evidence. Seal-0 contains 111 questions and is designed to expose weaknesses that easier factual benchmarks can miss.

Is a higher SealQA score better?+

Yes. The displayed figures are accuracy percentages: the share of evaluated questions graded CORRECT. Higher is better.

Why do the comparison rows have different sources?+

The table compiles explicitly attributed, publicly reported SealQA results. They were not all produced in one PropensityAI-controlled run, so retrieval implementations, evaluation dates and other settings may differ. The linked references are listed directly below the table.

What is PropensityAI's SealQA score?+

PropensityAI achieved 56.0% accuracy on the Seal-0 test split. All 111 responses were collected without errors, graded using the official GPT-4o-mini auto‑rater, and the final score is now complete.

How were PropensityAI's responses graded?+

The reproducible harness used the GPT-4o-mini auto-rater adapted from SimpleQA, returning the official labels CORRECT, INCORRECT and NOT_ATTEMPTED. Gold answers and benchmark metadata were withheld from the answer system until grading.

Try the search system under evaluation

Ask a difficult factual question and inspect PropensityAI's cited sources.

Open PropensityAI