AI search accuracy on the SealQA benchmark
A source-attributed view of published Seal-0 results on difficult factual questions with noisy and conflicting search evidence.
Published Seal-0 accuracy
Source-attributed results
Sourced comparison
SealQA results
Scored systems are ordered by published accuracy. PropensityAI leads with 56.0%.
| Rank | System | Accuracy | Status |
|---|---|---|---|
| #1 | PAI PropensityAI Grading complete · 56.0% accuracy | 56.0% | Scored |
| #2 | GPT GPT-5 Published by Parallel | 48.6% | Published |
| #3 | PPLX Perplexity Published in Search Evals | 40.5% | Published |
| #4 | PPLX Perplexity Deep Research Published by Parallel | 38.7% | Published |
| #5 | CLD Claude Opus 4.5 Reported in the cited arXiv paper | 26.6% | Published |
| #6 | BRV Brave Published in Search Evals | 22.5% | Published |
| #7 | GRK Grok 4 Reported by SealQA | 20.7% | Published |
| #8 | GEM Gemini 2.5 Pro Reported by SealQA | 19.8% | Published |
| #9 | O3 o3 Seal-0 with-search result | 17.1% | Published |
| #10 | YOU You.com Published by Tavily | 14.9% | Published |
| #11 | DS DeepSeek-R1-671B Seal-0 with-search result | 1.8% | Published |
Sources: SealQA, Parallel, Search Evals, Claude Opus 4.5 arXiv paper, and Tavily.
This is a source-attributed comparison, not a claim that every system was rerun under one controlled configuration. Retrieval providers, search depth, prompts, judge implementations and evaluation dates may differ; consult each linked source for its methodology.
Methodology
Testing search in difficult conditions
SealQA targets questions where useful evidence can be difficult to retrieve and web results may be noisy or contradictory. PropensityAI's harness runs the complete Seal-0 test split through its normal retrieval-and-answer pipeline and stores both answers and retrieval evidence.
Complete split
The pinned Seal-0 test set is validated before collection.
Leakage control
Gold answers, URLs, metadata and canary fields are withheld from the answer system.
Auditable collection
Responses, sources, retries and timings are retained for review.
Official-style labels
The judge returns CORRECT, INCORRECT or NOT_ATTEMPTED.
FAQ
SealQA benchmark questions
How to interpret the scores, sources and PropensityAI evaluation status.
What is the SealQA benchmark?+
SealQA evaluates search-augmented language models on difficult factual questions drawn from conflicting, noisy and hard-to-retrieve web evidence. Seal-0 contains 111 questions and is designed to expose weaknesses that easier factual benchmarks can miss.
Is a higher SealQA score better?+
Yes. The displayed figures are accuracy percentages: the share of evaluated questions graded CORRECT. Higher is better.
Why do the comparison rows have different sources?+
The table compiles explicitly attributed, publicly reported SealQA results. They were not all produced in one PropensityAI-controlled run, so retrieval implementations, evaluation dates and other settings may differ. The linked references are listed directly below the table.
What is PropensityAI's SealQA score?+
PropensityAI achieved 56.0% accuracy on the Seal-0 test split. All 111 responses were collected without errors, graded using the official GPT-4o-mini auto‑rater, and the final score is now complete.
How were PropensityAI's responses graded?+
The reproducible harness used the GPT-4o-mini auto-rater adapted from SimpleQA, returning the official labels CORRECT, INCORRECT and NOT_ATTEMPTED. Gold answers and benchmark metadata were withheld from the answer system until grading.
Try the search system under evaluation
Ask a difficult factual question and inspect PropensityAI's cited sources.