XL-DocBench

XL-DocBench

Benchmarking evidence-grounded extra-long document understanding.

Paper Code Dataset

Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡, Bei Liu2,*, Yifan Yang2, Qi Dai2, Ruichun Ma2, Kai Qiu2, Yunsheng Li2, Dongdong Chen2, Chong Luo2, Zhenzhong Chen1, Baining Guo2

1 Wuhan University2 Microsoft† Equal contribution‡ Work done during an internship at MSRA* Project leader

A human-verified benchmark for evidence-grounded reasoning across extra-long documents

Most document benchmarks test whether a model can answer. XL-DocBench also records where the answer comes from, what rule verifies it, and whether the available evidence is sufficient.

1,519

Questions verified end to end

Every retained item is re-answered by experts with full access to the source documents.

2,303

Professional-scale contexts

The median context is 211 pages, while the longest multi-document context spans 2,303 pages.

72.6%

Evidence beyond one page

1,103 questions require evidence distributed across multiple pages.

12

Diagnostic reasoning types

Twelve labels separate retrieval, set tracking, rule following, and abstention failures.

Six professional domains, one evidence standard

Select a domain to inspect its coverage in the final human-verified benchmark.

311questions

Legal & Regulatory

Why page-level evidence matters

The same wrong answer can come from very different system failures.

A model may miss the required page, read the right page but ignore a table, combine only part of the support, or answer even though a required condition is absent. XL-DocBench stores evidence pages, supporting snippets, answer formats, and typed verification rules so these failures can be distinguished.

The evidence-grounding loop
1

Locate

Search the complete document context.

2

Inspect

Read text, tables, charts, and figures.

3

Combine

Track all required support and exceptions.

4

Verify

Apply the typed rule or abstain.

From one page to 2,303 pages

Document QA has grown longer, but XL-DocBench also increases evidence dispersion, multimodal support, and cross-document scope.

AVERAGE CONTEXT PAGES2020 → 2026
DocVQA1.0MP-DocVQA8.3MMLongBench47.5LongDocURL85.6XL-DocBench297.1
297.1average context pages
2,303longest evaluation context
51questions over 1,000 pages

Tree-guided synthesis, verified by 194 human experts

Models propose difficult candidates from document branches; people establish the final answer, evidence pages, supporting quotes, and verification rule.

XL-DocBench tree-guided construction and human verification pipeline
XL-DocBench construction pipeline. Page-level evidence in the released benchmark comes from human verification.
01

Structure documents

Parse pages, text, tables, figures, and headings into a navigable tree.

02

Test branch dependency

Reject questions that remain answerable after a required branch is removed.

03

Refine and filter

Remove world-knowledge shortcuts, metadata leakage, weak support, and malformed answers.

04

Verify every item

Experts re-answer, mark exact pages and quotes, resolve ambiguity, and check the rule.

Current systems still fail on more than half of XL-DocBench

The strongest evaluated pipeline reaches 44.0% overall accuracy. Long windows help, but evidence selection and set tracking remain binding constraints.

44.0%best overall accuracy
1

A long window is not enough

Models with the same 1M-token window still differ at the same OCR budget.

2

Agents are not automatically better

With GPT-5.4, SimpleDoc reaches 44.0%, MDocAgent 35.0%, and DeepRead 32.2%.

3

Set tracking remains difficult

Best ranking, coverage, and set-difference accuracy stays below 37%.

SimpleDoc + GPT-5.444.0
SimpleDoc + Opus 4.640.6
Claude Opus 4.6 OCR39.8
SimpleDoc + GPT-5.238.3
GPT-5.4 OCR37.3

Same final accuracy, different failure modes

Breakdowns by context length, evidence count, and evidence span reveal whether a system loses information because the input is long or because the required support is dispersed.

Accuracy by context length, evidence-page count, and evidence span
Accuracy under context and evidence pressure for all evaluated systems.
01

Context pressure

Direct readers degrade as irrelevant pages grow, even before hard context limits bind.

02

Evidence pressure

Questions requiring several support pages expose incomplete retrieval and set tracking.

03

Span pressure

Evidence separated by hundreds of pages is harder to keep active and reconcile.

Retrieval is not the same as a correct answer

Evidence-hit curves and answer-yield curves separate page-search failures from evidence-use and rule-following failures.

Agent answer yield over retrieval rounds
Agent evidence hit over retrieval rounds

Inside a human-verified example

Each case exposes the question, evidence scope, reasoning type, and expert-verified support rather than showing only a final answer.

Multiple PDFsTemporal

Cross-document temporal reasoning

Cross-document temporal reasoning case study

Benchmark leaderboard across 27 evaluated systems

Filter by system family or inspect the full twelve-type reasoning breakdown. All scores come from the same deterministic evaluator and expert-verified ground truth.

Best values in the selected view are highlighted in red; second-best values are blue. Failed, missing, and unparsable predictions count as incorrect.

XL-DOCBENCH · 2026

Evidence-grounded reasoning across thousands of pages.

A fully human-verified benchmark for finding, combining, and validating evidence in extra-long professional documents.

Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡, Bei Liu2,*, Yifan Yang2, Qi Dai2, Ruichun Ma2, Kai Qiu2, Yunsheng Li2, Dongdong Chen2, Chong Luo2, Zhenzhong Chen1, Baining Guo2

1 Wuhan University2 Microsoft† Equal contribution* Project leader

CONTEXT SCALE2,303

pages in the longest evaluation context

1 page2,303 pages

A BENCHMARK BUILT FOR REAL DOCUMENT WORK

Long context is not enough.
Systems must find the right evidence.

Professional questions depend on policies, exceptions, tables, charts, and repeated measurements scattered across long files. XL-DocBench makes every answer auditable with expert-annotated evidence pages and typed verification rules.

1,519

Human-verified questions

Every retained item is re-answered and grounded by experts.

72.6%

Cross-page reasoning

1,103 questions require evidence from multiple pages.

36.6%

Multimodal evidence

Tables, charts, and figures are required, not decorative.

194

Human experts

Page support, answer rules, and ambiguity are checked manually.

WHY EVIDENCE-GROUNDED EVALUATION MATTERS

One answer can hide
four different failures.

A system can miss the right page, ignore part of the evidence, apply the wrong rule, or answer when support is absent. Final-answer accuracy alone cannot tell these failures apart.

THE EVIDENCE-GROUNDING LOOPREPEATS ACROSS PAGES AND DOCUMENTS
QUESTION

Which entity satisfies every requirement and exception?

TYPED RULE

All required conditions must hold; explicit exceptions override defaults.

VERIFIEDEntity A

3 evidence pages · rule satisfied

DOCUMENT QA, AT PROFESSIONAL SCALE

From one page
to hundreds.

Prior benchmarks made document QA longer. XL-DocBench moves to professional contexts while also increasing cross-page, cross-document, multimodal, and unanswerable coverage.

AVERAGE CONTEXT PAGES0100200300
DocVQA20201.0
MP-DocVQA20228.3
MMLongBench-Doc202447.5
LongDocURL202585.6
XL-DocBench2026297.1

3.5× longer on average than the closest prior benchmark · 2,303 pages at maximum

Benchmark positioning

Page-level evidence for
extra-long document reasoning.

XL-DocBench evaluates long-document systems under realistic document length, evidence dispersion, multimodal support, and explicit answer-verification conditions.

DOCUMENTS

331 public professional documents

Legal, finance, technical, medical, scientific, and narrative sources with contexts ranging from local documents to multi-PDF series.

ANNOTATION

Expert-verified evidence records

Every retained item includes answer support from annotated pages and quotes, together with an answer format and typed verification rule.

TASK CONDITIONS

Evidence spread and answer support

Multi-page, multimodal, cross-document, and None-answer examples expose distinct long-document failure modes.

72.6%multi-page evidence
36.6%table, chart, or image evidence
10.9%cross-document questions
14.4%None-answer cases
6professional domains
12diagnostic reasoning types

TABLE 01 · BENCHMARK LANDSCAPE

Scale, evidence, and scope.

XL-DocBench jointly increases context scale, evidence spread, cross-document scope, and answer-support diagnostics.

Benchmark Release Pages Tokens Cross-page Cross-doc. Unans. Source Avg. evidence pages
DocVQA2020-071.0151.5×××TXT/L/C/TAB/I1.0
ChartQA2022-031.0236.9×××C1.0
InfoVQA2021-041.2288.0×××L/C/TAB/I1.0
TAT-QA2022-071.1577.0×××TXT/TAB1.0
VisualWebBench2024-041.0452.4×××L/I1.0
MP-DocVQA2022-128.32,026.6×××TXT/L/C/TAB/I1.0
DUDE2023-055.71,831.5✓ (*)×✓ (*)TXT/L/C/TAB/I
SlideVQA2023-0120.02,030.5✓ (*)××TXT/L/C/TAB/I
MMLongBench-Doc2024-0747.521,214.1✓ (33.0%)×TXT/L/C/TAB/I1.88
LongDocURL2025-0785.643,622.6✓ (52.9%)×TXT/L/TAB/I1.53
XL-DocBench2026-04297.1227,463.2✓ (72.6%)TXT/TAB/C/I2.30

TXT/L/C/TAB/I: text, layout, chart, table, and image. (*) indicates the benchmark includes the capability but does not report the corresponding percentage.

Construction methodology

Tree-guided construction
and human verification.

Candidate generation operates over document-tree branches; human experts establish the final page-level evidence and verification labels.

FIG. 01Construction pipeline
XL-DocBench construction pipeline
  1. 01

    Structure documents

    Parse PDFs into pages, text, tables, figures, and hierarchical trees.

  2. 02

    Test branch dependency

    Keep candidates only when removing a selected branch removes necessary support.

  3. 03

    Refine and filter

    Reject local shortcuts, metadata leakage, malformed answers, and weak evidence.

  4. 04

    Verify with experts

    Experts re-answer, mark pages and quotes, and check the typed rule.

Dataset design and annotation

Reasoning labels, annotation records,
and dataset statistics.

Each example combines a reasoning label, evidence-grounded verification record, and context-level statistics for diagnostic evaluation.

2.1 REASONING TAXONOMY

T1

Core

Comparison · Reference chain

T2

Structural

Ranking · Coverage · Reconciliation

T3

Advanced

Set difference · Unanswerable · Temporal · Compliance · Counterfactual · Aggregation · Consistency

The labels separate evidence localization, set tracking, rule following, and abstention errors.

FIG. 02Reasoning labels and evidence structure
Distribution of reasoning labels and evidence structure across XL-DocBench

2.2 ANNOTATION RECORD

Each retained example pairs its answer with evidence pages, evidence snippets, an answer format, and a typed verification rule.

EVIDENCEPages 83 · 217 · 891

Traceable sources reveal whether a model found all required support.

RULEReconcile values under a stated condition.

Verification reflects the reasoning operation, not only the answer surface.

ANSWERVerified and normalized

Unsupported answers are measured explicitly with None-answer cases.

2.3 DATASET STATISTICS

Context length, evidence span, evidence count, modality, and answer format describe distinct forms of document-level difficulty.

FIG. 03Dataset composition
XL-DocBench context, evidence span, modality, and answer-format composition

XL-DocBench leaderboard

Evaluation results and
system comparison.

Direct long-context readers and PDF-based agents are evaluated against the same expert-verified answers and typed verification rules.

BEST OVERALL ACCURACY44.0%

SimpleDoc + GPT-5.4 still fails on more than half of the human-verified questions.

01

A long window is not enough.

GPT-5.4 and Claude Opus 4.6 both support 1M tokens, yet differ at the same OCR budget. Systems must still suppress irrelevant pages and keep the right evidence active.

02

Agents are not automatically better.

With GPT-5.4, SimpleDoc reaches 44.0%, while MDocAgent reaches 35.0% and DeepRead 32.2%. Retrieval quality determines whether the interface helps.

03

Set tracking remains the bottleneck.

Best ranking, coverage, and set-difference accuracy remains only 35.8%, 30.8%, and 36.4%.

PRIMARY METRIC · RULE-BASED ACCURACY

Scores are percentages. Failed, missing, and unparsable predictions are counted as incorrect.

MODEL FAMILY
INPUT

DIAGNOSTIC VIEW

Aggregate rank alone hides context-length, evidence-count, and evidence-span failure modes.

FIG. 04Accuracy under context and evidence pressure
Accuracy diagnostics by context length, evidence-page count, and evidence span

Retrieval analysis

Evidence retrieval and
answer-yield dynamics.

Evidence-hit dynamics separate page-search failures from evidence-use and rule-following failures.

FIG. 05Agent answer yield
Agent answer yield as retrieval rounds increase
FIG. 06Evidence-hit dynamics
Evidence hit rate as the number of retrieved and read page rounds increases

Case studies

Human-verified evidence
case studies.

These records make retrieval, evidence-use, and rule-following distinctions directly inspectable.

SELECTED EXAMPLES

Cross-document temporal reasoning

This example requires linking evidence across related documents rather than extracting a local phrase from a single page.

Evidence scope
Multiple PDFs
Reasoning type
Temporal
Ground truth
Expert verified
EXAMPLECross-document temporal
Cross-document temporal reasoning case study