XL-DocBench construction and human-verification pipeline Benchmark · 2026

01 / BENCHMARK

XL-DocBench

Evidence-grounded reasoning across hundreds or thousands of pages, with 1,519 human-verified questions and exact page-level support.

2,303
maximum context pages
72.6%
multi-page evidence
194
human experts
View project
DocAtlas mutable-state document interaction framework Agent system · 2026

02 / AGENT SYSTEM

DocAtlas

Long-document understanding as mutable-state interaction through Search, Read, Note, Review, and end-to-end reinforcement learning.

71.4%
MMLongBench-Doc
63.7%
RL-trained 4B policy
4
stateful tools
View project