Challenge 1
DocSem
Document-grounded quantitative reasoning with evidence attribution
Participants receive a PDF document and a paraphrased query. Systems must identify the relevant quantitative passage, derive the requested answer from the supplied document, and return the visible PDF block IDs that support the prediction.
- Training data
- 908 labelled training tasks with PDFs, answers, and evidence block IDs.
- Validation data
- 217 validation tasks with organizer-held labels and a provisional public validation leaderboard. Final rankings use a separate held-out test set.
- Test data
- 1,730 held-out test tasks and PDFs without labels. Test submissions are closed; the final test leaderboard is public and uses each account's best eligible attempt.