What is DocInsights?

DocInsights is the Workshop on Document Intelligence and Understanding, co-located with EMNLP 2026. It focuses on document-centered NLP beyond plain text: structure-aware modeling, multimodal evidence, grounding, retrieval, evaluation, and deployment.

What submission types are accepted?

We welcome long papers, short papers, position and theory papers, benchmark and dataset papers, survey papers, system demos, reproducibility studies, and negative or diagnostic analyses.

Can I submit non-archival work?

Yes. DocInsights accepts direct non-archival submissions. Non-archival work will not appear in proceedings and may be submitted elsewhere later, subject to the policies of those venues.

How do ARR commitments work?

Eligible papers with completed ACL Rolling Review reviews and meta-review may be committed through the DocInsights ARR Commitment OpenReview group. Decisions are based on ARR reviews, meta-review, relevance, and fit to the workshop.

Are submissions double-blind?

Direct submissions must be anonymized for double-blind review. The review process will use OpenReview and at least three reviews plus meta-review.

What are the page limits?

Long papers may include up to 8 pages of content. Short papers may include up to 4 pages of content. References and appendices are unlimited.

What challenges are running?

DocInsights hosts two distinct challenges: DocSem for document-grounded quantitative reasoning with evidence attribution, and Dr.DocBench for expert-level parsing of complex documents.

When does the competition run?

DocSem ran from August 3 through September 10, 2026, and its final test leaderboard is now public. Dr.DocBench runs from August 10 through October 10, 2026, with submissions closing at 12:59 PM UTC. Each official portal is the system of record for exact submission status.

What is the challenge prize pool?

The combined prize pool will exceed USD 5,000. Track allocations, eligibility requirements, team limits, and award conditions will be announced with the final rules.

Can challenge participants present their work?

Yes. Participants can still submit concise system papers for workshop consideration. Selected contributions will be invited to present at DocInsights 2026.

How do I submit a shared-task system paper?

Participants in DocSem or Dr.DocBench can still submit a system paper through the shared-task OpenReview venue. The deadline is September 15, 2026 at 11:59 PM UTC. Authors may indicate an archival or non-archival preference; the review committee will make the final archival/non-archival decision.

Was the DocSem dataset updated?

Yes. The training split was updated on August 31, 2026 to correct seven annotation inconsistencies following community feedback. The changes affect only training data; the task definition and data format are unchanged. Use the latest public dataset release.

Were DocSem validation labels corrected?

Yes. Three organizer-only validation ground-truth labels have now been corrected, most recently on September 3, 2026, following additional data review. All existing submissions were rescored, and the leaderboard now reflects the updated results. Public validation inputs, task definition, and data format are unchanged.

Where are the final DocSem test results?

The DocSem portal now opens on the Final test leaderboard. It publishes each Hugging Face account's best eligible attempt from up to three accepted test submissions, with the public columns Rank, Team, Submission name, Selected attempt, Total attempts, and Joint Exact Accuracy. The portal has separate Validation leaderboard and Final test leaderboard views.

The public test dataset contains 1,730 held-out tasks and PDFs without labels. Submissions could contain any non-empty subset of task IDs. Omitted tasks count as incorrect against the full split, and "answer": null with "evidence": [] represents an abstention.

Will the workshop support remote participation?

Participation details will follow EMNLP guidance and organizer confirmation. The site will be updated when room, schedule, and participation modes are final.

When will speakers and the program be announced?

Confirmed speakers, accepted papers, and the final program will be posted after speaker confirmations, review decisions, and EMNLP schedule details are public-ready.