Deniable

Document category · long-form documents and ebooks

Long-form document and ebook datasets for context-heavy testing.

Long documents create a different testing problem from short examples. The answer may depend on a definition, table, or qualification several sections away. Long-form test documents help you inspect that context deliberately.

Browse dataset packsBack to use cases

The testing job

Test the distance between a question and its evidence.

Long-form retrieval and summarisation can fail through lost headings, weak chunk boundaries, or a conclusion that ignores an earlier exception. A document dataset gives you repeatable material for checking those failure modes.

  • Definitions introduced before they are used
  • Sections with similar language but different scope
  • Tables, references, and qualifications around the main text
  • Questions that require more than one passage

Useful checks

Review context, structure, and citations together.

Use long-form documents for retrieval evaluation, section-aware summaries, document search, question answering, and citation review. The source structure should remain visible enough for a reviewer to follow the answer back to its evidence.

  • Did the workflow preserve section boundaries?
  • Did it retrieve the passage with the needed qualification?
  • Does the summary distinguish findings from background?
  • Can a reviewer verify each important claim?

Deniable approach

Designed scenarios make long context testable.

Human input and creative scenario design decide which relationships matter in each document. The aim is to provide a useful internal test condition, not to imply that a synthetic long-form document represents a publisher, course, or knowledge base.

Catalogue note

Check the document scope before ingestion.

Catalogue pages will state the exact length, structure, and metadata for each release. Fixed packs are delivered as ZIP files, while authenticated API retrieval uses prepaid credits for workflows that fetch selected documents.