Deniable

Testing & evaluation

Benchmarking datasets for repeatable testing and evaluation.

A useful benchmark is more than a score. It gives a team repeatable inputs, a clear testing question, and enough context to understand what changed between two runs.

Browse dataset packsBack to use cases

The testing job

Turn a change into a comparable run.

Benchmarking datasets help you compare a parser, classifier, retriever, or answer workflow across a known set of inputs. The point is to make a change visible and reviewable, not to suggest that one number captures every production risk.

  • Regression checks before a release
  • Prompt, model, or parser comparisons
  • Documented edge-case suites for QA
  • Evaluation runs that need the same inputs again

Useful checks

Keep the test question close to the result.

Use evaluation cases with your own labels, review process, and acceptance criteria. A benchmark becomes more useful when a failed case points to a concrete behaviour such as missed context, incorrect routing, or an unsupported answer.

  • Which cases changed between runs?
  • Can a reviewer inspect the input and output together?
  • Are changes separated from random variation?
  • Does the suite cover the workflow you actually operate?

Deniable approach

Human input gives a benchmark its purpose.

Deniable combines human review with creative scenario design so cases are built around a failure mode or evaluation question. The result is not a ranking promise and does not replace your own production acceptance tests.

Catalogue note

Start with the evaluation path you need.

Fixed packs will be delivered as ZIP files, and authenticated API retrieval will use prepaid credits for selected documents. The catalogue will state the available metadata, labels, and release scope before purchase or use.