core/evals/generalization/manifests/openbookqa.yaml
Shay 34dc0d64c5 feat(evals): add generalization benchmark manifests + policy (PR-1 of 2)
Adds policy document and sealed manifest records for the first 8
external audit datasets. No data is vendored. Local cache paths are
gitignored. Fetch/verify scripts and smoke fixtures come in PR-2.

Datasets: GSM1K, ASDiv, SVAMP, PARA-MAWPS, ARC-Easy, ARC-Challenge,
          OpenBookQA, CLUTRR
2026-06-23 05:59:36 -07:00

19 lines
788 B
YAML

dataset: OpenBookQA
purpose: sealed_audit_not_training
description: >
Elementary open-book science QA: ~6,000 questions over 1,329 core
science facts. Requires combining a retrieved fact with broader
commonsense/world knowledge. See arXiv:1809.02789.
source: hf://allenai/openbookqa
license: Apache-2.0 # confirm; Allenai typically Apache-2.0
version: pinned_release_or_commit
split: test
sha256: TODO_AFTER_DOWNLOAD
local_cache: .data/benchmarks/openbookqa/
repo_policy: manifest_only
inspection_policy: aggregate_reports_only
mutation_policy: no_direct_pack_policy_operator_mutation
smoke_fixture: null # PR-2: small smoke slice
notes: |
Evidence + commonsense audit. 5,957 questions, 1,329 core science facts.
Good probe for fact retrieval + inference composition in CORE.