Skip to content

TPJEP v0 eight-task runner

Kind: reference.

Parent: #106, #131 (records), #130 (pins).

Loads eight vendored JevBench rows, maps them to native Noul / Choice / Score questions, and runs them through JudgmentPort. Gold expected values never enter model inputs.

Modules

Piece Location
Loader + pins typevet_evals.tpjep.loader
Runner + receipt metadata typevet_evals.tpjep.runner
Answer → record scoring typevet_evals.tpjep.outcome
Eight-row fixture tests/fixtures/tpjep/eight_task_smoke.jsonl
Provenance tests/fixtures/tpjep/PROVENANCE.md

Run metadata records both dataset hashes from #130: manifest (typellm_manifest_sha256) on each attempt row, plus manifest and local concat hashes on TpjepRunMetadata.

Commands

Offline:

uv run pytest evals/tests/unit/test_eval_tpjep_loader.py evals/tests/contract/test_tpjep_runner_offline.py -q

Live (requires local llama.cpp router; skips are not passes):

uv run pytest evals/tests/live/test_tpjep_smoke_live.py -m live -q

Live JSONL and summary are written under scratchpad/tpjep/ (gitignored).