Skip to content

Testing pyramid and markers

Kind: reference.

This page is the lookup for pytest markers, default commands, and what each layer can verify. Law lives in CLAUDE.md and AGENTS.md. Narrative and fixture labeling sit in Verified evidence and inferred claims.

Three layers

Layer Marker Path Default CI Coverage counted
Unit unit tests/unit/, evals/tests/unit/ yes yes
Contract contract tests/contract/, evals/tests/contract/ yes yes
Live live tests/live/, evals/tests/live/ no no

Default pytest excludes live (-m "not live" in pyproject.toml). The default suite must keep ≥ 90 coverage (tool.coverage.report.fail_under). A live pass does not replace unit or contract proof.

testpaths holds tests and evals/tests, so one uv run pytest runs both. The evals/tests/ layers test the typevet-evals workspace member.

What each layer proves

Unit. Pure domain rules, inbound wiring with fakes, adapter edge paths with controlled inputs. No real network and no live model weights.

Contract. The offline fake generation adapter and LlamaCppGenerationAdapter (sync) or AsyncLlamaCppGenerationAdapter (async) behave the same on shared fixtures under tests/fixtures/generation_contract.py. HTTP is replayed with httpx.MockTransport. Agreement verifies adapter compatibility for those labeled cases only.

Live. One exercised call against the configured router and model. Opt in with pytest -m live. See Run Gemma 4 on llama.cpp.

For release evidence, set TYPEVET_REQUIRE_LIVE=1 so missing router or model configuration fails live tests instead of skipping. See live eval runner.

The live eval runner adds an optional BoolQ/Banking77 slice with attempted / schema-valid / gold-match counters over GenerationPort. That is task accuracy on gold for a tiny limit, not ECE.

Valid JSON shape for a run is not the same as correct judgment. Do not infer calibration or task accuracy from pyramid passes alone unless the run used the loader eval runner and you report its gold-match counter explicitly.

Shared GenerationPort fixtures (judgevet shape)

Contract fixtures are synthetic: the test author defines the request, fake value or failure, and mocked chat-completion body. Each fixture has a name and label field for scope reporting on issues.

Sync parity: tests/contract/test_outbound.py. Async parity: tests/contract/test_async_outbound.py.

Add new port behaviour to the fixture list first, then extend fakes and the llama.cpp adapter until both sides agree. Do not weaken the default coverage floor or add a fourth pyramid layer to do it.

Shared JudgmentPort fixtures (judgevet shape)

Offline judgment contract fixtures live in tests/fixtures/judgment_contract.py. They exercise ContractJudgmentFake against labeled success and error cases. There is no live llama.cpp judgment adapter in the default suite yet.

Suite: tests/contract/test_judgment_port.py.

Scoring-backed judgment adapter fixtures live in tests/fixtures/judgment_scoring_contract.py. Suite: tests/contract/test_judgment_scoring_adapter.py.

Shared CandidateScoringPort fixtures

Offline scoring contract fixtures live in tests/fixtures/scoring_contract.py. ContractScoringFake there aliases public typevet.testing.ScriptedScoringFake. They exercise that fake against labeled success and fail-closed error cases (missing candidates, non-finite logprobs, unsupported stage). There is no live llama.cpp scoring adapter in the default suite yet.

Suite: tests/contract/test_scoring_port.py.

Commands

uv run pytest -m "unit or contract"
uv run pytest -m contract
uv run pytest -m live   # opt-in; not default CI
uv run pytest evals/tests/live/test_vllm_acceptance_live.py -m live -q -s
uv run pytest --cov=typevet --cov=typevet_evals --cov-report=term-missing
uv run mkdocs build --strict   # site build; also a unit test

Pre-commit runs unit and contract via the configured pytest hook. Import-linter contracts in pyproject.toml are unrelated to pytest contract markers.