Skip to content

Receipt blocks shared by the eval runs

Kind: reference. The serving-metrics, server_args, wording parts, wording_digests and wording judge identity receipt blocks. The face, signature and check runs write the first two. The throughput sweeps write the server_args block. The wording receipts write the wording parts and the wording_digests block. Issue #347 moved the first two sections to this page.

Piece Module
/metrics read, changes and gauges typevet_evals.serving_metrics
server_args block typevet_evals.throughput.server_args
Wording parts block typevet_evals.wording.digests and typevet_evals.wording.parts
wording_digests block typevet_evals.wording.digests

Serving metrics block

The server key of the image-run throughput block holds the vLLM /metrics changes over the run. It is null for llama.cpp. The LFW reference lists the other throughput keys.

The server block holds the e2e, queue and prefill histogram changes, prefix_cache_hit_rate, counters and gauges. The counters are the changes of vllm:prefix_cache_hits, vllm:prefix_cache_queries, vllm:mm_cache_hits, vllm:mm_cache_queries, vllm:prompt_tokens, vllm:prompt_tokens_cached, vllm:generation_tokens, vllm:request_success and vllm:num_preemptions. The gauges are vllm:kv_cache_usage_perc, vllm:num_requests_running and vllm:num_requests_waiting, read before and after the run. A gauge is a point value, not a peak. A missing series is unknown. The two /metrics reads are outside the wall time. Requests from other clients of the same server also change the counters. The names come from vLLM v0.30.0.

Server arguments block

Issue #341 adds a server_args key to the receipts that record serving metrics. These are the face, signature and check receipts and the throughput sweep receipts. The key tells a reader which vLLM server configuration made the receipt. typevet_evals.throughput.server_args builds it.

Key Meaning
cache_config Labels of the vLLM gauge vllm:cache_config_info, as strings; null when no reading holds the gauge
cache_config_source Always metrics: the server reported cache_config on /metrics
caller_stated Value of TYPEVET_VLLM_SERVER_ARGS, unchanged; null when the variable is not set

The cache_config labels are the vLLM CacheConfig fields. Examples are block_size, enable_prefix_caching and gpu_memory_utilization. The block also keeps the engine label that the vLLM Prometheus logger adds. llama.cpp has no such gauge, so its cache_config is null. A gauge sample can carry a Prometheus timestamp after its value. The parser ignores the timestamp.

/metrics does not show the scheduler and model flags. Examples are --max-num-seqs, --max-num-batched-tokens and --logprobs-mode. Only caller_stated can hold them. It is what the caller stated, not what the server reported. typevet does not parse or check it.

typevet does not read /server_info. That route needs VLLM_SERVER_DEV_MODE=1, and the vLLM security guidance does not allow that mode in production.

An image run takes cache_config from the /metrics reading before the run. When that read fails, it uses the reading after the run. A throughput sweep uses the first level reading that holds the gauge. Neither run adds a /metrics read. Receipts written before #341 have no server_args key.

Wording parts block

Issue #363 adds the wording parts to the held-out receipt (#309), the comparison receipt (#329) and the evolution artifact. A wording candidate is a mapping of part name to text. The judgment text parts page lists the part names.

Key Meaning
seed_text The seed instructions text, verbatim
evolved_text The evolved instructions text, verbatim
components The part names the run evolved, in selection order
seed_parts The full seed mapping, part name to text
evolved_parts The full evolved mapping, part name to text

The two mappings have the same keys. A part that is not in components is frozen, so it has the same text in both mappings. In evolved_parts, each selected part is the gepa-adk evolved_components text, unchanged. seed_text and evolved_text stay for the #252 comparison.

A held-out or comparison receipt takes the parts from run.parts. score_held_out sets them from the WordingParts that it takes as evolved. WordingParts.instructions_only gives the parts of an evolved instructions text over the full seed mapping, with or without criteria. A run without parts records instructions only. components is then ["instructions"]. The receipt refuses parts whose instructions texts differ from seed_text or evolved_text.

The evolution artifact also records length_cap. It maps each evolved part to its cap: the floor of 1.5 times the length of the seed part. Before #363, length_cap was one number.

Issue #369 adds a part_table key to the evolution artifact and the held-out receipt. The key comes after seed_parts. The comparison receipt does not have it.

Key Meaning
part_table Each criterion part name to its Choice label, Noul key or Score level
dataset The dataset of the run, in a pubmedqa evolution artifact only

For a Noul seed with criteria, part_table maps criteria_true to true. It also maps criteria_false to false. A seed without criteria gives an empty part_table. For the PubMedQA Choice seed, each criteria_<label> maps to its label.

The evolution artifact of a DIFrauD run has no dataset key. Its dataset_revision is the pinned DIFrauD revision. A pubmedqa artifact sets dataset to qiaojin/PubMedQA pqa_labeled, and its dataset_revision is null. The held-out receipt names its dataset in pins.dataset. Receipts written before #369 have no part_table key.

Wording digests block

Issue #362 adds a wording_digests key to the held-out receipt (#309) and the comparison receipt (#329). Issue #363 adds it to the evolution artifact. A reader matches a receipt to a candidate by digest. A reader does not need to compare the full text. The receipts keep seed_text and evolved_text verbatim.

Each receipt digests two mappings: seed is seed_parts, and evolved is evolved_parts. Each arm holds these keys:

Key Meaning
components Part name to the SHA-256 hex digest of its UTF-8 text, for every part
mapping SHA-256 hex digest of the canonical JSON of the full mapping
gepa_candidate_id gepa-adk Candidate.id of the selected parts only

The shape, for a run that evolves instructions only:

{
  "seed": {
    "components": {"instructions": "<64 hex>"},
    "mapping": "<64 hex>",
    "gepa_candidate_id": "<12 hex>"
  },
  "evolved": {
    "components": {"instructions": "<64 hex>"},
    "mapping": "<64 hex>",
    "gepa_candidate_id": "<12 hex>"
  }
}

The digest rules, in Python. m is the full mapping, and s holds the parts named in components:

Digest Rule
component sha256(text.encode("utf-8")).hexdigest()
mapping sha256(json.dumps(m, sort_keys=True, separators=(",", ":")).encode("utf-8")).hexdigest()
gepa_candidate_id sha256(json.dumps(s, sort_keys=True, ensure_ascii=False).encode()).hexdigest()[:12]

The mapping rule uses the Python default and escapes non-ASCII text as \uXXXX. A tool that recomputes it outside Python must do the same. The mapping digest is not equal to gepa-adk Candidate.id. The gepa-adk rule keeps the default JSON spaces and does not escape non-ASCII text. It also keeps only 12 hex characters. The block therefore records both values. gepa-adk Candidate.id is in gepa_adk/domain/models.py.

A gepa-adk candidate holds only the parts that the run evolves. The wording run names each gepa-adk component by its part name. Thus gepa_candidate_id equals the candidate_id that the gepa-adk engine logs for the same candidate. When components names every part, it equals Candidate(components=evolved_parts).id. A unit test checks both cases against the installed gepa-adk.

Receipts written before #362 have no wording_digests key. Receipts written before #363 have no components, seed_parts or evolved_parts key. Their gepa_candidate_id does not match the engine log, because those runs named the gepa-adk component by the question key, for example is_scam.

Wording judge identity block

The wording evolution artifact records judge_provider and a judge_identity block. The live test evals/tests/live/test_wording_evolution_live.py writes them. TYPEVET_WORDING_JUDGE_PROVIDER selects the judge.

judge_provider judge_identity keys
gemma (default) requested_model, reported_models
jev requested_model, reported_models, base_url
ollama requested_model, reported_models, base_url, ollama_version, model_digest

Issue #333 adds the ollama judge. It uses the judgevet HTTPSystemOneAdapter with a placeholder key. It has no spend cap and no key check.

Variable Default Use
TYPEVET_OLLAMA_BASE http://localhost:11434 Ollama server URL
TYPEVET_WORDING_JUDGE nimble Judge model name
TYPEVET_OLLAMA_TIMEOUT 600 Read timeout, in seconds
TYPEVET_WORDING_CONCURRENCY 1 gepa-adk max_concurrent_evals

The test reads ollama_version from GET /api/version before the first judge call. It reads model_digest from the GET /api/tags entry of the model. A model name without a tag also matches the :latest tag. The test fails when /api/tags does not list the model. The failure message does not show the listed names.

With ollama, the held-out and comparison tests use the same judge. They need an evolution artifact whose judge_provider is ollama. Their receipts record backend ollama and the server facts. Their pins add a judge_identity key with the keys of the table above. The served_template pin is null. The run identity takes ollama_version as the server build. Receipts of the other backends have no judge_identity pin.