Banking77 proxy label and Choice vs Noul metrics¶
Kind: reference. This page states the six-intent proxy label, how finvet routes questions in calibrate vs evolve, and which metric claims typevet may make on Banking77.
Parent: #51, #53. Loader work is #58. Partner policy: eval-partner-data-policy.md.
finvet is the source corpus for labels and questions. See finvet
The Banking77 fraud label,
Why finvet evolves on Banking77,
and finvet.data.banking77 (FRAUD_INTENTS in
banking77.py).
What Banking77 actually labels¶
PolyAI Banking77 (CC BY 4.0) assigns each short banking query one of 77
intent names. It does not label fraud. typevet and finvet both use a
proxy binary: fraud or not_fraud.
The six-intent proxy (finvet choice)¶
finvet maps six intent names to fraud. Every other intent is not_fraud.
The dataset card does not group intents. finvet made this choice. A different
group gives different agreement numbers.
Intent (Banking77 category) |
Proxy label |
|---|---|
card_payment_not_recognised |
fraud |
cash_withdrawal_not_recognised |
fraud |
direct_debit_payment_not_recognised |
fraud |
transaction_charged_twice |
fraud |
compromised_card |
fraud |
extra_charge_on_statement |
fraud |
| any other intent | not_fraud |
Two fraud intents do not literally say “unauthorized transaction”.
compromised_card and extra_charge_on_statement still count as fraud under
this rule. finvet documents that mismatch in
why-banking77.
typevet v1 adopts the same six-intent collapse as finvet unless a judgment issue overrides it (#53 acceptance).
Jev questions finvet uses on banking text¶
finvet defines three fraud-triage questions in FRAUD_QUESTIONS plus
is_scam from SCAM_QUESTIONS for the evolve agent only
(questions.py):
| Question name | Type | Role on Banking77 |
|---|---|---|
reports_unauthorized |
Noul | “Did the customer report a transaction they did not authorize?” |
fraud_type |
Choice (six fraud kinds + not_fraud + unclear) |
Finer fraud kind, not the 77 intents |
is_scam |
Noul | Scam/phishing wording; weak fit on bank-service queries |
The proxy gold label is still six-intent fraud / not_fraud. None of
the questions relabel the 77 intents directly.
Calibrate vs evolve routing (finvet)¶
These commands measure different things. Do not merge their tables without stating the command and the question.
finvet calibrate --dataset banking77¶
- Asks one Noul:
reports_unauthorizedon each message. - Compares Jev’s yes-probability to the proxy label (
fraud= positive). - Documented in finvet Run a calibration.
typevet does not copy finvet calibration ECE figures into eval claims until verification policy #50 accepts them. This page defines label and question policy, not calibration scores.
finvet evolve --dataset banking77¶
- The ADK agent calls
ask_jevonce per message. - The agent picks one of
is_scam,reports_unauthorized, orfraud_type(seeAGENT_QUESTIONSinagent.py). - The seed instruction lists all three names. gepa-adk may rewrite the
instruction to route by topic (for example unauthorized charges →
reports_unauthorized, exchange rates →fraud_type). finvet records an example in why-banking77.
Probability for the evolved Decision depends on the question:
| Question used | Positive (fraud) probability rule |
|---|---|
reports_unauthorized or is_scam |
Jev Noul yes-probability |
fraud_type |
Sum of Choice label probabilities except not_fraud and unclear (fraud_probability in finvet agent.py) |
Held-out agreement uses the same 0.5 threshold on that probability against
the proxy label. Question mix affects the score. Report question_used
distribution when you compare two evolve results.
Choice vs Noul — metric claims typevet may make¶
Accepted direction from #53: Noul-primary on Banking77, with optional
fraud_type Choice as a secondary probe—not a second gold standard.
| Claim | Allowed when |
|---|---|
Agreement (or log-loss) of reports_unauthorized Noul vs six-intent proxy |
Primary v1 Banking77 metric. State proxy label explicitly. |
Agreement of fraud_type Choice vs proxy or vs collapsed fraud probability |
Secondary. State that Choice labels are not the 77 intents and gold is still the proxy. |
| “Calibrated on Banking77” / ECE as typevet eval headline | Not yet. finvet live ECE is research context only until #50. |
| Direct numeric comparison of calibrate ECE (Noul-only) to evolve agreement (routed Noul+Choice) | No. Different commands, questions, and aggregation rules. |
Using is_scam as the default Banking77 metric |
No for v1. Bank queries are not scam SMS. finvet keeps is_scam in evolve for template parity with DIFrauD. |
When typevet exposes System One judgments (#11),
compile Decisions that mirror finvet names: Noul for reports_unauthorized,
Choice for fraud_type. JevBench (#24)
remains the protocol peer. Banking77 regression tracks finvet domain choices,
not JevBench substitution.
typevet loader (#58)¶
The loader is typevet_evals.datasets.banking77. It has the
six-intent FRAUD_INTENTS collapse from finvet, test-split CSV parsing, optional
balanced_sample / load_test_split(..., balanced=True), and
REPORTS_UNAUTHORIZED_NOUL_SCHEMA for the primary v1 Noul fixture. CI uses
checked-in CSV under tests/fixtures/banking77/ (no Hugging Face Hub). Optional
fraud_type Choice mapping stays documented here and in finvet; not required in
the first loader revision.
Sampling note¶
finvet banking77.load balances fraud and not_fraud rows up to --limit
(banking77.py).
Report limit, seed, and balance when you publish agreement numbers. A
balanced sample is not the raw intent distribution.
Related typevet pages¶
- CLINC150 domain shard map — complementary Choice
shards; no label mapping to Banking77 or
FRAUD_INTENTS(#66). - Eval partner data policy — public Banking77 vs partner NBA exclusion.
- TypeLLM, Jev and judgevet — Noul, Choice, and probability shape for a future judgment port.
- Glossary — Banking77 proxy label.