Every claim the repository makes, with its source statement, its evidence, and the limitation that bounds it. Statuses come from the repository's own vocabulary — a Planned or Unsupported claim keeps its evidence absence visible.
Measured TF-IDF baseline: Accuracy = 0.8653 (measured on the untouched test split) ml
Measured DistilBERT: Accuracy = 0.6955 (measured on the untouched test split) ml
Measured TF-IDF baseline: Macro-F1 = 0.8654 (measured on the untouched test split) ml
Measured DistilBERT: Macro-F1 = 0.6620 (measured on the untouched test split) ml
Measured TF-IDF baseline: Coverage at threshold = 0.7500 (measured on the untouched test split) ml
Measured DistilBERT: Coverage at threshold = 0.6799 (measured on the untouched test split) ml
Measured TF-IDF baseline: Accepted accuracy = 0.9333 (measured on the untouched test split) ml
Measured DistilBERT: Accepted accuracy = 0.8161 (measured on the untouched test split) ml
Measured TF-IDF baseline: Selective risk = 0.0667 (measured on the untouched test split) ml
Measured DistilBERT: Selective risk = 0.1839 (measured on the untouched test split) ml
Measured TF-IDF baseline: Expected calibration error = 0.4883 (measured on the untouched test split) ml
Measured DistilBERT: Expected calibration error = 0.4697 (measured on the untouched test split) ml
Measured DistilBERT (validation-selected): Persisted threshold = 0.16841767053420467 (measured on the untouched test split) ml
Implemented U01 — Project foundation and reproducibility: Implemented ml
Implemented U02 — BANKING77 data contract: Implemented ml
Implemented U03 — TF-IDF logistic-regression baseline: Implemented ml
Implemented U04 — DistilBERT training and immutable artifact: Implemented ml
Implemented U05 — Comparative evaluation and selective prediction: Implemented ml
Implemented U06 — FastAPI inference and real-artifact demo: Implemented ml
Implemented U07 — Full validation and acceptance gate: Implemented ml
Implemented U08 — Documentation and recruiter-ready delivery: Implemented ml
Measured the fine-tuned transformer loses to the lexical baseline ml
Measured sealed transformer bundle that this repository does not track ml
Measured A validation-only threshold ml
Measured Evaluation on the untouched test split ml
Measured The baseline wins by 0.2034 macro-F1 ml
Limited A confidence score here is a ranking signal for abstention, not a probability of correctness ml
Limited The curated unsupported-request check is a behavioral check, not an out-of-distribution benchmark ml
Limited Descriptive for that machine, not a service level ml
Measured 43 passed, 1 not evidenced, 0 blocked, with an empty cause list ml
Measured That is precisely why the promotion came last ml
Measured That CI run is the first clean-runner audit of a tree where U08 is `Implemented` ml
Measured in 19.14s, and the acceptance audit at the verdict above. The audit recipe ml
Measured : the step logs contain no `##[error]`, no `make: ml
Measured The CI acceptance step is now enforced ml
Measured NFR-001's RAM clause passes as `Measured` ml
Measured NFR-001's GPU clause is `not_applicable`, not an unmeasured MUST ml
Measured The 585/563 contradiction is resolved above ml
Limited The curated unsupported-request fixture is a behavioral check, not an OOD benchmark ml
Limited Neither model's confidence is a probability of correctness ml
Limited The demo is local-only; the acceptance audit is evidenced by a fully evidenced local run and a degraded CI run — by declaration, not as a deficiency awaiting a fix ml
Limited The CI acceptance step is now enforced, and a step conclusion is still not a substitute for its logs ml
Limited [Makefile:38: acceptance] Error 1` followed by `##[error]Process completed with exit code 2`, because `make` exits 2 when a recipe fails while `scripts/validate_acceptance.py` itself exits 1 ml
Limited Latency is descriptive, not a service level ml
Planned Docker is POST-WEEKEND. operations
Measured Strict-MVP verdict: PASS — issued by `make acceptance`, which exits zero only when no cause remains ml
Limited BANKING77 is pinned to `1fb62b1bb4635df59a8e1b2f2bc5e0643b2856c8`; data preparation validates its source split contract and records local provenance. ml
Limited The DistilBERT base model is pinned to `12040accade4e8a0f71eabdb258fecc2e7e948be` and one CPU fine-tune has been run. ml
Limited Nothing was tuned in response to that result ml
Limited The transformer's confidence never exceeded 0.6000 on any test example, leaving six of the fifteen bins empty ml
Limited Each model's ECE happens to equal its aggregate confidence-accuracy gap here because every occupied bin errs in the same direction ml
Limited Reported ECE is only comparable against another ECE computed under the same binning — 15 equal-width bins, left-closed and right-open with the final bin closed, empty bins excluded from the average. ml
Limited The audit refuses to overwrite a fully-evidenced `reports/acceptance.json` with a degraded one ml
Limited Measured metrics are the baseline's and transformer's test accuracy, macro-F1, weighted-F1, and per-class figures; the transformer's validation metrics and selected threshold; both models' test cover… ml
Limited Test macro-F1 and weighted-F1 agree to reported precision for both models because the test split is exactly balanced at 40 examples per class, so the weights are uniform ml
Limited Both models were evaluated at the same threshold, which was selected from the transformer's validation confidences ml
Unsupported CUDA and GPU compatibility are unverified ml
Limited A synthetic dataset, frozen embeddings, deterministic API predictor, or one-epoch training path would be degraded evidence and cannot satisfy a violated MUST requirement. ml
Planned Docker and all deployment work are POST-WEEKEND. ml