Model cards

What each model does — and where it fails

None of these models writes a published number. They reduce manual work and surface problems; the scoring itself stays deterministic.

Document input extractor

Purpose
Reads supplier declarations and certificates to propose material, mass and packaging values.
Basis
Vision-language model prompted against a fixed schema. No fine-tuning on partner data.
Evaluation
Field-level precision measured on a held-out set of partner documents.

Known limits

  • Handwritten declarations degrade sharply.
  • Cannot verify a claim, only transcribe it — output always enters as supplier-declared at best.
  • Every field requires human confirmation before scoring.

Category classifier

Purpose
Assigns an incoming SKU to a scoring category so the correct baseline applies.
Basis
Text classification over product title, description and material fields.
Evaluation
Top-1 accuracy against human-assigned categories, reviewed each methodology release.

Known limits

  • Novel or hybrid product types default to the broadest matching category.
  • A misclassification changes the baseline, so category is human-reviewed before publication.

Anomaly flagger

Purpose
Flags submitted inputs that sit far outside category norms for manual review.
Basis
Statistical outlier detection over the historical input distribution per category.
Evaluation
Flag precision and reviewer agreement rate.

Known limits

  • Sparse categories produce noisy thresholds.
  • Flags are advisory: nothing is rejected automatically.

The deterministic half of the pipeline is described under platform overview.