Model cards
What each model does — and where it fails
None of these models writes a published number. They reduce manual work and surface problems; the scoring itself stays deterministic.
Document input extractor
- Purpose
- Reads supplier declarations and certificates to propose material, mass and packaging values.
- Basis
- Vision-language model prompted against a fixed schema. No fine-tuning on partner data.
- Evaluation
- Field-level precision measured on a held-out set of partner documents.
Known limits
- Handwritten declarations degrade sharply.
- Cannot verify a claim, only transcribe it — output always enters as supplier-declared at best.
- Every field requires human confirmation before scoring.
Category classifier
- Purpose
- Assigns an incoming SKU to a scoring category so the correct baseline applies.
- Basis
- Text classification over product title, description and material fields.
- Evaluation
- Top-1 accuracy against human-assigned categories, reviewed each methodology release.
Known limits
- Novel or hybrid product types default to the broadest matching category.
- A misclassification changes the baseline, so category is human-reviewed before publication.
Anomaly flagger
- Purpose
- Flags submitted inputs that sit far outside category norms for manual review.
- Basis
- Statistical outlier detection over the historical input distribution per category.
- Evaluation
- Flag precision and reviewer agreement rate.
Known limits
- Sparse categories produce noisy thresholds.
- Flags are advisory: nothing is rejected automatically.
The deterministic half of the pipeline is described under platform overview.