A visual inspection lab

See the defect.
Know the limits.

A useful classifier needs more than a prediction. Inspect real photographs, tune the review threshold, and see exactly where AI gets it wrong.

Explore the lab
From image to decision
01
Inspect the evidence

Real photos. Real model estimates.

02
Choose what needs a person

Clear, flag, or ask for human review.

03
Measure on unseen examples

Two policies. One held-out test set.

The inspection bench

Look closer. Decide better.

VisA photograph collection
Product family

Saved real API output · Oct 9, 2026, 4:23 a.m. UTC

Calibration / 01

Candle under inspection

VisA dataset
Candle selected for visual inspectionOriginal photograph

Known normal references

Two normal references used by policy v2
Separate normal candle reference 1
Reference 1
Separate normal candle reference 2
Reference 2

Separate reference photos. Never included in calibration or holdout metrics.

Inspection resultRecorded run
Anomaly suspected

The estimated anomaly probability is at or above the flag threshold.

Anomaly probability98.0%
Photo usability97.0%
Detected familyCandle
Family confidence100.0%
Visual severity2.00 / 3

Model estimates, not calibrated certainty. Severity describes visible appearance only.

Run details
Model
gpt-6-luna
Policy
Reference guided
Captured
Oct 9, 2026, 4:23 a.m. UTC
API latency
0.70 seconds
Input tokens
2,473

Severity distribution (none / subtle / clear / extensive): 0.0% · 0.0% · 100.0% · 0.0%

Your review

Your label is kept separately. It does not change dataset truth or retrain the model.

A clear result means no visible anomaly was detected under this policy. It does not certify function or safety.

Built for scrutiny

Trust starts with
knowing what it is.

01 / The photographs

A real, attributed dataset

A selected VisA subset from Amazon Science. Separate reference, calibration and holdout splits. Dataset labels remain separate from model outputs.

Dataset & CC BY 4.0 license
02 / The intelligence

Decisions, not a chatbot

OpenAI Decisions estimates anomaly and usability probabilities, identifies the product family, and scores visible severity. Code applies the routing rules.

Explore the API
03 / The boundary

Evidence before automation

Human corrections are review records, not automatic retraining. Saved API results are labelled. Live uploads remain in memory, with no demo-side image storage.

Read the implementation
Webytex for Business

What could your team
stop checking by hand?

We turn a specific operational problem into working software, with clear evidence of where AI helps and where a person still belongs.

Let’s discuss your workflow