- Cultural correctness100
- Safety95
- Text legibility95
- Aesthetic quality95
- Forward appeal85
- Brand fit·
- Prompt adherence·
Passed everything. The publish gate can ship these.
Something's off — wrong text, missing motif. Waits in the inbox for a person.
Cultural or safety fail. Code keeps it locked; only an audited human override opens it.
rehearsed flow — every score and verdict is from a real QA run
It doesn't check everything — just enough to be sure.
Every dot is a waiting greeting; the colour says which brief it came from. The QA Agent picks a fair mix from every brief — never-checked ones first — and just enough of them to trust the answer. It's the same math pollsters use.
- population
- 61
- strata
- 14 briefs
- confidence
- 95%
- margin
- ±10%
- sample
- n = 31
n = z²p(1−p)/e² · FPC · z=1.96 · p=0.2 · e=0.10 · N=61
A stricter answer needs a bigger sample, but never more than everything — slide the dial and watch n move.
The veto is code, not opinion.
A judge model weighs all seven scores and can be generous. But if a card fails on culture or safety, code locks it — the judge can't lift it, and no good average can wash it out.
This is the actual guard. The judge's verdict is an input; the veto is computed from the two veto metrics and cannot be averaged away.
What ships, what stays back.
Publishing is one decision on the whole run: the checked-and-passed ship, the unchecked ship with them — that's what the sample vouched for — and the flagged and vetoed stay back. One click rolls it all back.
On staging this run evaluated 31 of 61 (⌀ 86/100): 2 shipped clean, 29 were flagged — most for the missing-overlay-text defect the judge caught across the batch — and 1 was vetoed for frightening imagery. That inbox is real. See the real run in the console ↗
Which one got flagged?
Three real greetings from the same run. One of them was flagged — and not for anything you'd see at a glance. Pick.
How it's really built
Population = every staged greeting; a candidate set is frozen per run; the sample is stratified by brief, risk-weighted (never-scored first), sized by Wilson with FPC. Each metric evaluator is a versioned Langfuse prompt run against the image + its intended context; the LLM judge integrates; the veto (cultural-correctness, safety) is enforced in code; a run-scoped publish gate ships what passed and holds what didn't, with rollback.
Why it exists
Pretty images that don't forward are worthless; wrong images that do forward are worse. You can't eyeball thousands of cards every day, and you can't let a model be the last word on culture — so: sample honestly, measure explicitly, judge transparently, veto in code, and keep a human able to overrule. That is the whole QA philosophy in one sentence.