Namaste Ji by Ayushman Dash

Exhibit · The QA theatre

Every greeting faces the judge — watch the line run.

Nothing reaches a phone unchecked. Cards queue up, the QA Agent reads seven things off each one, and a stamp comes down — ship, flag or veto. Every score and verdict on this line is real, from a real check-run on staging. Click any card to read why.

〜 the line never stops ✓ real scorecards ⊙ beat the QA Agent below
The honest sample

It doesn't check everything — just enough to be sure.

Every dot is a waiting greeting; the colour says which brief it came from. The QA Agent picks a fair mix from every brief — never-checked ones first — and just enough of them to trust the answer. It's the same math pollsters use.

population
61
strata
14 briefs
confidence
95%
margin
±10%
sample
n = 31

n = z²p(1−p)/e² · FPC · z=1.96 · p=0.2 · e=0.10 · N=61

A stricter answer needs a bigger sample, but never more than everything — slide the dial and watch n move.

The one rule that can't be argued with

The veto is code, not opinion.

A judge model weighs all seven scores and can be generous. But if a card fails on culture or safety, code locks it — the judge can't lift it, and no good average can wash it out.

This is the actual guard. The judge's verdict is an input; the veto is computed from the two veto metrics and cannot be averaged away.

The gate

What ships, what stays back.

Publishing is one decision on the whole run: the checked-and-passed ship, the unchecked ship with them — that's what the sample vouched for — and the flagged and vetoed stay back. One click rolls it all back.

32would ship2 passed + 30 vouched for by the sample
29held back28 flag · 1 veto
nothing has moved

On staging this run evaluated 31 of 61 (⌀ 86/100): 2 shipped clean, 29 were flagged — most for the missing-overlay-text defect the judge caught across the batch — and 1 was vetoed for frightening imagery. That inbox is real. See the real run in the console

Beat the QA Agent

Which one got flagged?

Three real greetings from the same run. One of them was flagged — and not for anything you'd see at a glance. Pick.

How it's really built

Population = every staged greeting; a candidate set is frozen per run; the sample is stratified by brief, risk-weighted (never-scored first), sized by Wilson with FPC. Each metric evaluator is a versioned Langfuse prompt run against the image + its intended context; the LLM judge integrates; the veto (cultural-correctness, safety) is enforced in code; a run-scoped publish gate ships what passed and holds what didn't, with rollback.

Why it exists

Pretty images that don't forward are worthless; wrong images that do forward are worse. You can't eyeball thousands of cards every day, and you can't let a model be the last word on culture — so: sample honestly, measure explicitly, judge transparently, veto in code, and keep a human able to overrule. That is the whole QA philosophy in one sentence.