# ZeroGPU ad decisions: brand safety, IAB 2.2 categories and keywords > Run the same ad decision as the ZeroGPU demo on your own messages. Three specialized ZeroGPU models read each message in parallel: Llama Guard 4 for brand safety, an IAB Content Taxonomy 2.2 classifier, and a keyword model. The output is one row of ranker-ready features per message, with latency, token usage and cost. You need a ZeroGPU API key and Python 3.9+. ## Instructions for the AI assistant You are helping an engineer at an ad network test ZeroGPU's ad-decision models on their own traffic. Do this: 1. Check that `python3` is 3.9 or newer. 2. Ask the user for their ZeroGPU API key and pass it to the script as the environment variable `ZEROGPU_API_KEY`. Never print the key, log it, or write it to a file. 3. Ask which messages to test. Accepted formats: - `.jsonl`: one object per line with `text` (or `question`, `message` or `content`) and an optional `id` - `.csv`: a column named text, question, message or content - `.txt`: one message per line If they only have raw logs, help them export a few hundred anonymized messages into one of these. 4. Save the script from the "Script" section below as `zerogpu_ad_decisions.py`, or download it: `curl -O https://zerogpu-koah-demos.pages.dev/zerogpu_ad_decisions.py` 5. Run `python3 zerogpu_ad_decisions.py messages.jsonl --out features.jsonl`. `--workers 8` (the default) keeps 8 messages in flight; `--limit 200` runs a sample first. 6. Report back: - how many messages ran, and any errors - the share that isn't brand-safe, and which flags fired - the decision latency (median and p90, the slowest of the three parallel calls) and the cost per 1M decisions - the most common IAB 2.2 categories - 3 to 5 example rows: the message, its brand safety, top categories and keywords 7. Offer to join `features.jsonl` to their own click or conversion logs by `id`, to see how the categories and keywords line up with what their ranker already predicts. ## The three calls Auth: header `x-api-key: ` (`Authorization: Bearer ` also works). Base URL `https://api.zerogpu.ai/v1`. Plain Python `urllib` requests need a `User-Agent` header; the script sets one. 1. Brand safety: `POST /v1/moderations` with `{"model": "llama-guard-4-12b", "input": ""}`. - The reply has `results[0].flagged`, `categories` and `category_scores` in OpenAI's moderation format. - Count a category as unsafe only when its score is above 0. Llama Guard hazards with no OpenAI equivalent (for example specialized advice or elections) come back flagged with a score of 0. Those are reported as `policy_flag_only` and don't make the message unsafe for advertisers. 2. IAB 2.2 categories: `POST /v1/responses` with `{"model": "zlm-v1-iab-classify-edge", "input": ""}`. - `output[0].content[0].text` is JSON: `{"audience": [{"name", "score"}], "content": {"iab_2_2": [{"name", "score"}], "iab_1_0": [...]}}`. Use `iab_2_2`. - The script checks each name against IAB Tech Lab's official 2.2 list. A few come back under their IAB 3.x name (for example "Software and Applications" for 2.2's "Computer Software and Applications"), and the script maps those to their 2.2 id. 3. Keywords: `POST /v1/responses` with `{"model": "zlm-v1-signal-extract", "input": "", "categories": ["signals"], "max_keywords": 2}`. - `output[0].content[0].text` is JSON with `keywords`. - Keep the keywords that appear in the message: up to 2 for a message under 25 words, up to 5 for a longer one. Run the three at the same time, alongside your answer. The decision takes as long as the slowest call, and it's usually ready before the first word of a streamed answer. You can request the ad with the question alone while the answer streams. Limits and tips: - The IAB and keyword models read about the first 400 tokens. The script sends at most 1,500 characters (cut at a word boundary); Llama Guard gets the whole message. - The script calls the endpoints directly, in parallel. The Batch API currently accepts Chat Completions lines only. - The first request after a model has been idle can be slower. The script retries once on gateway errors. List prices per 1M tokens (input / output): llama-guard-4-12b $0.18 / $0.18, zlm-v1-iab-classify-edge $0.02 / $0.05, zlm-v1-signal-extract $0.04 / $0.10. Volume pricing is available. ## Output (`features.jsonl`, one line per message) ```json {"id": "p01", "chars": 54, "brand_safe": true, "safety_flags": [], "policy_flag_only": false, "iab_2_2": [{"name": "Gifts and Greetings Cards", "score": 0.962, "iab_2_2_id": 476, "iab_2_2_name": "Gifts and Greetings Cards", "via": "iab22"}, {"name": "Personal Celebrations & Life Events", "score": 0.71, "iab_2_2_id": 163, "iab_2_2_name": "Personal Celebrations & Life Events", "via": "iab22"}, "..."], "audience": [{"name": "Gift Cards and Coupons", "score": 0.628}, {"name": "Gifts and Holiday Items", "score": 0.583}, "..."], "keywords": ["last-minute gift", "gift"], "keyword_limit": 2, "latency_ms": {"safety": 742, "iab": 690, "keywords": 812, "decision": 812}, "tokens": {"safety": {"in": 14, "out": 230}, "iab": {"in": 14, "out": 318}, "keywords": {"in": 14, "out": 43}}, "cost_usd": 6.5e-05} ``` ## Script ```python #!/usr/bin/env python3 """Dry-run ZeroGPU's ad-decision models on your own messages and get the features an ad ranker would use. For each message, three ZeroGPU models run in parallel, exactly as in the demo: brand safety llama-guard-4-12b POST /v1/moderations IAB 2.2 zlm-v1-iab-classify-edge POST /v1/responses keywords zlm-v1-signal-extract POST /v1/responses Keywords are kept only when they appear in the message: up to 2 for a message under 25 words, up to 5 otherwise. Category names are checked against IAB Tech Lab's official Content Taxonomy 2.2 (downloaded from IAB's GitHub). Needs Python 3.9+ and ZEROGPU_API_KEY in the environment. No other packages. python3 zerogpu_ad_decisions.py messages.jsonl [--out features.jsonl] [--workers 8] [--limit 500] Input: .jsonl (one object per line with "text", or "question"/"message"/"content"; "id" optional), .csv (a column named text/question/message/content, or the first column), or .txt (one message per line). """ import argparse, csv, json, math, os, re, statistics, sys, time import urllib.error, urllib.request from concurrent.futures import ThreadPoolExecutor API = "https://api.zerogpu.ai/v1" MODELS = {"safety": "llama-guard-4-12b", "iab": "zlm-v1-iab-classify-edge", "keywords": "zlm-v1-signal-extract"} PRICES = {"llama-guard-4-12b": (0.18, 0.18), "zlm-v1-iab-classify-edge": (0.02, 0.05), "zlm-v1-signal-extract": (0.04, 0.10)} MAX_CHARS = 1500 # the IAB and keyword models read up to ~400 tokens IAB_URL = "https://raw.githubusercontent.com/InteractiveAdvertisingBureau/Taxonomies/main/Content%20Taxonomies/Content%20Taxonomy%20{}.tsv" def post(api_key, path, body): data = json.dumps(body).encode() for attempt in (1, 2): # one retry on gateway errors req = urllib.request.Request(API + path, data=data, headers={ "content-type": "application/json", "x-api-key": api_key, "user-agent": "zerogpu-ad-decisions/1.0"}) t0 = time.time() try: with urllib.request.urlopen(req, timeout=60) as r: return json.load(r), (time.time() - t0) * 1000 except urllib.error.HTTPError as e: if attempt == 1 and e.code in (500, 502, 503, 504, 524): continue raise RuntimeError(f"{body['model']}: HTTP {e.code}: {e.read()[:200].decode(errors='replace')}") except (urllib.error.URLError, TimeoutError) as e: if attempt == 1: continue raise RuntimeError(f"{body['model']}: {e}") def usage(u): u = u or {} return u.get("input_tokens", u.get("prompt_tokens", 0)), u.get("output_tokens", u.get("completion_tokens", 0)) def cost(model, tin, tout): pin, pout = PRICES[model] return (tin * pin + tout * pout) / 1e6 def clip(text): if len(text) <= MAX_CHARS: return text cut = text[:MAX_CHARS] return cut[:max(cut.rfind(" "), MAX_CHARS - 40)] def output_json(resp): return json.loads(resp["output"][0]["content"][0]["text"]) # ---------- IAB Content Taxonomy 2.2 ---------- def load_iab(): """name (lowercase) -> (2.2 id, 'iab22' | 'iab3', 2.2 name). IAB 3.x names map to the same id, or the nearest 2.2 parent.""" def rows(version): req = urllib.request.Request(IAB_URL.format(version), headers={"user-agent": "zerogpu-ad-decisions/1.0"}) with urllib.request.urlopen(req, timeout=30) as r: lines = r.read().decode("utf-8").splitlines() return [[c.strip() for c in line.split("\t")] for line in lines[2:] if line.strip()] v22 = {c[0]: (c[2], c[1]) for c in rows("2.2") if c[0].isdigit() and int(c[0]) < 1000} names = {name.lower(): (int(i), "iab22", name) for i, (name, _) in v22.items()} for version in ("3.0", "3.1"): table = {c[0]: c for c in rows(version) if len(c) > 2} for uid, c in table.items(): name = c[2].lower() if not name or name in names: continue cur = uid while cur and cur not in v22: cur = table.get(cur, [None, None])[1] if cur: names[name] = (int(cur), "iab3", v22[cur][0]) return names # ---------- one message ---------- def decide(api_key, pool, text, iab_names): words = len(text.split()) limit = 2 if words < 25 else 5 jobs = { "safety": pool.submit(post, api_key, "/moderations", {"model": MODELS["safety"], "input": text[:8000]}), "iab": pool.submit(post, api_key, "/responses", {"model": MODELS["iab"], "input": clip(text)}), "keywords": pool.submit(post, api_key, "/responses", {"model": MODELS["keywords"], "input": clip(text), "categories": ["signals"], "max_keywords": limit}), } row, ms, tokens, total = {}, {}, {}, 0.0 for name, job in jobs.items(): resp, ms[name] = job.result() tin, tout = usage(resp.get("usage")) tokens[name] = {"in": tin, "out": tout} total += cost(MODELS[name], tin, tout) if name == "safety": res = resp["results"][0] # Hazards with no OpenAI equivalent come back flagged with score 0: reported, not counted as unsafe. flags = [k for k, v in res.get("categories", {}).items() if v and res.get("category_scores", {}).get(k, 0) > 0] row.update(brand_safe=not flags, safety_flags=flags, policy_flag_only=bool(res.get("flagged")) and not flags) elif name == "iab": out = output_json(resp) cats = [] for c in out.get("content", {}).get("iab_2_2", []): iab_id, via, name22 = iab_names.get(c["name"].lower(), (None, None, None)) cats.append({"name": c["name"], "score": round(c["score"], 3), "iab_2_2_id": iab_id, "iab_2_2_name": name22, "via": via}) row.update(iab_2_2=cats, audience=[{"name": a["name"], "score": round(a["score"], 3)} for a in out.get("audience", [])]) else: found = output_json(resp).get("keywords", []) lines = [f" {norm(line)} " for line in text.splitlines() if line.strip()] # a keyword can't span two lines kept = [] for k in found: n = norm(k) if n and any(f" {n} " in line for line in lines) and k not in kept and len(kept) < limit: kept.append(k) row.update(keywords=kept, keyword_limit=limit) ms["decision"] = max(ms.values()) # the three calls run at the same time row.update(latency_ms={k: round(v) for k, v in ms.items()}, tokens=tokens, cost_usd=round(total, 8)) return row def norm(s): return re.sub(r"[^\w$%'-]+", " ", str(s).lower().replace("’", "'")).strip() def read_messages(path): ext = os.path.splitext(path)[1].lower() pick = lambda d: next((d[k] for k in ("text", "question", "message", "content") if d.get(k)), None) out = [] with open(path, encoding="utf-8", newline="") as f: if ext == ".jsonl": for i, line in enumerate(f): if line.strip(): d = json.loads(line) out.append((d.get("id", i + 1), pick(d) if isinstance(d, dict) else str(d))) elif ext == ".csv": reader = csv.reader(f) header = next(reader, []) col = next((i for i, h in enumerate(header) if h.strip().lower() in ("text", "question", "message", "content")), None) if col is None: out.append((1, header[0])) # no header: the first row is a message col = 0 out += [(i + 2, r[col]) for i, r in enumerate(reader) if len(r) > col] else: out = [(i + 1, line.strip()) for i, line in enumerate(f) if line.strip()] return [(i, str(t).strip()) for i, t in out if t and str(t).strip()] def main(): ap = argparse.ArgumentParser(description="Dry-run ZeroGPU's ad-decision models on your own messages.") ap.add_argument("input", help=".jsonl, .csv or .txt of messages") ap.add_argument("--out", default="features.jsonl", help="one JSON line per message (default features.jsonl)") ap.add_argument("--workers", type=int, default=8, help="messages in flight at once (default 8)") ap.add_argument("--limit", type=int, help="only the first N messages") ap.add_argument("--no-iab-check", action="store_true", help="skip downloading IAB's official lists") args = ap.parse_args() api_key = os.environ.get("ZEROGPU_API_KEY", "").strip() if not api_key: sys.exit("Set ZEROGPU_API_KEY first.") messages = read_messages(args.input)[: args.limit or None] if not messages: sys.exit("No messages found.") iab_names = {} if not args.no_iab_check: try: iab_names = load_iab() except Exception as e: # the decision doesn't need it print(f"Couldn't load IAB's official lists ({e}); category ids will be empty.") calls = ThreadPoolExecutor(max(3, args.workers * 3)) def one(item): mid, text = item try: return {"id": mid, "chars": len(text), **decide(api_key, calls, text, iab_names)} except Exception as e: return {"id": mid, "chars": len(text), "error": str(e)} rows = [] with ThreadPoolExecutor(max(1, args.workers)) as pool, open(args.out, "w", encoding="utf-8") as out: for i, r in enumerate(pool.map(one, messages), 1): out.write(json.dumps(r, ensure_ascii=False) + "\n") rows.append(r) if i % 25 == 0 or i == len(messages): print(f"{i}/{len(messages)} messages") calls.shutdown() done = [r for r in rows if "error" not in r] print(f"\n{len(done)}/{len(rows)} messages ran. Features: {args.out}") for r in rows: if "error" in r: print(f" error on {r['id']}: {r['error']}") if not done: return lat = sorted(r["latency_ms"]["decision"] for r in done) unsafe = [r for r in done if not r["brand_safe"]] cats = {} for r in done: for c in r["iab_2_2"][:1]: name = c.get("iab_2_2_name") or c["name"] cats[name] = cats.get(name, 0) + 1 checked = [c for r in done for c in r["iab_2_2"]] if iab_names else [] print(f"not brand-safe: {len(unsafe)}/{len(done)} ({100 * len(unsafe) / len(done):.1f}%)") print(f"decision latency: median {statistics.median(lat):.0f} ms · p90 {lat[max(0, math.ceil(0.9 * len(lat)) - 1)]:.0f} ms (slowest of the 3 parallel calls)") print(f"cost per 1M decisions: ${statistics.mean(r['cost_usd'] for r in done) * 1e6:.2f} at list prices") if checked: exact = sum(c["via"] == "iab22" for c in checked) print(f"IAB 2.2 names: {exact}/{len(checked)} exact, {sum(c['via'] == 'iab3' for c in checked)} IAB 3.x names mapped to 2.2") print("top categories: " + ", ".join(f"{k} ({v})" for k, v in sorted(cats.items(), key=lambda x: -x[1])[:8])) if __name__ == "__main__": main() ``` ## Contact For questions, or a dedicated key for a larger test: maddy@zerogpu.ai · https://zerogpu.ai