SimplyCodes publishes promo codes and tells shoppers which ones work. A machine now makes most of that call. It gathers evidence from real checkout attempts, decides a verdict, scores every code, and sends codes back for re-testing. Human editors still hold the last word on the codes the machine is not yet trusted with.

Our verification rulebook says the machine cannot be trusted further without an instrument. It caps the machine's authority near half of all verdicts until an answer key exists, a curated set of codes with known outcomes, and the machine scores 95 of 100 against it. The rulebook adds the rule that makes this a mission rather than an internal task: the answer key must be written by people outside our systems. Nobody who builds, runs, or grades the machine may write the key it is graded against.

You take about 400 SimplyCodes promo codes to real merchant checkouts, record what each one did with a screenshot, and hand us the sealed key.

Our goals for this mission

The key does not exist. It was due on 2026-07-10 as the rulebook's committed first move. What exists instead is a drift audit that compares two models to each other, which read 93.8% agreement when it was last verified on 2026-07-09, and an older 499-row graded panel drawn on 2026-03-10 that predates the current verdict logic and cannot grade the machine as it runs today. A separate project inside the company builds the machine-side comparison harness. This mission builds the independent key that harness is checked against.

Two of our published outcomes lean on this key. The code-accuracy outcome promises verdict accuracy of 95% or better and today has no way to measure it. The coverage outcome watches machine coverage climbing from 56.6% of active codes toward 80%, and the rulebook caps the machine's share of verdicts near half until the key exists.

This is also the shape of work we pay outsiders for. It is labor a person has to do, a real cart at a real checkout with a real code, and a machine cannot do it for us without becoming the thing being graded. It needs no access to our code, our warehouse, or our internal tools. Its output is the answer key, not software.

What you would do

  • Kickoff, in Santa Monica. The protocol walkthrough, a 30-code pilot labeled together (15 Shopify merchants, 15 not), every disagreement resolved into the method page before anyone labels alone, and the final row count set from the pilot's share of codes that can be verified before payment. We hand over five things: the seeded draw as a spreadsheet, a cart-building guide written from each code's stated terms so a small cart never falsely fails a code with a minimum, an evidence template, the machine's failure-class vocabulary, and a shared folder that lives outside every engineering system.
  • Labeling. Evidence on every row, a daily drop into the shared folder. 20% of rows are labeled twice by different people; those rows are drawn at random up front and spread across the whole window, and neither labeler knows which rows they are. Midway, we run an agreement read on the rows so far, between labelers only, never accuracy.
  • Adjudication and sealing. Adjudication of disagreements, the method page, the handoff, and sealing.

The codes are drawn by us, not chosen by you. The draw has two parts. The core, 300 rows, is a proportional random sample of the codes the machine rules on, drawn with a seed we record, held to at least 30% non-Shopify merchants because the machine's error on non-Shopify checkouts is about four times its Shopify error, and limited to merchants the machine has a verdict on record for. The supplement, 100 rows, over-represents codes near the machine's decision boundary. The two are labeled identically and reported separately. The file you receive carries none of the machine's fields. You never learn what the machine thinks.

Labor, stated honestly: a gradeable row takes 12 to 25 minutes once the cart, any account, and the required second attempt on a rejection are counted. 400 gradeable rows plus 20% double-labeling is 100 to 165 hours before retries. That is roughly two people working full time, so we contract two labelers for this mission. The agreement sets the window, the kickoff pilot sets the final count, and 300 gradeable core rows is the floor.

What success looks like

  • The deliverable is a sealed answer key. Each row records the merchant, the code, the exact host and checkout URL we asked you to test, the UTC timestamp of the final attempt, the cart that was built and its subtotal, the device, region and connection type, the outcome class, the failure class where there was one, the discount observed against the discount the code page stated, the merchant's message word for word where there was one, a link to the screenshot, the protocol version, and the labeler.
  • Five outcome classes, and nothing else. APPLIED: the discount showed on the order summary before payment. APPLIED_WITH_CONDITION: it applied only once a named condition was met (a minimum, a category, a first order, a membership); the condition is written down. REJECTED: the checkout refused it; the message is quoted. UNVERIFIABLE_BEFORE_PAYMENT: the merchant only validates at payment and no purchase was made. CHECKOUT_UNREACHABLE: a login wall, a region block, or a bot wall stopped the attempt. Only the first three classes are gradeable. The last two are reported as a share, never as evidence that a code is dead.
  • 400 gradeable rows with evidence at delivery: the 300-row core plus the 100-row hard-case supplement, never pooled.
  • The double-labeled rows agree at 0.80 or better on Gwet's AC1 (an agreement score that corrects for chance), with the interval's lower bound above 0.70, or the affected rows are re-labeled under the final protocol.
  • We compute the accuracy read. You never do. Within a week of delivery we publish the first read against the key, with the non-gradeable share and worst-case bounds beside it. At 300 rows an observed 95% carries a margin of roughly 92% to 97%.
  • Zero key rows present in any training source, verified by us at sealing.

Ground rules

  • No access to our repositories, warehouse, dashboards, or internal tools. Inputs arrive as files; outputs return as files.
  • You never see the machine's verdicts, confidence, or scores for any row.
  • Reach every merchant by typing its URL. Never click out from a SimplyCodes, Knoji, or Dealspotr page; those links fire affiliate clicks and pollute our revenue data. All stated terms come from the handed spreadsheet, not the live code page.
  • A clean browser profile with zero extensions of any kind, a fresh cookie jar per merchant, a residential connection, no VPN. The SimplyCodes extension is never installed on the labeling device.
  • Test the exact host in the row. A host mismatch is a void row, not a disagreement.
  • A row without a screenshot and a timestamp does not exist. Screenshots are redacted of personal data before delivery.
  • No more than a handful of attempts per merchant per day. One account per merchant, using an identity dedicated to this work. No scripted or automated attempts of any kind.
  • No purchases unless a purchase budget is set in writing before kickoff. The default is none.
  • Single-use codes are excluded from the draw.
  • The sealed key never enters a training source. We verify that at sealing.
  • The accuracy read is computed by us and published by us. You deliver labels, not a score.
  • Scope, timing, and pay are set in the contractor agreement, at your stated rate, with an NDA that covers merchant terms of service, data handling, and indemnity. The codes are public; the key is not.

What happens after you send a proposal

3 steps, and a person reads every proposal.

  1. You send a short written proposal. It covers how you understand this mission and how you would approach it, and you attach your resume or a LinkedIn profile.
  2. We read it and reach out with questions. There is no test and no exercise.
  3. If there’s a fit, the mission begins under a contractor agreement and an NDA, paid at your stated rate, with kickoff in Santa Monica.